A HUMAN-IN-THE-LOOP VALIDATION INTERFACE FOR SCALABLE QUALITY ASSURANCE OF OMOP DATABASES
Author(s)
Lucas Sterckx, PhD, Narges Farokhshad, MSc, Stephanie Vandeput, MSc, Siddharth Agarwal, MSc, Dries Hens, MD.
LynxCare, Leuven, Belgium.
LynxCare, Leuven, Belgium.
OBJECTIVES: Quality assurance of NLP-extracted clinical concepts for OMOP-compliant databases is time-consuming and resource-intensive. Every processing risks introducing new errors, the effort to scan for mistakes remains constant even as extraction quality improves, and many workflows emphasize errors while failing to quantify overall review effort or accuracy. We developed an expert-validator interface, augmented by intelligent sampling to reduce validation time while producing statistically sound quality metrics.
METHODS: A validation interface serves validators batches of multi-lingual NLP extractions, one datapoint at a time, rather than one exhaustive dashboard. A validation service tracks per-datapoint progress and selects samples using multiple configurable strategies (volume-optimized, uncertainty-based, and embedding-based diversity sampling). Reviewers accept or reject extractions and provide additional feedback on common errors like underspecified codes or wrong contextual disambiguation, as well as adjust attributes (negated, historical, hypothetical, experiencer). A context-expansion toggle (sentence, list-item, section, document) provides additional context when needed. All approvals and corrections feed a centralized annotation store and analytics dashboard showing acceptance, rejection, and NLP precision per datapoint.
RESULTS: Validators reviewed 86 concepts/hour using the interface versus 16 concepts/hour with a prior generic dashboard focused only on error flagging (a 5-fold increase). Because clinical text is highly redundant, the system propagates a single accept/reject decision to semantically similar concepts within clusters, reducing annotations needed for representative coverage by 70%. Approvals are logged so each extraction is validated only once, focusing review on new or changed concepts across reprocessing runs. Also capturing the true-positive concepts, not just errors, mitigates training bias toward edge cases and yields a reusable benchmark set. Metrics enable reporting and progress demonstration.
CONCLUSIONS: Combining a purpose-built validation interface with intelligent sampling and approval logging transforms ad-hoc error-hunting into measurable, scalable quality assurance. This makes continuous quality monitoring feasible for resource-constrained OMOP implementations and supports reliable observational research, particularly in multilingual data networks.
METHODS: A validation interface serves validators batches of multi-lingual NLP extractions, one datapoint at a time, rather than one exhaustive dashboard. A validation service tracks per-datapoint progress and selects samples using multiple configurable strategies (volume-optimized, uncertainty-based, and embedding-based diversity sampling). Reviewers accept or reject extractions and provide additional feedback on common errors like underspecified codes or wrong contextual disambiguation, as well as adjust attributes (negated, historical, hypothetical, experiencer). A context-expansion toggle (sentence, list-item, section, document) provides additional context when needed. All approvals and corrections feed a centralized annotation store and analytics dashboard showing acceptance, rejection, and NLP precision per datapoint.
RESULTS: Validators reviewed 86 concepts/hour using the interface versus 16 concepts/hour with a prior generic dashboard focused only on error flagging (a 5-fold increase). Because clinical text is highly redundant, the system propagates a single accept/reject decision to semantically similar concepts within clusters, reducing annotations needed for representative coverage by 70%. Approvals are logged so each extraction is validated only once, focusing review on new or changed concepts across reprocessing runs. Also capturing the true-positive concepts, not just errors, mitigates training bias toward edge cases and yields a reusable benchmark set. Metrics enable reporting and progress demonstration.
CONCLUSIONS: Combining a purpose-built validation interface with intelligent sampling and approval logging transforms ad-hoc error-hunting into measurable, scalable quality assurance. This makes continuous quality monitoring feasible for resource-constrained OMOP implementations and supports reliable observational research, particularly in multilingual data networks.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
RWD188
Topic
Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Data Protection, Integrity, & Quality Assurance
Disease
No Additional Disease & Conditions/Specialized Treatment Areas