BAYESIAN LATENT CLASS MODELS FOR VALIDATING PHENOTYPING ALGORITHMS IN MEDICO-ADMINISTRATIVE DATABASES: A SYSTEMATIC REVIEW WITH FOCUS ON CASE-ONLY GOLD STANDARDS

Author(s)

Léo Petitdidier, BSc, Clara Bouvard, MSc, Nathanael SEDMAK, MSc, Tiphaine Porte, MSc, Morgane Swital, PhD, Audrey Lajoinie, PharmD, PhD.
RCTs, Lyon, France.
OBJECTIVES: Algorithms identifying patients from routinely collected health data use indirect indicators of disease status, leading to misclassification and require validation. This evaluation is constrained by gold standards restricted to confirmed cases, leaving non-cases unidentified and limiting estimation of key metrics, particularly specificity and negative predictive value. This review assesses the use of Bayesian latent class models (BLCMs) for algorithm validation in medico-administrative databases, focusing on settings with only positive cases (e.g., registries, clinical cohorts).
METHODS: A structured literature search was conducted on PubMed to identify studies using BLCMs to validate phenotyping algorithms in medico-administrative databases with gold standards restricted to confirmed cases. The equation combined keywords (BLCMs, healthcare database, gold standard, algorithms). Titles and abstracts were screened, followed by full-text review. Performance measures (e.g., sensitivity, specificity) were extracted.
RESULTS: Six studies using BLCMs in administrative health databases were found, with incomplete or absent gold standards focusing on case definitions over strict algorithm validation. Three applications for BLCMs emerged: (i) direct inference of latent disease status (n=4), (ii) construction of a latent standard for model evaluation (n=1), and (iii) evaluation of predefined algorithms using BLCM-derived reference (n=6). In one study, BLCMs improved sensitivity (95.9% vs 91.9%) and specificity (99.7% vs 90.8%) compared to rule-based approaches, with more accurate prevalence (3.5% vs 0.3%-1.5%). Another study reported higher prevalence (13.6% vs ~4%-5%) using BLCMs, improving case identification by integrating multiple data sources.
CONCLUSIONS: BLCMs provide a framework for validating phenotyping algorithms with incomplete gold standards, improving performance by accounting for misclassification assuming conditional independence between data sources. Although BLCMs have been applied in medico-administrative databases primarily in the absence of gold standards, their use is broader in diagnostic test evaluation across diverse settings. Validation of phenotyping algorithms with BLCMs in medico-administrative databases requires more studies, particularly with case-only gold standards.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

SA41

Topic

Study Approaches

Topic Subcategory

Literature Review & Synthesis

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×