THE SENSITIVITY ANALYSIS PARADOX IN TARGET TRIAL EMULATION: AN EMPIRICAL TAXONOMY OF BIAS MITIGATION AND PROTOCOL DESIGN IN 208 STUDIES
Author(s)
ALBAN FABRE, PhD, MPH, MSc1, Sarah Rosen, MSc2, Alice ROULEAU, MASc, MBA, MPH, PharmD2.
1Thermo Fisher Scientific, Barcelona, Spain, 2Thermo Fisher Scientific, Ivry-sur-Seine, France.
1Thermo Fisher Scientific, Barcelona, Spain, 2Thermo Fisher Scientific, Ivry-sur-Seine, France.
OBJECTIVES: Health Technology Assessments (HTAs) increasingly rely on Target Trial Emulations (TTEs) to populate economic models with real-world effectiveness estimates. While guidelines emphasize reporting completeness, validity depends on robust operational execution. This study establishes an empirical taxonomy of sensitivity practices and protocol designs across 208 TTEs to evaluate whether reported sensitivity testing safeguards HTA decisions.
METHODS: We conducted a systematic literature review and structural audit of 208 peer-reviewed TTE studies (2018-2025). Categories were derived from causal-inference bias domains and assigned by the reported function of each sensitivity analysis: Patient-Level Confounding Adjustments for balancing measured baseline covariates, Structural Timeline Validations for validating timeline/eligibility structure, and Unmeasured Bias Quantifiers for quantifying residual/unmeasured bias. Structural execution errors were defined as mismatches between target-trial specification and observational implementation in time zero, eligibility, treatment assignment, or follow-up. Testing strategies and design components, including estimand type, database type, and protocol mapping tables, were cross-tabulated against audited protocol-to-emulation fidelity.
RESULTS: Overall, 93.3% (194/208) of TTEs reported bias-focused sensitivity testing. However, a profound paradox emerged: 61.1% (127/208) relied exclusively on baseline patient-level matching or alternative cohort definitions, approaches that do not directly evaluate structural execution errors. Only 26.0% (54/208) applied unmeasured bias quantifiers. In protocol design, pure Intention-to-Treat (ITT) strategies exhibited higher structural error rates than Per-Protocol or dual-estimand designs (61.9% vs. 34.7%), driven by eligibility-alignment drift. A formal protocol-to-emulation mapping table was associated with lower structural error rates, reducing errors from 60.6% to 38.7%.
CONCLUSIONS: High procedural compliance in TTE sensitivity testing masks a conceptual blind spot. Most reported sensitivity analyses adjust for baseline patient selection rather than validating the structural integrity of the data emulation timeline. For HTA bodies, assertions of “sensitivity testing” may overstate robustness; payers can optimize appraisal by moving beyond binary checklists to prioritize dossiers utilizing explicit side-by-side protocol mapping tables and dual-estimand designs.
METHODS: We conducted a systematic literature review and structural audit of 208 peer-reviewed TTE studies (2018-2025). Categories were derived from causal-inference bias domains and assigned by the reported function of each sensitivity analysis: Patient-Level Confounding Adjustments for balancing measured baseline covariates, Structural Timeline Validations for validating timeline/eligibility structure, and Unmeasured Bias Quantifiers for quantifying residual/unmeasured bias. Structural execution errors were defined as mismatches between target-trial specification and observational implementation in time zero, eligibility, treatment assignment, or follow-up. Testing strategies and design components, including estimand type, database type, and protocol mapping tables, were cross-tabulated against audited protocol-to-emulation fidelity.
RESULTS: Overall, 93.3% (194/208) of TTEs reported bias-focused sensitivity testing. However, a profound paradox emerged: 61.1% (127/208) relied exclusively on baseline patient-level matching or alternative cohort definitions, approaches that do not directly evaluate structural execution errors. Only 26.0% (54/208) applied unmeasured bias quantifiers. In protocol design, pure Intention-to-Treat (ITT) strategies exhibited higher structural error rates than Per-Protocol or dual-estimand designs (61.9% vs. 34.7%), driven by eligibility-alignment drift. A formal protocol-to-emulation mapping table was associated with lower structural error rates, reducing errors from 60.6% to 38.7%.
CONCLUSIONS: High procedural compliance in TTE sensitivity testing masks a conceptual blind spot. Most reported sensitivity analyses adjust for baseline patient selection rather than validating the structural integrity of the data emulation timeline. For HTA bodies, assertions of “sensitivity testing” may overstate robustness; payers can optimize appraisal by moving beyond binary checklists to prioritize dossiers utilizing explicit side-by-side protocol mapping tables and dual-estimand designs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR157
Topic
Methodological & Statistical Research
Topic Subcategory
Confounding, Selection Bias Correction, Causal Inference
Disease
No Additional Disease & Conditions/Specialized Treatment Areas