NEUTRAL COMPARISON STUDY OF CAUSAL MACHINE LEARNING METHODS FOR ESTIMATING INDIVIDUALISED TREATMENT EFFECTS IN ECONOMIC EVALUATIONS
Author(s)
Aisha Moolla, MD, MPH.
MD | Health Economics | Causal Machine Learning, ScHARR, The University of Sheffield, Sheffield, United Kingdom.
MD | Health Economics | Causal Machine Learning, ScHARR, The University of Sheffield, Sheffield, United Kingdom.
OBJECTIVES: Economic evaluations underpinning health technology assessment (HTA) typically report average treatment effects, masking heterogeneity that can render a cost-ineffective intervention favourable in some subgroups and harmful in others. Causal machine learning (ML) methods can estimate individual treatment effects (ITEs) without strong parametric assumptions. However, their comparative performance in HTA-relevant settings, with skewed cost and utility outcomes and a need for uncertainty estimation, remains largely unevaluated. This study performs a neutral comparison of 11 ITE estimation methods identified from recent systematic reviews.
METHODS: A modular data-generating process, extending an established cost-effectiveness (CEA) forest simulation framework, generated synthetic patient-level QALY and cost data across seven scenarios varying covariate dimensionality, functional complexity, noise, and collider structures, crossed with four sample sizes (500-100,000) and five confounding strengths. Methods were assessed against oracle ITEs for QALYs, costs, and net monetary benefit (NMB) using precision in estimating heterogeneous effects (PEHE), which quantifies the mean squared error between estimated and true ITEs; bias, sign (decision) accuracy; 95% confidence-interval coverage; and variable importance.
RESULTS: Benchmarking is complete for CEA forest, the most widely used ITE method in cost-effectiveness contexts. PEHE ranged from £6,748 (Scenario 1, base case) to £12,763 (Scenario 2, quadratic cost), and declined sharply with sample size for every scenario, falling from a mean of £16,395 (n = 500) to £4,332 (n = 100,000). QALY sign accuracy was near 100% except in the scenario containing a genuinely harmed subgroup. NMB decision accuracy was statistically significant only at n≥10,000. Across scenarios, coverage ranged from 63-89%, indicating systematic undercoverage. Full cross-method results will be available at the time of presentation.
CONCLUSIONS: Simulation benchmarking reveals sample-size-dependent failure points in cost-effectiveness classification and interval calibration that are directly relevant to HTA decision-making. This comparison provides evidence-based guidance for selecting ITE estimation methods for economic evaluation.
METHODS: A modular data-generating process, extending an established cost-effectiveness (CEA) forest simulation framework, generated synthetic patient-level QALY and cost data across seven scenarios varying covariate dimensionality, functional complexity, noise, and collider structures, crossed with four sample sizes (500-100,000) and five confounding strengths. Methods were assessed against oracle ITEs for QALYs, costs, and net monetary benefit (NMB) using precision in estimating heterogeneous effects (PEHE), which quantifies the mean squared error between estimated and true ITEs; bias, sign (decision) accuracy; 95% confidence-interval coverage; and variable importance.
RESULTS: Benchmarking is complete for CEA forest, the most widely used ITE method in cost-effectiveness contexts. PEHE ranged from £6,748 (Scenario 1, base case) to £12,763 (Scenario 2, quadratic cost), and declined sharply with sample size for every scenario, falling from a mean of £16,395 (n = 500) to £4,332 (n = 100,000). QALY sign accuracy was near 100% except in the scenario containing a genuinely harmed subgroup. NMB decision accuracy was statistically significant only at n≥10,000. Across scenarios, coverage ranged from 63-89%, indicating systematic undercoverage. Full cross-method results will be available at the time of presentation.
CONCLUSIONS: Simulation benchmarking reveals sample-size-dependent failure points in cost-effectiveness classification and interval calibration that are directly relevant to HTA decision-making. This comparison provides evidence-based guidance for selecting ITE estimation methods for economic evaluation.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR100
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Confounding, Selection Bias Correction, Causal Inference
Disease
No Additional Disease & Conditions/Specialized Treatment Areas