USING CROSS-VALIDATION TO ASSESS MODEL SUITABILITY FOR NEW DECISION PROBLEMS: A CASE STUDY IN PREDIABETES
Author(s)
James Altunkaya, DPhil, Philip Clarke, PhD, Amanda Adler, PhD, MD, José Leal, DPhil.
University of Oxford, Oxford, United Kingdom.
University of Oxford, Oxford, United Kingdom.
OBJECTIVES: Health economic models are frequently reused beyond populations in which risk equations were estimated. This creates uncertainty about whether an available model is fit for a target decision problem, particularly in areas of substantial clinical heterogeneity. We present a case study in prediabetes examining how cross-validation can support assessment of population suitability for value-assessment within disease-specific reference models.
METHODS: Two new individual-level reference models of prediabetes were derived from DPP/DPPOS (n=3,081; 21-year follow-up) and NAVIGATOR data (n=9,306; 6-year follow-up). Both models used a common structure and coding implementation. Random-effects equations predicted annual risk-factor progression, whilst parametric time-to-event equations predicted diabetes, cardiovascular events, and mortality using lagged risk factors and time-updated event histories. A stochastic microsimulation operationalised equations to estimate patient costs and outcomes. Internal validation assessed calibration and discrimination, whilst a combination of cross-validation and further external data was used to diagnose whether lack of transportability reflected differences in baseline populations, risk-factor trajectories, or event-risk relationships.
RESULTS: Derivation and validation populations differed substantially in age, sex, ethnicity, body mass index, cardiovascular history, follow-up, and observed event rates. Whilst common model structures improved transparency and comparability, this was insufficient to ensure equivalent external performance. Cross-validation distinguished whether model prediction error arose from differences in risk-factor trajectories or event-risk relationships, identifying areas where calibration could improve model generalisability.
CONCLUSIONS: Universal disease reference models remain a worthwhile but challenging aspiration in HTA. In the interim, the suitability of individual models for each decision problem could readily be made more explicit. Structured cross-validation can indicate when existing models are credible for a target population, and where recalibration, re-estimation, or alternative model selections are needed to best inform reimbursement decisions. Our analysis suggests model suitability is best judged by actively testing transportability of models’ linked prediction mechanisms, rather than assessing overlap in target and model derivation populations alone.
METHODS: Two new individual-level reference models of prediabetes were derived from DPP/DPPOS (n=3,081; 21-year follow-up) and NAVIGATOR data (n=9,306; 6-year follow-up). Both models used a common structure and coding implementation. Random-effects equations predicted annual risk-factor progression, whilst parametric time-to-event equations predicted diabetes, cardiovascular events, and mortality using lagged risk factors and time-updated event histories. A stochastic microsimulation operationalised equations to estimate patient costs and outcomes. Internal validation assessed calibration and discrimination, whilst a combination of cross-validation and further external data was used to diagnose whether lack of transportability reflected differences in baseline populations, risk-factor trajectories, or event-risk relationships.
RESULTS: Derivation and validation populations differed substantially in age, sex, ethnicity, body mass index, cardiovascular history, follow-up, and observed event rates. Whilst common model structures improved transparency and comparability, this was insufficient to ensure equivalent external performance. Cross-validation distinguished whether model prediction error arose from differences in risk-factor trajectories or event-risk relationships, identifying areas where calibration could improve model generalisability.
CONCLUSIONS: Universal disease reference models remain a worthwhile but challenging aspiration in HTA. In the interim, the suitability of individual models for each decision problem could readily be made more explicit. Structured cross-validation can indicate when existing models are credible for a target population, and where recalibration, re-estimation, or alternative model selections are needed to best inform reimbursement decisions. Our analysis suggests model suitability is best judged by actively testing transportability of models’ linked prediction mechanisms, rather than assessing overlap in target and model derivation populations alone.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA164
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Decision & Deliberative Processes
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Diabetes/Endocrine/Metabolic Disorders (including obesity), No Additional Disease & Conditions/Specialized Treatment Areas