RAPID AND HIGH-FIDELITY REPLICATION OF ECONOMIC MODELS USING AN LLM-ENABLED MODEL BUILDING PLATFORM
Author(s)
Jordan Amdahl, BS1, Reece Grindley, MSc2.
1Director of Product, Avalere Health, London, United Kingdom, 2Avalere Health, London, United Kingdom.
1Director of Product, Avalere Health, London, United Kingdom, 2Avalere Health, London, United Kingdom.
OBJECTIVES: Developing cost-effectiveness models (CEMs) in R already provides great efficiencies over traditional Excel-based models. Now, with the integration of large language models (LLMs) and artificial intelligence (AI) agents into R modelling workflows, CEMs can be developed faster to enhance analyses. This study evaluated the performance of SensePredict(tm), a proprietary AI tool for cost-effectiveness modelling.
METHODS: A total of seven economic models were used in the replication study. Three were selected from a non-systematic literature review based on quality of reporting and overall replicability. An additional four example models were created to cover specific challenging cases. Of the seven models, four were Markov cohort models (MCMs) and three were partitioned survival models (PSMs). For each model, a text prompt was created for the AI tool, containing only the minimal information needed to replicate. Results for total life-years (LYs), total quality-adjusted life-years (QALYs), total costs, and net monetary benefit (NMB) from replicated models were compared with the original published results to derive a percentage error rate.
RESULTS: For LYs and QALYs, error rates were low across all replicated models with an average percentage error of 0.1% (range: 0.0% - 1.1%). Total costs showed a slightly higher percentage error of 0.6% (range: 0.0% - 5.9%). For NMB, the average percentage error was 2.0% (range: 0.0% - 19.0%). Replication of NMB was more accurate for MCMs than PSMs (0.4% vs. 3.7%), though this difference was driven entirely by a single outlier result.
CONCLUSIONS: These results demonstrate the ability of a proprietary LLM-enabled workflow to accurately support CEM development. As with all AI-assisted activities in HEOR, the direction and supervision of the expert is foundational. Further research is needed to explore the performance on a larger dataset of published models of varying complexity and to support tasks adjacent to model programming such as estimation and reporting.
METHODS: A total of seven economic models were used in the replication study. Three were selected from a non-systematic literature review based on quality of reporting and overall replicability. An additional four example models were created to cover specific challenging cases. Of the seven models, four were Markov cohort models (MCMs) and three were partitioned survival models (PSMs). For each model, a text prompt was created for the AI tool, containing only the minimal information needed to replicate. Results for total life-years (LYs), total quality-adjusted life-years (QALYs), total costs, and net monetary benefit (NMB) from replicated models were compared with the original published results to derive a percentage error rate.
RESULTS: For LYs and QALYs, error rates were low across all replicated models with an average percentage error of 0.1% (range: 0.0% - 1.1%). Total costs showed a slightly higher percentage error of 0.6% (range: 0.0% - 5.9%). For NMB, the average percentage error was 2.0% (range: 0.0% - 19.0%). Replication of NMB was more accurate for MCMs than PSMs (0.4% vs. 3.7%), though this difference was driven entirely by a single outlier result.
CONCLUSIONS: These results demonstrate the ability of a proprietary LLM-enabled workflow to accurately support CEM development. As with all AI-assisted activities in HEOR, the direction and supervision of the expert is foundational. Further research is needed to explore the performance on a larger dataset of published models of varying complexity and to support tasks adjacent to model programming such as estimation and reporting.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EE608
Topic
Economic Evaluation
Disease
No Additional Disease & Conditions/Specialized Treatment Areas