RAPID AND HIGH-FIDELITY REPLICATION OF ECONOMIC MODELS USING AN LLM-ENABLED MODEL BUILDING PLATFORM

Author(s)

Jordan Amdahl, BS1, Reece Grindley, MSc2.
1Director of Product, Avalere Health, London, United Kingdom, 2Avalere Health, London, United Kingdom.
OBJECTIVES: Developing cost-effectiveness models (CEMs) in R already provides great efficiencies over traditional Excel-based models. Now, with the integration of large language models (LLMs) and artificial intelligence (AI) agents into R modelling workflows, CEMs can be developed faster to enhance analyses. This study evaluated the performance of SensePredict(tm), a proprietary AI tool for cost-effectiveness modelling.
METHODS: A total of seven economic models were used in the replication study. Three were selected from a non-systematic literature review based on quality of reporting and overall replicability. An additional four example models were created to cover specific challenging cases. Of the seven models, four were Markov cohort models (MCMs) and three were partitioned survival models (PSMs). For each model, a text prompt was created for the AI tool, containing only the minimal information needed to replicate. Results for total life-years (LYs), total quality-adjusted life-years (QALYs), total costs, and net monetary benefit (NMB) from replicated models were compared with the original published results to derive a percentage error rate.
RESULTS: For LYs and QALYs, error rates were low across all replicated models with an average percentage error of 0.1% (range: 0.0% - 1.1%). Total costs showed a slightly higher percentage error of 0.6% (range: 0.0% - 5.9%). For NMB, the average percentage error was 2.0% (range: 0.0% - 19.0%). Replication of NMB was more accurate for MCMs than PSMs (0.4% vs. 3.7%), though this difference was driven entirely by a single outlier result.
CONCLUSIONS: These results demonstrate the ability of a proprietary LLM-enabled workflow to accurately support CEM development. As with all AI-assisted activities in HEOR, the direction and supervision of the expert is foundational. Further research is needed to explore the performance on a larger dataset of published models of varying complexity and to support tasks adjacent to model programming such as estimation and reporting.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

EE608

Topic

Economic Evaluation

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×