IMPORTANT CONSIDERATIONS IN THE SELECTION OF AI TOOLS TO SUPPORT HEALTH ECONOMIC MODEL REPLICATION

Author(s)

Alex Hirst, BSc, MSc, Laith Yakob, Doctorate, MSc.
Adelphi Values PROVE, Bollington, United Kingdom.
OBJECTIVES: The use of AI to support programming is well established; however, the choice of AI model remains an evolving decision. This research compared the performance of two leading large language models (LLMs) in replicating published economic models, to understand the role model selection may play in reproduction accuracy.
METHODS: Two published economic models were selected to span differing structures and therapeutic contexts: a decision tree model in pain treatment, and an oncology partitioned survival model (PSM) with comparison to three-state and five-state state-transition model (STM) structures. Each publication was provided to two LLMs (Claude Opus 4.8 and OpenAI GPT-5.5) using an identical prompt, with clear instruction on context and intended modelling approach.
RESULTS: Both models produced running implementations. GPT-5.5 introduced one programming error (a recursive issue) that initially prevented the oncology model from running, but resolved it on prompting. As anticipated, both LLMs replicated the simpler decision tree model more accurately than the oncology model. Total QALYs matched the published values almost exactly (~100%) for both LLMs across models. In the pain model, Claude reproduced published costs to within ~0.1% (a difference of under £1 per patient), whereas GPT-5.5 underestimated costs by approximately 10-18% (tramadol 90.3%, buprenorphine 81.7% of published) due to incorrectly interpreting inputs for the duration of therapy. For the oncology model, replication of incremental QALYs varied by structure: under the PSM, Claude achieved 97.7% versus 87.3% for GPT-5.5; under the three-state STM, 105.8% versus 82.5%; and under the five-state STM, 98.0% versus 103.7%.
CONCLUSIONS: Both LLMs reproduced published models with reasonable accuracy, but performance diverged with model complexity, particularly for cost estimation and incremental outcomes in the oncology model. Findings suggest that AI model selection is a material consideration when using LLMs to support health economic model replication and that outputs should be validated.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

EE515

Topic

Economic Evaluation

Disease

Musculoskeletal Disorders (Arthritis, Bone Disorders, Osteoporosis, Other Musculoskeletal), Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×