VALIDATING GENERATIVE AI (CLAUDE-IN-EXCEL ADD-IN) FOR BUILDING HEALTH ECONOMIC MODELS IN EXCEL FROM PUBLISHED ARTICLES
Author(s)
Subhajit Gupta, MSc1, Craig van Rensburg, MComm Economics2, Marc Bollée, BA (Math)2, Anns Thomas, MSc3, Satarupa Talukdar, MSc3, Ray Gani, PhD4.
1PharmaQuant, Bangalore, India, 2PharmaQuant, Dublin, Ireland, 3PharmaQuant, Kolkata, India, 4PharmaQuant, London, United Kingdom.
1PharmaQuant, Bangalore, India, 2PharmaQuant, Dublin, Ireland, 3PharmaQuant, Kolkata, India, 4PharmaQuant, London, United Kingdom.
OBJECTIVES: Generative AI (GenAI) is increasingly proposed for cost-effectiveness model (CEM) construction, yet its accuracy and reliability remain unestablished. Here we develop a framework to evaluate GenAI CEM rebuilds and use it on the Claude-in-Excel add-in, to see how reliably Claude rebuilds published CEMs in Excel using manuscripts and supplementary materials.
METHODS: Four published CEMs spanning different therapeutic areas and Markov/partitioned survival structures were rebuilt in Excel using Claude. For each, the primary publication and supplements were uploaded and a fixed prompt instructed Claude to extract all parameters and build the model. This was repeated three times per model to assess reproducibility. The CEM was then refined using additional prompts until marginal improvements were less than manual programming alone by an experienced health economist. Each version was assessed using a framework that scored extraction and build accuracy.
RESULTS: Claude rebuilt models modularly and generally produced sound, traceable formulas. However, structural and calculation choices (e.g., cycle length, half-cycle correction, discounting and state-transition logic) were applied inconsistently. Input extraction accuracy exceeded 90% but critical errors occurred (e.g., survival parameters, misuse of inputs). Claude did independently identify reporting omissions in source publications. Output accuracy varied with model complexity and data availability — up to 66% for incremental QALYs. Where reporting was incomplete, Claude approximated reasonable assumptions following user prompting. Marginal improvements plateaued after around 6 complex prompts, requiring 4-8 hours time from an experienced modeler from start to finish.
CONCLUSIONS: The proposed evaluation framework demonstrated that Claude-in-Excel is a capable assistant for building CEMs using GenAI, with the potential to significantly reduce model build timelines. However, it is not autonomous, and an experienced health economist is required for supervision and to maximize its utility. Over-reliance on GenAI can introduce significant errors with detailed expert review and refinement remaining critical throughout the process.
METHODS: Four published CEMs spanning different therapeutic areas and Markov/partitioned survival structures were rebuilt in Excel using Claude. For each, the primary publication and supplements were uploaded and a fixed prompt instructed Claude to extract all parameters and build the model. This was repeated three times per model to assess reproducibility. The CEM was then refined using additional prompts until marginal improvements were less than manual programming alone by an experienced health economist. Each version was assessed using a framework that scored extraction and build accuracy.
RESULTS: Claude rebuilt models modularly and generally produced sound, traceable formulas. However, structural and calculation choices (e.g., cycle length, half-cycle correction, discounting and state-transition logic) were applied inconsistently. Input extraction accuracy exceeded 90% but critical errors occurred (e.g., survival parameters, misuse of inputs). Claude did independently identify reporting omissions in source publications. Output accuracy varied with model complexity and data availability — up to 66% for incremental QALYs. Where reporting was incomplete, Claude approximated reasonable assumptions following user prompting. Marginal improvements plateaued after around 6 complex prompts, requiring 4-8 hours time from an experienced modeler from start to finish.
CONCLUSIONS: The proposed evaluation framework demonstrated that Claude-in-Excel is a capable assistant for building CEMs using GenAI, with the potential to significantly reduce model build timelines. However, it is not autonomous, and an experienced health economist is required for supervision and to maximize its utility. Over-reliance on GenAI can introduce significant errors with detailed expert review and refinement remaining critical throughout the process.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EE656
Topic
Economic Evaluation, Methodological & Statistical Research, Study Approaches