CAN A LARGE LANGUAGE MODEL QUALITY CONTROL A HEALTH-ECONOMIC MODEL? RELIABILITY AND FAILURE MODES OF A LARGE LANGUAGE MODEL AS A QUALITY-CONTROL REVIEWER OF HEALTH-ECONOMIC MODELS
Author(s)
Daniel Ribes, MSc1, Agota Szende, MSc, PhD2, Frank Cinfio, BSc3, Alasdair D. Henry, PhD2, Marko Zivkovic, PhD4.
1Genesis Research Group, Barcelona, Spain, 2Genesis Research Group, London, United Kingdom, 3Genesis Research Group, Indianapolis, IN, USA, 4Genesis Research, LLC, Hoboken, NJ, USA.
1Genesis Research Group, Barcelona, Spain, 2Genesis Research Group, London, United Kingdom, 3Genesis Research Group, Indianapolis, IN, USA, 4Genesis Research, LLC, Hoboken, NJ, USA.
OBJECTIVES: Large-language-models (LLMs) are increasingly proposed for quality control (QC) of cost-effectiveness models, yet their reliability and blind spots are poorly characterized. Rather than asking whether an LLM can QC a model, we characterized its detection accuracy, reproducibility, and failure modes.
METHODS: We seeded 72 errors (syntactic/structural, hardcoding, referencing, calculation logic, results/interpretation, transcription/units, behavioral invariants) into a real, de-identified partitioned-survival oncology model (24 sheets, ~91,000 formula cells), one per workbook plus one containing all 72. A general-purpose LLM (Claude Opus 4.8) reviewed each complete model in three independent rounds using one model with plain prompts without additional tuning or tooling to show baseline performance. We profiled recall by category, reproducibility, and failure structure.
RESULTS: Reviewing one error per model, the LLM detected 58/72 by majority of three rounds (81%) and 64/72 by union (89%); single-round recall ranged 56-60/72 (78-83%), and 59/72 faults received identical verdicts across all three rounds. Detection was near-complete for calculation-logic, syntactic and hardcoding faults and weakest for transcription/units. The 14 undetected faults were systematic, not random: in every case the injected cell was a valid formula returning a believable number - source-dependent values, signal-free label/unit errors, peerless isolated cells, and locally plausible wrong-reference results. Each review returned a median of 8 flagged cells. Under dense review, recall fell to 33/72 (46%) by majority and 51/72 (71%) by union with 7 false positives.
CONCLUSIONS: We demonstrated that a general-purpose LLM, used simply and without tuning, can identify most seeded errors. However, it missed a predictable few - valid formulas returning believable numbers, verifiable only against an external source. Although giving the model the source documents or using more advanced setups would narrow the gap, human QC remains essential.
METHODS: We seeded 72 errors (syntactic/structural, hardcoding, referencing, calculation logic, results/interpretation, transcription/units, behavioral invariants) into a real, de-identified partitioned-survival oncology model (24 sheets, ~91,000 formula cells), one per workbook plus one containing all 72. A general-purpose LLM (Claude Opus 4.8) reviewed each complete model in three independent rounds using one model with plain prompts without additional tuning or tooling to show baseline performance. We profiled recall by category, reproducibility, and failure structure.
RESULTS: Reviewing one error per model, the LLM detected 58/72 by majority of three rounds (81%) and 64/72 by union (89%); single-round recall ranged 56-60/72 (78-83%), and 59/72 faults received identical verdicts across all three rounds. Detection was near-complete for calculation-logic, syntactic and hardcoding faults and weakest for transcription/units. The 14 undetected faults were systematic, not random: in every case the injected cell was a valid formula returning a believable number - source-dependent values, signal-free label/unit errors, peerless isolated cells, and locally plausible wrong-reference results. Each review returned a median of 8 flagged cells. Under dense review, recall fell to 33/72 (46%) by majority and 51/72 (71%) by union with 7 false positives.
CONCLUSIONS: We demonstrated that a general-purpose LLM, used simply and without tuning, can identify most seeded errors. However, it missed a predictable few - valid formulas returning believable numbers, verifiable only against an external source. Although giving the model the source documents or using more advanced setups would narrow the gap, human QC remains essential.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR43
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas