CAN A LARGE LANGUAGE MODEL QUALITY CONTROL A HEALTH-ECONOMIC MODEL? RELIABILITY AND FAILURE MODES OF A LARGE LANGUAGE MODEL AS A QUALITY-CONTROL REVIEWER OF HEALTH-ECONOMIC MODELS

Author(s)

Daniel Ribes, MSc1, Agota Szende, MSc, PhD2, Frank Cinfio, BSc3, Alasdair D. Henry, PhD2, Marko Zivkovic, PhD4.
1Genesis Research Group, Barcelona, Spain, 2Genesis Research Group, London, United Kingdom, 3Genesis Research Group, Indianapolis, IN, USA, 4Genesis Research, LLC, Hoboken, NJ, USA.
OBJECTIVES: Large-language-models (LLMs) are increasingly proposed for quality control (QC) of cost-effectiveness models, yet their reliability and blind spots are poorly characterized. Rather than asking whether an LLM can QC a model, we characterized its detection accuracy, reproducibility, and failure modes.
METHODS: We seeded 72 errors (syntactic/structural, hardcoding, referencing, calculation logic, results/interpretation, transcription/units, behavioral invariants) into a real, de-identified partitioned-survival oncology model (24 sheets, ~91,000 formula cells), one per workbook plus one containing all 72. A general-purpose LLM (Claude Opus 4.8) reviewed each complete model in three independent rounds using one model with plain prompts without additional tuning or tooling to show baseline performance. We profiled recall by category, reproducibility, and failure structure.
RESULTS: Reviewing one error per model, the LLM detected 58/72 by majority of three rounds (81%) and 64/72 by union (89%); single-round recall ranged 56-60/72 (78-83%), and 59/72 faults received identical verdicts across all three rounds. Detection was near-complete for calculation-logic, syntactic and hardcoding faults and weakest for transcription/units. The 14 undetected faults were systematic, not random: in every case the injected cell was a valid formula returning a believable number - source-dependent values, signal-free label/unit errors, peerless isolated cells, and locally plausible wrong-reference results. Each review returned a median of 8 flagged cells. Under dense review, recall fell to 33/72 (46%) by majority and 51/72 (71%) by union with 7 false positives.
CONCLUSIONS: We demonstrated that a general-purpose LLM, used simply and without tuning, can identify most seeded errors. However, it missed a predictable few - valid formulas returning believable numbers, verifiable only against an external source. Although giving the model the source documents or using more advanced setups would narrow the gap, human QC remains essential.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR43

Topic

Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×