OFF-THE-SHELF LLM, HUMAN EXPERT, OR PURPOSE-BUILT AGENT? A CONTROLLED VALIDATION HACKATHON BENCHMARKING THREE APPROACHES TO QUALITY CONTROL OF HEALTH ECONOMIC MODELS

Author(s)

Tushar Srivastava, MSc, Hanan Irfan, MSc, Thaison Tong, MSc, Shilpi Swami, MSc.
ConnectHEOR, London, United Kingdom.
OBJECTIVES: QC of Excel-based CEMs underpins HTA credibility but remains manual and slow. General-purpose LLMs are increasingly proposed for this. We ran a controlled hackathon comparing three QC approaches on identical models with seeded errors.
METHODS: Two production-grade models were tested: a rare-disease Markov model (~150,000 cells) and a metastatic endometrial cancer partitioned-survival model (~120,000 cells). Forty errors (20 per model) were seeded across four types: numeric or transcription, broken references, recalculation-dependent logic, and hygiene. Three arms used identical workbooks: (1) an off-the-shelf frontier LLM with no execution layer; (2) two senior health economists reviewing independently then reconciling; and (3) an in-house developed AI validator agent. The five-hour limit reflects a realistic working session and deliberately constrains the human arm, whose accuracy would be expected to rise with the days available for manual QC. A blinded panel scored every flag for sensitivity, precision, and time.
RESULTS: Within the time-window the agent found 38 of 40 errors (95%), the human experts 22 (55%), and the off-the-shelf LLM 18 (45%); the human total was bounded by the available time, not by capability. The gap was widest for recalculation-dependent logic errors (agent 9 of 10, humans 5 of 10, LLM 2 of 10); only the agent caught issues around treatment waning calculation in subgroups and survival-input calculation. The LLM produced the most false positives, including hallucinated references (precision 31%, versus 85% for the agent). The agent completed its pass in about two hours per model.
CONCLUSIONS: Within a fixed time window, a purpose-built validator outperformed both an off-the-shelf LLM and human experts, especially on logic errors. The human result reflects time pressure, not reviewer skill, the realistic constraint HTA QC operates under. An off-the-shelf LLM alone is not enough for HTA-grade validation. Best results paired the agent as a fast first-line layer with senior human adjudication.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR181

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×