WHEN AI GETS HEOR WRONG: A TAXONOMY OF SILENT FAILURE MODES IN AI-ASSISTED HEALTH ECONOMIC WORK AND HOW TO DETECT THEM
Author(s)
Tushar Srivastava, MSc1, Hanan Irfan, MSc2, Kunal Swami, MSc3, Shilpi Swami, MSc1.
1ConnectHEOR, London, United Kingdom, 2ConnectHEOR, London, India, 3ConnectHEOR Ltd, Delhi, India.
1ConnectHEOR, London, United Kingdom, 2ConnectHEOR, London, India, 3ConnectHEOR Ltd, Delhi, India.
OBJECTIVES: Most published work on AI in HEOR reports what AI can do; far less report how it fails, and the failures that matter most are silent: plausible outputs that pass a quick read but are wrong. We synthesised failure cases from controlled studies of AI-assisted HEOR into a structured taxonomy, each mode paired with its detection signature and the control that catches it.
METHODS: We pooled failure cases from completed in-house studies spanning the AI-assisted HEOR workflow: cost-effectiveness model validation, a hackathon comparing off-the-shelf LLM, human expert, and purpose-built agent, model interrogation for assessors, and SLR screening and extraction. Each failure was classified by mechanism, by whether it survived a standard face-validity check, by where in the workflow it arose, and by its consequence for a submission. Each mode was then mapped to the detection method and the human or automated control that reliably catches it.
RESULTS: Failures grouped into recurring modes. Fabrication: plausible but non-existent references, comparators, or values, most damaging because fluent and confident. Recalculation-dependent logic errors: structurally invalid results (for example, transition probabilities exceeding one under a subgroup) that pass static inspection and surface only on forced recalculation. Confident mis-extraction: correct-looking but wrong inputs pulled from source documents. Drift: outputs that change as the underlying model updates, breaking earlier validation. Automation complacency: errors missed because a fluent output discouraged scrutiny. The highest-risk modes shared one property: they survive face validity, so they are caught not by reading the output but by independent recomputation, source tracing, and pre-specified checks.
CONCLUSIONS: AI-assisted HEOR fails in identifiable, recurring ways, and the most dangerous failures are the ones that look right. A shared failure-mode taxonomy turns scattered anecdotes into a managed risk surface, lets teams target validation where it matters, and gives HTA bodies a vocabulary for what to ask of AI-assisted submissions.
METHODS: We pooled failure cases from completed in-house studies spanning the AI-assisted HEOR workflow: cost-effectiveness model validation, a hackathon comparing off-the-shelf LLM, human expert, and purpose-built agent, model interrogation for assessors, and SLR screening and extraction. Each failure was classified by mechanism, by whether it survived a standard face-validity check, by where in the workflow it arose, and by its consequence for a submission. Each mode was then mapped to the detection method and the human or automated control that reliably catches it.
RESULTS: Failures grouped into recurring modes. Fabrication: plausible but non-existent references, comparators, or values, most damaging because fluent and confident. Recalculation-dependent logic errors: structurally invalid results (for example, transition probabilities exceeding one under a subgroup) that pass static inspection and surface only on forced recalculation. Confident mis-extraction: correct-looking but wrong inputs pulled from source documents. Drift: outputs that change as the underlying model updates, breaking earlier validation. Automation complacency: errors missed because a fluent output discouraged scrutiny. The highest-risk modes shared one property: they survive face validity, so they are caught not by reading the output but by independent recomputation, source tracing, and pre-specified checks.
CONCLUSIONS: AI-assisted HEOR fails in identifiable, recurring ways, and the most dangerous failures are the ones that look right. A shared failure-mode taxonomy turns scattered anecdotes into a managed risk surface, lets teams target validation where it matters, and gives HTA bodies a vocabulary for what to ask of AI-assisted submissions.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR47
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas