AUTOMATED QUALITY CONTROL OF HEALTH ECONOMIC EXCEL MODELS USING A LARGE LANGUAGE MODEL AGENT: A MULTI-MODEL VALIDATION STUDY
Author(s)
Andre Verhoek, MSc1, Suzan Serip, MSc2, Yiduo Zhang, BA, MA, PhD2.
1HMS Lead, AstraZeneca, Barcelona, Spain, 2AstraZeneca, Barcelona, Spain.
1HMS Lead, AstraZeneca, Barcelona, Spain, 2AstraZeneca, Barcelona, Spain.
OBJECTIVES: No published method exists for automated quality control (QC) of Excel-based health economic models. Manual QC using established frameworks (TECH-VER, CDA-AMC) is time-intensive and resource-constrained. We developed and validated an agentic large language model (LLM) workflow that operationalizes the combined TECH-VER and CDA-AMC checklists to perform structured, reproducible QC of cost-effectiveness models (CEMs) and budget impact models (BIMs).
METHODS: The TECH-VER 5-stage verification process and CDA-AMC 53-item validation tool were distilled into an 8-phase QC protocol executable by an LLM agent (Claude, Anthropic): conceptual validation, clinical effectiveness, transparency checks, black-box testing, white-box testing, Excel-specific concerns, overall validation, and documentation. The agent inspects Excel models programmatically — reading formulas, named ranges, VBA macros, and cell values via static code inspection — and applies the protocol sequentially. We validated the approach on four industry models of increasing complexity: a 1-year biologic switching analysis in severe eosinophilic asthma, a Markov CEM in idiopathic pulmonary fibrosis, a 54-sheet partitioned-survival CEM in asthma, and a 5-year BIM in resistant hypertension. AI-generated QC reports were compared head-to-head against independent human QC of the same models.
RESULTS: The LLM agent detected 100% of major and 100% of moderate flags identified by human reviewers across all four models, with 95% concordance on minor flags; the agent additionally identified issues missed by human reviewers (e.g., out-of-range PSA distribution bounds). QC completion times ranged from 28 to 116 minutes depending on model complexity, at $1-$15 per model. Each run produced a structured report documenting every test, expected versus actual results, and severity-graded findings.
CONCLUSIONS: This is the first demonstration of an LLM agent performing systematic QC of health economic Excel models against established verification frameworks. Full concordance on major and moderate findings, combined with reduction of QC time from days to under two hours, establishes feasibility for production deployment in HTA-grade modelling workflows.
METHODS: The TECH-VER 5-stage verification process and CDA-AMC 53-item validation tool were distilled into an 8-phase QC protocol executable by an LLM agent (Claude, Anthropic): conceptual validation, clinical effectiveness, transparency checks, black-box testing, white-box testing, Excel-specific concerns, overall validation, and documentation. The agent inspects Excel models programmatically — reading formulas, named ranges, VBA macros, and cell values via static code inspection — and applies the protocol sequentially. We validated the approach on four industry models of increasing complexity: a 1-year biologic switching analysis in severe eosinophilic asthma, a Markov CEM in idiopathic pulmonary fibrosis, a 54-sheet partitioned-survival CEM in asthma, and a 5-year BIM in resistant hypertension. AI-generated QC reports were compared head-to-head against independent human QC of the same models.
RESULTS: The LLM agent detected 100% of major and 100% of moderate flags identified by human reviewers across all four models, with 95% concordance on minor flags; the agent additionally identified issues missed by human reviewers (e.g., out-of-range PSA distribution bounds). QC completion times ranged from 28 to 116 minutes depending on model complexity, at $1-$15 per model. Each run produced a structured report documenting every test, expected versus actual results, and severity-graded findings.
CONCLUSIONS: This is the first demonstration of an LLM agent performing systematic QC of health economic Excel models against established verification frameworks. Full concordance on major and moderate findings, combined with reduction of QC time from days to under two hours, establishes feasibility for production deployment in HTA-grade modelling workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR101
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas