AN AGENTIC AI MODEL VALIDATOR FOR AUTOMATED QUALITY CHECK OF EXCEL-BASED BUDGET IMPACT MODELS
Author(s)
Tushar Srivastava, MSc1, Hanan Irfan, MSc1, Nishtha Neeraj2, Syed Salleh, PhD1.
1ConnectHEOR, London, United Kingdom, 2India.
1ConnectHEOR, London, United Kingdom, 2India.
OBJECTIVES: QC of Excel-based budget impact models (BIMs) remains manual and inconsistently documented, leaving errors that can materially distort projected expenditure. We evaluated an AI Model Validator Agent that runs a structured QC suite over Excel BIMs, benchmarked against senior human validators.
METHODS: The agent interprets each QC check, infers the expected behaviour, determines applicability, and plans validation, with value extraction, independent recalculation, and scenario toggling. An executor layer runs static and dynamic checks returning pass or fail, observations, severity (critical, moderate, minor), and recommendations into a reviewable, overridable model score. The agent was tested on two production-grade anonymized BIMs: 5 year based endometrial cancer model (around 10 worksheets, around 60,000 populated cells) and a 5-year COPD model using a mixed prevalent and incidence population (around 13 worksheets, around 80,000 populated cells). Twenty-four errors of graded difficulty were seeded across both, spanning population funnels, market-share displacement, discounting, annualisation, cost-year indexing, and scenario-toggle logic.
RESULTS: The agent detected 23 of 24 (96%) seeded errors across the two models. Examples of critical findings that surfaced only on recalculation: a scenario-toggle fault where identical inputs across scenarios still returned a non-zero incremental budget impact, and an uptake-linkage error where new-treatment uptake set to 0% failed to zero projected expenditure. A population double-counting error in the incidence funnel overstated 5-year budget impact by about GBP 8.4 million (around 14%) against a base case near GBP 62 million. It also flagged market share failing to sum to 100% under specific toggles. The single miss was an inflation-indexing inconsistency in an appendix table; a few false positives needed adjudication. Automated QC completed in under 90 minutes per model, over 90% faster than manual review.
CONCLUSIONS: An agentic, reasoning-based validator delivered accurate, rapid, and auditable first-pass QC for Excel BIMs, complementing rather than replacing senior expert review within governed HTA workflows.
METHODS: The agent interprets each QC check, infers the expected behaviour, determines applicability, and plans validation, with value extraction, independent recalculation, and scenario toggling. An executor layer runs static and dynamic checks returning pass or fail, observations, severity (critical, moderate, minor), and recommendations into a reviewable, overridable model score. The agent was tested on two production-grade anonymized BIMs: 5 year based endometrial cancer model (around 10 worksheets, around 60,000 populated cells) and a 5-year COPD model using a mixed prevalent and incidence population (around 13 worksheets, around 80,000 populated cells). Twenty-four errors of graded difficulty were seeded across both, spanning population funnels, market-share displacement, discounting, annualisation, cost-year indexing, and scenario-toggle logic.
RESULTS: The agent detected 23 of 24 (96%) seeded errors across the two models. Examples of critical findings that surfaced only on recalculation: a scenario-toggle fault where identical inputs across scenarios still returned a non-zero incremental budget impact, and an uptake-linkage error where new-treatment uptake set to 0% failed to zero projected expenditure. A population double-counting error in the incidence funnel overstated 5-year budget impact by about GBP 8.4 million (around 14%) against a base case near GBP 62 million. It also flagged market share failing to sum to 100% under specific toggles. The single miss was an inflation-indexing inconsistency in an appendix table; a few false positives needed adjudication. Automated QC completed in under 90 minutes per model, over 90% faster than manual review.
CONCLUSIONS: An agentic, reasoning-based validator delivered accurate, rapid, and auditable first-pass QC for Excel BIMs, complementing rather than replacing senior expert review within governed HTA workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR235
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas