AUTOMATED TECHNICAL VERIFICATION OF HEALTH ECONOMIC MODELS USING A GENERATIVE-AI AGENT: CONCORDANCE WITH EXPERT REVIEW ACROSS SPREADSHEET AND PROGRAMMATIC IMPLEMENTATIONS

Author(s)

Shubhram Pandey, MSc1, Rajdeep Kaur, PhD2, Barinder Singh, RPh3.
1Senior Consultant and Head, Modeling and Advanced Analytics, Pharmacoevidence Pvt. Ltd., SAS Nagar, Mohali, India, 2Pharmacoevidence Pvt. Ltd., Mohali, India, 3Pharmacoevidence Pvt. Ltd., SAS Nagar Mohali, India.
OBJECTIVES: Technical verification of health economic models is labour-intensive, inconsistently performed, and a recognised source of error and credibility risk. We developed and evaluated a generative-AI (GenAI) agent that verifies both spreadsheet (Excel) and script-based (R/Python) models against TECH-VER and complementary checklists, benchmarked against expert reviewers.
METHODS: The agent pairs a retrieval-augmented large language model with deterministic parsers: Excel workbooks are decomposed into cells, formulas, named ranges and dependency graphs; R/Python models undergo static code analysis plus instrumented execution. Each component is mapped to TECH-VER black-box and white-box tests and to AdViSHE validation-status items, with automated checks for discounting, half-cycle correction, transition-probability and cohort conservation, extreme-value/null-parameter behaviour, and dimensional consistency. Outputs are a severity-graded error log and an auto-populated verification report. Evaluation used 25 models (15 Excel, 10 R/Python) seeded with 120 plausible errors of known type; agent detections were compared against independent verification by three experienced modellers (reference standard).
RESULTS: The agent achieved 88% overall detection sensitivity and 96% specificity versus expert review (Cohen's κ=0.82), with a 6% false-positive rate. Detection was higher for spreadsheet than script-based models (91% vs 84%). Mean verification time fell from 6.2 hours (expert) to 9 minutes (agent), a ~41× reduction. Errors requiring structural judgement - inappropriate survival extrapolation, mis-specified comparator pathways - were most frequently missed.
CONCLUSIONS: A GenAI agent delivers fast, reproducible, checklist-anchored technical verification across both spreadsheet and programmatic models, materially reducing reviewer burden while standardising documentation. The approach augments rather than replaces expert judgement; conceptual validity, structural assumptions and external validation remain human responsibilities.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR294

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×