WHY ARE HTA REVIEWERS STILL HUNTING FOR CODING ERRORS? AN AI VALIDATOR AGENT AS A FIRST-LINE QA LAYER FOR SUBMITTED MODELS
Author(s)
Tushar Srivastava, MSc, Hanan Irfan, MSc, Thaison Tong, PhD, Shilpi Swami, MSc.
ConnectHEOR, London, United Kingdom.
ConnectHEOR, London, United Kingdom.
OBJECTIVES: HTA bodies and their assessment teams face rising submission volumes and finite reviewer capacity, with much reviewer time spent identifying computational and logical errors in economic models before methodological appraisal can begin. We tested whether an AI model validator agent, deployed as a first-line quality-control (QC) layer, could speed up verification of a company-submitted model while preserving reviewer control.
METHODS: We ran a proof-of-concept simulation of an HTA body’s first line verification of a company submitted model, with the agent standing in for the technical verification an assessor performs before methodological appraisal.The agent ingests an Excel model and documentation and produces sheet-by-sheet summary for reviewer approval.. The test case was a submission quality six-state Markov model in early triple-negative breast cancer. To reflect the errors an assessor identified during first line QC, we seeded 40 graded errors across transition logic, discounting, state-utility mapping, cohort accounting and cross-referencing.
RESULTS: The agent identified 39 of 40 seeded errors (97.5%), including every critical structural error. First-pass verification took under four hours per model, against multi-day manual checking, cutting time to a severity-ranked issue list by roughly 80%. False positives were rare (positive predictive value 94%), the single miss was a face-validity judgement rather than a computational error, and simulated reviewers retained the agent's classification on 92% of flags.
CONCLUSIONS: As a proof of concept, the validator acts as a first-line QC enabler for HTA reviewers, not a replacement for their critique. It lets a reviewer verify several models in parallel, resolves technical errors early so assessors can concentrate on methodological judgement, and leaves a line-by-line audit trail supporting transparency. These results come from a single simulated model not yet evaluated within an HTA body; the next step is prospective evaluation with an HTA reviewer, using naturally occurring rather than seeded errors.
METHODS: We ran a proof-of-concept simulation of an HTA body’s first line verification of a company submitted model, with the agent standing in for the technical verification an assessor performs before methodological appraisal.The agent ingests an Excel model and documentation and produces sheet-by-sheet summary for reviewer approval.. The test case was a submission quality six-state Markov model in early triple-negative breast cancer. To reflect the errors an assessor identified during first line QC, we seeded 40 graded errors across transition logic, discounting, state-utility mapping, cohort accounting and cross-referencing.
RESULTS: The agent identified 39 of 40 seeded errors (97.5%), including every critical structural error. First-pass verification took under four hours per model, against multi-day manual checking, cutting time to a severity-ranked issue list by roughly 80%. False positives were rare (positive predictive value 94%), the single miss was a face-validity judgement rather than a computational error, and simulated reviewers retained the agent's classification on 92% of flags.
CONCLUSIONS: As a proof of concept, the validator acts as a first-line QC enabler for HTA reviewers, not a replacement for their critique. It lets a reviewer verify several models in parallel, resolves technical errors early so assessors can concentrate on methodological judgement, and leaves a line-by-line audit trail supporting transparency. These results come from a single simulated model not yet evaluated within an HTA body; the next step is prospective evaluation with an HTA reviewer, using naturally occurring rather than seeded errors.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR124
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas