Can Agentic AI Deliver HTA-Ready Health Economic Models? Governance, Validation, and Trust
Moderator
Andrew Briggs, DPhil, London School of Hygiene & Tropical Medicine, London, United Kingdom
Speakers
Andre Verhoek, MSc, AstraZeneca, Barcelona, Spain; Jag Chhatwal, PhD, Harvard Medical School / Massachusetts General Hospital, Boston, MA, United States; Xiaoyan Wang, PhD, New Orleans, United States
ISSUE:
Large language model (LLM) agents can now construct Excel-based health economic models, execute multi-phase quality control protocols, generate HTA-compliant technical reports, and produce structured bibliographies—all with validated accuracy. Yet no consensus exists on whether AI-generated outputs that pass identical verification standards to human-generated work should be treated equivalently in HTA submissions. The field faces a governance gap: proof-of-concept demonstrations have outpaced frameworks for responsible production deployment. This panel debates whether current validation standards suffice for LLM-generated health economic outputs, or whether new governance is required.
OVERVIEW:
Three speakers from different sectors present empirical perspectives within a 60-minute session. The industry speaker (10 min) presents results from a multi-model validation programme: automated QC achieving concordance with human reviewers across 10+ models, technical reports reaching zero expert revision, functional cost-effectiveness models built in under 3 hours, and reference management with full structural accuracy. The consultancy speaker (10 min) addresses scalability and trust, reporting significant timeline reductions but arguing that the critical success factor is decomposing modelling tasks into auditable sub-steps, with governance calibrated to use case—from early asset valuation to submission-grade models. The academic speaker (10 min) examines where current reporting frameworks—including ELEVATE-GenAI—fall short for autonomous agent workflows, proposing adapted criteria including specification completeness scoring and assumption provenance tracking. A 15-minute moderated discussion addresses: (1) Is an LLM-generated model that passes ISPOR-SMDM equivalent to a human-built model that passes ISPOR-SMDM? (2) What frameworks must HTA agencies implement to audit AI involvement, and where is the threshold between AI-assisted and AI-generated modeling? (3) Where is human oversight essential versus performative? This benefits health economic modellers, HTA assessors, pharmaceutical submission teams, and HEOR consultants navigating AI adoption.
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research