RETROSPECTIVE VALIDATION OF A SOURCE-GROUNDED AI-ENABLED EVIDENCE-RISK TRIAGE WORKFLOW FOR ANTICIPATING EVIDENCE GAPS IN EUROPEAN HTA APPRAISALS
Author(s)
Ankit Dhaundiyal, MS1, Jagriti Prasad, MPH2, Tejesh S, MPH2, Mona Thangamma, MSc Health Economics2, Megha Tharad, PhD3.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
OBJECTIVES: To evaluate whether a source-grounded AI-enabled evidence-risk triage workflow could retrospectively anticipate material concerns later raised by European HTA bodies and support earlier dossier-readiness assessment.
METHODS: Retrospective validation was conducted across 10 oncology product-indication cases using public regulatory, clinical, and HTA documents from NICE (England), HAS (France), and G-BA/IQWiG (Germany). Cases were purposively selected for available pre- and post-appraisal reference documents, rather than to represent all European, EU, or EU5 agencies. During prediction, the workflow was restricted to pre-appraisal sources, including EMA assessment materials, pivotal publications, and trial registries; final HTA appraisal documents were withheld as the reference standard. Task-specialised modules supported PICOTS decomposition, source retrieval/extraction, comparator/endpoint alignment, indirect-comparison critique, PRO/HRQoL assessment, uncertainty characterisation, gap classification, citation/contradiction checking, and severity scoring. Predicted concerns were mapped to a prespecified taxonomy and matched against agency-raised concerns using strict and broad rules.
RESULTS: Across 10 cases, reference documents identified 70 material HTA evidence concerns, and the workflow generated 133 predicted evidence-gap rows. Strict matching identified 51/70 concerns (72.9% recall; exact 95% CI: 60.9%-82.8%; p<0.001 versus 50% no-signal benchmark). Broad matching identified 68/70 concerns (97.1% recall; exact 95% CI: 90.1%-99.7%). Strict and broad precision were 38.3% and 64.7%, supporting recall-oriented early triage rather than autonomous final classification. Source-grounding was confirmed for 70.7%; 29.3% were plausible but insufficiently supported. No direct contradiction with cited evidence was observed.
CONCLUSIONS: The workflow showed proof-of-concept value for surfacing likely HTA scrutiny areas before submission-critical use. Rather than replacing expert judgement, it may focus earlier review on comparator choice, endpoint relevance, evidence maturity, indirect-comparison limitations, PRO/HRQoL evidence, and uncertainty. Generalisability is limited by the small oncology-focused purposive sample, selected agency scope, moderate precision, plausible-only outputs, and absence of a manual gap-analysis or baseline NLP comparator. Broader prospective validation across agencies, therapy areas, and workflow-efficiency endpoints is required.
METHODS: Retrospective validation was conducted across 10 oncology product-indication cases using public regulatory, clinical, and HTA documents from NICE (England), HAS (France), and G-BA/IQWiG (Germany). Cases were purposively selected for available pre- and post-appraisal reference documents, rather than to represent all European, EU, or EU5 agencies. During prediction, the workflow was restricted to pre-appraisal sources, including EMA assessment materials, pivotal publications, and trial registries; final HTA appraisal documents were withheld as the reference standard. Task-specialised modules supported PICOTS decomposition, source retrieval/extraction, comparator/endpoint alignment, indirect-comparison critique, PRO/HRQoL assessment, uncertainty characterisation, gap classification, citation/contradiction checking, and severity scoring. Predicted concerns were mapped to a prespecified taxonomy and matched against agency-raised concerns using strict and broad rules.
RESULTS: Across 10 cases, reference documents identified 70 material HTA evidence concerns, and the workflow generated 133 predicted evidence-gap rows. Strict matching identified 51/70 concerns (72.9% recall; exact 95% CI: 60.9%-82.8%; p<0.001 versus 50% no-signal benchmark). Broad matching identified 68/70 concerns (97.1% recall; exact 95% CI: 90.1%-99.7%). Strict and broad precision were 38.3% and 64.7%, supporting recall-oriented early triage rather than autonomous final classification. Source-grounding was confirmed for 70.7%; 29.3% were plausible but insufficiently supported. No direct contradiction with cited evidence was observed.
CONCLUSIONS: The workflow showed proof-of-concept value for surfacing likely HTA scrutiny areas before submission-critical use. Rather than replacing expert judgement, it may focus earlier review on comparator choice, endpoint relevance, evidence maturity, indirect-comparison limitations, PRO/HRQoL evidence, and uncertainty. Generalisability is limited by the small oncology-focused purposive sample, selected agency scope, moderate precision, plausible-only outputs, and absence of a manual gap-analysis or baseline NLP comparator. Broader prospective validation across agencies, therapy areas, and workflow-efficiency endpoints is required.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR40
Topic
Health Policy & Regulatory, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas