EVIDENCE-INFORMED VALIDATION FRAMEWORK FOR GENERATIVE AI-ASSISTED SYSTEMATIC LITERATURE REVIEWS IN HEOR: IMPLICATIONS FOR EUROPEAN HTA AND JCA

Author(s)

Geetika Sharma, MS.
Vedara IQ, Pune, India.
OBJECTIVES: To develop an evidence-informed validation framework for generative AI-assisted systematic literature reviews (SLRs) in HEOR and assess whether current performance is sufficient for European HTA and Joint Clinical Assessment (JCA)-relevant use.
METHODS: We conducted a targeted review of peer-reviewed comparative studies published 2020-2026, supplemented by ISPOR trend reports, PRISMA/Cochrane/RAISE reporting and governance guidance, NICE and European HTA/JCA documents, and major tool documentation. We extracted recall/sensitivity, precision or specificity, F1 where available, time savings, inter-rater agreement or reproducibility, and reported error modes. We then built an evidence-informed human-in-the-loop validation framework and simulated a representative HEOR screening scenario using conservative ranges reported in the literature.
RESULTS: Across comparative studies, title/abstract recall or sensitivity was generally high (0.85-1.00), whereas precision varied substantially (0.03-0.49) because of topic complexity and low inclusion prevalence; reported F1 ranged from 0.05 to approximately 0.64 when calculable. Workload reduction ranged from 47% to 75%, and repeat-run agreement or consistency ranged from 95.4% to 98.9%; one validation study reported kappa=0.83 for repeated title/abstract screening runs. In an evidence-informed HEOR scenario, a human-supervised workflow using AI for prioritisation and second-review support was projected to reduce manual title/abstract effort by 55%-68% while preserving a recall target of at least 0.95. Principal failure modes were over-inclusion, false negatives in sparse or ambiguous abstracts, prompt/model drift, and weaker performance for full-text or complex judgement tasks. Minimum requirements for defensible use were locked prompts and model versions, protocol registration, database and registry coverage, audit trails, dual-review adjudication of exclusions, and stage-specific performance reporting.
CONCLUSIONS: Generative AI can materially accelerate HEOR SLRs, but current evidence supports augmented rather than autonomous use. For European HTA/JCA settings, acceptability is most likely when AI is prospectively validated, transparently documented, and embedded within accountable human review.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

SA105

Topic

Health Technology Assessment, Study Approaches

Topic Subcategory

Literature Review & Synthesis

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×