VALIDATING GENERATIVE AI-ASSISTED SYSTEMATIC LITERATURE REVIEW WORKFLOWS IN HEOR: ACCURACY, EFFICIENCY, AND GOVERNANCE IMPLICATIONS

Author(s)

Geetika Sharma, Masters of Science(MS)1, Monica Verma, MPH2, Anand Jha, MBA3.
1Ansea Consultants Pte Ltd, Pune, India, 2Associate Director, Ansea Consultants Pte Ltd, Singapore, Singapore, 3Ansea Consultants Pte Ltd, Singapore, Singapore.
OBJECTIVES: Systematic literature reviews (SLRs) underpin HEOR, HTA, and payer decision making but remain resource intensive. Generative artificial intelligence (AI) and related machine-learning tools may improve efficiency across screening and data extraction; however, their validity, reproducibility, and governance requirements for HEOR use remain unclear. We synthesized recent evidence to identify where AI-assisted SLR workflows are sufficiently robust for implementation and where human verification remains essential.
METHODS: A targeted methodological review was performed using peer-reviewed studies, conference evaluations, and official guidance published between 2021 and 2026. Sources included PubMed-indexed articles, Value in Health/ISPOR presentations, Research Synthesis Methods, Systematic Reviews, official ISPOR instructions, and guidance from Cochrane/RAISE, PRISMA-related reporting frameworks, EMA, and the European AI governance sources. Evidence was synthesized across five domains: title/abstract screening, full-text screening, data extraction, operational efficiency, and governance.
RESULTS: Evidence was concentrated in screening rather than extraction. A scoping review identified 123 automation studies, of which 72.4% addressed record screening and 10.6% addressed data extraction. In comparative screening studies, tuned large language model workflows achieved sensitivity from 0.93 to 1.00 in selected datasets, with substantial workload reduction; however, precision decreased markedly in low-prevalence settings. In data extraction studies, accuracy was high for structured study characteristics (for example, 92.4% to 96.3%) but lower for complex efficacy, safety, and numeric result fields. Recent case studies reported time savings ranging from 23% to greater than 90%, depending on task and workflow design. Governance sources consistently required human oversight, transparent reporting, source verification, and explicit management of hallucinations, privacy, and auditability risks.
CONCLUSIONS: AI-assisted SLR workflows are increasingly viable for HEOR when implemented as validated, human-in-the-loop systems. The strongest near-term use cases are screening prioritization and first-pass extraction of structured study descriptors. For HTA-grade evidence generation, local validation against dual-human reference standards, task-specific performance thresholds, and full audit trails should be treated as minimum requirements.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA41

Topic

Health Technology Assessment

Topic Subcategory

Decision & Deliberative Processes, Systems & Structure

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×