Generative AI for Systematic Literature Reviews: ISPOR Good Practices Task Force Findings
Moderator
Dalia Dawoud, BSc, MSc, PhD, Cytel, London, United Kingdom
Speakers
Jag Chhatwal, PhD, Harvard Medical School / Massachusetts General Hospital, Boston, MA, United States; Raphael Sonabend-Friend, NICE, London, United Kingdom; Sven L Klijn, MSc, Bristol Myers Squibb, Princeton, NJ, United States
PURPOSE: Systematic literature reviews (SLRs) are foundational to evidence-based medicine, health technology assessment (HTA), yet they remain labor-intensive and resource-demanding.
Generative AI (GenAI) tools are increasingly being applied across SLR work?ows, with early studies demonstrating promise for screening, data extraction, and report drafting. However, key concerns persist regarding hallucinated outputs, reproducibility, prompt sensitivity, and variable performance across tasks. This workshop will share the ?ndings and recommendations of the ISPOR Task Force on GenAI for SLRs, equipping attendees to responsibly integrate GenAI into evidence synthesis while preserving research integrity. DESCRIPTION: This forum presents the ?nal ?ndings and recommendations of the ISPOR Good Practices Task Force on GenAI for SLRs. The Task Force conducted a rapid evidence assessment identifying 115 studies (November 2022--July 2025) and applied a structured assessment framework across eight evaluation domains to seven core SLR tasks: search strategy development, title/abstract screening, full-text screening, data extraction, risk of bias assessment, qualitative summarization, and report writing. Task Force members will present task-level performance summaries, recommended accuracy metrics, and consensus-based good practice recommendations. Key ?ndings indicate that GenAI can augment--but not replace--human expertise. The strongest evidence supports title/abstract screening and structured data extraction within recall-oriented, human-in-the-loop work?ows. Interpretive tasks such as risk of bias assessment demonstrated the lowest readiness for GenAI integration. Autonomous deployment was not supported for any SLR task. The forum will present the task force recommendations addressing: (1) appropriate and unsupported uses of GenAI by SLR task; (2) required human oversight and accountability safeguards; (3) recommended performance metrics for each task; and (4) reporting and transparency standards aligned with the ELEVATE- GenAI framework. The session will close with a facilitated audience discussion.
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches