Is There a Consensus on the Framework for Evaluating Artificial Intelligence (AI)-Assisted Systematic Review Tools in HEOR?

Moderator

C. Daniel Mullins, PhD, University of Maryland School of Medicine, Baltimore, MD, United States

Speakers

Lizheng Shi, PhD, Tulane University School of Public Health and Tropical Medicine, New Orleans, LA, United States; Yiduo Zhang, BA, MA, PhD, AstraZeneca, Barcelona, Spain; Mei Yang, PhD, NouStarX, Short Hills, NJ, United States

Issue: AI-assisted systematic literature reviews (AI-SLRs) have rapidly expanded, with dozens of commercial and open-source tools now available. However, the HEOR and HTA community lacks shared standards for evaluating tool performance, ensuring reproducibility, documenting risk, and benchmarking AI-SLR output against human-conducted reviews. As evidence synthesis teams increasingly adopt AI-SLRs for HEOR and HTA projects, a possible solution is to establish an evaluation framework from the perspectives of multiple stakeholders. A case competition (a.k.a. challenge) using AI-SLR tools could be used to test the common evaluation framework. This session will discuss the need for a common evaluation framework and a process to conduct the HEOR/HTA community-driven benchmarking challenge. Overview: This panel will bring together perspectives from ISPOR leadership, academia, AI methodology, and the pharmaceutical industry to co-create a path forward. An overview provided by Dr. Mullins will outline the current landscape and the role of ISPOR in shaping good-practice guidance for AI in evidence synthesis (5 min). A senior evidence generation lead (Dr. Zhang) will share real-world insights on the challenges and opportunities of AI-SLR adoption, highlighting vendor assessment, validation requirements, compliance considerations, quality control, and the balance between efficiency gains and decision-grade reliability (15 min). Dr. Shi will introduce the concept of case competitions, discussing the feasibility of shared reference datasets, “gold-standard” systematic reviews, and open evaluation protocols to support innovation and methodological rigor (10 min). Dr. Yang will then present a structured framework for evaluating AI-SLR tools, covering reference standard construction, accuracy metrics, reproducibility, workload reduction, hallucination/error types, and governance considerations aligned with GenAI reporting expectations (15 min). The session will conclude with interactive polling and collaborative consensus to identify priority actions, including whether HEOR/HTA communities should form a working group, issue guidance, or sponsor an AI-SLR challenge.

Topic

Health Technology Assessment, Methodological & Statistical Research

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×