AI AS A SECOND REVIEWER: ASSESSING A LANGUAGE MODEL'S SUITABILITY FOR ABSTRACT SCREENING IN SYSTEMATIC AND SCOPING REVIEWS

Author(s)

Allan Tran, PhD.
University of Toronto, Toronto, ON, Canada.
OBJECTIVES: To evaluate whether a large language model (GPT-4o) can effectively serve as a second reviewer during the abstract screening stage of systematic and scoping reviews, by comparing its performance to human reviewers in terms of agreement, classification accuracy, and simulated consensus decisions.
METHODS: A retrospective analysis was conducted using data from a previously completed scoping review comprising 2879 abstracts, each screened by two of three trained reviewers. Reviewers labelled each abstract as “Include”, “Exclude”, or Unclear”. GPT-4o was prompted with identical screening criteria and asked to classify each abstract. Agreement was assessed using Cohen’s Kappa and proportionate agreement. Model performance was evaluated against final consensus decisions using accuracy, precision, recall, and F1-score. Simulated reviewer replacement scenarios were conducted to assess the impact of substituting a human reviewer with the model.
RESULTS: Cohen’s Kappa between the AI and human reviewers ranged from 0.19 to 0.44, within the range of inter-reviewer variability (0.30-0.55). The model achieved high recall (0.80) but lower precision (0.24) for “Include” decisions, reflecting a sensitivity-focused approach. In simulated replacement, AI-human reviewer pairs matched the original consensus in over 96% of cases.
CONCLUSIONS: GPT-4o demonstrated agreement and classification performance comparable to human reviewers, particularly in identifying relevant studies. Its high recall suggests it may act as a safeguard against false exclusions. However, its lower precision underscores the need for human oversight in hybrid workflows. Overall, GPT-4o shows promise as a conservative second reviewer for abstract screening in systematic and scoping reviews, potentially reducing workload while preserving sensitivity in evidence selection.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MT45

Topic

Medical Technologies, Methodological & Statistical Research, Real World Data & Information Systems

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×