AI AS A SECOND REVIEWER: ASSESSING A LANGUAGE MODEL'S SUITABILITY FOR ABSTRACT SCREENING IN SYSTEMATIC AND SCOPING REVIEWS
Author(s)
Allan Tran, PhD.
University of Toronto, Toronto, ON, Canada.
University of Toronto, Toronto, ON, Canada.
OBJECTIVES: To evaluate whether a large language model (GPT-4o) can effectively serve as a second reviewer during the abstract screening stage of systematic and scoping reviews, by comparing its performance to human reviewers in terms of agreement, classification accuracy, and simulated consensus decisions.
METHODS: A retrospective analysis was conducted using data from a previously completed scoping review comprising 2879 abstracts, each screened by two of three trained reviewers. Reviewers labelled each abstract as “Include”, “Exclude”, or Unclear”. GPT-4o was prompted with identical screening criteria and asked to classify each abstract. Agreement was assessed using Cohen’s Kappa and proportionate agreement. Model performance was evaluated against final consensus decisions using accuracy, precision, recall, and F1-score. Simulated reviewer replacement scenarios were conducted to assess the impact of substituting a human reviewer with the model.
RESULTS: Cohen’s Kappa between the AI and human reviewers ranged from 0.19 to 0.44, within the range of inter-reviewer variability (0.30-0.55). The model achieved high recall (0.80) but lower precision (0.24) for “Include” decisions, reflecting a sensitivity-focused approach. In simulated replacement, AI-human reviewer pairs matched the original consensus in over 96% of cases.
CONCLUSIONS: GPT-4o demonstrated agreement and classification performance comparable to human reviewers, particularly in identifying relevant studies. Its high recall suggests it may act as a safeguard against false exclusions. However, its lower precision underscores the need for human oversight in hybrid workflows. Overall, GPT-4o shows promise as a conservative second reviewer for abstract screening in systematic and scoping reviews, potentially reducing workload while preserving sensitivity in evidence selection.
METHODS: A retrospective analysis was conducted using data from a previously completed scoping review comprising 2879 abstracts, each screened by two of three trained reviewers. Reviewers labelled each abstract as “Include”, “Exclude”, or Unclear”. GPT-4o was prompted with identical screening criteria and asked to classify each abstract. Agreement was assessed using Cohen’s Kappa and proportionate agreement. Model performance was evaluated against final consensus decisions using accuracy, precision, recall, and F1-score. Simulated reviewer replacement scenarios were conducted to assess the impact of substituting a human reviewer with the model.
RESULTS: Cohen’s Kappa between the AI and human reviewers ranged from 0.19 to 0.44, within the range of inter-reviewer variability (0.30-0.55). The model achieved high recall (0.80) but lower precision (0.24) for “Include” decisions, reflecting a sensitivity-focused approach. In simulated replacement, AI-human reviewer pairs matched the original consensus in over 96% of cases.
CONCLUSIONS: GPT-4o demonstrated agreement and classification performance comparable to human reviewers, particularly in identifying relevant studies. Its high recall suggests it may act as a safeguard against false exclusions. However, its lower precision underscores the need for human oversight in hybrid workflows. Overall, GPT-4o shows promise as a conservative second reviewer for abstract screening in systematic and scoping reviews, potentially reducing workload while preserving sensitivity in evidence selection.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MT45
Topic
Medical Technologies, Methodological & Statistical Research, Real World Data & Information Systems
Disease
No Additional Disease & Conditions/Specialized Treatment Areas