CAN ARTIFICIAL INTELLIGENCE RELIABLY PERFORM TITLE-ABSTRACT SCREENING FOR SYSTEMATIC LITERATURE REVIEWS?
Author(s)
Shainki Sharma, MPharm1, Geetank Kamboj, MPharm1, Abhishek Malik, MSc2, Hemant Rathi, MSc2.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
OBJECTIVES: To evaluate the performance of artificial intelligence (AI) for title-abstract screening by comparing AI-generated screening decisions with those of human reviewers across multiple systematic literature reviews (SLRs).
METHODS: Three previously completed SLRs across different research areas were included in the evaluation. Title-abstract screening was performed using the OpenAI GPT-5 mini model within the EasySLR™ platform. For each review, identical predefined eligibility criteria were applied by both the AI and human reviewers. Human reviewer decisions served as the reference standard against which AI decisions were evaluated. AI performance was assessed separately for each SLR using sensitivity (the proportion of eligible records correctly included), specificity (the proportion of ineligible records correctly excluded), precision (the proportion of AI-included records that were truly eligible), and accuracy (the overall proportion of correctly classified records).
RESULTS: Across the three SLRs, the number of citations screened ranged from 262 to 559. AI demonstrated consistently high sensitivity (96.4-100%), specificity (88.7-94.0%), and accuracy (92.1-94.4%) were consistently high, whereas precision was moderate (65.9-82.6%). High sensitivity was observed across all three reviews, indicating that AI identified nearly all studies deemed eligible by human reviewers. Similarly, high specificity and accuracy demonstrated reliable exclusion of ineligible records. However, the moderate precision reflected a variable number of false-positive inclusions.
CONCLUSIONS: Overall, title-abstract screening using AI demonstrated high sensitivity, specificity, and accuracy with moderate precision; however, human oversight remains necessary to ensure methodological rigour. These findings should be interpreted cautiously, as performance may vary across research questions. Future research should evaluate AI performance using larger datasets, diverse therapeutic areas, and refine screening protocols to improve AI interpretation of eligibility criteria.
METHODS: Three previously completed SLRs across different research areas were included in the evaluation. Title-abstract screening was performed using the OpenAI GPT-5 mini model within the EasySLR™ platform. For each review, identical predefined eligibility criteria were applied by both the AI and human reviewers. Human reviewer decisions served as the reference standard against which AI decisions were evaluated. AI performance was assessed separately for each SLR using sensitivity (the proportion of eligible records correctly included), specificity (the proportion of ineligible records correctly excluded), precision (the proportion of AI-included records that were truly eligible), and accuracy (the overall proportion of correctly classified records).
RESULTS: Across the three SLRs, the number of citations screened ranged from 262 to 559. AI demonstrated consistently high sensitivity (96.4-100%), specificity (88.7-94.0%), and accuracy (92.1-94.4%) were consistently high, whereas precision was moderate (65.9-82.6%). High sensitivity was observed across all three reviews, indicating that AI identified nearly all studies deemed eligible by human reviewers. Similarly, high specificity and accuracy demonstrated reliable exclusion of ineligible records. However, the moderate precision reflected a variable number of false-positive inclusions.
CONCLUSIONS: Overall, title-abstract screening using AI demonstrated high sensitivity, specificity, and accuracy with moderate precision; however, human oversight remains necessary to ensure methodological rigour. These findings should be interpreted cautiously, as performance may vary across research questions. Future research should evaluate AI performance using larger datasets, diverse therapeutic areas, and refine screening protocols to improve AI interpretation of eligibility criteria.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR110
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas