ASSESSING THE PERFORMANCE OF ARTIFICIAL INTELLIGENCE FOR TITLE-ABSTRACT AND FULL-TEXT SCREENING IN SYSTEMATIC LITERATURE REVIEWS...
Author(s)
Geetank Kamboj, MPharm1, Shainki Sharma, MPharm1, Abhishek Malik, MSc2, Hemant Rathi, MSc2.
1Skyward Analytics Private Limited, Gurugram, India, 2EasySLR, Gurugram, India.
1Skyward Analytics Private Limited, Gurugram, India, 2EasySLR, Gurugram, India.
OBJECTIVES: The use of artificial intelligence (AI) in systematic literature reviews (SLRs) has grown substantially in recent years due to its potential to improve review timelines and optimize screening processes. Despite this increasing adoption, there remains limited evidence on the effectiveness of AI when independently performing study screening tasks. Therefore, evaluating AI performance at both the title-abstract and full-text screening stages is important to understand its potential role in supporting evidence synthesis. To assess the performance of AI for title-abstract and full-text screening using the EasySLR platform compared with human reviewers.
METHODS: The performance of AI screening was evaluated across four SLRs using identical screening protocols for both AI and human reviewers. Human decisions served as the reference standard for comparison. Accuracy, sensitivity, and specificity were calculated separately for each project. OpenAI GPT-5.0 mini was used for title-abstract screening, while GPT-5.2 was used for full-text screening.
RESULTS: Across the four projects, 268 to 948 citations were screened. During title-abstract screening, sensitivity was high (80-100%), whereas specificity (60.9-88.5%) and accuracy (63.0-90.1%) ranged from moderate to high. These findings suggest that AI effectively identified relevant studies but generated variable false positives. Full-text screening showed improved and more consistent performance, with sensitivity ranging from 83.3% to 100%, specificity of 100%, and accuracy between 91.7% and 100%. Overall, strong agreement between AI and human screening decisions was observed during full-text screening.
CONCLUSIONS: The findings indicate that AI may support SLR workflows by improving screening efficiency while maintaining strong agreement with human reviewers, particularly during full-text evaluation. Although AI demonstrated high sensitivity, human oversight remains essential to ensure methodological rigor. These findings should be interpreted cautiously, as performance may vary across research questions. Future research should evaluate AI performance on larger datasets and refine screening protocol for better understanding by AI.
METHODS: The performance of AI screening was evaluated across four SLRs using identical screening protocols for both AI and human reviewers. Human decisions served as the reference standard for comparison. Accuracy, sensitivity, and specificity were calculated separately for each project. OpenAI GPT-5.0 mini was used for title-abstract screening, while GPT-5.2 was used for full-text screening.
RESULTS: Across the four projects, 268 to 948 citations were screened. During title-abstract screening, sensitivity was high (80-100%), whereas specificity (60.9-88.5%) and accuracy (63.0-90.1%) ranged from moderate to high. These findings suggest that AI effectively identified relevant studies but generated variable false positives. Full-text screening showed improved and more consistent performance, with sensitivity ranging from 83.3% to 100%, specificity of 100%, and accuracy between 91.7% and 100%. Overall, strong agreement between AI and human screening decisions was observed during full-text screening.
CONCLUSIONS: The findings indicate that AI may support SLR workflows by improving screening efficiency while maintaining strong agreement with human reviewers, particularly during full-text evaluation. Although AI demonstrated high sensitivity, human oversight remains essential to ensure methodological rigor. These findings should be interpreted cautiously, as performance may vary across research questions. Future research should evaluate AI performance on larger datasets and refine screening protocol for better understanding by AI.
Conference/Value in Health Info
2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand
Value in Health, Volume 55, Issue S1
Code
SA5
Topic
Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
No Additional Disease & Conditions/Specialized Treatment Areas