HUMAN, AI, OR BOTH? EVALUATING THE PERFORMANCE OF AI-ASSISTED ABSTRACT SCREENING IN SYSTEMATIC LITERATURE REVIEWS
Author(s)
Nadine D. Younan, MSc, PhD1, Jiumei Gao, MSc2, Chaitanyamayee Kalakota, MSc2, Robin Philip, MSci, MSc2.
1Senior Analyst, Evimed Solutions Ltd, Amersham, United Kingdom, 2Evimed Solutions Ltd, Amersham, United Kingdom.
1Senior Analyst, Evimed Solutions Ltd, Amersham, United Kingdom, 2Evimed Solutions Ltd, Amersham, United Kingdom.
OBJECTIVES: Artificial intelligence (AI) is increasingly being integrated into systematic literature reviews (SLRs), particularly for labour-intensive stages such as title or abstract screening. While AI-assisted approaches may reduce reviewer burden and improve efficiency, concerns remain regarding their accuracy and transparency. This study evaluated the accuracy and efficiency of dual-human, human-AI, and AI-only abstract-screening workflows within an SLR.
METHODS: An SLR compared three abstract-screening workflows: dual-human screening; a hybrid human-AI screening; and AI-only screening. All workflows were followed by an independent human adjudication. Sensitivity, specificity, and screening times were assessed.
RESULTS: A total of 4,252 records underwent title and abstract screening. Following adjudication, sensitivity and specificity were: 64.8% and 99.8% for dual-human screening; 14.3% and 99.7% for human-AI screening; and 39.4% and 99.9% for the AI-only workflow. Specificity remained consistently high (>99%) across all workflows, indicating few false-positive inclusions. AI-assisted workflows substantially reduced screening time. Dual-human screening required approximately 1 minute per abstract, compared with 0.15-0.25 minutes for the human-AI workflow and 0.011 minutes for the AI-only workflow. The AI-only workflow demonstrated the greatest efficiency gains while maintaining high specificity, despite the considerable upfront effort required for implementation. Furthermore, both AI-assisted approaches showed lower sensitivity than dual-human screening, necessitating additional human rescreening of excluded studies to ensure that important information was captured, underscoring the continued importance of human oversight in abstract screening.
CONCLUSIONS: AI-assisted abstract screening substantially improved efficiency while maintaining high specificity; however, sensitivity was lower than that with dual-human screening. This suggests AI may reduce reviewer burden and support evidence synthesis workflows but should augment rather than replace human reviewers. Even in fully AI-driven workflows, human review remains necessary to prevent inadvertent exclusion of relevant articles during full-text screening and data extraction. Further optimisation of AI-assisted workflows is needed to improve sensitivity while preserving efficiency gains.
METHODS: An SLR compared three abstract-screening workflows: dual-human screening; a hybrid human-AI screening; and AI-only screening. All workflows were followed by an independent human adjudication. Sensitivity, specificity, and screening times were assessed.
RESULTS: A total of 4,252 records underwent title and abstract screening. Following adjudication, sensitivity and specificity were: 64.8% and 99.8% for dual-human screening; 14.3% and 99.7% for human-AI screening; and 39.4% and 99.9% for the AI-only workflow. Specificity remained consistently high (>99%) across all workflows, indicating few false-positive inclusions. AI-assisted workflows substantially reduced screening time. Dual-human screening required approximately 1 minute per abstract, compared with 0.15-0.25 minutes for the human-AI workflow and 0.011 minutes for the AI-only workflow. The AI-only workflow demonstrated the greatest efficiency gains while maintaining high specificity, despite the considerable upfront effort required for implementation. Furthermore, both AI-assisted approaches showed lower sensitivity than dual-human screening, necessitating additional human rescreening of excluded studies to ensure that important information was captured, underscoring the continued importance of human oversight in abstract screening.
CONCLUSIONS: AI-assisted abstract screening substantially improved efficiency while maintaining high specificity; however, sensitivity was lower than that with dual-human screening. This suggests AI may reduce reviewer burden and support evidence synthesis workflows but should augment rather than replace human reviewers. Even in fully AI-driven workflows, human review remains necessary to prevent inadvertent exclusion of relevant articles during full-text screening and data extraction. Further optimisation of AI-assisted workflows is needed to improve sensitivity while preserving efficiency gains.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA20
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
No Additional Disease & Conditions/Specialized Treatment Areas