AI-ASSISTED VS. MANUAL SYSTEMATIC LITERATURE REVIEW IN MULTIPLE SCLEROSIS MORTALITY: A COMPARATIVE ASSESSMENT
Author(s)
Pavan Ghunawat, M. Pharm1, Sirish Cholasamudram, M. Pharm1, Lita Araujo, BSc, MSc, PhD2, Hugo Dubucq, MPH3.
1Sanofi Healthcare India Private Limited, Hyderabad, India, 2Sanofi, Cambridge, MA, USA, 3Sanofi, Barcelona, Spain.
1Sanofi Healthcare India Private Limited, Hyderabad, India, 2Sanofi, Cambridge, MA, USA, 3Sanofi, Barcelona, Spain.
OBJECTIVES: Systematic literature reviews (SLRs) are resource-intensive, requiring substantial manual effort across screening, data extraction, and quality control. AI-assisted tools offer potential to accelerate this process; however, their performance relative to manual review remains insufficiently characterized. This study evaluated an AI-assisted tool against a manual SLR in a mortality in multiple sclerosis (MS) review, and assessing time efficiency.
METHODS: A database search yielded 5,058 records. The AI tool and manual review were applied in parallel across deduplication, Title/Abstract (TIAB) screening, and full-text (FT) screening. Sensitivity and specificity were calculated against manual review as the reference standard.
RESULTS: Deduplication was completed in comparable time by both approaches. TIAB sensitivity was 69%, improving to 74% after prompt refinement, with specificity declining from 91% to 86%. FT sensitivity reached 96% after refinement. Although the AI-tool reduced TIAB processing from 10 to 3 days and FT screening from 8 to 2 days, this validation was conducted against a completed manual review, facilitating prompt refinement through clear reference findings. Without this, time savings may be offset by quality control demands, and prompt deficiency identification. Recurrent limitations included reproducibility inconsistencies, publication type misclassifications, population and outcome detection failures, and inability to de-merge uploaded PDFs underscoring that human-in-the-loop validation is essential at every stage. Furthermore, additional manual steps including conference abstract searching and PDF downloading remain outside the scope of AI-assisted automation, indicating that time savings are not feasible across all SLR processes.
CONCLUSIONS: In this SLR, the AI-assisted tool demonstrated meaningful time savings and acceptable sensitivity; however, human oversight was indispensable. Notably, limitations observed for a review assessing straightforward mortality outcome suggest that the performance, accuracy, and true time savings of AI tools in more complex, multidimensional SLRs encompassing heterogeneous outcomes, diverse study designs, and nuanced inclusion criteria warrants rigorous further examination in evidence synthesis workflows.
METHODS: A database search yielded 5,058 records. The AI tool and manual review were applied in parallel across deduplication, Title/Abstract (TIAB) screening, and full-text (FT) screening. Sensitivity and specificity were calculated against manual review as the reference standard.
RESULTS: Deduplication was completed in comparable time by both approaches. TIAB sensitivity was 69%, improving to 74% after prompt refinement, with specificity declining from 91% to 86%. FT sensitivity reached 96% after refinement. Although the AI-tool reduced TIAB processing from 10 to 3 days and FT screening from 8 to 2 days, this validation was conducted against a completed manual review, facilitating prompt refinement through clear reference findings. Without this, time savings may be offset by quality control demands, and prompt deficiency identification. Recurrent limitations included reproducibility inconsistencies, publication type misclassifications, population and outcome detection failures, and inability to de-merge uploaded PDFs underscoring that human-in-the-loop validation is essential at every stage. Furthermore, additional manual steps including conference abstract searching and PDF downloading remain outside the scope of AI-assisted automation, indicating that time savings are not feasible across all SLR processes.
CONCLUSIONS: In this SLR, the AI-assisted tool demonstrated meaningful time savings and acceptable sensitivity; however, human oversight was indispensable. Notably, limitations observed for a review assessing straightforward mortality outcome suggest that the performance, accuracy, and true time savings of AI tools in more complex, multidimensional SLRs encompassing heterogeneous outcomes, diverse study designs, and nuanced inclusion criteria warrants rigorous further examination in evidence synthesis workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EPH128
Topic
Epidemiology & Public Health, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Safety & Pharmacoepidemiology
Disease
Neurological Disorders, No Additional Disease & Conditions/Specialized Treatment Areas