EVALUATING AI-ASSISTED SYSTEMATIC LITERATURE REVIEW: A COCHRANE BASED COMPARISON ACROSS SEARCH, SCREENING, AND EXTRACTION
Author(s)
Wing Yu Tang, MPH1, Joseph C. Cappelleri, MPH, MS, PhD2, Haitao Chu, PhD, MD3, Jarjieh Fang, MPH4, Pradyumna Shyama Prasad, .5, Ben Rachbach, BA5, Hamsa Pillai, BS5, Tylar Tannenbaum, BA5.
1Pfizer, New York, NY, USA, 2Pfizer, Newington, CT, USA, 3Pfizer, Edina, MN, USA, 4Pfizer, Brooklyn, NY, USA, 5Elicit, Covina, CA, USA.
1Pfizer, New York, NY, USA, 2Pfizer, Newington, CT, USA, 3Pfizer, Edina, MN, USA, 4Pfizer, Brooklyn, NY, USA, 5Elicit, Covina, CA, USA.
OBJECTIVES: Transparent, stage-specific validation is important to AI-assisted systematic literature review workflows. We evaluated the performance of an AI tool workflow against Cochrane reviews across search, abstract screening, full-text screening, and extraction.
METHODS: Reference datasets were drawn from open-access Cochrane reviews spanning 12 MeSH areas, using published decisions on inclusion, exclusion, and data extraction as the standard for comparison. For search, review titles were used as semantic-search queries, with results evaluated against the included studies from each review that had a resolvable DOI [digital object identifier] (7,532 studies; 888 reviews). Abstract screening was tested by reconstructing MEDLINE searches to produce candidate papers, the workflow classified papers against each review’s eligibility criteria (108 reviews; 931 positives, 5,162 negatives). For full-text screening, open-access PDFs were classified against eligibility criteria at both the paper and individual-criterion level (74 reviews; 377 papers). Lastly, extraction was evaluated by reconstructing review-specific questions for methods, participants, and interventions, comparing outputs to published study characteristics tables (28 reviews; 319 tasks). Search metrics were averaged across reviews; screening and extraction metrics were pooled across studies.
RESULTS: At the search stage, the mean recall was 95.0%, with the majority of reviews able to achieve full recall. Abstract screening performed strongly across all metrics: 96.9% sensitivity, 92.5% specificity, 70.1% positive predictive value (PPV), 99.4% negative predictive value (NPV), and 93.2% accuracy. Full-text screening achieved 99.5% paper-level sensitivity, 70.1% specificity, 81.2% PPV, 99.1% NPV, and 86.7% accuracy; criterion-level accuracy was 94.8%. Extraction accuracy was 95.6%.
CONCLUSIONS: This AI-assisted workflow demonstrated strong performance across retrieval, screening, and extraction. Evaluations like this one can help establish a practical benchmark for measuring AI-assisted literature reviews against established methods and for assessing performance across different AI solutions. Future work should continue standardizing AI evaluation frameworks and exploring broader implications for systematic reviews and meta-analyses.
METHODS: Reference datasets were drawn from open-access Cochrane reviews spanning 12 MeSH areas, using published decisions on inclusion, exclusion, and data extraction as the standard for comparison. For search, review titles were used as semantic-search queries, with results evaluated against the included studies from each review that had a resolvable DOI [digital object identifier] (7,532 studies; 888 reviews). Abstract screening was tested by reconstructing MEDLINE searches to produce candidate papers, the workflow classified papers against each review’s eligibility criteria (108 reviews; 931 positives, 5,162 negatives). For full-text screening, open-access PDFs were classified against eligibility criteria at both the paper and individual-criterion level (74 reviews; 377 papers). Lastly, extraction was evaluated by reconstructing review-specific questions for methods, participants, and interventions, comparing outputs to published study characteristics tables (28 reviews; 319 tasks). Search metrics were averaged across reviews; screening and extraction metrics were pooled across studies.
RESULTS: At the search stage, the mean recall was 95.0%, with the majority of reviews able to achieve full recall. Abstract screening performed strongly across all metrics: 96.9% sensitivity, 92.5% specificity, 70.1% positive predictive value (PPV), 99.4% negative predictive value (NPV), and 93.2% accuracy. Full-text screening achieved 99.5% paper-level sensitivity, 70.1% specificity, 81.2% PPV, 99.1% NPV, and 86.7% accuracy; criterion-level accuracy was 94.8%. Extraction accuracy was 95.6%.
CONCLUSIONS: This AI-assisted workflow demonstrated strong performance across retrieval, screening, and extraction. Evaluations like this one can help establish a practical benchmark for measuring AI-assisted literature reviews against established methods and for assessing performance across different AI solutions. Future work should continue standardizing AI evaluation frameworks and exploring broader implications for systematic reviews and meta-analyses.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR126
Topic
Methodological & Statistical Research, Organizational Practices, Real World Data & Information Systems
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas