EVALUATING AI-ASSISTED SYSTEMATIC LITERATURE REVIEW: A COCHRANE BASED COMPARISON ACROSS SEARCH, SCREENING, AND EXTRACTION

Author(s)

Wing Yu Tang, MPH1, Joseph C. Cappelleri, MPH, MS, PhD2, Haitao Chu, PhD, MD3, Jarjieh Fang, MPH4, Pradyumna Shyama Prasad, .5, Ben Rachbach, BA5, Hamsa Pillai, BS5, Tylar Tannenbaum, BA5.
1Pfizer, New York, NY, USA, 2Pfizer, Newington, CT, USA, 3Pfizer, Edina, MN, USA, 4Pfizer, Brooklyn, NY, USA, 5Elicit, Covina, CA, USA.
OBJECTIVES: Transparent, stage-specific validation is important to AI-assisted systematic literature review workflows. We evaluated the performance of an AI tool workflow against Cochrane reviews across search, abstract screening, full-text screening, and extraction.
METHODS: Reference datasets were drawn from open-access Cochrane reviews spanning 12 MeSH areas, using published decisions on inclusion, exclusion, and data extraction as the standard for comparison. For search, review titles were used as semantic-search queries, with results evaluated against the included studies from each review that had a resolvable DOI [digital object identifier] (7,532 studies; 888 reviews). Abstract screening was tested by reconstructing MEDLINE searches to produce candidate papers, the workflow classified papers against each review’s eligibility criteria (108 reviews; 931 positives, 5,162 negatives). For full-text screening, open-access PDFs were classified against eligibility criteria at both the paper and individual-criterion level (74 reviews; 377 papers). Lastly, extraction was evaluated by reconstructing review-specific questions for methods, participants, and interventions, comparing outputs to published study characteristics tables (28 reviews; 319 tasks). Search metrics were averaged across reviews; screening and extraction metrics were pooled across studies.
RESULTS: At the search stage, the mean recall was 95.0%, with the majority of reviews able to achieve full recall. Abstract screening performed strongly across all metrics: 96.9% sensitivity, 92.5% specificity, 70.1% positive predictive value (PPV), 99.4% negative predictive value (NPV), and 93.2% accuracy. Full-text screening achieved 99.5% paper-level sensitivity, 70.1% specificity, 81.2% PPV, 99.1% NPV, and 86.7% accuracy; criterion-level accuracy was 94.8%. Extraction accuracy was 95.6%.
CONCLUSIONS: This AI-assisted workflow demonstrated strong performance across retrieval, screening, and extraction. Evaluations like this one can help establish a practical benchmark for measuring AI-assisted literature reviews against established methods and for assessing performance across different AI solutions. Future work should continue standardizing AI evaluation frameworks and exploring broader implications for systematic reviews and meta-analyses.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR126

Topic

Methodological & Statistical Research, Organizational Practices, Real World Data & Information Systems

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×