EVALUATING THE RELIABILITY OF GENERATIVE AI FOR LITERATURE SCREENING: A COMPARATIVE ANALYSIS OF MANUAL, HYBRID, AND FULLY AUTOMATED APPROACHES

Author(s)

Keziah Dutt1, Julie Boelen, BSc, MPhil2, Sarah Dewilde, PhD3.
1Statistician, Services in Health Economics, Belgium, 2Services in Health Economics BV, Brussels, Belgium, 3SHE Bv, Brussels, Belgium.
OBJECTIVES: Artificial Intelligence is increasingly applied to labor-intensive tasks including systematic literature reviews. However, widespread adoption requires rigorous, scientific evaluation to ensure reliability and safeguard research quality. This study compares manual, hybrid, and fully AI-driven screening approaches to assess performance and potential risks.
METHODS: A previously completed manual literature review was used as reference. Literature screenings based on identical study design and research questions were implemented using an AI tool in two use cases. In the “AI-only” approach, the tool autonomously generated search queries, in- and exclusion criteria, and conducted abstract and full-text screening. In the “hybrid” approach, AI-generated queries and criteria were refined by a researcher, and included studies underwent additional manual validation. Both approaches enabled the AI to automatically exclude articles that did not meet all defined eligibility conditions. Manual and hybrid data extractions were then compared for 17 predefined data elements, including study details and rates of adverse events. It was assumed that the manual data extraction was the gold standard and was complete.
RESULTS: Across all approaches, 58 studies were identified, 41 of which were manually verified as relevant. The sensitivity of the manual, hybrid, and AI-only approaches are 87.5%, 67.5%, and 30% respectively, and the AI-only approach included 17 non-relevant studies. Overall, the hybrid method correctly extracted 52.9% of data. The number of patients without side effects and follow-up time were correctly extracted by the hybrid approach in 23.9% and 25.4% of studies, whereas extraction accuracy for selected adverse events ranged from 44.8% to 61.2% for the hybrid method.
CONCLUSIONS: The manual approach demonstrated the highest sensitivity, identifying the greatest number of relevant studies not captured by AI-assisted methods. While AI approaches substantially reduce workload and time, they do introduce errors and omissions that impact evidence, completeness, and validity. Hybrid approaches may offer a balance between efficiency and oversight.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR286

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×