EVALUATING THE RELIABILITY OF GENERATIVE AI FOR LITERATURE SCREENING: A COMPARATIVE ANALYSIS OF MANUAL, HYBRID, AND FULLY AUTOMATED APPROACHES
Author(s)
Keziah Dutt1, Julie Boelen, BSc, MPhil2, Sarah Dewilde, PhD3.
1Statistician, Services in Health Economics, Belgium, 2Services in Health Economics BV, Brussels, Belgium, 3SHE Bv, Brussels, Belgium.
1Statistician, Services in Health Economics, Belgium, 2Services in Health Economics BV, Brussels, Belgium, 3SHE Bv, Brussels, Belgium.
OBJECTIVES: Artificial Intelligence is increasingly applied to labor-intensive tasks including systematic literature reviews. However, widespread adoption requires rigorous, scientific evaluation to ensure reliability and safeguard research quality. This study compares manual, hybrid, and fully AI-driven screening approaches to assess performance and potential risks.
METHODS: A previously completed manual literature review was used as reference. Literature screenings based on identical study design and research questions were implemented using an AI tool in two use cases. In the “AI-only” approach, the tool autonomously generated search queries, in- and exclusion criteria, and conducted abstract and full-text screening. In the “hybrid” approach, AI-generated queries and criteria were refined by a researcher, and included studies underwent additional manual validation. Both approaches enabled the AI to automatically exclude articles that did not meet all defined eligibility conditions. Manual and hybrid data extractions were then compared for 17 predefined data elements, including study details and rates of adverse events. It was assumed that the manual data extraction was the gold standard and was complete.
RESULTS: Across all approaches, 58 studies were identified, 41 of which were manually verified as relevant. The sensitivity of the manual, hybrid, and AI-only approaches are 87.5%, 67.5%, and 30% respectively, and the AI-only approach included 17 non-relevant studies. Overall, the hybrid method correctly extracted 52.9% of data. The number of patients without side effects and follow-up time were correctly extracted by the hybrid approach in 23.9% and 25.4% of studies, whereas extraction accuracy for selected adverse events ranged from 44.8% to 61.2% for the hybrid method.
CONCLUSIONS: The manual approach demonstrated the highest sensitivity, identifying the greatest number of relevant studies not captured by AI-assisted methods. While AI approaches substantially reduce workload and time, they do introduce errors and omissions that impact evidence, completeness, and validity. Hybrid approaches may offer a balance between efficiency and oversight.
METHODS: A previously completed manual literature review was used as reference. Literature screenings based on identical study design and research questions were implemented using an AI tool in two use cases. In the “AI-only” approach, the tool autonomously generated search queries, in- and exclusion criteria, and conducted abstract and full-text screening. In the “hybrid” approach, AI-generated queries and criteria were refined by a researcher, and included studies underwent additional manual validation. Both approaches enabled the AI to automatically exclude articles that did not meet all defined eligibility conditions. Manual and hybrid data extractions were then compared for 17 predefined data elements, including study details and rates of adverse events. It was assumed that the manual data extraction was the gold standard and was complete.
RESULTS: Across all approaches, 58 studies were identified, 41 of which were manually verified as relevant. The sensitivity of the manual, hybrid, and AI-only approaches are 87.5%, 67.5%, and 30% respectively, and the AI-only approach included 17 non-relevant studies. Overall, the hybrid method correctly extracted 52.9% of data. The number of patients without side effects and follow-up time were correctly extracted by the hybrid approach in 23.9% and 25.4% of studies, whereas extraction accuracy for selected adverse events ranged from 44.8% to 61.2% for the hybrid method.
CONCLUSIONS: The manual approach demonstrated the highest sensitivity, identifying the greatest number of relevant studies not captured by AI-assisted methods. While AI approaches substantially reduce workload and time, they do introduce errors and omissions that impact evidence, completeness, and validity. Hybrid approaches may offer a balance between efficiency and oversight.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR286
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas