HUMAN VERSUS ARTIFICIAL INTELLIGENCE FOR DATA EXTRACTION IN SYSTEMATIC LITERATURE REVIEWS: A REVIEW OF ACCURACY, EFFICIENCY, AND THE EVOLVING ROLE OF HUMAN REVIEWERS
Author(s)
Nitendra K. Mishra, M.S1, Sonam Vats, M Pharm1, Xuan Wang, MD2.
1ICON, Bengaluru, India, 2ICON, Taby, Sweden.
1ICON, Bengaluru, India, 2ICON, Taby, Sweden.
OBJECTIVES: Data extraction is a resource-intensive step in systematic literature reviews (SLRs). Advances in artificial intelligence (AI), particularly large language models (LLMs), have enabled automated extraction; however, concerns remain regarding reliability and the need for human validation. This review evaluated AI-assisted data extraction and the role of human oversight in evidence synthesis.
METHODS: A targeted literature review identified studies evaluating AI-based data extraction in SLRs. MEDLINE, EMBASE, and the Cochrane Database of Systematic Reviews were searched on 20 May 2026 for publications from January 2025 onward to include latest evidence. Studies evaluating AI-assisted or automated extraction against human review or validation were included. Outcomes included extraction accuracy, precision, recall, concordance, and time efficiency. Findings were synthesized narratively.
RESULTS: The search identified 537 records; 80 studies underwent full-text review, and 7 met the inclusion criteria. ChatGPT/OpenAI models were evaluated in four studies, Gemini in three, and Elicit and DeepSeek in one each. AI consistently improved efficiency while maintaining generally high accuracy. One study reported fully correct extraction in 82% of fields, with at least one model correctly extracting 95% of fields in 27-36 minutes. AI-assisted extraction achieved similar accuracy to human-only extraction (91.0% vs. 89.0%) while reducing extraction time by a median of 41 minutes per study. Another study reported precision, recall, and F1-scores of 98.2%, 96.6%, and 97.4%, respectively, reducing extraction time from approximately 240 minutes to 4.5 minutes per article. However, accuracy varied across studies and data types, ranging from 37.0% to 66.7% in regulatory and clinical trial extraction tasks. Hallucinations, missing data, and errors in complex outcomes remained common.
CONCLUSIONS: AI-assisted data extraction improves SLR efficiency while achieving accuracy comparable to human extraction. However, performance variability and persistent errors indicate that human oversight remains essential. Human-in-the-loop workflows may best balance efficiency, accuracy, and methodological rigor.
METHODS: A targeted literature review identified studies evaluating AI-based data extraction in SLRs. MEDLINE, EMBASE, and the Cochrane Database of Systematic Reviews were searched on 20 May 2026 for publications from January 2025 onward to include latest evidence. Studies evaluating AI-assisted or automated extraction against human review or validation were included. Outcomes included extraction accuracy, precision, recall, concordance, and time efficiency. Findings were synthesized narratively.
RESULTS: The search identified 537 records; 80 studies underwent full-text review, and 7 met the inclusion criteria. ChatGPT/OpenAI models were evaluated in four studies, Gemini in three, and Elicit and DeepSeek in one each. AI consistently improved efficiency while maintaining generally high accuracy. One study reported fully correct extraction in 82% of fields, with at least one model correctly extracting 95% of fields in 27-36 minutes. AI-assisted extraction achieved similar accuracy to human-only extraction (91.0% vs. 89.0%) while reducing extraction time by a median of 41 minutes per study. Another study reported precision, recall, and F1-scores of 98.2%, 96.6%, and 97.4%, respectively, reducing extraction time from approximately 240 minutes to 4.5 minutes per article. However, accuracy varied across studies and data types, ranging from 37.0% to 66.7% in regulatory and clinical trial extraction tasks. Hallucinations, missing data, and errors in complex outcomes remained common.
CONCLUSIONS: AI-assisted data extraction improves SLR efficiency while achieving accuracy comparable to human extraction. However, performance variability and persistent errors indicate that human oversight remains essential. Human-in-the-loop workflows may best balance efficiency, accuracy, and methodological rigor.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR183
Topic
Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas