HUMAN VERSUS ARTIFICIAL INTELLIGENCE FOR DATA EXTRACTION IN SYSTEMATIC LITERATURE REVIEWS: A REVIEW OF ACCURACY, EFFICIENCY, AND THE EVOLVING ROLE OF HUMAN REVIEWERS

Author(s)

Nitendra K. Mishra, M.S1, Sonam Vats, M Pharm1, Xuan Wang, MD2.
1ICON, Bengaluru, India, 2ICON, Taby, Sweden.
OBJECTIVES: Data extraction is a resource-intensive step in systematic literature reviews (SLRs). Advances in artificial intelligence (AI), particularly large language models (LLMs), have enabled automated extraction; however, concerns remain regarding reliability and the need for human validation. This review evaluated AI-assisted data extraction and the role of human oversight in evidence synthesis.
METHODS: A targeted literature review identified studies evaluating AI-based data extraction in SLRs. MEDLINE, EMBASE, and the Cochrane Database of Systematic Reviews were searched on 20 May 2026 for publications from January 2025 onward to include latest evidence. Studies evaluating AI-assisted or automated extraction against human review or validation were included. Outcomes included extraction accuracy, precision, recall, concordance, and time efficiency. Findings were synthesized narratively.
RESULTS: The search identified 537 records; 80 studies underwent full-text review, and 7 met the inclusion criteria. ChatGPT/OpenAI models were evaluated in four studies, Gemini in three, and Elicit and DeepSeek in one each. AI consistently improved efficiency while maintaining generally high accuracy. One study reported fully correct extraction in 82% of fields, with at least one model correctly extracting 95% of fields in 27-36 minutes. AI-assisted extraction achieved similar accuracy to human-only extraction (91.0% vs. 89.0%) while reducing extraction time by a median of 41 minutes per study. Another study reported precision, recall, and F1-scores of 98.2%, 96.6%, and 97.4%, respectively, reducing extraction time from approximately 240 minutes to 4.5 minutes per article. However, accuracy varied across studies and data types, ranging from 37.0% to 66.7% in regulatory and clinical trial extraction tasks. Hallucinations, missing data, and errors in complex outcomes remained common.
CONCLUSIONS: AI-assisted data extraction improves SLR efficiency while achieving accuracy comparable to human extraction. However, performance variability and persistent errors indicate that human oversight remains essential. Human-in-the-loop workflows may best balance efficiency, accuracy, and methodological rigor.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR183

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×