ARTIFICIAL INTELLIGENCE FOR DATA EXTRACTION IN LIVE HEOR LITERATURE REVIEWS: A REAL-WORLD ANALYSIS ACROSS MORE THAN 20 REVIEWS

Author(s)

Saifuddin M. Kharawala1, Divyanshu Jindal, BCom MBA2, Neeti Chana, MPharm3, Ashish Som, BCA MBA2, Payal Rana, PhD3, Harveen Baxi, MPharm2, Monique Martin, MSc MBA4, Paul Gandhi, MD MBA LLM5.
1Bridge Medical Consulting Ltd, Tonbridge, United Kingdom, 2Red Nucleus, Delhi, India, 3RedNucleus, Delhi, India, 4Red Nucleus, London, United Kingdom, 5Red Nucleus, Yardley, PA, USA.
OBJECTIVES: To evaluate the pre-quality-control (pre-QC) performance of generative artificial intelligence (AI) for data extraction, the most resource-intensive stage of health economics and outcomes research (HEOR) literature reviews.
METHODS: Since 2023, AI-supported extraction workflows were applied across more than 20 live systematic and targeted literature reviews, including clinical trial and observational/real-world evidence reviews. AI was used for initial extraction/categorization and full data extraction into Excel extraction grids and Word data tables. OpenAI GPT (generative pre-trained transformer) and Google Gemini models were used, with model versions evolving over time. AI processing was restricted to publicly available or free-to-access full texts. Prompting approaches were refined through extensive R&D testing and live project implementation, with further optimization for the specific requirements of each review. Each AI output underwent 100% human QC, enabling direct identification of pre-QC AI errors. Completeness was defined as the proportion of required elements extracted by AI, and accuracy as exact match to the human-QC value and unit.
RESULTS: Across more than 20 live projects, AI performance was consistently high before human QC. Initial extraction/categorization achieved median sensitivity of approximately 96% (range: approximately 92-100%) and median accuracy of approximately 98% (range: approximately 94-99%). Full extraction achieved median completeness of approximately 96% (range: approximately 80-100%) and median accuracy of approximately 98% (range: approximately 90-100%). Completeness was more variable than accuracy, with lower values observed in selected high-complexity extraction tasks.
CONCLUSIONS: In routine HEOR literature-review projects, generative AI demonstrated high pre-QC completeness and accuracy for extraction across multiple review types, output formats, and workflow stages. Performance should be interpreted as reflecting both model capability and accumulated experience in review-specific prompt optimization. The findings support AI as a scalable augmentation tool for data-intensive SLR extraction tasks, while reinforcing the need for expert-led 100% QC before outputs are used for decision-grade evidence synthesis.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

CO109

Topic

Clinical Outcomes, Economic Evaluation, Methodological & Statistical Research

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×