ARTIFICIAL INTELLIGENCE FOR DATA EXTRACTION IN LIVE HEOR LITERATURE REVIEWS: A REAL-WORLD ANALYSIS ACROSS MORE THAN 20 REVIEWS
Author(s)
Saifuddin M. Kharawala1, Divyanshu Jindal, BCom MBA2, Neeti Chana, MPharm3, Ashish Som, BCA MBA2, Payal Rana, PhD3, Harveen Baxi, MPharm2, Monique Martin, MSc MBA4, Paul Gandhi, MD MBA LLM5.
1Bridge Medical Consulting Ltd, Tonbridge, United Kingdom, 2Red Nucleus, Delhi, India, 3RedNucleus, Delhi, India, 4Red Nucleus, London, United Kingdom, 5Red Nucleus, Yardley, PA, USA.
1Bridge Medical Consulting Ltd, Tonbridge, United Kingdom, 2Red Nucleus, Delhi, India, 3RedNucleus, Delhi, India, 4Red Nucleus, London, United Kingdom, 5Red Nucleus, Yardley, PA, USA.
OBJECTIVES: To evaluate the pre-quality-control (pre-QC) performance of generative artificial intelligence (AI) for data extraction, the most resource-intensive stage of health economics and outcomes research (HEOR) literature reviews.
METHODS: Since 2023, AI-supported extraction workflows were applied across more than 20 live systematic and targeted literature reviews, including clinical trial and observational/real-world evidence reviews. AI was used for initial extraction/categorization and full data extraction into Excel extraction grids and Word data tables. OpenAI GPT (generative pre-trained transformer) and Google Gemini models were used, with model versions evolving over time. AI processing was restricted to publicly available or free-to-access full texts. Prompting approaches were refined through extensive R&D testing and live project implementation, with further optimization for the specific requirements of each review. Each AI output underwent 100% human QC, enabling direct identification of pre-QC AI errors. Completeness was defined as the proportion of required elements extracted by AI, and accuracy as exact match to the human-QC value and unit.
RESULTS: Across more than 20 live projects, AI performance was consistently high before human QC. Initial extraction/categorization achieved median sensitivity of approximately 96% (range: approximately 92-100%) and median accuracy of approximately 98% (range: approximately 94-99%). Full extraction achieved median completeness of approximately 96% (range: approximately 80-100%) and median accuracy of approximately 98% (range: approximately 90-100%). Completeness was more variable than accuracy, with lower values observed in selected high-complexity extraction tasks.
CONCLUSIONS: In routine HEOR literature-review projects, generative AI demonstrated high pre-QC completeness and accuracy for extraction across multiple review types, output formats, and workflow stages. Performance should be interpreted as reflecting both model capability and accumulated experience in review-specific prompt optimization. The findings support AI as a scalable augmentation tool for data-intensive SLR extraction tasks, while reinforcing the need for expert-led 100% QC before outputs are used for decision-grade evidence synthesis.
METHODS: Since 2023, AI-supported extraction workflows were applied across more than 20 live systematic and targeted literature reviews, including clinical trial and observational/real-world evidence reviews. AI was used for initial extraction/categorization and full data extraction into Excel extraction grids and Word data tables. OpenAI GPT (generative pre-trained transformer) and Google Gemini models were used, with model versions evolving over time. AI processing was restricted to publicly available or free-to-access full texts. Prompting approaches were refined through extensive R&D testing and live project implementation, with further optimization for the specific requirements of each review. Each AI output underwent 100% human QC, enabling direct identification of pre-QC AI errors. Completeness was defined as the proportion of required elements extracted by AI, and accuracy as exact match to the human-QC value and unit.
RESULTS: Across more than 20 live projects, AI performance was consistently high before human QC. Initial extraction/categorization achieved median sensitivity of approximately 96% (range: approximately 92-100%) and median accuracy of approximately 98% (range: approximately 94-99%). Full extraction achieved median completeness of approximately 96% (range: approximately 80-100%) and median accuracy of approximately 98% (range: approximately 90-100%). Completeness was more variable than accuracy, with lower values observed in selected high-complexity extraction tasks.
CONCLUSIONS: In routine HEOR literature-review projects, generative AI demonstrated high pre-QC completeness and accuracy for extraction across multiple review types, output formats, and workflow stages. Performance should be interpreted as reflecting both model capability and accumulated experience in review-specific prompt optimization. The findings support AI as a scalable augmentation tool for data-intensive SLR extraction tasks, while reinforcing the need for expert-led 100% QC before outputs are used for decision-grade evidence synthesis.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
CO109
Topic
Clinical Outcomes, Economic Evaluation, Methodological & Statistical Research
Disease
No Additional Disease & Conditions/Specialized Treatment Areas