DEVELOPMENT AND VALIDATION OF AN AI-ENABLED PLATFORM FOR AUTOMATING SYSTEMATIC LITERATURE REVIEWS: MULTI-AGENT ACCURACY ASSESSMENT ACROSS SCREENING AND DATA EXTRACTION
Author(s)
Jagriti Prasad, MPH1, Tejesh S, MPH1, Mona Thangamma, MSc Health Economics1, Ullas Ulahannan, MPH1, Prabhu Dutta Shaw, MPH, MBA1, Ankit Dhaundiyal, MS2, Megha Tharad, PhD3.
1Evalueserve Pvt Ltd, Bengaluru, India, 2Evalueserve GmbH, Rheinbach, Germany, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
1Evalueserve Pvt Ltd, Bengaluru, India, 2Evalueserve GmbH, Rheinbach, Germany, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
OBJECTIVES: To develop and validate an artificial intelligence (AI)-enabled platform for automating key systematic literature review (SLR) activities and evaluate the accuracy and safety of integrated screening and data extraction agents against human-reviewed benchmarks across multiple disease indications.
METHODS: An AI-enabled SLR platform was developed comprising two task-specific agents: (1) a title/abstract screening agent for first-pass study selection and (2) a data extraction agent for structured field-level extraction. Screening performance was evaluated across five indications using 989 studies. AI first-pass decisions were compared with final human-reviewed outcomes following full-text assessment. Outcomes were categorized as concordant exclusions, concordant inclusions, AI over-inclusions subsequently excluded by humans, and AI exclusions later included by humans (critical misses). Re-adjusted accuracy reflected alignment with final human-reviewed outcomes, while miss rate represented the proportion of critical misses. Data extraction accuracy was assessed through manual quality control across 27 variables by comparing AI outputs with source publications.
RESULTS: Across 989 studies, the screening agent achieved 97.8% alignment with final human-reviewed outcomes and a critical miss rate of 2.22%. Most discordance was attributable to conservative AI over-inclusion rather than exclusion of relevant evidence. Miss rates varied across indications, reflecting relatively small indication-level sample sizes. The data extraction agent achieved 91.9% accuracy across 27 structured variables. Errors were concentrated in complex or context-dependent fields, whereas core epidemiologic, study-design, and population characteristics demonstrated high extraction accuracy.
CONCLUSIONS: The AI-enabled SLR platform demonstrated high alignment with final human-reviewed outcomes, a low critical miss rate, and robust data extraction performance. Findings support a human-in-the-loop, AI-augmented approach to accelerate evidence synthesis while maintaining methodological rigor for HTA and evidence-generation workflows. Prospective evaluations should assess performance across additional indications, review designs, and real-world operational settings.
METHODS: An AI-enabled SLR platform was developed comprising two task-specific agents: (1) a title/abstract screening agent for first-pass study selection and (2) a data extraction agent for structured field-level extraction. Screening performance was evaluated across five indications using 989 studies. AI first-pass decisions were compared with final human-reviewed outcomes following full-text assessment. Outcomes were categorized as concordant exclusions, concordant inclusions, AI over-inclusions subsequently excluded by humans, and AI exclusions later included by humans (critical misses). Re-adjusted accuracy reflected alignment with final human-reviewed outcomes, while miss rate represented the proportion of critical misses. Data extraction accuracy was assessed through manual quality control across 27 variables by comparing AI outputs with source publications.
RESULTS: Across 989 studies, the screening agent achieved 97.8% alignment with final human-reviewed outcomes and a critical miss rate of 2.22%. Most discordance was attributable to conservative AI over-inclusion rather than exclusion of relevant evidence. Miss rates varied across indications, reflecting relatively small indication-level sample sizes. The data extraction agent achieved 91.9% accuracy across 27 structured variables. Errors were concentrated in complex or context-dependent fields, whereas core epidemiologic, study-design, and population characteristics demonstrated high extraction accuracy.
CONCLUSIONS: The AI-enabled SLR platform demonstrated high alignment with final human-reviewed outcomes, a low critical miss rate, and robust data extraction performance. Findings support a human-in-the-loop, AI-augmented approach to accelerate evidence synthesis while maintaining methodological rigor for HTA and evidence-generation workflows. Prospective evaluations should assess performance across additional indications, review designs, and real-world operational settings.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR222
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas