AI-ASSISTED VS MANUAL TARGETED LITERATURE REVIEWS: A CASE STUDY EVALUATING EFFICIENCY, ACCURACY, AND COMPLETENESS
Author(s)
Kota Vidyasagar, M.Pharm1, Amitkumar Pagada, M.Pharm1, Lisa SUTOUR, PharmD2, Nupur Greene, PhD2, Hugo Dubucq, MPH3.
1Sanofi Healthcare India Private Limited, Hyderabad, India, 2Sanofi US Services Inc., Cambridge, MA, USA, 3Sanofi-Aventis SA (Spain), Barcelona, Spain.
1Sanofi Healthcare India Private Limited, Hyderabad, India, 2Sanofi US Services Inc., Cambridge, MA, USA, 3Sanofi-Aventis SA (Spain), Barcelona, Spain.
OBJECTIVES: Traditional targeted literature reviews (TLRs) are resource-intensive, creating opportunities for artificial intelligence (AI) to improve efficiency, reduce manual effort, and accelerate evidence synthesis. This study compared AI-assisted and manual TLR approaches in terms of screening performance, data extraction accuracy, and overall efficiency in a case study evaluating the efficacy and safety of a targeted product in a selected disease.
METHODS: Leveraging a protocol-derived prompt, the generative AI tool screened publications retrieved from key electronic databases. The AI extracted predefined data on study characteristics, baseline demographics, and key efficacy and safety outcomes. Inclusion criteria were defined a priori and included relevant study design, population and outcome of interest. AI performance was benchmarked against an independent manual review using pre-defined performance metrics. Sensitivity (relevant studies correctly identified) and specificity (irrelevant studies correctly excluded) were measured against manual review as the reference. All AI-generated screening decisions and extracted data underwent 100% human quality validation by expert reviewers.
RESULTS: Among 603 publications, AI-assisted screening achieved 69% sensitivity and 78% specificity across title/abstract and full-text stages. Lower sensitivity indicates a risk of missed relevant studies without human validation. The combined AI-assisted screening and expert validation process identified 40 publications from 25 studies in 7 days, compared to 9 days with manual screening. For data extraction, AI reduced expert review time by 34% (8 vs. 12 days), with 70% completeness across 28 key data variables and 75% accuracy within the completed fields. These findings highlight a trade-off between efficiency gains and completeness and accuracy of extracted evidence. Key limitations included inability to extract data from figures, difficulties in processing conference proceedings and inconsistent structuring of extracted outputs, requiring additional manual intervention.
CONCLUSIONS: The AI-assisted workflows can improve efficiency in TLRs. However, moderate sensitivity and limitations in data extraction underscore the need for human oversight, particularly in complex evidence landscapes.
METHODS: Leveraging a protocol-derived prompt, the generative AI tool screened publications retrieved from key electronic databases. The AI extracted predefined data on study characteristics, baseline demographics, and key efficacy and safety outcomes. Inclusion criteria were defined a priori and included relevant study design, population and outcome of interest. AI performance was benchmarked against an independent manual review using pre-defined performance metrics. Sensitivity (relevant studies correctly identified) and specificity (irrelevant studies correctly excluded) were measured against manual review as the reference. All AI-generated screening decisions and extracted data underwent 100% human quality validation by expert reviewers.
RESULTS: Among 603 publications, AI-assisted screening achieved 69% sensitivity and 78% specificity across title/abstract and full-text stages. Lower sensitivity indicates a risk of missed relevant studies without human validation. The combined AI-assisted screening and expert validation process identified 40 publications from 25 studies in 7 days, compared to 9 days with manual screening. For data extraction, AI reduced expert review time by 34% (8 vs. 12 days), with 70% completeness across 28 key data variables and 75% accuracy within the completed fields. These findings highlight a trade-off between efficiency gains and completeness and accuracy of extracted evidence. Key limitations included inability to extract data from figures, difficulties in processing conference proceedings and inconsistent structuring of extracted outputs, requiring additional manual intervention.
CONCLUSIONS: The AI-assisted workflows can improve efficiency in TLRs. However, moderate sensitivity and limitations in data extraction underscore the need for human oversight, particularly in complex evidence landscapes.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA38
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
No Additional Disease & Conditions/Specialized Treatment Areas