EVALUATING AI-ASSISTED LITERATURE REVIEW PLATFORMS AT TASK-LEVEL AGAINST HUMAN-LED BASELINES
Author(s)
Rajshree Pandey, MPH, PhD1, Shobhit Sood, PharmD2, Hussein Jeafar, PharmD2, Nikhil Bhatia, MBA, PharmD3, Arpita Nag, PhD, MBA, MS, PharmD1, Yiduo Zhang, BA, MA, PhD4.
1Alexion AstraZeneca Rare Disease, Boston, MA, USA, 2Rutgers University/ AstraZeneca, Gaithersburg, MD, USA, 3AstraZeneca, Gaithersburg, MD, USA, 4AstraZeneca, Barcelona, Spain.
1Alexion AstraZeneca Rare Disease, Boston, MA, USA, 2Rutgers University/ AstraZeneca, Gaithersburg, MD, USA, 3AstraZeneca, Gaithersburg, MD, USA, 4AstraZeneca, Barcelona, Spain.
OBJECTIVES: Literature reviews (LR; targeted or systematic) are labor intensive to conduct for health technology developers. Artificial intelligence (AI)-enabled platforms lack evidence on probabilistic outputs or task-level performance against human standards. Past AI evaluations assessed single LR steps; none tested AI across all stages of LR. We developed a scalable, task-level framework evaluating AI-enabled LR platform accuracy and efficiency against human-led baseline.
METHODS: An eligibility criterion was created, which guided platform selection; chosen platforms were binary-scored against 29 literature-validated features. AI-enabled results for title/abstract screening, full-text screening, and data extraction were benchmarked against three human-led LRs, with sensitivity, specificity, accuracy, precision, F1 (composite balance of sensitivity and precision), and error rate measured at task-level. Sensitivity was prioritized as the most critical measure, as missing a relevant publication represents the highest methodological risk in LRs. Efficiency by time saved versus human dual review was also assessed by AI-facing task.
RESULTS: Ten commercial AI-enabled LR platforms were included using a single decision path: platform eligibility screening (pass/fail gate), feature scoring (differentiation), and task-level metrics (benchmarking against human gold standard). An initial evaluation demonstrated 33%-47% time savings but required human oversight. Performance was task-specific: title and abstract screening sensitivity ranged from 46%-100%, full-text review sensitivity from 47%-98% (F1 up to 0.89), and data extraction accuracy from 30%-81%. However, low specificity across tasks meant high false positives and requiring human adjudication of data before use. These metrics, agnostic to review type, suggest performance transferability across diverse LR types. Detailed results will be presented in the poster.
CONCLUSIONS: AI delivered substantial time savings, but sensitivity performance varied 46%-100%. This further demonstrates that human-in-the-loop oversight is non-negotiable for LR: platforms required human adjudication due to low specificity. Time efficiency gains are real when paired with mandatory human quality assurance to ensure methodological rigor.
METHODS: An eligibility criterion was created, which guided platform selection; chosen platforms were binary-scored against 29 literature-validated features. AI-enabled results for title/abstract screening, full-text screening, and data extraction were benchmarked against three human-led LRs, with sensitivity, specificity, accuracy, precision, F1 (composite balance of sensitivity and precision), and error rate measured at task-level. Sensitivity was prioritized as the most critical measure, as missing a relevant publication represents the highest methodological risk in LRs. Efficiency by time saved versus human dual review was also assessed by AI-facing task.
RESULTS: Ten commercial AI-enabled LR platforms were included using a single decision path: platform eligibility screening (pass/fail gate), feature scoring (differentiation), and task-level metrics (benchmarking against human gold standard). An initial evaluation demonstrated 33%-47% time savings but required human oversight. Performance was task-specific: title and abstract screening sensitivity ranged from 46%-100%, full-text review sensitivity from 47%-98% (F1 up to 0.89), and data extraction accuracy from 30%-81%. However, low specificity across tasks meant high false positives and requiring human adjudication of data before use. These metrics, agnostic to review type, suggest performance transferability across diverse LR types. Detailed results will be presented in the poster.
CONCLUSIONS: AI delivered substantial time savings, but sensitivity performance varied 46%-100%. This further demonstrates that human-in-the-loop oversight is non-negotiable for LR: platforms required human adjudication due to low specificity. Time efficiency gains are real when paired with mandatory human quality assurance to ensure methodological rigor.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR237
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas