EVALUATING AI-ASSISTED LITERATURE REVIEW PLATFORMS AT TASK-LEVEL AGAINST HUMAN-LED BASELINES

Author(s)

Rajshree Pandey, MPH, PhD1, Shobhit Sood, PharmD2, Hussein Jeafar, PharmD2, Nikhil Bhatia, MBA, PharmD3, Arpita Nag, PhD, MBA, MS, PharmD1, Yiduo Zhang, BA, MA, PhD4.
1Alexion AstraZeneca Rare Disease, Boston, MA, USA, 2Rutgers University/ AstraZeneca, Gaithersburg, MD, USA, 3AstraZeneca, Gaithersburg, MD, USA, 4AstraZeneca, Barcelona, Spain.
OBJECTIVES: Literature reviews (LR; targeted or systematic) are labor intensive to conduct for health technology developers. Artificial intelligence (AI)-enabled platforms lack evidence on probabilistic outputs or task-level performance against human standards. Past AI evaluations assessed single LR steps; none tested AI across all stages of LR. We developed a scalable, task-level framework evaluating AI-enabled LR platform accuracy and efficiency against human-led baseline.
METHODS: An eligibility criterion was created, which guided platform selection; chosen platforms were binary-scored against 29 literature-validated features. AI-enabled results for title/abstract screening, full-text screening, and data extraction were benchmarked against three human-led LRs, with sensitivity, specificity, accuracy, precision, F1 (composite balance of sensitivity and precision), and error rate measured at task-level. Sensitivity was prioritized as the most critical measure, as missing a relevant publication represents the highest methodological risk in LRs. Efficiency by time saved versus human dual review was also assessed by AI-facing task.
RESULTS: Ten commercial AI-enabled LR platforms were included using a single decision path: platform eligibility screening (pass/fail gate), feature scoring (differentiation), and task-level metrics (benchmarking against human gold standard). An initial evaluation demonstrated 33%-47% time savings but required human oversight. Performance was task-specific: title and abstract screening sensitivity ranged from 46%-100%, full-text review sensitivity from 47%-98% (F1 up to 0.89), and data extraction accuracy from 30%-81%. However, low specificity across tasks meant high false positives and requiring human adjudication of data before use. These metrics, agnostic to review type, suggest performance transferability across diverse LR types. Detailed results will be presented in the poster.
CONCLUSIONS: AI delivered substantial time savings, but sensitivity performance varied 46%-100%. This further demonstrates that human-in-the-loop oversight is non-negotiable for LR: platforms required human adjudication due to low specificity. Time efficiency gains are real when paired with mandatory human quality assurance to ensure methodological rigor.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR237

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×