ACCURACY OF CLOUD-BASED LARGE LANGUAGE MODELS TO SCREEN ABSTRACTS FOR TARGETED LITERATURE REVIEWS VARIES BY TOPIC

Author(s)

Alison Martin, MSc, MD1, Michal Witkowski, MSc2, Jay Bilimoria, PhD2, Holly Gould, MSc2, Raymond Hugo Henderson, BSc, MSc, PhD2, Heritage Kristilere, MPH, MD2, Tahera Patel, MSc2, Hannah Rice, BSc2.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Abstract screening for literature reviews is labour-intensive and error-prone. Large language models (LLMs) could save time, especially for targeted reviews, but the best way to use them to maximise accuracy is still unclear.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by a senior SLR expert.
RESULTS: Refinement of the protocol research questions and PICOS criteria was necessary to minimize ambiguity and maximise the amount of information provided to the LLM. Using a “safe” mean score threshold of >1 for inclusion, recall (sensitivity) of the LLM ranged from 90% for a burden of illness (BOI) review on rare infections to 98% for real-world efficacy studies in cancer. Precision/specificity was 15% for a review to support a global value dossier (GVD) for diagnostic tests for rare genetic diseases to 63% for an efficacy review in chronic liver disease. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining ≥95% recall. For BOI/GVD reviews, no safe threshold existed. Between 0.4% -22.1% of human decisions were changed after reviewing papers with AI scores under 1 and over 4.
CONCLUSIONS: AI is poor at interpreting standard SLR protocols for BOI/GVD reviews, especially where relevant data is not reported in journal abstracts. The combination of human and AI screening maximises accuracy and can reduce screening time, but optimizing AI performance for broad reviews may require a new SLR paradigm that goes beyond a single PICOS construct.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR250

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×