ACCURACY OF CLOUD-BASED LARGE LANGUAGE MODELS TO SCREEN ABSTRACTS FOR TARGETED LITERATURE REVIEWS VARIES BY TOPIC
Author(s)
Alison Martin, MSc, MD1, Michal Witkowski, MSc2, Jay Bilimoria, PhD2, Holly Gould, MSc2, Raymond Hugo Henderson, BSc, MSc, PhD2, Heritage Kristilere, MPH, MD2, Tahera Patel, MSc2, Hannah Rice, BSc2.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Abstract screening for literature reviews is labour-intensive and error-prone. Large language models (LLMs) could save time, especially for targeted reviews, but the best way to use them to maximise accuracy is still unclear.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by a senior SLR expert.
RESULTS: Refinement of the protocol research questions and PICOS criteria was necessary to minimize ambiguity and maximise the amount of information provided to the LLM. Using a “safe” mean score threshold of >1 for inclusion, recall (sensitivity) of the LLM ranged from 90% for a burden of illness (BOI) review on rare infections to 98% for real-world efficacy studies in cancer. Precision/specificity was 15% for a review to support a global value dossier (GVD) for diagnostic tests for rare genetic diseases to 63% for an efficacy review in chronic liver disease. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining ≥95% recall. For BOI/GVD reviews, no safe threshold existed. Between 0.4% -22.1% of human decisions were changed after reviewing papers with AI scores under 1 and over 4.
CONCLUSIONS: AI is poor at interpreting standard SLR protocols for BOI/GVD reviews, especially where relevant data is not reported in journal abstracts. The combination of human and AI screening maximises accuracy and can reduce screening time, but optimizing AI performance for broad reviews may require a new SLR paradigm that goes beyond a single PICOS construct.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by a senior SLR expert.
RESULTS: Refinement of the protocol research questions and PICOS criteria was necessary to minimize ambiguity and maximise the amount of information provided to the LLM. Using a “safe” mean score threshold of >1 for inclusion, recall (sensitivity) of the LLM ranged from 90% for a burden of illness (BOI) review on rare infections to 98% for real-world efficacy studies in cancer. Precision/specificity was 15% for a review to support a global value dossier (GVD) for diagnostic tests for rare genetic diseases to 63% for an efficacy review in chronic liver disease. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining ≥95% recall. For BOI/GVD reviews, no safe threshold existed. Between 0.4% -22.1% of human decisions were changed after reviewing papers with AI scores under 1 and over 4.
CONCLUSIONS: AI is poor at interpreting standard SLR protocols for BOI/GVD reviews, especially where relevant data is not reported in journal abstracts. The combination of human and AI screening maximises accuracy and can reduce screening time, but optimizing AI performance for broad reviews may require a new SLR paradigm that goes beyond a single PICOS construct.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR250
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases