CLOUD-BASED LARGE LANGUAGE MODELS TO SCREEN ABSTRACTS FOR TARGETED LITERATURE REVIEWS - TIME SAVED AND ACCURACY IMPROVEMENT
Author(s)
Alison Martin, MSc, MD1, Michal Witkowski, MSc2, Jay Bilimoria, PhD2, Holly Gould, MSc2, Raymond Hugo Henderson, BSc, MSc, PhD2, Heritage Kristilere, MPH, MD2, Tahera Patel, MSc2, Hannah Rice, BSc2.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Abstract screening for literature reviews is labour-intensive and error-prone. Large language models (LLMs) could save time and improve accuracy, especially for targeted reviews, but the best way to use them is still unclear.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by an SLR expert.
RESULTS: The proportion of human decisions that were changed after reviewing papers with AI scores under 1 and over 4 was 0.4% for a global value dossier (GVD) review for diagnostic tests in cancer; 2.5% for a GVD review in rare diseases; 2.6% for a burden of illness (BOI) review of rare infections; 3.7% for an efficacy study in chronic liver disease; and 22.1% for an efficacy review of real-world studies in cancer. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining at least 95% recall. This equates to a reduction in screening time of 17-49 hours per project. For BOI/GVD reviews, no safe exclusion threshold existed, even after refining the research questions and PICOS criteria to make them more AI-appropriate.
CONCLUSIONS: Using AI as a second or third screener can improve accuracy of targeted reviews but cannot always be relied on to reduce the number of abstracts needing human screening. Further research is needed to find the optimal way of interacting with AI for more complex reviews where a single PICOS construct does not easily represent the range of evidence being sought.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by an SLR expert.
RESULTS: The proportion of human decisions that were changed after reviewing papers with AI scores under 1 and over 4 was 0.4% for a global value dossier (GVD) review for diagnostic tests in cancer; 2.5% for a GVD review in rare diseases; 2.6% for a burden of illness (BOI) review of rare infections; 3.7% for an efficacy study in chronic liver disease; and 22.1% for an efficacy review of real-world studies in cancer. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining at least 95% recall. This equates to a reduction in screening time of 17-49 hours per project. For BOI/GVD reviews, no safe exclusion threshold existed, even after refining the research questions and PICOS criteria to make them more AI-appropriate.
CONCLUSIONS: Using AI as a second or third screener can improve accuracy of targeted reviews but cannot always be relied on to reduce the number of abstracts needing human screening. Further research is needed to find the optimal way of interacting with AI for more complex reviews where a single PICOS construct does not easily represent the range of evidence being sought.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR32
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases