CLOUD-BASED LARGE LANGUAGE MODELS TO SCREEN ABSTRACTS FOR TARGETED LITERATURE REVIEWS - TIME SAVED AND ACCURACY IMPROVEMENT

Author(s)

Alison Martin, MSc, MD1, Michal Witkowski, MSc2, Jay Bilimoria, PhD2, Holly Gould, MSc2, Raymond Hugo Henderson, BSc, MSc, PhD2, Heritage Kristilere, MPH, MD2, Tahera Patel, MSc2, Hannah Rice, BSc2.
1Crystallise Ltd, Stanford le Hope, United Kingdom, 2Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Abstract screening for literature reviews is labour-intensive and error-prone. Large language models (LLMs) could save time and improve accuracy, especially for targeted reviews, but the best way to use them is still unclear.
METHODS: We retrospectively evaluated the accuracy of gpt-5-nano for screening abstracts for 5 targeted reviews. The LLM screened each abstract 5 times based on the research protocol, with each abstract scored 1 (definite exclude) to 5 (definite include) and mean scores calculated. AI scores were compared with human screening decisions, with the gold standard determined by an SLR expert.
RESULTS: The proportion of human decisions that were changed after reviewing papers with AI scores under 1 and over 4 was 0.4% for a global value dossier (GVD) review for diagnostic tests in cancer; 2.5% for a GVD review in rare diseases; 2.6% for a burden of illness (BOI) review of rare infections; 3.7% for an efficacy study in chronic liver disease; and 22.1% for an efficacy review of real-world studies in cancer. For efficacy reviews in liver disease and cancer, the optimal cut-off score for exclusion was 1.0 and 1.2, which reduced the human screening workload by 40% to 43% respectively while maintaining at least 95% recall. This equates to a reduction in screening time of 17-49 hours per project. For BOI/GVD reviews, no safe exclusion threshold existed, even after refining the research questions and PICOS criteria to make them more AI-appropriate.
CONCLUSIONS: Using AI as a second or third screener can improve accuracy of targeted reviews but cannot always be relied on to reduce the number of abstracts needing human screening. Further research is needed to find the optimal way of interacting with AI for more complex reviews where a single PICOS construct does not easily represent the range of evidence being sought.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR32

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×