EVERY RECORD OR RANK AND STOP? RECALL SAFETY OF AN LLM THAT DECIDES ON ALL RECORDS VERSUS ACTIVE-LEARNING TRIAGE IN SYSTEMATIC REVIEW SCREENING
Author(s)
Farzana Malik, Ph.D.1, Mark Goodlad, BSc2.
1Managing Partner, Cogience, London, United Kingdom, 2Cogience, London, United Kingdom.
1Managing Partner, Cogience, London, United Kingdom, 2Cogience, London, United Kingdom.
OBJECTIVES: AI screening tools for systematic literature review (SLR), typically rank citations by predicted relevance and stop once a rule judges recall sufficient, cutting human screening by 30 to 90 percent. The saving comes from an unscreened tail the model never labels, in which any missed study is invisible. For a regulated HTA submission, one missed study is the worst error possible. We tested whether a large language model can instead decide include or exclude on every record, removing the tail, while keeping recall
METHODS: Each title and abstract were tagged against the protocol PICO , to which recall-first logic was applied. Every decision was logged with a reason, and no record was left unscreened. We validated against two publicly available datasets: the CLEF 2019 TAR benchmark (4 Cochrane reviews, 1,268 records, every record human-judged) and a 38-topic Cochrane run. This analysis was conducted using Anthropic's Claude Opus 4.8. model.
RESULTS: On CLEF, screening recall was 100 percent at both abstract and full-text level across all four topics. All 100 human-kept abstracts and all 49 full-text includes were retained, with zero missed includes, and every disagreement was over-inclusion. Across the 38-topic run, screening recall was 94.6 percent (191 of 202) of relevant records reaching the screen; the few misses were records whose abstract lacked the deciding term.
CONCLUSIONS: An LLM can decide on every record at recall matching or beating the best published active-learning runs, without the unscreened-tail risk those tools accept. For HTA reviews, where a missed study can change a conclusion, this is a defensible alternative to rank-and-stop triage. Model capability is improving year on year and cost is falling fast, so screening every record, infeasible until recently, is now practical and may significantly improve the quality of SLRs.
METHODS: Each title and abstract were tagged against the protocol PICO , to which recall-first logic was applied. Every decision was logged with a reason, and no record was left unscreened. We validated against two publicly available datasets: the CLEF 2019 TAR benchmark (4 Cochrane reviews, 1,268 records, every record human-judged) and a 38-topic Cochrane run. This analysis was conducted using Anthropic's Claude Opus 4.8. model.
RESULTS: On CLEF, screening recall was 100 percent at both abstract and full-text level across all four topics. All 100 human-kept abstracts and all 49 full-text includes were retained, with zero missed includes, and every disagreement was over-inclusion. Across the 38-topic run, screening recall was 94.6 percent (191 of 202) of relevant records reaching the screen; the few misses were records whose abstract lacked the deciding term.
CONCLUSIONS: An LLM can decide on every record at recall matching or beating the best published active-learning runs, without the unscreened-tail risk those tools accept. For HTA reviews, where a missed study can change a conclusion, this is a defensible alternative to rank-and-stop triage. Model capability is improving year on year and cost is falling fast, so screening every record, infeasible until recently, is now practical and may significantly improve the quality of SLRs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA33
Topic
Clinical Outcomes, Epidemiology & Public Health, Health Technology Assessment
Topic Subcategory
Value Frameworks & Dossier Format
Disease
Rare & Orphan Diseases