EVERY RECORD OR RANK AND STOP? RECALL SAFETY OF AN LLM THAT DECIDES ON ALL RECORDS VERSUS ACTIVE-LEARNING TRIAGE IN SYSTEMATIC REVIEW SCREENING

Author(s)

Farzana Malik, Ph.D.1, Mark Goodlad, BSc2.
1Managing Partner, Cogience, London, United Kingdom, 2Cogience, London, United Kingdom.
OBJECTIVES: AI screening tools for systematic literature review (SLR), typically rank citations by predicted relevance and stop once a rule judges recall sufficient, cutting human screening by 30 to 90 percent. The saving comes from an unscreened tail the model never labels, in which any missed study is invisible. For a regulated HTA submission, one missed study is the worst error possible. We tested whether a large language model can instead decide include or exclude on every record, removing the tail, while keeping recall
METHODS: Each title and abstract were tagged against the protocol PICO , to which recall-first logic was applied. Every decision was logged with a reason, and no record was left unscreened. We validated against two publicly available datasets: the CLEF 2019 TAR benchmark (4 Cochrane reviews, 1,268 records, every record human-judged) and a 38-topic Cochrane run. This analysis was conducted using Anthropic's Claude Opus 4.8. model.
RESULTS: On CLEF, screening recall was 100 percent at both abstract and full-text level across all four topics. All 100 human-kept abstracts and all 49 full-text includes were retained, with zero missed includes, and every disagreement was over-inclusion. Across the 38-topic run, screening recall was 94.6 percent (191 of 202) of relevant records reaching the screen; the few misses were records whose abstract lacked the deciding term.
CONCLUSIONS: An LLM can decide on every record at recall matching or beating the best published active-learning runs, without the unscreened-tail risk those tools accept. For HTA reviews, where a missed study can change a conclusion, this is a defensible alternative to rank-and-stop triage. Model capability is improving year on year and cost is falling fast, so screening every record, infeasible until recently, is now practical and may significantly improve the quality of SLRs.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA33

Topic

Clinical Outcomes, Epidemiology & Public Health, Health Technology Assessment

Topic Subcategory

Value Frameworks & Dossier Format

Disease

Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×