TARGETED SINGLE-RULE EXCLUSION SCREENERS FOR SYSTEMATIC AND TARGETED LITERATURE REVIEWS: SMALL LLMS REACH HIGH PRECISION ON RECURRING ELIGIBILITY CRITERIA
Author(s)
Michal Witkowski, MSc, Alison Martin, MSc, MD, Jay Bilimoria, PhD, Holly Gould, MSc, Raymond Hugo Henderson, BSc, MSc, PhD, Heritage Kristilere, MPH, MD, Tahera Patel, MSc, Hannah Rice, BSc.
Crystallise Ltd, Colchester, United Kingdom.
Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: One of the major limiting factors for AI abstract screening is precision - whether an AI exclusion decision can be trusted. Many eligibility rules recur across reviews (e.g. PICO). We assessed whether small LLMs restricted to a single rule can deliver AI exclusions safe enough to automate at trustworthy precision.
METHODS: We evaluated two small LLMs run in-house (qwen3.5:4b, gemma4:12b) for single-rule exclusion precision against a frontier cloud model (gpt-5-nano) and human decisions. Five recurring rules (study design, population, disease, intervention, study size) were assessed across five reviews (26,778 records, include-rate 12-57%): two well-specified reviews (clinical-efficacy review in chronic liver disease [A] and real-world-evidence [RWE] efficacy review in cancer [B]) and three broader-scope reviews (global-value-dossier [GVD] review of diagnostic tests in rare genetic diseases [C], RWE efficacy in cancer [D], and burden-of-illness [BOI] review in rare infections [E]). We also compared the frontier and small models' per-rule reasons, assessing agreement on the cited rule.
RESULTS: Single-rule exclusion precision reached 0.93-1.00 for study design, population and disease on most reviews, but the pattern was rule- and model-specific. On the two well-specified reviews gemma4:12b held strong precision across every rule (0.91-1.00), and qwen3.5:4b was comparable except on the intervention rule, where it collapsed in review A (0.48). The broader-scope reviews diverged sharply: precision held across all rules in the GVD review C (0.90-1.00, both models), was moderate in the BOI review E (0.87-0.92), but collapsed in the cancer review D (most rules <0.80). Compared with the frontier model, the small LLMs cited the same rule for 78-95% of shared exclusions.
CONCLUSIONS: Recurring eligibility rules are good candidates for targeted, single-rule AI screeners deployed only where precision is demonstrably high. Because precision is rule-, model- and review-specific, safe-to-automate rules should be selected per model and per review following a small calibration sample.
METHODS: We evaluated two small LLMs run in-house (qwen3.5:4b, gemma4:12b) for single-rule exclusion precision against a frontier cloud model (gpt-5-nano) and human decisions. Five recurring rules (study design, population, disease, intervention, study size) were assessed across five reviews (26,778 records, include-rate 12-57%): two well-specified reviews (clinical-efficacy review in chronic liver disease [A] and real-world-evidence [RWE] efficacy review in cancer [B]) and three broader-scope reviews (global-value-dossier [GVD] review of diagnostic tests in rare genetic diseases [C], RWE efficacy in cancer [D], and burden-of-illness [BOI] review in rare infections [E]). We also compared the frontier and small models' per-rule reasons, assessing agreement on the cited rule.
RESULTS: Single-rule exclusion precision reached 0.93-1.00 for study design, population and disease on most reviews, but the pattern was rule- and model-specific. On the two well-specified reviews gemma4:12b held strong precision across every rule (0.91-1.00), and qwen3.5:4b was comparable except on the intervention rule, where it collapsed in review A (0.48). The broader-scope reviews diverged sharply: precision held across all rules in the GVD review C (0.90-1.00, both models), was moderate in the BOI review E (0.87-0.92), but collapsed in the cancer review D (most rules <0.80). Compared with the frontier model, the small LLMs cited the same rule for 78-95% of shared exclusions.
CONCLUSIONS: Recurring eligibility rules are good candidates for targeted, single-rule AI screeners deployed only where precision is demonstrably high. Because precision is rule-, model- and review-specific, safe-to-automate rules should be selected per model and per review following a small calibration sample.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR201
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases