OPEN-WEIGHT LLMS RUN LOCALLY MATCH AND SOMETIMES BEAT FRONTIER MODELS FOR SYSTEMATIC AND TARGETED LITERATURE REVIEW ABSTRACT SCREENING

Author(s)

Michal Witkowski, MSc, Alison Martin, MSc, MD, Jay Bilimoria, PhD, Holly Gould, MSc, Raymond Hugo Henderson, BSc, MSc, PhD, Heritage Kristilere, MPH, MD, Tahera Patel, MSc, Hannah Rice, BSc.
Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: AI title/abstract screening typically uses frontier cloud models that send abstracts to a third-party vendor, raising data-privacy, auditability and access concerns. We explored whether a smaller LLM run locally can match frontier-model accuracy across five real literature reviews.
METHODS: AI screening was evaluated on 26,778 records from five reviews (include-rate 12-57%): two well-specified reviews (a clinical-efficacy review in chronic liver disease [A] and an RWE efficacy review in cancer [B]) and three broader-scope reviews (a GVD review of diagnostic tests in rare genetic diseases [C], an RWE genomic-diagnostic review in cancer [D], and a burden-of-illness review in rare infections [E]). We compared two open-weight models (qwen3.5:4b, gemma4:12b) run locally against a frontier cloud model (gpt-5-nano) and human decisions. For each small model we applied a recall-first approach, excluding only on unanimous agreement. We measured recall and work-saved-at-95%-recall (WSS@95).
RESULTS: No model was best everywhere. On the well-specified chronic-liver-disease review (A), gemma4:12b beat the frontier model on recall (0.972 vs. 0.963) and workload (WSS@95: 0.36 vs. 0.23); qwen3.5:4b outperformed all models on the cancer genomic-diagnostic review (D; recall 0.966 vs. 0.941). The frontier model led the other three reviews (B, C, E), including the two reviews (C, E) where no model reached 95% recall.
CONCLUSIONS: On well-specified reviews, locally run open models rival or beat a frontier cloud model on recall and workload while supporting confidential data handling and maintaining a full audit trail. No model was universally best, and on two reviews no model reached the 95%-recall safety threshold with meaningful work saved (a criteria-quality limitation rather than a model limitation), so model choice should be made per review from a small calibration sample. Locally run open models are a credible and privacy-preserving option where the review criteria are well specified.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR131

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×