MULTI-AI MODEL TRIAGE IN LITERATURE REVIEWS CONCENTRATES HUMAN REVIEW ON THE UNCERTAIN RECORDS AND RECOVERS THE FRONTIER MODEL'S SILENT MISSES

Author(s)

Michal Witkowski, MSc, Alison Martin, MSc, MD, Jay Bilimoria, PhD, Holly Gould, MSc, Raymond Hugo Henderson, BSc, MSc, PhD, Heritage Kristilere, MPH, MD, Tahera Patel, MSc, Hannah Rice, BSc.
Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Reliance on a single AI abstract screening model risks silent exclusion of relevant studies. A multi-model approach may improve screening reliability by increasing confidence where models agree and flagging uncertain records for human review. We evaluated a human-in-the-loop triage workflow, with a smaller LLMs cross-checking exclusions made by a frontier model, across five real reviews.
METHODS: Three AI models - two small LLMs (qwen3.5:4b, gemma4:12b) and one frontier model (gpt-5-nano) - screened 26,778 records from five reviews (inclusion rates 12-57%): two well-specified reviews (clinical efficacy in chronic liver disease [A]; RWE efficacy in cancer [B]) and three broader-scope reviews (RWE genomic-diagnostic review in cancer [C], a burden-of-illness review in rare infections [D], and a GVD review of diagnostic tests in rare genetic diseases [E]). Model outputs were compared with human decisions. Unanimous exclusions were auto-excluded, unanimous keeps auto-retained, and disagreements escalated for human review. We measured the proportion auto-resolved, AI-exclusion precision, and workflow recall. Separately, all frontier-model exclusions were re-screened by a small local LLM to estimate false exclusions recovery.
RESULTS: The workflow auto-resolved 48-79% of records. In the two well-specified reviews it resolved 57-62% with recall >=0.99 and AI-exclusion precision 0.99-1.00. In the burden-of-illness review (D), 32% of records were AI-excluded at 0.94 precision. The high-prevalence cancer review (C) was the main exception: automatic resolutions were mostly keeps and AI-exclusion precision was lower (0.62). Cross-checking recovered 60-75% of the frontier model's missed studies when models failed differently, improving recall from 0.963 to 0.989 in the chronic-liver-disease review (A) and from 0.941 to 0.973 in the cancer review (C). Recovery was lower (5-26%) in reviews where the small LLM usually repeated, rather than corrected, the frontier model's false exclusions.
CONCLUSIONS: A human-in-the-loop multi-model triage workflow can concentrate reviewer effort on uncertain records while reducing silent exclusions of relevant studies.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR243

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×