MULTI-AI MODEL TRIAGE IN LITERATURE REVIEWS CONCENTRATES HUMAN REVIEW ON THE UNCERTAIN RECORDS AND RECOVERS THE FRONTIER MODEL'S SILENT MISSES
Author(s)
Michal Witkowski, MSc, Alison Martin, MSc, MD, Jay Bilimoria, PhD, Holly Gould, MSc, Raymond Hugo Henderson, BSc, MSc, PhD, Heritage Kristilere, MPH, MD, Tahera Patel, MSc, Hannah Rice, BSc.
Crystallise Ltd, Colchester, United Kingdom.
Crystallise Ltd, Colchester, United Kingdom.
OBJECTIVES: Reliance on a single AI abstract screening model risks silent exclusion of relevant studies. A multi-model approach may improve screening reliability by increasing confidence where models agree and flagging uncertain records for human review. We evaluated a human-in-the-loop triage workflow, with a smaller LLMs cross-checking exclusions made by a frontier model, across five real reviews.
METHODS: Three AI models - two small LLMs (qwen3.5:4b, gemma4:12b) and one frontier model (gpt-5-nano) - screened 26,778 records from five reviews (inclusion rates 12-57%): two well-specified reviews (clinical efficacy in chronic liver disease [A]; RWE efficacy in cancer [B]) and three broader-scope reviews (RWE genomic-diagnostic review in cancer [C], a burden-of-illness review in rare infections [D], and a GVD review of diagnostic tests in rare genetic diseases [E]). Model outputs were compared with human decisions. Unanimous exclusions were auto-excluded, unanimous keeps auto-retained, and disagreements escalated for human review. We measured the proportion auto-resolved, AI-exclusion precision, and workflow recall. Separately, all frontier-model exclusions were re-screened by a small local LLM to estimate false exclusions recovery.
RESULTS: The workflow auto-resolved 48-79% of records. In the two well-specified reviews it resolved 57-62% with recall >=0.99 and AI-exclusion precision 0.99-1.00. In the burden-of-illness review (D), 32% of records were AI-excluded at 0.94 precision. The high-prevalence cancer review (C) was the main exception: automatic resolutions were mostly keeps and AI-exclusion precision was lower (0.62). Cross-checking recovered 60-75% of the frontier model's missed studies when models failed differently, improving recall from 0.963 to 0.989 in the chronic-liver-disease review (A) and from 0.941 to 0.973 in the cancer review (C). Recovery was lower (5-26%) in reviews where the small LLM usually repeated, rather than corrected, the frontier model's false exclusions.
CONCLUSIONS: A human-in-the-loop multi-model triage workflow can concentrate reviewer effort on uncertain records while reducing silent exclusions of relevant studies.
METHODS: Three AI models - two small LLMs (qwen3.5:4b, gemma4:12b) and one frontier model (gpt-5-nano) - screened 26,778 records from five reviews (inclusion rates 12-57%): two well-specified reviews (clinical efficacy in chronic liver disease [A]; RWE efficacy in cancer [B]) and three broader-scope reviews (RWE genomic-diagnostic review in cancer [C], a burden-of-illness review in rare infections [D], and a GVD review of diagnostic tests in rare genetic diseases [E]). Model outputs were compared with human decisions. Unanimous exclusions were auto-excluded, unanimous keeps auto-retained, and disagreements escalated for human review. We measured the proportion auto-resolved, AI-exclusion precision, and workflow recall. Separately, all frontier-model exclusions were re-screened by a small local LLM to estimate false exclusions recovery.
RESULTS: The workflow auto-resolved 48-79% of records. In the two well-specified reviews it resolved 57-62% with recall >=0.99 and AI-exclusion precision 0.99-1.00. In the burden-of-illness review (D), 32% of records were AI-excluded at 0.94 precision. The high-prevalence cancer review (C) was the main exception: automatic resolutions were mostly keeps and AI-exclusion precision was lower (0.62). Cross-checking recovered 60-75% of the frontier model's missed studies when models failed differently, improving recall from 0.963 to 0.989 in the chronic-liver-disease review (A) and from 0.941 to 0.973 in the cancer review (C). Recovery was lower (5-26%) in reviews where the small LLM usually repeated, rather than corrected, the frontier model's false exclusions.
CONCLUSIONS: A human-in-the-loop multi-model triage workflow can concentrate reviewer effort on uncertain records while reducing silent exclusions of relevant studies.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR243
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Gastrointestinal Disorders, Infectious Disease (non-vaccine), Oncology, Rare & Orphan Diseases