DEFINING HUMAN-IN-THE-LOOP IN AI-AUGMENTED SCREENING: A MULTI-LARGE LANGUAGE MODEL COMPARATIVE STUDY IN RARE DISEASE
Author(s)
April Betts, PhD1, Shubhram Pandey, MSc2, Robert Kidd, MSc3, Barinder Singh, RPh2.
1UCB, Slough, United Kingdom, 2Pharmacoevidence, Mohali, India, 3UCB, Copenhagen, Denmark.
1UCB, Slough, United Kingdom, 2Pharmacoevidence, Mohali, India, 3UCB, Copenhagen, Denmark.
OBJECTIVES: The first artificial-intelligence (AI)-assisted health technology assessment submission, accepted by the NICE UK, demonstrated that using AI as a second reviewer for title/abstract screening can reduce systematic literature review (SLR) time and costs by ~50%. The present study evaluates the feasibility of fully automated screening for a rare indication by leveraging a multi-agentic approach within a human-in-the-loop framework to drive further efficiency gains.
METHODS: Makhija et al. 2025 [GID-TA11540] previously applied an AI-assisted screening methodology in a NICE submission, with AI deployed as a second reviewer. Building on Makhija et al.'s methodology, this study investigates the feasibility of a fully automated screening approach. Pharmacoevidence AI SLR tool (metaSLR) was used to conduct automated title/abstract screening using multiple large language models (LLMs). Records with conflicts between multiple AI agents or low-confidence matrices were escalated for manual review.
RESULTS: Overall, 378 citations were screened using the proposed multi-LLM framework. All three LLMs achieved an accuracy of ~99%. The level of disagreement across different LLMs was minimal (~1%), and these instances required human review to finalize decisions. Only four citations were escalated for human review and quality assurance confirmed that no relevant citations were excluded by any of the AI models. The models demonstrated a modest over-inclusion rate of ~1%. Using a human-in-the-loop approach, time and cost savings were around 50% with semi-automated benchmark (AI as a second reviewer) to more than 90% with multiple-LLM approach.
CONCLUSIONS: This study demonstrates that a multi-LLM approach can generate high-quality SLR outputs up to ten times faster than traditional workflows. Human-in-the-loop could be operationalised across a spectrum, from a single human reviewer with AI to advanced multi-agent configurations. By combining targeted human oversight with a multi-agent architecture, this study highlights the potential of optimised human-AI collaboration to accelerate evidence synthesis and timely decision-making.
METHODS: Makhija et al. 2025 [GID-TA11540] previously applied an AI-assisted screening methodology in a NICE submission, with AI deployed as a second reviewer. Building on Makhija et al.'s methodology, this study investigates the feasibility of a fully automated screening approach. Pharmacoevidence AI SLR tool (metaSLR) was used to conduct automated title/abstract screening using multiple large language models (LLMs). Records with conflicts between multiple AI agents or low-confidence matrices were escalated for manual review.
RESULTS: Overall, 378 citations were screened using the proposed multi-LLM framework. All three LLMs achieved an accuracy of ~99%. The level of disagreement across different LLMs was minimal (~1%), and these instances required human review to finalize decisions. Only four citations were escalated for human review and quality assurance confirmed that no relevant citations were excluded by any of the AI models. The models demonstrated a modest over-inclusion rate of ~1%. Using a human-in-the-loop approach, time and cost savings were around 50% with semi-automated benchmark (AI as a second reviewer) to more than 90% with multiple-LLM approach.
CONCLUSIONS: This study demonstrates that a multi-LLM approach can generate high-quality SLR outputs up to ten times faster than traditional workflows. Human-in-the-loop could be operationalised across a spectrum, from a single human reviewer with AI to advanced multi-agent configurations. By combining targeted human oversight with a multi-agent architecture, this study highlights the potential of optimised human-AI collaboration to accelerate evidence synthesis and timely decision-making.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR248
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Neurological Disorders, Rare & Orphan Diseases