DEFINING HUMAN-IN-THE-LOOP IN AI-AUGMENTED SCREENING: A MULTI-LARGE LANGUAGE MODEL COMPARATIVE STUDY IN RARE DISEASE

Author(s)

April Betts, PhD1, Shubhram Pandey, MSc2, Robert Kidd, MSc3, Barinder Singh, RPh2.
1UCB, Slough, United Kingdom, 2Pharmacoevidence, Mohali, India, 3UCB, Copenhagen, Denmark.
OBJECTIVES: The first artificial-intelligence (AI)-assisted health technology assessment submission, accepted by the NICE UK, demonstrated that using AI as a second reviewer for title/abstract screening can reduce systematic literature review (SLR) time and costs by ~50%. The present study evaluates the feasibility of fully automated screening for a rare indication by leveraging a multi-agentic approach within a human-in-the-loop framework to drive further efficiency gains.
METHODS: Makhija et al. 2025 [GID-TA11540] previously applied an AI-assisted screening methodology in a NICE submission, with AI deployed as a second reviewer. Building on Makhija et al.'s methodology, this study investigates the feasibility of a fully automated screening approach. Pharmacoevidence AI SLR tool (metaSLR) was used to conduct automated title/abstract screening using multiple large language models (LLMs). Records with conflicts between multiple AI agents or low-confidence matrices were escalated for manual review.
RESULTS: Overall, 378 citations were screened using the proposed multi-LLM framework. All three LLMs achieved an accuracy of ~99%. The level of disagreement across different LLMs was minimal (~1%), and these instances required human review to finalize decisions. Only four citations were escalated for human review and quality assurance confirmed that no relevant citations were excluded by any of the AI models. The models demonstrated a modest over-inclusion rate of ~1%. Using a human-in-the-loop approach, time and cost savings were around 50% with semi-automated benchmark (AI as a second reviewer) to more than 90% with multiple-LLM approach.
CONCLUSIONS: This study demonstrates that a multi-LLM approach can generate high-quality SLR outputs up to ten times faster than traditional workflows. Human-in-the-loop could be operationalised across a spectrum, from a single human reviewer with AI to advanced multi-agent configurations. By combining targeted human oversight with a multi-agent architecture, this study highlights the potential of optimised human-AI collaboration to accelerate evidence synthesis and timely decision-making.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR248

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Neurological Disorders, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×