TRUST BEYOND EFFICIENCY: HUMAN-AI CONCORDANCE AS A KEY VALIDATION METRIC FOR AI-ENABLED SYSTEMATIC REVIEWS

Author(s)

Viji Queen V, PharmD1, George Alisha, B.E.2, Angeline Babitha Dhas, BS3, Swathirajan C R, Ph.D2, Revanth M, B.E.2, Meghan Oates-Zalesky, MSc4.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: Artificial intelligence (AI) is increasingly used in systematic review workflows to improve efficiency. However, adoption in evidence synthesis requires demonstrating alignment between AI-enabled and expert reviewer decisions while maintaining study identification performance. Human-AI concordance, alongside sensitivity and specificity, may provide an important validation metric for AI-enabled systematic reviews. This study evaluated the agreement between AI-enabled and human literature-screening decisions using the MadeAi-LR platform.
METHODS: A retrospective evaluation was conducted using four completed literature reviews across different therapeutic areas. A total of 556 citations retrieved from systematic database searches were screened using the MadeAi-LR Platform, which employs an AI-enabled, human-in-the-loop workflow to prioritize and classify records according to predefined eligibility criteria. Final screening decisions by experienced reviewers served as the reference standard. Agreement was assessed using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), overall concordance, and Cohen’s kappa (κ). Potential reductions in manual screening effort were also evaluated.
RESULTS: Of 556 citations screened, 188 studies were included following human review. AI-enabled screening achieved a sensitivity of 80.8%, specificity of 92.2%, PPV of 85.1%, and NPV of 89.7%. Overall concordance between AI-enabled and human decisions was 88.1%, corresponding to substantial agreement beyond chance (Cohen’s Kappa=0.74). The platform correctly identified 160 studies included by human reviewers, while remaining studies were identified through the human-in-the-lead review process. AI-enabled prioritization reduced citations requiring full manual review by 66.2%, corresponding to an estimated savings of 38 reviewer hours.
CONCLUSIONS: Substantial agreement between AI-enabled and human screening decisions suggests that concordance can serve as an important complement to traditional performance metrics when evaluating AI-enabled literature screening tools. Together with sensitivity and specificity, concordance may help establish confidence in the reliability and appropriate use of AI within evidence synthesis workflows.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR226

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×