TRUST BEYOND EFFICIENCY: HUMAN-AI CONCORDANCE AS A KEY VALIDATION METRIC FOR AI-ENABLED SYSTEMATIC REVIEWS
Author(s)
Viji Queen V, PharmD1, George Alisha, B.E.2, Angeline Babitha Dhas, BS3, Swathirajan C R, Ph.D2, Revanth M, B.E.2, Meghan Oates-Zalesky, MSc4.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: Artificial intelligence (AI) is increasingly used in systematic review workflows to improve efficiency. However, adoption in evidence synthesis requires demonstrating alignment between AI-enabled and expert reviewer decisions while maintaining study identification performance. Human-AI concordance, alongside sensitivity and specificity, may provide an important validation metric for AI-enabled systematic reviews. This study evaluated the agreement between AI-enabled and human literature-screening decisions using the MadeAi-LR platform.
METHODS: A retrospective evaluation was conducted using four completed literature reviews across different therapeutic areas. A total of 556 citations retrieved from systematic database searches were screened using the MadeAi-LR Platform, which employs an AI-enabled, human-in-the-loop workflow to prioritize and classify records according to predefined eligibility criteria. Final screening decisions by experienced reviewers served as the reference standard. Agreement was assessed using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), overall concordance, and Cohen’s kappa (κ). Potential reductions in manual screening effort were also evaluated.
RESULTS: Of 556 citations screened, 188 studies were included following human review. AI-enabled screening achieved a sensitivity of 80.8%, specificity of 92.2%, PPV of 85.1%, and NPV of 89.7%. Overall concordance between AI-enabled and human decisions was 88.1%, corresponding to substantial agreement beyond chance (Cohen’s Kappa=0.74). The platform correctly identified 160 studies included by human reviewers, while remaining studies were identified through the human-in-the-lead review process. AI-enabled prioritization reduced citations requiring full manual review by 66.2%, corresponding to an estimated savings of 38 reviewer hours.
CONCLUSIONS: Substantial agreement between AI-enabled and human screening decisions suggests that concordance can serve as an important complement to traditional performance metrics when evaluating AI-enabled literature screening tools. Together with sensitivity and specificity, concordance may help establish confidence in the reliability and appropriate use of AI within evidence synthesis workflows.
METHODS: A retrospective evaluation was conducted using four completed literature reviews across different therapeutic areas. A total of 556 citations retrieved from systematic database searches were screened using the MadeAi-LR Platform, which employs an AI-enabled, human-in-the-loop workflow to prioritize and classify records according to predefined eligibility criteria. Final screening decisions by experienced reviewers served as the reference standard. Agreement was assessed using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), overall concordance, and Cohen’s kappa (κ). Potential reductions in manual screening effort were also evaluated.
RESULTS: Of 556 citations screened, 188 studies were included following human review. AI-enabled screening achieved a sensitivity of 80.8%, specificity of 92.2%, PPV of 85.1%, and NPV of 89.7%. Overall concordance between AI-enabled and human decisions was 88.1%, corresponding to substantial agreement beyond chance (Cohen’s Kappa=0.74). The platform correctly identified 160 studies included by human reviewers, while remaining studies were identified through the human-in-the-lead review process. AI-enabled prioritization reduced citations requiring full manual review by 66.2%, corresponding to an estimated savings of 38 reviewer hours.
CONCLUSIONS: Substantial agreement between AI-enabled and human screening decisions suggests that concordance can serve as an important complement to traditional performance metrics when evaluating AI-enabled literature screening tools. Together with sensitivity and specificity, concordance may help establish confidence in the reliability and appropriate use of AI within evidence synthesis workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR226
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas