STRONG ON THE OBJECTIVE, DIVERGENT ON THE SUBJECTIVE: ARTIFICIAL INTELLIGENCE (AI) VS. HUMAN CRITICAL APPRAISAL OF RANDOMIZED CONTROLLED TRIALS (RCTS)
Author(s)
Muhammed Musthafa KK, M. Pharm.1, Aditi Priyadarshini, PharmD1, Ikksheta Sharma, M.Pharm.1, Martina Giulietti Smith, BSc2, Grace E. Fox, PhD3.
1OPEN Health HEOR & Market Access, Bengaluru, India, 2OPEN Health HEOR & Market Access, London, United Kingdom, 3OPEN Health HEOR & Market Access, New York, NY, USA.
1OPEN Health HEOR & Market Access, Bengaluru, India, 2OPEN Health HEOR & Market Access, London, United Kingdom, 3OPEN Health HEOR & Market Access, New York, NY, USA.
OBJECTIVES: Critical appraisal undertaken by humans is often resource-intensive and subject to inter-reviewer variability. AI has been suggested as a way to improve efficiency and consistency while supporting evidence synthesis. This study compares the efficiency (completion time) and agreement of AI-assisted critical appraisal with human reviewers in the appraisal of RCTs using the Cochrane Risk of Bias 1 (RoB-1) checklist.
METHODS: Primary publications of 25 RCTs in non-small cell lung cancer (NSCLC) were identified from PubMed and critically appraised independently by 2 human reviewers plus an AI-assisted SLR platform (Nested Knowledge, version 1.111.4) using the RoB-1 checklist (Low/High/Unclear risk). Agreement between AI assessments and consensus human assessments was evaluated using percentage agreement and Cohen's kappa (κ) statistics. In addition, completion times for appraisal were recorded.
RESULTS: The overall agreement between AI and consensus human assessment was 90% based on RoB-1 checklist. Domain-level agreement varied according to domains with higher levels of agreement in domains related to objective methodological reporting (Performance bias, 100%; Selection bias, 92%; Attrition bias, 96%), and lower agreement in domains requiring greater subjective interpretation (Reporting bias, 64%; Detection bias, 88%; Other bias, 88%). Cohen's kappa (κ) analysis demonstrated strong overall agreement (κ=0.83; 95% CI, 0.75-0.90), with the confidence interval spanning substantial to almost perfect agreement. In addition, as anticipated, the use of AI substantially reduced appraisal completion time compared with human reviewers (80 minutes vs 310 minutes) even allowing for independent human verification of each publication.
CONCLUSIONS: The use of AI substantially improved efficiency of critical appraisal while demonstrating almost perfect overall agreement with human appraisers. These findings suggest that AI can serve as an effective tool to support the critical appraisal of RCTs and improve review efficiency. However, human oversight remains essential, particularly for domains requiring nuanced methodological judgment.
METHODS: Primary publications of 25 RCTs in non-small cell lung cancer (NSCLC) were identified from PubMed and critically appraised independently by 2 human reviewers plus an AI-assisted SLR platform (Nested Knowledge, version 1.111.4) using the RoB-1 checklist (Low/High/Unclear risk). Agreement between AI assessments and consensus human assessments was evaluated using percentage agreement and Cohen's kappa (κ) statistics. In addition, completion times for appraisal were recorded.
RESULTS: The overall agreement between AI and consensus human assessment was 90% based on RoB-1 checklist. Domain-level agreement varied according to domains with higher levels of agreement in domains related to objective methodological reporting (Performance bias, 100%; Selection bias, 92%; Attrition bias, 96%), and lower agreement in domains requiring greater subjective interpretation (Reporting bias, 64%; Detection bias, 88%; Other bias, 88%). Cohen's kappa (κ) analysis demonstrated strong overall agreement (κ=0.83; 95% CI, 0.75-0.90), with the confidence interval spanning substantial to almost perfect agreement. In addition, as anticipated, the use of AI substantially reduced appraisal completion time compared with human reviewers (80 minutes vs 310 minutes) even allowing for independent human verification of each publication.
CONCLUSIONS: The use of AI substantially improved efficiency of critical appraisal while demonstrating almost perfect overall agreement with human appraisers. These findings suggest that AI can serve as an effective tool to support the critical appraisal of RCTs and improve review efficiency. However, human oversight remains essential, particularly for domains requiring nuanced methodological judgment.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR147
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology