CAN AI MATCH THE HUMAN EYE? VALIDATION OF A GENERATIVE AI TOOL FOR NEWCASTLE-OTTAWA QUALITY APPRAISAL IN SYSTEMATIC LITERATURE REVIEWS
Author(s)
Ritesh Dubey, PharmD1, Shivom Prajapati, M.Pharm1, Sukriti Sharma, MSc1, Adarsh Kumar, PharmD1, Rajdeep Kaur, PhD1, Sunil Kumar, M.Pharm1, Barinder Singh, RPh2, Pankaj Rai, MS1.
1Pharmacoevidence, Mohali, India, 2Pharmacoevidence, London, United Kingdom.
1Pharmacoevidence, Mohali, India, 2Pharmacoevidence, London, United Kingdom.
OBJECTIVES: Quality appraisal in systematic literature reviews (SLRs) is time-consuming and often decisions varies between reviewers. With the increasing use of multi-agent generative AI tools in evidence synthesis, their reliability for risk-of-bias assessment remains debated. This study aimed to compare and validate the capabilities of the Pharmacoevidence MetaSLR platform, powered by Claude Sonnet 4.6, with human reviewers for the risk of bias assessment using the Newcastle-Ottawa Scale (NOS) module for observational studies.
METHODS: NOS checklists comprising eight items across the Selection, Comparability, and Outcome domains, with a maximum score of 9 stars, were generated for 31 real‑world studies within the pulmonology therapeutic area. Each AI-generated checklist was compared with a checklist independently completed by a blinded human expert; any disagreement/misalignment was cross-checked by a independent subject-matter expert (SME). Agreements were assessed at both the item and total-score levels using per cent agreement, Cohen's kappa, intraclass correlation coefficient (ICC), and Bland-Altman analysis.
RESULTS: Overall item-level agreement was 81% (200/248 items). Total scores showed moderate agreement (ICC = 0.41; weighted kappa = 0.40), with minimal systematic bias (mean difference: +0.06 stars). The highest agreement was observed for objective items, including representativeness (94%) and exposure assessment (94%), whereas lower agreement was observed for follow-up adequacy (71%) and follow-up duration (61%). A second independent SME cross-checked all disagreements and confirmed that they concentrated in subjective domains (outcome assessment and comparability), whereas missing data in some conference abstracts also constrained the AI's appraisal.
CONCLUSIONS: LLMs achieved moderate, unbiased agreement with human reviewers on NOS appraisal, performing reliably on objective items but diverging on subjective domains requiring methodological judgment. LLMs may enable efficient screening process, though human oversight remains essential for nuanced assessment. Future work should test LLM performance across larger and more diverse evidence bases.
METHODS: NOS checklists comprising eight items across the Selection, Comparability, and Outcome domains, with a maximum score of 9 stars, were generated for 31 real‑world studies within the pulmonology therapeutic area. Each AI-generated checklist was compared with a checklist independently completed by a blinded human expert; any disagreement/misalignment was cross-checked by a independent subject-matter expert (SME). Agreements were assessed at both the item and total-score levels using per cent agreement, Cohen's kappa, intraclass correlation coefficient (ICC), and Bland-Altman analysis.
RESULTS: Overall item-level agreement was 81% (200/248 items). Total scores showed moderate agreement (ICC = 0.41; weighted kappa = 0.40), with minimal systematic bias (mean difference: +0.06 stars). The highest agreement was observed for objective items, including representativeness (94%) and exposure assessment (94%), whereas lower agreement was observed for follow-up adequacy (71%) and follow-up duration (61%). A second independent SME cross-checked all disagreements and confirmed that they concentrated in subjective domains (outcome assessment and comparability), whereas missing data in some conference abstracts also constrained the AI's appraisal.
CONCLUSIONS: LLMs achieved moderate, unbiased agreement with human reviewers on NOS appraisal, performing reliably on objective items but diverging on subjective domains requiring methodological judgment. LLMs may enable efficient screening process, though human oversight remains essential for nuanced assessment. Future work should test LLM performance across larger and more diverse evidence bases.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR155
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Confounding, Selection Bias Correction, Causal Inference
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Respiratory-Related Disorders (Allergy, Asthma, Smoking, Other Respiratory)