CAN AI MATCH THE HUMAN EYE? VALIDATION OF A GENERATIVE AI TOOL FOR NEWCASTLE-OTTAWA QUALITY APPRAISAL IN SYSTEMATIC LITERATURE REVIEWS

Author(s)

Ritesh Dubey, PharmD1, Shivom Prajapati, M.Pharm1, Sukriti Sharma, MSc1, Adarsh Kumar, PharmD1, Rajdeep Kaur, PhD1, Sunil Kumar, M.Pharm1, Barinder Singh, RPh2, Pankaj Rai, MS1.
1Pharmacoevidence, Mohali, India, 2Pharmacoevidence, London, United Kingdom.
OBJECTIVES: Quality appraisal in systematic literature reviews (SLRs) is time-consuming and often decisions varies between reviewers. With the increasing use of multi-agent generative AI tools in evidence synthesis, their reliability for risk-of-bias assessment remains debated. This study aimed to compare and validate the capabilities of the Pharmacoevidence MetaSLR platform, powered by Claude Sonnet 4.6, with human reviewers for the risk of bias assessment using the Newcastle-Ottawa Scale (NOS) module for observational studies.
METHODS: NOS checklists comprising eight items across the Selection, Comparability, and Outcome domains, with a maximum score of 9 stars, were generated for 31 real‑world studies within the pulmonology therapeutic area. Each AI-generated checklist was compared with a checklist independently completed by a blinded human expert; any disagreement/misalignment was cross-checked by a independent subject-matter expert (SME). Agreements were assessed at both the item and total-score levels using per cent agreement, Cohen's kappa, intraclass correlation coefficient (ICC), and Bland-Altman analysis.
RESULTS: Overall item-level agreement was 81% (200/248 items). Total scores showed moderate agreement (ICC = 0.41; weighted kappa = 0.40), with minimal systematic bias (mean difference: +0.06 stars). The highest agreement was observed for objective items, including representativeness (94%) and exposure assessment (94%), whereas lower agreement was observed for follow-up adequacy (71%) and follow-up duration (61%). A second independent SME cross-checked all disagreements and confirmed that they concentrated in subjective domains (outcome assessment and comparability), whereas missing data in some conference abstracts also constrained the AI's appraisal.
CONCLUSIONS: LLMs achieved moderate, unbiased agreement with human reviewers on NOS appraisal, performing reliably on objective items but diverging on subjective domains requiring methodological judgment. LLMs may enable efficient screening process, though human oversight remains essential for nuanced assessment. Future work should test LLM performance across larger and more diverse evidence bases.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR155

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Confounding, Selection Bias Correction, Causal Inference

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Respiratory-Related Disorders (Allergy, Asthma, Smoking, Other Respiratory)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×