AUTOMATING QUALITY ASSESSMENT: AGREEMENT BETWEEN AI AND HUMAN CRITICAL APPRAISAL USING THE NEWCASTLE-OTTAWA SCALE

Author(s)

Vicki Pierre, MSc, Marcia Reinhart, DPhil.
Tantalus Medical, Victoria, BC, Canada.
OBJECTIVES: To evaluate agreement between artificial intelligence (AI) and human reviewers in the critical appraisal of cohort studies using the Newcastle-Ottawa Scale (NOS) across multiple systematic literature reviews (SLRs).
METHODS: Cohort studies from three SLRs (covering rare disease, ophthalmology, and mental health topics) were identified where the same human reviewer had performed or validated all NOS-based appraisals. To address copyright constraints related to AI use, studies were filtered to include publications with AI reuse rights only. AI-generated assessments from the Nested Knowledge Smart Critical Appraisal function were compared with human ratings at the item level. Agreements were scored using scales from 0 to 1, including precision, recall, F1 (harmonic mean of precision and recall), and accuracy.
RESULTS: A total of 170 cohort studies were initially identified from the SLRs; however, fewer than 15% (24 studies) permitted AI reuse rights and were included in this analysis. Across all NOS items, recall was high (median 1.00), as the tool was able to answer the appraisal questions for nearly all studies. Median precision and accuracy scores were both 0.75, with variability driven primarily by false positives. The median F1 score was 0.86, with over 90% of studies exceeding a threshold of 0.70. Agreement between human and AI assessments was highest for the NOS ‘selection’ items (representativeness of the exposed cohort and selection of the non-exposed cohort) and lowest for the ‘outcome’ item assessing adequacy of follow-up. One outlier scored 0 on all assessments, showing both missed and overclassified items.
CONCLUSIONS: AI-based critical appraisal continues to evolve, allowing for improved SLR efficiency when used alongside human oversight and review. However, restricting analyses to studies permitting AI reuse substantially limited the usable dataset, underscoring that copyright barriers remain an ongoing challenge for scaling AI use in literature reviews.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR15

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Mental Health (including addiction), Rare & Orphan Diseases, Sensory System Disorders (Ear, Eye, Dental, Skin)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×