AGREEMENT BETWEEN ARTIFICIAL INTELLIGENCE AND HUMAN ASSESSMENT IN QUALITY APPRAISAL OF ECONOMIC EVALUATIONS USING THE DRUMMOND CHECKLIST

Author(s)

Nguyen Thi Nhan Phan, MPH, Michaela Lunan, PhD.
RTI Health Solutions, Manchester, United Kingdom.
OBJECTIVES: Artificial intelligence (AI) may support automation of systematic literature review (SLR) processes; however, its application to quality assessment (QA) of economic evaluations (EEs) remains minimally evaluated. QA of EEs is essential for assessing evidence robustness, and evaluating AI in this context is necessary to inform its appropriate integration into review workflows. This study evaluated structured AI-assisted QA within an economic SLR by assessing AI-human agreement using Cohen’s kappa (κ) and examining the impact of prompt refinement on performance.
METHODS: QA was conducted in Nested Knowledge using AI-assisted Adaptive Smart Tags. Prompts covered 36 questions from the expanded 36-item Drummond checklist, grouped into three categories: study design, data collection, and analysis/interpretation of results. QA response options were yes, no, unclear, and not appropriate. Ten economic evaluations from an SLR of chronic lymphocytic leukaemia were assessed. AI assessments were compared with human assessments. The prompt was refined by revising instructions to address AI-human disagreements observed with the initial prompt.
RESULTS: Overall observed agreement between AI and human assessment was 76.67% with the initial prompt and increased to 81.67% with the refined prompt. Across the three categories, observed agreement was highest for study design with both prompts (97.14% and 97.14%), followed by data collection (74.00% and 78.57%) and analysis/interpretation of results (69.29% and 77.33%). The refined prompt achieved moderate AI-human agreement (κ = 0.5431), compared with fair agreement using the initial prompt (κ = 0.3744). Of the 36 items, the greatest discrepancies were observed in judgement-based questions, including effectiveness-source design, subjects of benefit valuation, discount-rate rationale, statistical tests for stochastic data, and generalisability issues.
CONCLUSIONS: AI demonstrated moderate agreement with human assessments, suggesting a potential role in supporting quality appraisal of EEs in economic SLRs. Although prompt refinement improved agreement, further research is required to enhance AI performance for judgement-based QA criteria.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR202

Topic

Economic Evaluation, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×