AGREEMENT BETWEEN ARTIFICIAL INTELLIGENCE AND HUMAN ASSESSMENT IN QUALITY APPRAISAL OF ECONOMIC EVALUATIONS USING THE DRUMMOND CHECKLIST
Author(s)
Nguyen Thi Nhan Phan, MPH, Michaela Lunan, PhD.
RTI Health Solutions, Manchester, United Kingdom.
RTI Health Solutions, Manchester, United Kingdom.
OBJECTIVES: Artificial intelligence (AI) may support automation of systematic literature review (SLR) processes; however, its application to quality assessment (QA) of economic evaluations (EEs) remains minimally evaluated. QA of EEs is essential for assessing evidence robustness, and evaluating AI in this context is necessary to inform its appropriate integration into review workflows. This study evaluated structured AI-assisted QA within an economic SLR by assessing AI-human agreement using Cohen’s kappa (κ) and examining the impact of prompt refinement on performance.
METHODS: QA was conducted in Nested Knowledge using AI-assisted Adaptive Smart Tags. Prompts covered 36 questions from the expanded 36-item Drummond checklist, grouped into three categories: study design, data collection, and analysis/interpretation of results. QA response options were yes, no, unclear, and not appropriate. Ten economic evaluations from an SLR of chronic lymphocytic leukaemia were assessed. AI assessments were compared with human assessments. The prompt was refined by revising instructions to address AI-human disagreements observed with the initial prompt.
RESULTS: Overall observed agreement between AI and human assessment was 76.67% with the initial prompt and increased to 81.67% with the refined prompt. Across the three categories, observed agreement was highest for study design with both prompts (97.14% and 97.14%), followed by data collection (74.00% and 78.57%) and analysis/interpretation of results (69.29% and 77.33%). The refined prompt achieved moderate AI-human agreement (κ = 0.5431), compared with fair agreement using the initial prompt (κ = 0.3744). Of the 36 items, the greatest discrepancies were observed in judgement-based questions, including effectiveness-source design, subjects of benefit valuation, discount-rate rationale, statistical tests for stochastic data, and generalisability issues.
CONCLUSIONS: AI demonstrated moderate agreement with human assessments, suggesting a potential role in supporting quality appraisal of EEs in economic SLRs. Although prompt refinement improved agreement, further research is required to enhance AI performance for judgement-based QA criteria.
METHODS: QA was conducted in Nested Knowledge using AI-assisted Adaptive Smart Tags. Prompts covered 36 questions from the expanded 36-item Drummond checklist, grouped into three categories: study design, data collection, and analysis/interpretation of results. QA response options were yes, no, unclear, and not appropriate. Ten economic evaluations from an SLR of chronic lymphocytic leukaemia were assessed. AI assessments were compared with human assessments. The prompt was refined by revising instructions to address AI-human disagreements observed with the initial prompt.
RESULTS: Overall observed agreement between AI and human assessment was 76.67% with the initial prompt and increased to 81.67% with the refined prompt. Across the three categories, observed agreement was highest for study design with both prompts (97.14% and 97.14%), followed by data collection (74.00% and 78.57%) and analysis/interpretation of results (69.29% and 77.33%). The refined prompt achieved moderate AI-human agreement (κ = 0.5431), compared with fair agreement using the initial prompt (κ = 0.3744). Of the 36 items, the greatest discrepancies were observed in judgement-based questions, including effectiveness-source design, subjects of benefit valuation, discount-rate rationale, statistical tests for stochastic data, and generalisability issues.
CONCLUSIONS: AI demonstrated moderate agreement with human assessments, suggesting a potential role in supporting quality appraisal of EEs in economic SLRs. Although prompt refinement improved agreement, further research is required to enhance AI performance for judgement-based QA criteria.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR202
Topic
Economic Evaluation, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology