Evaluating ChatGPT's Efficacy and Efficiency in the Risk of Bias Assessment in Health Economic Evaluations: A Comparative Analysis Using the Drummond Checklist

Author(s)

Ohra S1, C RR1, Siddiqui MT1, Gupta J2, Siddiqui MK3
1EBM Health Consultants, New Delhi, DL, India, 2EBM Health, Cleckheaton, West Yorkshire, UK, 3EBM Health, New Delhi, DL, India

OBJECTIVES: Drummond checklist is a detailed tool for evaluating the methodological rigor of health economic (HE) studies, requiring a time-consuming review process. The use of large language models (LLMs) such as ChatGPT4.0 can significantly reduce the efforts required by a human reviewer. This study aimed at evaluating the performance of ChatGPT4.0 versus a trained human reviewer in assessing the quality of HE studies using the Drummond checklist.

METHODS: A list of HE studies was randomly selected from an existing review and each study was anonymized to ensure an unbiased evaluation by ChatGPT4.0 and the human reviewer. We developed standardized prompts using an iterative process by providing explicit instructions on the Drummond checklist, which encompasses 36 questions across study design and data collection methods, measurement and evaluation of outcomes, analysis and interpretation of results. These prompts included an appraisal guidance document, a response template and the study documents to be assessed for quality appraisal. The primary outcome of the analysis was the level of agreement assessed by Kappa statistics between ChatGPT4.0 and the human reviewer.

RESULTS: We piloted the prompt using 10 studies and post-standardization further expanded the assessment to another 27 studies. After standardization of the prompt, meantime in minutes to complete the checklist was significantly shorter with ChatGPT4.0 compared to the human reviewer (8 mins/study vs 25 mins/study, a 68% reduction). Our assessment indicated that the level of agreement between ChatGPT4.0 and the expert reviewer across the three domains ranged between 2.7% to 94.5%. Differences were noted in specific areas such as the generalizability of the study results which are more subjective and require human judgement.

CONCLUSIONS: LLMs such as ChatGPT4.0 can expedite the quality appraisal of health economic studies, but it should be utilized only as an adjunct to human expertise rather than a standalone reviewer, particularly for domains requiring subjective evaluation.

Conference/Value in Health Info

2024-05, ISPOR 2024, Atlanta, GA, USA

Value in Health, Volume 27, Issue 6, S1 (June 2024)

Acceptance Code

P54

Topic

Methodological & Statistical Research, Organizational Practices, Study Approaches

Topic Subcategory

Academic & Educational, Artificial Intelligence, Machine Learning, Predictive Analytics, Best Research Practices, Literature Review & Synthesis

Disease

no-additional-disease-conditions-specialized-treatment-areas, Oncology

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×