AI-Enabled Risk of Bias Assessment of RCTs in Systematic Reviews: A Case Study
Author(s)
Langham J1, Reason T2, Malcolm B3, Klijn S4, Gimblett A2
1Estima Scientific Ltd, London, LON, UK, 2Estima Scientific Ltd, London, UK, 3Bristol Myers Squibb, Middlesex, LON, UK, 4Bristol Myers Squibb, Lawrence Township, NJ, USA
Presentation Documents
OBJECTIVES: The task of assessing the risk of bias (RoB) of studies included in a systematic review is time-consuming and requires considerable expertise and judgement. The potential of large language models (LLMs), such as GPT-4 to assist and automate RoB assessment, particularly in randomised trials (RCTs) where reporting is standardised, remains unclear. This study evaluated the accuracy of GPT-4 in assessing the RoB of RCTs included in a published network meta-analysis (NMA).
METHODS: The RoB of each domain (randomization, blinding, missing outcome, outcome measurement, and selective outcome reporting) for ten RCTs was assessed by an experienced systematic reviewer and by GPT-4, using the revised Cochrane risk-of-bias tool for randomized trials (RoB 2). The risk was calculated using Cochrane algorithms based on answers to signalling questions. GPT-4 extracted text from the PDFs and from information downloaded from “ClinicalTrials.gov” to answer signalling questions via a set of prompts. The risk estimates provided by human reviewer and GPT-4 were compared.
RESULTS: The AI-enabled RoB assessment successfully extracted text and answered signalling questions, with very good agreement (70-100%) across questions. There was 80% agreement in the overall risk of bias judgement, 100% agreement in three domains, 90% in Domain 2 (masking) and 70% in Domain 1 (randomisation). Some discrepancy in the assessment of concealment of allocation and randomisation method was reflected in the overall risk assessment by domain.
CONCLUSIONS: This case study provides early evidence for the potential of LLMs to extract and summarise the relevant information required to assess RoB and to quickly deliver information in a concise form for assessing study quality. More work is required on how best to interact with a LLM to ensure that all relevant information is extracted and reported to help assist with quality assessment.
Conference/Value in Health Info
Value in Health, Volume 26, Issue 11, S2 (December 2023)
Code
MSR80
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Literature Review & Synthesis, Meta-Analysis & Indirect Comparisons
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology