AI-Enabled Risk of Bias Assessment of RCTs in Systematic Reviews: A Case Study

Author(s)

Langham J1, Reason T2, Malcolm B3, Klijn S4, Gimblett A2
1Estima Scientific Ltd, London, LON, UK, 2Estima Scientific Ltd, London, UK, 3Bristol Myers Squibb, Middlesex, LON, UK, 4Bristol Myers Squibb, Lawrence Township, NJ, USA

OBJECTIVES: The task of assessing the risk of bias (RoB) of studies included in a systematic review is time-consuming and requires considerable expertise and judgement. The potential of large language models (LLMs), such as GPT-4 to assist and automate RoB assessment, particularly in randomised trials (RCTs) where reporting is standardised, remains unclear. This study evaluated the accuracy of GPT-4 in assessing the RoB of RCTs included in a published network meta-analysis (NMA).

METHODS: The RoB of each domain (randomization, blinding, missing outcome, outcome measurement, and selective outcome reporting) for ten RCTs was assessed by an experienced systematic reviewer and by GPT-4, using the revised Cochrane risk-of-bias tool for randomized trials (RoB 2). The risk was calculated using Cochrane algorithms based on answers to signalling questions. GPT-4 extracted text from the PDFs and from information downloaded from “ClinicalTrials.gov” to answer signalling questions via a set of prompts. The risk estimates provided by human reviewer and GPT-4 were compared.

RESULTS: The AI-enabled RoB assessment successfully extracted text and answered signalling questions, with very good agreement (70-100%) across questions. There was 80% agreement in the overall risk of bias judgement, 100% agreement in three domains, 90% in Domain 2 (masking) and 70% in Domain 1 (randomisation). Some discrepancy in the assessment of concealment of allocation and randomisation method was reflected in the overall risk assessment by domain.

CONCLUSIONS: This case study provides early evidence for the potential of LLMs to extract and summarise the relevant information required to assess RoB and to quickly deliver information in a concise form for assessing study quality. More work is required on how best to interact with a LLM to ensure that all relevant information is extracted and reported to help assist with quality assessment.

Conference/Value in Health Info

2023-11, ISPOR Europe 2023, Copenhagen, Denmark

Value in Health, Volume 26, Issue 11, S2 (December 2023)

Code

MSR80

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Literature Review & Synthesis, Meta-Analysis & Indirect Comparisons

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Oncology

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×