Improving Efficiency of Living Systematic Literature Reviews (SLR) with Artificial Intelligence (AI): Assisted Extraction of Population, Intervention/Comparator, Outcome, and Study Design (P-I/C-O-S)

Author(s)

Liu R1, Jafar R2, Girard LA3, Thorlund K4, Rizzo M5, Forsythe A6
1Cytel Inc., Toronto, ON, Canada, 2Cytel Inc., Vancouver, BC, Canada, 3Cytel Inc., Montreal, QC, Canada, 4McMaster University, Hamilton, ON, Canada, 5Cytel, Kent, KEN, UK, 6Cytel, Waltham, MA, USA

OBJECTIVES: SLRs are crucial but time-consuming parts of any evidence-based research. SLR protocols’ modifications add further complexity. We developed an AI-assisted tool, LiveSTARTTM, which extracts P-I/C-O-S information from citations during SLR review.

METHODS: LiveSTARTTM utilizes a fine-tuned generative biomedical large-language model to identify P-I/C-O-S-relevant texts from citations (title+abstract) helping reviewers identify the information necessary for acceptance/rejection decision based on inclusion/exclusion criteria. The model was trained using in-house annotated datasets consisting of 1400 citations, spanning 18 oncology and 6 non-oncology indications from clinical, economic, and health-related-quality-of-life (HRQoL) SLRs.

The model accuracy was evaluated by two human reviewers who scored a separate test dataset with 350 LiveSTARTTM predictions against manually identified P-I/C-O-S using 3 categories: completely correct, completely incorrect, or partially correct. We defined two sets of evaluation metrics: 1) The strict score considered partially correct predictions as incorrect. 2) The lenient score considered partially correct predictions as correct. Additionally, we evaluated the time savings LiveSTARTTM P-I/C-O-S model provides when an SLR protocol changes midway through projects.

RESULTS: The strict accuracy for P-I/C-O-S extraction were 93%, 89%, 45%, and 97%, respectively. The lenient accuracy were 99%, 96%, 98%, and 98%, respectively. The average strict accuracy was 81%; the average lenient accuracy was 98%. The lowest score was related to the strict accuracy for Outcome extraction (45%). The model could generally predict most of outcome measures in citations (with a lenient accuracy of 98%) which we found to be sufficient to assist reviewers in decision-making. When one of the P-I/C-O-S inclusion/exclusion criteria was changed, it took a human reviewer 4 hours to re-review manually, compared to 1 hour with the help of LiveSTART P-I/C-O-S annotation, yielding 75% in time saving.

CONCLUSIONS: The utilization of LiveSTARTTM to extract P-I/C-O-S shows high accuracy and significant time savings while dealing with protocol changes.

Conference/Value in Health Info

2023-11, ISPOR Europe 2023, Copenhagen, Denmark

Value in Health, Volume 26, Issue 11, S2 (December 2023)

Acceptance Code

P24

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Literature Review & Synthesis

Disease

no-additional-disease-conditions-specialized-treatment-areas, Oncology

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×