Improving Efficiency of Living Systematic Literature Reviews (SLR) with Artificial Intelligence (AI): Assisted Extraction of Population, Intervention/Comparator, Outcome, and Study Design (P-I/C-O-S)
Author(s)
Liu R1, Jafar R2, Girard LA3, Thorlund K4, Rizzo M5, Forsythe A6
1Cytel Inc., Toronto, ON, Canada, 2Cytel Inc., Vancouver, BC, Canada, 3Cytel Inc., Montreal, QC, Canada, 4McMaster University, Hamilton, ON, Canada, 5Cytel, Kent, KEN, UK, 6Cytel, Waltham, MA, USA
OBJECTIVES: SLRs are crucial but time-consuming parts of any evidence-based research. SLR protocols’ modifications add further complexity. We developed an AI-assisted tool, LiveSTARTTM, which extracts P-I/C-O-S information from citations during SLR review.
METHODS: LiveSTARTTM utilizes a fine-tuned generative biomedical large-language model to identify P-I/C-O-S-relevant texts from citations (title+abstract) helping reviewers identify the information necessary for acceptance/rejection decision based on inclusion/exclusion criteria. The model was trained using in-house annotated datasets consisting of 1400 citations, spanning 18 oncology and 6 non-oncology indications from clinical, economic, and health-related-quality-of-life (HRQoL) SLRs.
The model accuracy was evaluated by two human reviewers who scored a separate test dataset with 350 LiveSTARTTM predictions against manually identified P-I/C-O-S using 3 categories: completely correct, completely incorrect, or partially correct. We defined two sets of evaluation metrics: 1) The strict score considered partially correct predictions as incorrect. 2) The lenient score considered partially correct predictions as correct. Additionally, we evaluated the time savings LiveSTARTTM P-I/C-O-S model provides when an SLR protocol changes midway through projects.RESULTS: The strict accuracy for P-I/C-O-S extraction were 93%, 89%, 45%, and 97%, respectively. The lenient accuracy were 99%, 96%, 98%, and 98%, respectively. The average strict accuracy was 81%; the average lenient accuracy was 98%. The lowest score was related to the strict accuracy for Outcome extraction (45%). The model could generally predict most of outcome measures in citations (with a lenient accuracy of 98%) which we found to be sufficient to assist reviewers in decision-making. When one of the P-I/C-O-S inclusion/exclusion criteria was changed, it took a human reviewer 4 hours to re-review manually, compared to 1 hour with the help of LiveSTART P-I/C-O-S annotation, yielding 75% in time saving.
CONCLUSIONS: The utilization of LiveSTARTTM to extract P-I/C-O-S shows high accuracy and significant time savings while dealing with protocol changes.
Conference/Value in Health Info
Value in Health, Volume 26, Issue 11, S2 (December 2023)
Acceptance Code
P24
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Literature Review & Synthesis
Disease
no-additional-disease-conditions-specialized-treatment-areas, Oncology