Adoption of Artificial Intelligence in Systematic Reviews
Author(s)
Bhagat A1, Moon D2, Khan H2, Kochar P2, Kanakagiri S2, Kaur R2, Bhalla Y2, Kaur K2, Singh R1, Goyal R1, Aggarwal A2
1IQVIA, Thane, MH, India, 2IQVIA, Gurugram, HR, India
Presentation Documents
OBJECTIVES: Evidence synthesis is necessary to make informed decisions regarding patients and health policies. HEOR organizations employ literature reviews to summarize all available evidence for answering specific clinical question. Researchers have become overloaded with literature searches due to the fast growth of medical literature. The use of artificial intelligence and machine learning (AI/ML) in semi-automation of the conventional manual literature review process could expedite the process. We explored the applicability and validated the AI classifier tool of the Distiller SR platform (Evidence Partners Inc. Ottawa, Canada) for title/abstract screening.
METHODS: In the double-review screening of a systematic review, the AI classifier tool was utilized as a second reviewer. Custom AI classifier was trained from a training set of already screened references and was validated on the same set of already reviewed references to get statistical insights, which included a Balanced Accuracy Score of 0.75 (+/- 0.04), Recall Score of 0.86 (+/- 0.09), and F1 Score of 0.71 (+/- 0.08), where 0 represented lowest accuracy and 1 represented highest accuracy. The trained classifier’s responses were compared to those of human reviewer. An independent human reviewer resolved any discrepancies between the decisions.
RESULTS: The trained and validated classifier reviewed 574 references and found 12.5% of them conflicting with the human reviewer’s responses. These conflicts were resolved by an independent human reviewer in around 2-3 hours and resulted in 2.8% false exclusions (considered as critical error) and 9.8% false inclusions (considered as non-critical error). The manual double-screening of the 574 records takes about 34-40 hours, while the AI classifier screening takes about 20-23 hours.
CONCLUSIONS: With AI classifier, as a second reviewer, the human effort in title-abstract screening was dramatically decreased, potentially saving up to 50-60% of human time. However, the quality issues and the lack of exclusion codes, limited its utility to internal or scoping reviews.
Conference/Value in Health Info
Value in Health, Volume 25, Issue 12S (December 2022)
Code
MSR42
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas