AUTOMATIC ABSTRACT SCREENING USING MACHINE LEARNING TECHNIQUES- ARE WE THERE YET AND HOW CAN WE MOVE FORWARD?
Author(s)
Rivolo S1, Marczell K2, Dillon-Murphy D3, Sarri G1, Benedict Á2
1Evidera, London, LON, UK, 2Evidera Inc., Budapest, Hungary, 3Evidera, London, UK
Background. Identification of evidence through systematic literature reviews (SLRs) is a key element in health technology assessments (HTAs). To meet HTA requirements by ensuring completeness and validity of conclusions from SLRs, the searches need to be recent and the screening to be thorough which can be time and resource consuming and skill demanding. To address this need machine learning techniques (MLTs) have been recently proposed to reduce the manual burden of conducting SLRs, and specifically with automatic AS (aAS). Our objective was to review the literature on recent publications (2016-2019) in Pubmed and ISPOR examining the role of MLTs for aAS and exposing specific challenges for their use in the context of HTA submissions. Approach. Preliminary analysis of their conclusions showed that there was (1) a large performance (precision) variability depending on the SLR topic and MLT algorithm used, (2) a lack of consistency in metrics used to assess MLTs performances (e.g. should the MLTs be more inclusive or more selective in the identified evidence?) restricting its further applicability. Furthermore, no benchmark datasets exist to standardise evaluation of MLTs for aAS that would facilitate the development of minimum requirements for aAS use in HTA submissions and create confidence in their applicability. Recommendation. We developed a proposal for standardisation of MLTs’ performance evaluation. Firstly, open-source datasets need to be made available to researcher for benchmarking MLTs performances in a systematic way. Secondly, performance metrics focusing on minimising false negatives should be favoured to avoid missing key citations at the expense of increasing full-text screening burden. Thirdly, transfer learning capabilities of MLTs across multiple datasets should be assessed to understand generalisability of MLTs performances. Fourthly, a panel of experts in performing SLRs for HTAs submission should be consulted to validate and expand these requirements for evaluating MLTs for aAS.
Conference/Value in Health Info
2019-11, ISPOR Europe 2019, Copenhagen, Denmark
Code
PNS32
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Specific Disease