PERFORMANCE OF AUTOMATED SCREENING OF CITATIONS COMPARED TO HUMAN REVIEWERS IN SYSTEMATIC LITERATURE REVIEWS- A SYSTEMATIC LITERATURE REVIEW

Author(s)

Lopes R1, Gauthier G2, Akhtar O3, Atanasov P1
1Amaris, Barcelona, Spain, 2Amaris, Toronto, ON, Canada, 3Amaris, London, UK

OBJECTIVES: Systematic literature reviews (SLRs) provide the highest level of evidence to inform healthcare decision-making. Citation screening in SLRs is a labour intensive process that limits timely dissemination of evidence. The use of machine-learning (ML) tools to automate citation screening for SLRs is becoming increasingly common. This SLR aims to identify evidence on the performance characteristics and labour reductions associated with automated and semi-automated ML tools compared to humans in facilitating the screening phase of SLRs.

METHODS: An SLR was conducted in PubMed and Embase to identify published literature related to the research question. Studies were assessed by two independent reviewers following the pre-specified selection criteria. Studies were included if they reported performance metrics comparing ML tools and the gold standard of two independent reviewers, in identifying relevant citations for SLRs.

RESULTS: A total of 801 unique citations studies were identified, with 15 meeting pre-specified inclusion criteria. Sensitivity of ML approaches was generally high (>80%; n=6), but demonstrated variability based on topic and methodology (range: 8%-100%). When both were reported (n=3), specificity was lower than sensitivity, although still high (typically 70%-90%). Accuracy was not commonly reported (n=2), but was consistently greater than 85%. Reduction of labour as measured by the WSS95 was the most commonly reported measure of performance (n=10). Labour reduction typically ranged between 20% and 50% reduction in effort.

CONCLUSIONS: The studies identified in this review demonstrated that ML technology shows favourable performance characteristics, and may result in considerable labour reductions. The performance of classifiers varied based on therapeutic area, and choice of ML model. Further model development and training with SLRs may improve the ability of ML algorithms to identify relevant studies, reduce researcher time commitment, and increase timely dissemination of SLR findings.

Conference/Value in Health Info

2018-11, ISPOR Europe 2018, Barcelona, Spain

Value in Health, Vol. 21, S3 (October 2018)

Code

PRM72

Topic

Real World Data & Information Systems

Topic Subcategory

Reproducibility & Replicability

Disease

Multiple Diseases

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×