HOW DOES MACHINE-LEARNING COMPARE TO AN INCOMING MEDICAL STUDENT IN EXTRACTING OUTCOMES RESULTS FROM ABSTRACTS?

Author(s)

Michelson M1, Ross M1, Minton S2
1Evid Science, El Segundo, CA, USA, 2Inferlink, El Segundo, CA, USA

Presentation Documents

OBJECTIVES: Literature analysis could benefit from machine-learning (ML) methods that parse medical text to extract reported results. Therefore, we benchmark one ML approach against an incoming medical school student in the same task, as a comparative study. METHODS: We ran an ML algorithm, using Named Entity Recognizers, against a corpus of PubMed abstracts from Randomized Controlled Trials. We then sampled 102 resulting sentences, limiting to those with at least two interventions compared to one another. The ML parsed the disease, end-points, interventions studied, and the reported numerator and denominator (e.g., the sentence, “The total response rate was 75.0% (6/8) in non-transplantation group and 37.5% (3/8) in transplantation group, respectively.” yields two results. The first has 6 as a numerator, 8 as a denominator, “total response rate” as the end-point, and “non-transplantation” as the intervention. The second result has 6, 8, “total response rate”, and “transplantation” for its attributes.) We then trained the student on the same task and gave him the 102 sentences to process. RESULTS: The student and the ML algorithm yielded the same end-points for 97/102 sentences (95%). We then compared triples of (numerator, denominator, intervention) across the sentences. The ML and student match on 87/102 (85%) of sentences (e.g., both pairs matched for each sentence 85% of the time between the ML and student). Accounting for errors associated with only the numerator and denominator (e.g., the interventions match), this number improves to 94/102 sentences (92%). These errors are mostly associated with incorrectly assigning one intervention’s quantities to the other. CONCLUSIONS: Although preliminary, we demonstrate that ML can be almost as effective as an incoming medical student in the task of turning written text into structured results.

Conference/Value in Health Info

2019-05, ISPOR 2019, New Orleans, LA, USA

Value in Health, Volume 22, Issue S1 (2019 May)

Code

PNS261

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Multiple Diseases, No Specific Disease

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×