DEVELOPMENT AND VALIDATION OF A METHOD TO EXTRACT LEFT VENTRICULAR EJECTION FRACTION DATA FROM EHR PHYSICIAN NOTES

Author(s)

Oguntuga A1, Overcash J2, Nguyen N2
1Veradigm Health, San Francisco, CA, USA, 2Veradigm Health, Raleigh, NC, USA

Presentation Documents

OBJECTIVES: Left Ventricular Ejection Fraction scores (LVEF) assess the pumping power of the left ventricular wall of the heart and it is an important factor to consider in retrospective analysis studies focused on Heart Failure (HF). We built a pipeline to extract LVEF results from clinical notes on a large cloud based Electronic Health Record (EHR) platform. A literature review showed that only rule-based approaches had been tried. We explore both a rule-based and machine learning approach.

METHODS: We manually annotated 4924 sentences containing LVEF results, pulled from de-identified SOAP notes. 1424 of the sentences were used to build the rule-based and machine learning pipelines, while 3500 sentences were used for validation.

RESULTS: Overall, the rule-based pipeline yielded an F-1 Score of 0.87 (Precision: 0.95, Recall: 0.81) and the machine learning based pipeline, a F-1 Score of 0.95 (Precision: 0.95, Recall: 0.94).

For the machine learning pipeline, results with ratio percentage value ( e.g. 60%) had a F-1 Score of 0.95 (Precision: 0.95, Recall: 0.93), results with interval percentage values (e.g. 35-50%) had a F1-Score of 0.94 (Precision: 0.96, Recall: 0.93) and relative percentage values (e.g. >40%) had a F1-Score of 0.93 (Precision: 0.96, Recall: 0.95). Results without dates had a 0.96 F-1 Score (Precision: 0.98, Recall: 0.94), results with MM-DD-YYYY dates had a 0.92 F-1 Score (Precision: 0.91, Recall: 0.94), results with dates MM-YYYY dates had a 0.79 F1-Score (Precision: 0.86, Recall: 0.76) and results with YYYY dates had a 0.96 F-1 Score (Precision: 0.98, Recall: 0.94). These results are based on the validation set.

CONCLUSIONS: The machine learning pipeline had a better recall performance than the rule-based pipeline, resulting in a much higher F1 accuracy score. This machine learning method can also be used to develop other extraction pipelines for other clinical data items found in EHR notes.

Conference/Value in Health Info

2020-05, ISPOR 2020, Orlando, FL, USA

Value in Health, Volume 23, Issue 5, S1 (May 2020)

Code

PCV82

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Cardiovascular Disorders

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×