PERFORMANCE OF MACHINE LEARNING ALGORITHMS IN PREDICTING 30 DAY HEART FAILURE READMISSIONS RISK USING AN ADMINISTRATIVE CLAIMS DATABASE

Author(s)

Dave C, Park H, Hartzema A
University of Florida, Gainesville, FL, USA

OBJECTIVES: To evaluate machine learning algorithms in modelling the risk of 30-day heart failure (HF) readmissions in a cohort of commercially insured patients in the US. METHODS: : We used Marketscan commercial claims data (2012-14) to identify a cohort of patients >=18 years admitted with a primary diagnosis of heart failure. Heart failure index admissions and 30 day readmissions were defined using the Centers for Medicare and Medicaid Services (CMS) definitions. Using a combination of CMS defined predictors and empirical analysis, we identified 146 predictors of hospital readmission. Predictors were assessed in the one year period prior to the index heart failure hospitalization. Study data were split into the training set (75%) and the test set (25%). We compared four commonly used machine learning algorithms for binary classification: elastic net regularized generalized linear models, random forests, gradient boosted machines, and support vector machines. We employed 10 fold cross validation on the training set to train each algorithm; predictive performance for each algorithm was assessed using the c-statistic [AUC] on the test set. RESULTS: : In a cohort of 17,631 patients with a qualifying index admission for heart failure, 2,830 patients had a readmission within 30 days. The mean age of the patients in the cohort was 55 years, and 60.2% were males. Based on the c-statistic on the test set, gradient boosted machines performed the best (AUC 0.66), followed by elastic net regularized logistic regression (AUC 0.65), random forests (AUC 0.65), and support vector machines (AUC 0.63). CONCLUSIONS: : Compared to previous models which report an AUC of 0.60, we were able to increase the AUC by 0.06, representing a 12% increase in predictive ability. However, this increase in AUC was largely driven by an increase in the number of predictors rather than any differences in performance between the machine learning algorithms themselves.

Conference/Value in Health Info

2017-05, ISPOR 2017, Boston, MA, USA

Value in Health, Vol. 20, No. 5 (May 2017)

Code

PRM101

Topic

Methodological & Statistical Research

Topic Subcategory

Confounding, Selection Bias Correction, Causal Inference, Modeling and simulation

Disease

Cardiovascular Disorders

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×