Predicting the Risk of Stroke Using Machine Learning on a Large Administrative Health Database
Author(s)
Ghiani M1, Maywald U2, Wilke T1
1IPAM, University of Wismar, Wismar, Germany, 2AOK PLUS, Dresden, Germany
OBJECTIVES: Stroke is a leading cause of death worldwide and understanding its risk factors is key for prevention. This study investigates the predictive performance of several machine learning classifiers using a large administrative database to predict individual stroke risk.
METHODS: We used data from AOK PLUS, a German sickness fund covering 3.6 million patients in Saxony and Thuringia. We identified all adult patients continuously insured between 01/01/2016-31/12/2017. The outcome variable was an indicator of whether the patient was hospitalized with a stroke diagnosis between 01/01/2017 (index date) and 31/12/2017. We collected 28 baseline characteristics over the period 01/01/2016-31/12/2016, including age, sex, prior stroke/TIA hospitalizations, comorbidities (including myocardial infarctions, hypertension, diabetes, and venous thromboembolism), prior treatments and procedures. The sample was split in 80% training and 20% testing and we evaluated the predictive performance of logistic model, lasso, ridge, and XGBoost. Measures of performance included sensitivity (true positive rate), specificity (true negative rate) and the area under the ROC curve (AUROC). We used random under-sampling to adjust for the high unbalance of the outcome across classes.
RESULTS: We included 2,543,965 adult patients continuously insured between 01/01/2016-31/12/2017 (53% females, mean age 54.1 years). In 2017, 0.75% of the population had a stroke hospitalization. The AUROC was 0.813 for logistic, lasso and ridge, and 0.814 for XGBoost. Sensitivity and specificity where, respectively, 0.784 and 0.707 for the logistic model and lasso, 0.791 and 0.698 for ridge, 0.795 and 0.698 for XGBoost. Features with highest importance included prior stroke or TIA, smoking and history of diabetes and hypertension.
CONCLUSIONS: Different algorithms displayed similar performance and our proposed logistic regression approach was on par with state-of-the-art ML algorithms for tabular data. The fitted logistic model could be deployed on new claims data to predict the risk of stroke, correctly identifying 78% of stroke patients and 71% of non-stroke patients.
Conference/Value in Health Info
Value in Health, Volume 25, Issue 12S (December 2022)
Acceptance Code
P58
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Health & Insurance Records Systems
Disease
sdc-cardiovascular-disorders-including-mi-stroke-circulatory