Predicting the Risk of Stroke Using Machine Learning on a Large Administrative Health Database

Author(s)

Ghiani M1, Maywald U2, Wilke T1
1IPAM, University of Wismar, Wismar, Germany, 2AOK PLUS, Dresden, Germany

OBJECTIVES: Stroke is a leading cause of death worldwide and understanding its risk factors is key for prevention. This study investigates the predictive performance of several machine learning classifiers using a large administrative database to predict individual stroke risk.

METHODS: We used data from AOK PLUS, a German sickness fund covering 3.6 million patients in Saxony and Thuringia. We identified all adult patients continuously insured between 01/01/2016-31/12/2017. The outcome variable was an indicator of whether the patient was hospitalized with a stroke diagnosis between 01/01/2017 (index date) and 31/12/2017. We collected 28 baseline characteristics over the period 01/01/2016-31/12/2016, including age, sex, prior stroke/TIA hospitalizations, comorbidities (including myocardial infarctions, hypertension, diabetes, and venous thromboembolism), prior treatments and procedures. The sample was split in 80% training and 20% testing and we evaluated the predictive performance of logistic model, lasso, ridge, and XGBoost. Measures of performance included sensitivity (true positive rate), specificity (true negative rate) and the area under the ROC curve (AUROC). We used random under-sampling to adjust for the high unbalance of the outcome across classes.

RESULTS: We included 2,543,965 adult patients continuously insured between 01/01/2016-31/12/2017 (53% females, mean age 54.1 years). In 2017, 0.75% of the population had a stroke hospitalization. The AUROC was 0.813 for logistic, lasso and ridge, and 0.814 for XGBoost. Sensitivity and specificity where, respectively, 0.784 and 0.707 for the logistic model and lasso, 0.791 and 0.698 for ridge, 0.795 and 0.698 for XGBoost. Features with highest importance included prior stroke or TIA, smoking and history of diabetes and hypertension.

CONCLUSIONS: Different algorithms displayed similar performance and our proposed logistic regression approach was on par with state-of-the-art ML algorithms for tabular data. The fitted logistic model could be deployed on new claims data to predict the risk of stroke, correctly identifying 78% of stroke patients and 71% of non-stroke patients.

Conference/Value in Health Info

2022-11, ISPOR Europe 2022, Vienna, Austria

Value in Health, Volume 25, Issue 12S (December 2022)

Acceptance Code

P58

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Health & Insurance Records Systems

Disease

sdc-cardiovascular-disorders-including-mi-stroke-circulatory

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×