PREDICTING THE 1-YEAR HOSPITALIZATION AMONG PATIENTS WITH OBESITY USING MACHINE LEARNING TECHNIQUES: IMPLICATIONS FROM A LARGE-SCALE HEALTH ADMINISTRATIVE DATABASE
Author(s)
Yao Xu, MSc1, Jiayue Tian, MSc1, Koki Idehara, PhD2, Fei Zhao, MSc1, Seok-Won Kim, PhD2, Sven Demiya, PhD2.
1RWS, IQVIA, Shanghai, China, 2RWS, IQVIA Solutions Japan G. K, Tokyo, Japan.
1RWS, IQVIA, Shanghai, China, 2RWS, IQVIA Solutions Japan G. K, Tokyo, Japan.
OBJECTIVES: While real-world data are increasingly used for clinical outcome prediction in observational studies, conventional regression approaches such as generalized linear models (GLM) may be sensitive to model misspecification, potentially leading to suboptimal predictive performance and limited generalizability. This study evaluated the accuracy of machine learning methods - Random Forest and XGBoost - in predicting 1-year all-cause hospital admission among patients with obesity, compared with a GLM baseline, using a large administrative claims database.
METHODS: Using the Japan IQVIA Claims Database, which covers all age group, 79,909 patients with obesity were identified. Baseline covariates included demographics, prevalent comorbidities, procedures, and laboratory tests. A GLM method was used as the baseline model and compared with Random Forest and XGBoost for predicting 1-year all-cause hospital admission. To mitigate overfitting, the dataset was randomly split into training (80%) and testing (20%) subsets. Model performance was evaluated using AUC.
RESULTS: Among the 79,909 patients with obesity, median age was 50 years (IQR, 45-60), 40,710 (51.0%) were male, median BMI was 32.1 (IQR, 30.8-34.3), 3,375 (4.22%) were diagnosed with obesity at baseline, 48,934 (61.24%) were having medication for obesity-related medications at baseline. Both GLM and Machine Learning models were applied to predict 1-year all-cause hospital admission during follow-up period. As a result, XGBoost achieved best prediction performance, with AUC (0.75 on training set, 0.68 on testing set), compared to Random Forest with AUC (0.80 on training set, 0.65 on testing set), and GLM with AUC (0.66 on training set, 0.64 on testing set). Features such as age, BMI group, diagnosis of dyslipidaemia / hypertension, and medications for type-2 diabetes explain most of the variability of 1-year all-cause hospital admission.
CONCLUSIONS: Machine Learning methods provide a robust strategy for outcome prediction in observational study, enhance the reliability of estimates using RWD for rigorous clinical and healthcare decision-making in real-world settings.
METHODS: Using the Japan IQVIA Claims Database, which covers all age group, 79,909 patients with obesity were identified. Baseline covariates included demographics, prevalent comorbidities, procedures, and laboratory tests. A GLM method was used as the baseline model and compared with Random Forest and XGBoost for predicting 1-year all-cause hospital admission. To mitigate overfitting, the dataset was randomly split into training (80%) and testing (20%) subsets. Model performance was evaluated using AUC.
RESULTS: Among the 79,909 patients with obesity, median age was 50 years (IQR, 45-60), 40,710 (51.0%) were male, median BMI was 32.1 (IQR, 30.8-34.3), 3,375 (4.22%) were diagnosed with obesity at baseline, 48,934 (61.24%) were having medication for obesity-related medications at baseline. Both GLM and Machine Learning models were applied to predict 1-year all-cause hospital admission during follow-up period. As a result, XGBoost achieved best prediction performance, with AUC (0.75 on training set, 0.68 on testing set), compared to Random Forest with AUC (0.80 on training set, 0.65 on testing set), and GLM with AUC (0.66 on training set, 0.64 on testing set). Features such as age, BMI group, diagnosis of dyslipidaemia / hypertension, and medications for type-2 diabetes explain most of the variability of 1-year all-cause hospital admission.
CONCLUSIONS: Machine Learning methods provide a robust strategy for outcome prediction in observational study, enhance the reliability of estimates using RWD for rigorous clinical and healthcare decision-making in real-world settings.
Conference/Value in Health Info
2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand
Value in Health, Volume 55, Issue S1
Code
MSR22
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, SDC: Diabetes/Endocrine/Metabolic Disorders (including obesity)