PREDICTING THE 1-YEAR HOSPITALIZATION AMONG PATIENTS WITH OBESITY USING MACHINE LEARNING TECHNIQUES: IMPLICATIONS FROM A LARGE-SCALE HEALTH ADMINISTRATIVE DATABASE

Author(s)

Yao Xu, MSc1, Jiayue Tian, MSc1, Koki Idehara, PhD2, Fei Zhao, MSc1, Seok-Won Kim, PhD2, Sven Demiya, PhD2.
1RWS, IQVIA, Shanghai, China, 2RWS, IQVIA Solutions Japan G. K, Tokyo, Japan.
OBJECTIVES: While real-world data are increasingly used for clinical outcome prediction in observational studies, conventional regression approaches such as generalized linear models (GLM) may be sensitive to model misspecification, potentially leading to suboptimal predictive performance and limited generalizability. This study evaluated the accuracy of machine learning methods - Random Forest and XGBoost - in predicting 1-year all-cause hospital admission among patients with obesity, compared with a GLM baseline, using a large administrative claims database.
METHODS: Using the Japan IQVIA Claims Database, which covers all age group, 79,909 patients with obesity were identified. Baseline covariates included demographics, prevalent comorbidities, procedures, and laboratory tests. A GLM method was used as the baseline model and compared with Random Forest and XGBoost for predicting 1-year all-cause hospital admission. To mitigate overfitting, the dataset was randomly split into training (80%) and testing (20%) subsets. Model performance was evaluated using AUC.
RESULTS: Among the 79,909 patients with obesity, median age was 50 years (IQR, 45-60), 40,710 (51.0%) were male, median BMI was 32.1 (IQR, 30.8-34.3), 3,375 (4.22%) were diagnosed with obesity at baseline, 48,934 (61.24%) were having medication for obesity-related medications at baseline. Both GLM and Machine Learning models were applied to predict 1-year all-cause hospital admission during follow-up period. As a result, XGBoost achieved best prediction performance, with AUC (0.75 on training set, 0.68 on testing set), compared to Random Forest with AUC (0.80 on training set, 0.65 on testing set), and GLM with AUC (0.66 on training set, 0.64 on testing set). Features such as age, BMI group, diagnosis of dyslipidaemia / hypertension, and medications for type-2 diabetes explain most of the variability of 1-year all-cause hospital admission.
CONCLUSIONS: Machine Learning methods provide a robust strategy for outcome prediction in observational study, enhance the reliability of estimates using RWD for rigorous clinical and healthcare decision-making in real-world settings.

Conference/Value in Health Info

2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand

Value in Health, Volume 55, Issue S1

Code

MSR22

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, SDC: Diabetes/Endocrine/Metabolic Disorders (including obesity)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×