AI-DRIVEN SYNTHETIC PATIENT GENERATION FRAMEWORK TO MODEL OBESITY USING THE CONSTANCES COHORT

Author(s)

Diane Vincent, MSc1, Louise Dry, MSc1, Nina Temam, PharmD1, Antoine Movschin, MSc1, Sofiane Kab, PharmD2, Antoine Duclos, MD, PhD3, Pauline Guilmin, MSc1.
1Quinten Health, Paris, France, 2Paris Cité University, Paris-Saclay University, UVSQ, Inserm, Epidemiological Population Cohorts Unit (UMS11), Villejuif, France, 3Paris Cité University, Paris-Saclay University, UVSQ, Inserm, Epidemiological Population Cohorts Unit (UMS11), Villejuif, France and Paris Cité University and Sorbonne Paris Nord University, Inserm, INRAE, Centre for Research in Epidemiology and Statistic, Paris, France.
OBJECTIVES: Obesity is a chronic, heterogeneous disease associated with substantial cardiovascular burden and diverse patient profiles, creating challenges for evidence generation and healthcare decision-making. This study aimed to develop an artificial intelligence (AI)-driven framework capable of generating and evaluating realistic synthetic real-world Obesity patient cohorts that preserve both patient heterogeneity and privacy.
METHODS: Data were obtained from the French CONSTANCES cohort linked to the national health claims database (SNDS). Adults with obesity, defined as body mass index (BMI) ≥30 kg/m² or overweight defined as BMI ≥27 kg/m² with weight-related comorbidities, were identified using BMI measurements recorded in CONSTANCES. Baseline characteristics included clinical variables, medical history, and lifestyle. A Cox proportional hazards (CPH) model was developed to predict time to first 5-point major adverse cardiovascular event (5P-MACE) and evaluated on an independent test set. Synthetic baseline patient characteristics were generated using a Gaussian copula approach, and longitudinal outcomes were simulated using the trained CPH model combined with inverse transform sampling. Synthetic data performance was assessed across three domains: fidelity through comparisons of generated and observed variable distributions and correlation structures; utility, through agreement of generated and observed Kaplan-Meier outcome trajectories; and privacy through disclosure-risk assessment. Fidelity and utility were assessed in the overall population and across clinically relevant subgroups defined by cardiovascular comorbidities, diabetes and sex.
RESULTS: The framework was developed using data from 29,046 patients with Obesity. High fidelity and utility were observed across the overall population and clinically relevant subgroups, with close agreement in baseline characteristics and 5P-MACE trajectories, while maintaining promising privacy-preserving properties.
CONCLUSIONS: This AI-driven framework generated realistic synthetic Obesity cohorts that may enable augmentation of under-represented subgroups or simulation of alternative therapeutic trajectories, thereby supporting research on progression, cardiovascular outcomes, and long-term prevention strategies to inform clinical and healthcare decision-making in Obesity.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR130

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Diabetes/Endocrine/Metabolic Disorders (including obesity)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×