SYNTHETICALLY GENERATED DATA ARE USEFUL FOR CLINICAL TASKS EVEN WHEN DATA ARE UNREALISTIC

Author(s)

Bret Nestor1, Running Yang, BSc1, Tanady Bryan, BSc1, Nirupama Tamvada, MSc2, Emanuel Krebs, MA3, Deirdre Weymann, MSc4, Xiaoxiao Li, PhD5, Dean A. Regier, BA, MA, PhD6.
1University of British Columbia, Vancouver, BC, Canada, 2BC Cancer research centre, Vancouver, BC, Canada, 3Cancer Control Research, BC Cancer, Vancouver, BC, Canada, 4Regulatory Science Lab, BC Cancer, Burnaby, BC, Canada, 5The University of British Columbia, Vancouver, BC, Canada, 6BC Cancer - ARCC - UBC, Burnaby, BC, Canada.
OBJECTIVES: Synthetic data promise to accelerate real-world evidence while safeguarding patient privacy. The use of synthetic data to uncover causal mechanisms requires data generation techniques that replicate real-world complexities missed by conventional simulation approaches. We investigate the diversity and plausibility of synthetic electronic health records (EHRs) generated from a large language model (LLM) and evaluate suitability for training clinical outcome models without access to individual patient data.
METHODS: We use public pretraining and private finetuning of a deep learning model designed for long sequences (Mamba architecture) to generate synthetic EHRs. Using the MIMIC-IV dataset derived from the Beth Israel Deaconess Medicine Center ICU between 2008 and 2022, we transform data into the medical event data standard for continued pretraining. Next, we use advanced cancer research program data from British Columbia, Canada collected between 2012-2026, to refine a site-specific model. We evaluate the synthetic records for: plausibility (proximity to real records using Wasserstein distance); diversity (perplexity); and ability to forecast clinical outcomes. Our clinical outcomes include probability of survival greater than one year, and progression after first line of cancer treatment. Classifiers are trained using either synthetic or real data, with performance measured using AUROC.
RESULTS: Drawing on data from 364,627 MIMIC-IV patients and 1,949 BC Cancer patients, we generated 20000 synthetic text EHRs. Synthetic data becomes less realistic (increasing Wasserstein distance) before becoming more diverse (increasing perplexity). Classification results diverged, with synthetic data underperforming real data for the survival task (0.61 vs. 0.73 AUROC) while outperforming models trained on real data for the progression task (0.65 vs. 0.61 AUROC).
CONCLUSIONS: Language modelling produces reliable clinical outcome models despite unrealistic synthetic data representations. This LLM model provides rapid evidence generation to inform access within small cohorts. Future research is necessary to mitigate the risk of synthetic data exposing real patient information.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR82

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×