SYNTHETICALLY GENERATED DATA ARE USEFUL FOR CLINICAL TASKS EVEN WHEN DATA ARE UNREALISTIC
Author(s)
Bret Nestor1, Running Yang, BSc1, Tanady Bryan, BSc1, Nirupama Tamvada, MSc2, Emanuel Krebs, MA3, Deirdre Weymann, MSc4, Xiaoxiao Li, PhD5, Dean A. Regier, BA, MA, PhD6.
1University of British Columbia, Vancouver, BC, Canada, 2BC Cancer research centre, Vancouver, BC, Canada, 3Cancer Control Research, BC Cancer, Vancouver, BC, Canada, 4Regulatory Science Lab, BC Cancer, Burnaby, BC, Canada, 5The University of British Columbia, Vancouver, BC, Canada, 6BC Cancer - ARCC - UBC, Burnaby, BC, Canada.
1University of British Columbia, Vancouver, BC, Canada, 2BC Cancer research centre, Vancouver, BC, Canada, 3Cancer Control Research, BC Cancer, Vancouver, BC, Canada, 4Regulatory Science Lab, BC Cancer, Burnaby, BC, Canada, 5The University of British Columbia, Vancouver, BC, Canada, 6BC Cancer - ARCC - UBC, Burnaby, BC, Canada.
OBJECTIVES: Synthetic data promise to accelerate real-world evidence while safeguarding patient privacy. The use of synthetic data to uncover causal mechanisms requires data generation techniques that replicate real-world complexities missed by conventional simulation approaches. We investigate the diversity and plausibility of synthetic electronic health records (EHRs) generated from a large language model (LLM) and evaluate suitability for training clinical outcome models without access to individual patient data.
METHODS: We use public pretraining and private finetuning of a deep learning model designed for long sequences (Mamba architecture) to generate synthetic EHRs. Using the MIMIC-IV dataset derived from the Beth Israel Deaconess Medicine Center ICU between 2008 and 2022, we transform data into the medical event data standard for continued pretraining. Next, we use advanced cancer research program data from British Columbia, Canada collected between 2012-2026, to refine a site-specific model. We evaluate the synthetic records for: plausibility (proximity to real records using Wasserstein distance); diversity (perplexity); and ability to forecast clinical outcomes. Our clinical outcomes include probability of survival greater than one year, and progression after first line of cancer treatment. Classifiers are trained using either synthetic or real data, with performance measured using AUROC.
RESULTS: Drawing on data from 364,627 MIMIC-IV patients and 1,949 BC Cancer patients, we generated 20000 synthetic text EHRs. Synthetic data becomes less realistic (increasing Wasserstein distance) before becoming more diverse (increasing perplexity). Classification results diverged, with synthetic data underperforming real data for the survival task (0.61 vs. 0.73 AUROC) while outperforming models trained on real data for the progression task (0.65 vs. 0.61 AUROC).
CONCLUSIONS: Language modelling produces reliable clinical outcome models despite unrealistic synthetic data representations. This LLM model provides rapid evidence generation to inform access within small cohorts. Future research is necessary to mitigate the risk of synthetic data exposing real patient information.
METHODS: We use public pretraining and private finetuning of a deep learning model designed for long sequences (Mamba architecture) to generate synthetic EHRs. Using the MIMIC-IV dataset derived from the Beth Israel Deaconess Medicine Center ICU between 2008 and 2022, we transform data into the medical event data standard for continued pretraining. Next, we use advanced cancer research program data from British Columbia, Canada collected between 2012-2026, to refine a site-specific model. We evaluate the synthetic records for: plausibility (proximity to real records using Wasserstein distance); diversity (perplexity); and ability to forecast clinical outcomes. Our clinical outcomes include probability of survival greater than one year, and progression after first line of cancer treatment. Classifiers are trained using either synthetic or real data, with performance measured using AUROC.
RESULTS: Drawing on data from 364,627 MIMIC-IV patients and 1,949 BC Cancer patients, we generated 20000 synthetic text EHRs. Synthetic data becomes less realistic (increasing Wasserstein distance) before becoming more diverse (increasing perplexity). Classification results diverged, with synthetic data underperforming real data for the survival task (0.61 vs. 0.73 AUROC) while outperforming models trained on real data for the progression task (0.65 vs. 0.61 AUROC).
CONCLUSIONS: Language modelling produces reliable clinical outcome models despite unrealistic synthetic data representations. This LLM model provides rapid evidence generation to inform access within small cohorts. Future research is necessary to mitigate the risk of synthetic data exposing real patient information.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR82
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology