SYNTHETIC SWEDISH CANCER REGISTRY DATA FOR HEALTH ECONOMIC RESEARCH AND INTERACTIVE DECISION SUPPORT
Author(s)
Anders Berglund, PhD, Hanna Vikman, MSc, Hampus Hållberg, BSc, Erik Lampa, PhD.
Epistat AB, Uppsala, Sweden.
Epistat AB, Uppsala, Sweden.
OBJECTIVES: Access to individual-level cancer registry data is often restricted due to privacy regulations, limiting health-economic and methodological research. Synthetic data may offer a privacy-preserving alternative while retaining key characteristics of real-world evidence. This study evaluated whether synthetic data generated from the Swedish cancer registry could reproduce patient characteristics and survival outcomes.
METHODS: Data from the Swedish Cancer Registry were used to identify a population-based cohort of 370,137 patients with breast, prostate, lung, or colorectal cancer between 2014-2024. Synthetic datasets were generated using the synthpop package in R. A full synthetic dataset and additional datasets based on 5% and 1% random samples of the original population were created. Demographic characteristics and overall survival were compared between original and synthetic data. Results were implemented in an interactive R Shiny dashboard.
RESULTS: The synthetic datasets closely reproduced distributions of age, sex, cancer type, disease stage, year of diagnosis, vital status, and survival time. Differences between original and synthetic datasets were minimal. Kaplan-Meier analyses showed strong agreement, with survival curves closely overlapping throughout most of the follow-up period, although some divergence was observed in the tails. Stratified analyses confirmed preservation of survival differences across cancer types, disease stages, and calendar years. Similar patterns were observed in datasets generated from smaller samples, although variability increased as sample size decreased. Findings were visualized in an interactive dashboard enabling transparent exploration of outcomes and clinically relevant subgroups.
CONCLUSIONS: Synthetic data generated from the Swedish Cancer Registry preserved demographical characteristics and survival patterns while reducing disclosure risk. Findings were consistent across synthesis approaches, supporting the use of synthetic data for real-world evidence generation, health-economic research, and methodological development when patient-level data are inaccessible. An interactive dashboard enables transparent exploration of outcomes and provides a platform for future decision-support applications.
METHODS: Data from the Swedish Cancer Registry were used to identify a population-based cohort of 370,137 patients with breast, prostate, lung, or colorectal cancer between 2014-2024. Synthetic datasets were generated using the synthpop package in R. A full synthetic dataset and additional datasets based on 5% and 1% random samples of the original population were created. Demographic characteristics and overall survival were compared between original and synthetic data. Results were implemented in an interactive R Shiny dashboard.
RESULTS: The synthetic datasets closely reproduced distributions of age, sex, cancer type, disease stage, year of diagnosis, vital status, and survival time. Differences between original and synthetic datasets were minimal. Kaplan-Meier analyses showed strong agreement, with survival curves closely overlapping throughout most of the follow-up period, although some divergence was observed in the tails. Stratified analyses confirmed preservation of survival differences across cancer types, disease stages, and calendar years. Similar patterns were observed in datasets generated from smaller samples, although variability increased as sample size decreased. Findings were visualized in an interactive dashboard enabling transparent exploration of outcomes and clinically relevant subgroups.
CONCLUSIONS: Synthetic data generated from the Swedish Cancer Registry preserved demographical characteristics and survival patterns while reducing disclosure risk. Findings were consistent across synthesis approaches, supporting the use of synthetic data for real-world evidence generation, health-economic research, and methodological development when patient-level data are inaccessible. An interactive dashboard enables transparent exploration of outcomes and provides a platform for future decision-support applications.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
RWD16
Topic
Clinical Outcomes, Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Data Protection, Integrity, & Quality Assurance, Distributed Data & Research Networks
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Genetic, Regenerative & Curative Therapies, Neurological Disorders, Oncology, Rare & Orphan Diseases