DISTRIBUTIONAL FIDELITY OF SYNTHETIC GERMAN CLAIMS DATA UNDER REAL-WORLD STUDY DESIGNS: IMPLICATIONS FOR HCRU AND COST ANALYSIS

Author(s)

Tobias Heidler, Staatsexamen Pharmazie1, Michael Schultze, Dr., Staatsexamen Tiermedizin2, George Kafatos, PhD, MSc3, Bagmeet Behera, PhD, MSc4, Caroline Lienau, MSc5, Alexander Franz Unger, Dr., Mag.5, Valentina Balko, DPhil5, Julius Brandenburg, PhD, Diplom Biologie5, Lea Grotenrath, MSc5, Zhenchen Wang, PhD6, Philipp Großer, MSc7, Adam Hilbert, MSc8, Nils Kossack, Dipl.-Math.1, Marc Pignot, PhD, MSc2.
1WIG2 GmbH – Scientific Institute for Health Economics and Health System Research, Leipzig, Germany, 2ZEG – Berlin Center for Epidemiology and Health Research GmbH, Berlin, Germany, 3Amgen Limited, Uxbridge, United Kingdom, 4Amgen Research (Munich) GmbH, Munich, Germany, 5AstraZeneca GmbH, Hamburg, Germany, 6Medicines and Healthcare products Regulatory Agency (MHRA), London, United Kingdom, 7Limebit GmbH, Berlin, Germany, 8ai4medicine UG, Berlin, Germany.
OBJECTIVES: Synthetic health data are proposed as privacy-preserving alternatives to real-world claims data for health economics research. Whether these datasets preserve distributional properties required for HCRU analyses - drug, outpatient, and inpatient costs and utilization - under real-world study designs remains poorly understood. We quantify the distributional fidelity of four synthetic German SHI (Statutory Health Insurance) datasets across major HCRU cost domains under both a raw population and an applied SLE (Systemic Lupus Erythematosus) cohort definition.
METHODS: Four synthetic datasets (DS-ARF, DS-BNN-WangTucker, DS-BNN-Kaur, DS-GAN) were compared against a German SHI ground truth (~6,743 patients; DS-WIG2) across drug, outpatient, and inpatient outcomes. Distributional similarity (1-KS-D), normalised mean difference (NMD), Wasserstein distance, and tail quantile ratios were assessed for full and conditional distributions under two scenarios: raw population and applied SLE cohort (SLE diagnosis, ≥1-year continuous pre-index insurance, age ≥18).
RESULTS: Cohort attrition exposed a 7× gap: DS-ARF and DS-BNN-WangTucker retained only ~8% of patients after cohort selection (reference: 60%), which substantially affected all downstream analyses. Raw scenario drug cost NMD ranged +12% (DS-ARF) to −68% (DS-GAN); outpatient NMD +3% to −87%; inpatient cost similarity 0.65-0.96. DS-BNN-Kaur showed a notable discrepancy: 60% outpatient zero-utilisation (reference: 10%) with +128% conditional cost inflation resulting in coincidental mean equality (€1,116 vs €1,114; KS-similarity=0.40). P99 drug costs ranged 21-113% of reference; comparable tail truncation was observed in outpatient and inpatient domains. Under the SLE cohort definition, fidelity deteriorated across generators and HCRU domains.
CONCLUSIONS: While synthetic datasets were able to reproduce some broad HCRU and cost patterns, important differences remained across all domains. These discrepancies could meaningfully affect estimates. Careful validation is therefore important before synthetic data are used for health economic analyses or decision-making.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

RWD149

Topic

Real World Data & Information Systems, Study Approaches

Topic Subcategory

Data Protection, Integrity, & Quality Assurance, Health & Insurance Records Systems

Disease

Systemic Disorders/Conditions (Anesthesia, Auto-Immune Disorders (n.e.c.), Hematological Disorders (non-oncologic), Pain)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×