Generate Synthetic Data in R for a Hypothetical Alzheimer's Disease Trial
Author(s)
Handels R1, Jönsson L2, Raket LL3
1Maastricht University, maastricht, LI, Netherlands, 2Karolinska Institutet, Stockholm, Sweden, 3Lund University, Lund, Sweden
Presentation Documents
OBJECTIVES: Representative data of recent Alzheimer’s Disease (AD) trials are difficult to obtain. We aimed to generate a synthetic version of an original real-world observational dataset, subsequently apply a plausible AD treatment effect, and make our method open-source available.
METHODS: Synthetic data was generated in the following steps: (1) Obtain real-world data from the ADNI study on demographic (age, sex, education), clinical (cognition: MMSE and ADAS; function: FAQ; composite cognition/function: CDR, ADCOMS) and biological (genetics: APOE4; cerebrospinal fluid: ABeta, Tau; imaging: PET-SUVR-centiloid) outcomes at baseline, 6, 12 and/or 18-month follow-up (35 variables), with missing data multiple-imputed to obtain 10 sets of 537 individuals. (2) Estimate (theoretical) minimum and maximum (all continuous variables) and proportions (all categorical variables). (3) Rescale to 0-1 range (continuous). (4) Estimate beta distribution shape parameters (method of moments; continuous). (5) Transform to cumulative density function (using shape parameters; continuous) and to cumulative probability (categorical). (6) Converted to a normal distribution. (7) Estimate variance-covariance matrix. (8) Generate random correlated normal data using Cholesky decomposition of variance-covariance. (9) Transform to cumulative density function. (10) Transform to inverse cumulative density function of beta distribution (using shape parameters; continuous). (11) Rescale to original range (using min/max and proportions from step 2). (12) Keep half as control arm, and half as intervention arm, and estimate change from baseline. (13) Multiply intervention change from baseline with self-defined hypothetical relative treatment effect. We assumed correlations on normalized scale were similar to correlations on original scale. Code is available on github: https://github.com/ronhandels/synthetic-correlated-data.
RESULTS: The synthetic distribution and mean over time showed large similarity to the original data (visually assessed). The absolute difference in pairwise correlations between original and synthetic data median was 0.02 (95th percentile=0.11, max=0.18).
CONCLUSIONS: We judged our method sufficiently valid to generate synthetic correlated plausible hypothetical trial results.
Conference/Value in Health Info
Value in Health, Volume 26, Issue 11, S2 (December 2023)
Code
MSR98
Topic
Clinical Outcomes, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Clinical Trials, Comparative Effectiveness or Efficacy, Decision Modeling & Simulation
Disease
Neurological Disorders, No Additional Disease & Conditions/Specialized Treatment Areas