THE GENERATION OF SYNTHETIC CLINICAL TRIAL DATA
Author(s)
Mosquera L
Replica Analytics, Vancouver, ON, Canada
Presentation Documents
OBJECTIVES: Making clinical trial data available for secondary analysis and building innovative AI/machine learning models requires addressing privacy concerns. These concerns exist whether the data is shared internal or external to the organization. An effective way to address these concerns is through the creation of synthetic data. METHODS: Synthetic data was generated for a trial that assessed the efficacy of a new adjuvant chemotherapy for patients with gastrointestinal stromal tumors (available via Project Data Sphere). A classification tree method was used to create a synthetic data that retains the distributions and dependencies in the trial data. A comprehensive analysis of the data utility of the synthesized data was performed to determine if similar conclusions can be drawn from the synthetic dataset as from the original dataset. Utility tests included: comparisons of distributions, descriptive statistics, complex multivariate models, including propensity scores and the model’s associated predictive error. RESULTS: The synthetic data had almost identical distributions as the original data. Absolute differences in bivariate correlations were very small across all comparisons. The differences in the area under the ROC curve across all possible multivariate models built using CART were less than 0.05, and the absolute relative differences for continuous outcomes had a median below 20%. Global propensity scores statistics indicate that there was no substantial difference between the original and synthetic data. CONCLUSIONS: The utility analysis indicated that the synthetic data preserved the relationships between variables from the original data, using a simple data generation method. Synthetic data methods should be considered at least for exploratory secondary analysis on clinical trial data as they are relatively efficient to operationalize while mitigating privacy concerns.
Conference/Value in Health Info
2019-11, ISPOR Europe 2019, Copenhagen, Denmark
Code
PCN429
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology