FROM PRIVATE SITES TO BIG DATA WITHOUT COMPROMISING PRIVACY- A CASE OF NEUROIMAGING DATA CLASSIFICATION
Author(s)
Plis S1, Sarwate A2, Turner J3, Arbabshirani M1, Calhoun V1
1The Mind Research Network, Albuquerque, NM, USA, 2Rutgers, The State University of New Jersey, Piscataway, NM, USA, 3Georgia State University, Atlanta, GA, USA
OBJECTIVES: Open data sharing for large-scale studies is an expensive and resource-wasteful approach that does not scale well with participating sites. Critically, some data simply cannot be shared due to privacy concerns and/or risk of re-identification. We pursue distributed computation that only shares data derivatives, such as statistical summaries, as a notable alternative allowing data holders to maintain control over data access. Privacy protection, if implemented via the differential privacy framework, can quantify and controllably reduce the risk of sharing the results of computations on private data. Unfortunately, it can impair the quality of statistical estimates at a single site. Using structural MRI data we present a distributed differentially private classifier that greatly improves the accuracy. METHODS: We used a combined MRI dataset from four separate schizophrenia studies conducted at Johns Hopkins University (JHU), the Maryland Psychiatric Research Center (MPRC), the Institute of Psychiatry, London, UK (IOP), and the Western Psychiatric Institute and Clinic at the University of Pittsburgh (WPIC). The sample comprised 198 schizophrenia patients and 191 matched healthy controls. We trained differentially private classifiers on these data using a method that minimizes a perturbed version of the SVM classifier. The outputs of these classifiers were used to train a logistic regression classifier. A combined SMVs plus logistic regression classifier was tested for accuracy. RESULTS: Using 100 random splits of data into 70% training and 30% validation groups followed by splitting the training set into 11 sites of 24/25 subjects each, we assessed classification accuracy at each site and of combination per split. Average site-accuracy was below 78%, while our combined classifier averaged to 96%: substantial and significant improvement with Bonferroni-corrected p-values below 1.8e-33. CONCLUSIONS: Our approach provides a way to enable use of distributed big data while still providing privacy and thus a much larger effective sample size.
Conference/Value in Health Info
2014-05, ISPOR 2014, Palais des Congres de Montreal
Value in Health, Vol. 17, No. 3 (May 2014)
Code
PRM52
Topic
Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Confounding, Selection Bias Correction, Causal Inference, Reproducibility & Replicability
Disease
Mental Health