FROM PRIVATE SITES TO BIG DATA WITHOUT COMPROMISING PRIVACY- A CASE OF NEUROIMAGING DATA CLASSIFICATION

Author(s)

Plis S1, Sarwate A2, Turner J3, Arbabshirani M1, Calhoun V1
1The Mind Research Network, Albuquerque, NM, USA, 2Rutgers, The State University of New Jersey, Piscataway, NM, USA, 3Georgia State University, Atlanta, GA, USA

OBJECTIVES: Open data sharing for large-scale studies is an expensive and resource-wasteful approach that does not scale well with participating sites.  Critically, some data simply cannot be shared due to privacy concerns and/or risk of re-identification. We pursue distributed computation that only shares data derivatives, such as statistical summaries, as a notable alternative allowing data holders to maintain control over data access. Privacy protection, if implemented via the differential privacy framework, can quantify and controllably reduce the risk of sharing the results of computations on private data. Unfortunately, it can impair the quality of statistical estimates at a single site.  Using structural MRI data we present a distributed differentially private classifier that greatly improves the accuracy. METHODS: We used a combined MRI dataset from four separate schizophrenia studies conducted at Johns Hopkins University (JHU), the Maryland Psychiatric Research Center (MPRC), the Institute of Psychiatry, London, UK (IOP), and the Western Psychiatric Institute and Clinic at the University of Pittsburgh (WPIC). The sample comprised 198 schizophrenia patients and 191 matched healthy controls. We trained  differentially private classifiers on these data using a method that minimizes a perturbed version of the SVM classifier.  The outputs of these classifiers were used to train a logistic regression classifier. A combined SMVs plus logistic regression classifier was tested for accuracy. RESULTS: Using 100 random splits of data into 70% training and 30% validation groups followed by splitting the training set into 11 sites of 24/25 subjects each, we assessed classification accuracy at each site and of combination per split. Average site-accuracy was below 78%, while our combined classifier averaged to 96%: substantial and significant improvement with Bonferroni-corrected p-values below 1.8e-33. CONCLUSIONS: Our approach provides a way to enable use of distributed big data while still providing privacy and thus a much larger effective sample size.

Conference/Value in Health Info

2014-05, ISPOR 2014, Palais des Congres de Montreal

Value in Health, Vol. 17, No. 3 (May 2014)

Code

PRM52

Topic

Methodological & Statistical Research, Real World Data & Information Systems

Topic Subcategory

Confounding, Selection Bias Correction, Causal Inference, Reproducibility & Replicability

Disease

Mental Health

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×