PREPARING MULTIMODAL REAL-WORLD ONCOLOGY DATA FOR FEDERATED RESEARCH: AN END-TO-END ANONYMIZATION AND RE-IDENTIFICATION RISK ASSESSMENT PIPELINE
Author(s)
Alice Andalo', Msc1, Valentina Danesi, PhD2, Paolo De Angelis, Msc2, Andrea Roncadori, BSc, MSc2, Ilaria Massa, BSc2, Nicola Gentili, Msc2.
1Data Scientist, IRCCS IRST, Meldola, Italy, 2IRCCS IRST, Meldola, Italy.
1Data Scientist, IRCCS IRST, Meldola, Italy, 2IRCCS IRST, Meldola, Italy.
OBJECTIVES: Federated research infrastructures support privacy-preserving secondary use of real-world data, but local datasets may still pose re-identification risks. This study aimed to develop and empirically validate an end-to-end anonymization pipeline for multimodal data in federated research, while preserving analytical utility and documenting residual risk.
METHODS: We conducted a case study within the FLUTE project at an Italian cancer center. The dataset included 285 retrospective prostate cancer cases, seven structured diagnostic variables, a binary endpoint for clinically significant prostate cancer, and matched MRI DICOM studies. The workflow combined governance safeguards, pseudonymization for clinical-imaging linkage, clinical data anonymization, metadata minimization, and empirical re-identification testing. Clinical data were anonymized with ARX using k-anonymity, l-diversity, and controlled quasi-identifier generalization. Imaging data were processed through a policy-driven pipeline using a strict allow-list, retaining only metadata needed for file validity and quantitative biomarker extraction. Residual risk was assessed through a simulated insider re-identification test involving internal professionals with role-based access profiles.
RESULTS: The pipeline generated a research-ready multimodal dataset linking anonymized clinical records with MRI studies for federated analysis. Structured data met predefined privacy constraints while retaining variables needed for predictive modeling. The pipeline removed or transformed identifiers, normalized temporal information, selected prostate-relevant series, and preserved acquisition-critical metadata for quantitative analysis. In empirical testing, no re-identification occurred through software interfaces. One linkage was possible only in a privileged backend scenario, requiring cross-system queries, free-text report review, probabilistic discrimination among candidates, and substantial manual effort.
CONCLUSIONS: A data protection-by-design anonymization pipeline can support secondary use of multimodal real-world data in federated research while preserving analytical utility and reducing re-identification risk. The findings show that federated learning and local anonymization are complementary safeguards, requiring governance controls, clinical data anonymization, metadata minimization, linkage-key destruction, and empirical validation under realistic access scenarios.
METHODS: We conducted a case study within the FLUTE project at an Italian cancer center. The dataset included 285 retrospective prostate cancer cases, seven structured diagnostic variables, a binary endpoint for clinically significant prostate cancer, and matched MRI DICOM studies. The workflow combined governance safeguards, pseudonymization for clinical-imaging linkage, clinical data anonymization, metadata minimization, and empirical re-identification testing. Clinical data were anonymized with ARX using k-anonymity, l-diversity, and controlled quasi-identifier generalization. Imaging data were processed through a policy-driven pipeline using a strict allow-list, retaining only metadata needed for file validity and quantitative biomarker extraction. Residual risk was assessed through a simulated insider re-identification test involving internal professionals with role-based access profiles.
RESULTS: The pipeline generated a research-ready multimodal dataset linking anonymized clinical records with MRI studies for federated analysis. Structured data met predefined privacy constraints while retaining variables needed for predictive modeling. The pipeline removed or transformed identifiers, normalized temporal information, selected prostate-relevant series, and preserved acquisition-critical metadata for quantitative analysis. In empirical testing, no re-identification occurred through software interfaces. One linkage was possible only in a privileged backend scenario, requiring cross-system queries, free-text report review, probabilistic discrimination among candidates, and substantial manual effort.
CONCLUSIONS: A data protection-by-design anonymization pipeline can support secondary use of multimodal real-world data in federated research while preserving analytical utility and reducing re-identification risk. The findings show that federated learning and local anonymization are complementary safeguards, requiring governance controls, clinical data anonymization, metadata minimization, linkage-key destruction, and empirical validation under realistic access scenarios.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
RWD92
Topic
Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Data Protection, Integrity, & Quality Assurance, Reproducibility & Replicability
Disease
Oncology