PREPARING MULTIMODAL REAL-WORLD ONCOLOGY DATA FOR FEDERATED RESEARCH: AN END-TO-END ANONYMIZATION AND RE-IDENTIFICATION RISK ASSESSMENT PIPELINE

Author(s)

Alice Andalo', Msc1, Valentina Danesi, PhD2, Paolo De Angelis, Msc2, Andrea Roncadori, BSc, MSc2, Ilaria Massa, BSc2, Nicola Gentili, Msc2.
1Data Scientist, IRCCS IRST, Meldola, Italy, 2IRCCS IRST, Meldola, Italy.
OBJECTIVES: Federated research infrastructures support privacy-preserving secondary use of real-world data, but local datasets may still pose re-identification risks. This study aimed to develop and empirically validate an end-to-end anonymization pipeline for multimodal data in federated research, while preserving analytical utility and documenting residual risk.
METHODS: We conducted a case study within the FLUTE project at an Italian cancer center. The dataset included 285 retrospective prostate cancer cases, seven structured diagnostic variables, a binary endpoint for clinically significant prostate cancer, and matched MRI DICOM studies. The workflow combined governance safeguards, pseudonymization for clinical-imaging linkage, clinical data anonymization, metadata minimization, and empirical re-identification testing. Clinical data were anonymized with ARX using k-anonymity, l-diversity, and controlled quasi-identifier generalization. Imaging data were processed through a policy-driven pipeline using a strict allow-list, retaining only metadata needed for file validity and quantitative biomarker extraction. Residual risk was assessed through a simulated insider re-identification test involving internal professionals with role-based access profiles.
RESULTS: The pipeline generated a research-ready multimodal dataset linking anonymized clinical records with MRI studies for federated analysis. Structured data met predefined privacy constraints while retaining variables needed for predictive modeling. The pipeline removed or transformed identifiers, normalized temporal information, selected prostate-relevant series, and preserved acquisition-critical metadata for quantitative analysis. In empirical testing, no re-identification occurred through software interfaces. One linkage was possible only in a privileged backend scenario, requiring cross-system queries, free-text report review, probabilistic discrimination among candidates, and substantial manual effort.
CONCLUSIONS: A data protection-by-design anonymization pipeline can support secondary use of multimodal real-world data in federated research while preserving analytical utility and reducing re-identification risk. The findings show that federated learning and local anonymization are complementary safeguards, requiring governance controls, clinical data anonymization, metadata minimization, linkage-key destruction, and empirical validation under realistic access scenarios.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

RWD92

Topic

Methodological & Statistical Research, Real World Data & Information Systems

Topic Subcategory

Data Protection, Integrity, & Quality Assurance, Reproducibility & Replicability

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×