P344: Privacy-Preserving Federated Analysis Enables International Rare Disease Evidence Generation
Poster Presenter
Roopal Bhatnagar
Senior Data Analyst
Critical Path Institute United States
Objectives
To demonstrate how privacy-preserving federated analysis of U.S. and European autosomal dominant polycystic kidney disease (ADPKD) cohorts on Rare Disease Cures Accelerator–Data and Analytics Platform (RDCA-DAP) enables cross-regional evidence generation for rare disease drug development.
Method
Shared variables in harmonized Critical Path Institute (C-Path) PKD Consortium and European Rare Kidney Disease Registry (ERKReg) data were compared via federated t-tests and chi-square on RDCA-DAP, with checks for distribution, variance, and outliers.
Results
The RDCA-DAP is a secure research environment enabling federated analysis of distributed rare disease datasets while maintaining institutional data control. Data from two cohorts harmonized to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) were analyzed on RDCA-DAP without transferring patient-level data. The hypothesis tested was that, after harmonization to OMOP, key ADPKD clinical measures (e.g., eGFR and MRI-derived TKV) would be broadly comparable between U.S. and European cohorts, while demographic distributions could differ.
The U.S. dataset from C-Path PKD Consortium included 2,595 data points from observational cohorts and the Tolvaptan Efficacy and Safety in Management of ADPKD and Its Outcomes (TAME-PKD) clinical trial. The European cohort from ERKReg included 4,571 data points.
Key variables analyzed included age, sex, race/ethnicity, body mass index (BMI), serum creatinine (SCR), estimated glomerular filtration rate (eGFR), and total kidney volume (TKV) measured by computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, with Mayo imaging classification. Comparisons were conducted overall and within adult and pediatric populations.
Significant differences were observed in anthropometric and demographic measures, including BMI and race/ethnicity distributions, and Mayo imaging class distributions (p<0.001). However, core indicators of kidney function and disease burden, i.e. SCR, eGFR, and MRI-derived TKV, were broadly comparable between U.S. and European cohorts, particularly in adults, supporting the study hypothesis.
These findings show federated analysis on RDCA-DAP can generate cross-regional insights without central data pooling while preserving data privacy and governance, enabling multinational rare disease research and evidence generation relevant for clinical trial design and drug development.
Conclusion
This study showcases the capabilities of the RDCA-DAP to enable secure, privacy-preserving federated analysis across geographically distributed rare disease datasets. By harmonizing data to the OMOP CDM and executing analyses within a trusted research environment, investigators were able to collaboratively analyze U.S. and European ADPKD cohorts without transferring patient-level data or compromising institutional data governance.
These findings highlight how federated analytical infrastructures can overcome longstanding barriers to multinational rare disease research, including regulatory constraints, data-sharing limitations, and fragmented datasets. Through the RDCA-DAP platform, distributed datasets can be analyzed collaboratively to identify clinically meaningful similarities and differences across populations while maintaining privacy, security, and institutional control of data.
Importantly, this approach provides a scalable framework for generating cross-regional real-world evidence in rare diseases. Such capabilities are critical for supporting natural history studies, informing clinical trial design, enabling external control cohort development, and strengthening evidence generation for therapeutic development and regulatory decision-making.
Collectively, this work demonstrates how federated data analysis platforms like RDCA-DAP can accelerate rare disease research by enabling global collaboration across distributed datasets, unlocking insights that would otherwise remain inaccessible due to data-sharing constraints.