Re-identification Risks and the Mosaic Effect : Are Health Data Ever Truly Anonymous ?

Insights from Panel 1 of “Data Privacy for Health” Summit, Brussels, 29 January 2026

At the DPH Summit in Brussels, the SRB case reshaped the debate on anonymization under the GDPR. For the Life Sciences sector, the question is no longer whether pseudonymized Data qualifies as Personal Data, but how contextual risk, governance responsibilities, and emerging technologies are redefining compliance and innovation.

🔐 When Pseudonymization Is Not Enough

Timmothy Dangeon opened the discussion with a technically grounded and deliberately provocative message: pseudonymization should never be assumed; it must be demonstrated and challenged. His presentation focused on MRI Data and the risk of facial reconstruction, illustrating how Data thought to be de-identified can still carry re-identification potential when analyzed with advanced techniques.

Imaging Data is often perceived as neutral once names and direct identifiers are removed. Yet MRI scans contain structural information that, under certain conditions, may allow reconstruction of facial features. The issue is not that such re-identification is routine or trivial, but that the possibility exists, and evolves with technological progress. What was once considered “remote” may become “reasonable” as tools become more accessible.

This intervention set the tone for the day: anonymization is not a static label but a moving target. Risk must be assessed dynamically, and claims of anonymity must be supported by evidence rather than assumptions. In this context, the SRB case does not eliminate responsibility; it increases the need for methodological rigor. If identifiability is contextual, then contextual analysis must be robust.

The implications for research are immediate. Clinical imaging Datasets, especially in rare diseases, cannot rely solely on the removal of direct identifiers. Risk assessment must consider reconstruction techniques, Data linkage possibilities and future technological developments. The message was clear: legal comfort without technical validation is fragile.

🧬 Synthetic Data as a Structural Response

Where Dangeon highlighted the limits of pseudonymization, Professor Pierre-Antoine Gourraud proposed a forward-looking alternative: augmented ( fit-for use) anonymous synthetic Data as a structural solution to the identifiability dilemma.

Grounding his analysis in the well-established anonymization criteria of the former Article 29 Working Party (“WP 29”), singling out, linkability and inference, he reminded the audience that anonymization must prevent all three. The challenge is that traditional anonymization techniques often degrade utility, making Datasets less valuable for research while still failing to eliminate residual risk. It is also strict application of GDPR minimization principle, if possible why risking re-identification ?

Synthetic Data generation offers a different approach. Rather than masking or suppressing original Data, it creates new Datasets that replicate statistical structure and analytical behavior without corresponding to real individuals. Properly designed Synthetic Datasets exhibit structural similarity, informational relevance and indistinguishability from the original Dataset in analytical terms.

For research, the potential is considerable. Synthetic Datasets can facilitate open science, allow Data sharing for peer review, support training environments for AI models, and enable sandbox experimentation without exposing Personal Data. In Clinical Research contexts where Data reuse is essential for validation and replication, this could represent a paradigm shift.

However, Gourraud’s intervention was not a call for deregulation. The compliance of Synthetic Data generation with the anonymization criteria of WP 29 must be documented. It must be auditable. It must ensure that rare patterns do not inadvertently allow inference. Moreover, synthetic representations may still carry ethical implications if misused or misinterpreted.

The underlying proposition is strategic rather than tactical: instead of continuously debating whether pseudonymized Data crosses the anonymity threshold, research ecosystems might redesign their Data architecture to reduce reliance on Personal Data in the first place whenever possible. No need to use Personal Data for non-personal application.

🔬 Anonymization in Rare and Ultra-Rare Diseases

Donovan Sheppard brought the discussion back to fundamentals but applied to one of the most sensitive domains: rare and ultra-rare diseases. His central question, “What if you are alone?”, captured the structural vulnerability of anonymization in small populations.

In rare disease research, even heavily pseudonymized Datasets may allow singling out simply because of uniqueness. When only a handful of patients worldwide share a specific condition, combinations of non-direct identifiers can quickly narrow down possibilities. In such contexts, the argument that “we do not hold the key” may provide limited reassurance.

Sheppard emphasized that anonymization evaluation through the WP 29 criteria must heavily rely on the real-world context of the Dataset. Legal disclaimers and formal distancing from identification capabilities do not substitute for substantive risk evaluation.

His intervention also underscored the governance dimension. In multinational pharmaceutical environments, anonymization assessments influence secondary use strategies, Data sharing agreements, AI development pipelines and regulatory positioning. Divergent interpretations between partners can create operational friction and ethical tension.

The rare disease perspective therefore reinforces a broader lesson: identifiability is relational. It depends not only on Data structure but on context, ecosystem, and the evolving landscape of available knowledge.

🎤 The Practitioner’s Reality: What the Audience Revealed

The interactive exchanges throughout the summit revealed that professionals are navigating these debates in real time. When asked about the difficulties encountered in practice, participants mentioned divergent role interpretations (controller versus processor), challenges in conducting joint re-identification assessments, prolonged contracting discussions, ethics committee disagreements and reputational risks.

When asked what arguments organizations rely on to claim anonymity, many cited the absence of access to identification keys, the lack of direct identifiers, or the inability to re-identify from their specific perspective. These arguments mirror the SRB reasoning, yet their operationalization remains complex.

Most participants saw potential in the evolving interpretation but expressed uncertainty regarding consequences and responsibilities. This ambivalence reflects maturity rather than confusion. The sector recognizes the opportunity to reduce unnecessary regulatory burden but also understands that premature conclusions may create longer-term instability.

📌 Conclusion: From Relative Anonymity to Structured Responsibility

The discussions at the Summit converge toward a central insight: the SRB case does not simplify the anonymization debate, it reframes it.

If anonymity is contextual, then contextual analysis must be rigorous. If identifiability depends on reasonable means, then those means must be documented and periodically reassessed. If Synthetic Data offers a pathway beyond Personal Data constraints, it must be validated and governed responsibly.

For the Life Sciences sector, the challenge is architectural. It concerns how Data Ecosystems are designed, how responsibilities are allocated, and how trust is preserved. Innovation and Data Protection are not opposing forces; they are interdependent. Scientific progress depends on public trust that Personal Data will not be mishandled or prematurely declared anonymous.

The SRB reasoning invites organizations to move beyond binary thinking. Pseudonymization is not anonymity. Absence of a key is not absence of risk. Synthetic Data is not a magic solution. What is required is structured governance, transparent documentation and interdisciplinary collaboration between legal, technical and ethical experts.

In the end, anonymization is not merely a legal status, it is a commitment to protecting individuals while enabling collective benefit. The future of Health Research will depend less on semantic debates and more on our capacity to build Data frameworks that are robust, auditable and worthy of trust.

Author: Winnie Dongbou

Prev post
Next post
Powered by MyData-TRUST

Want to subscribe to our newsletter ?

Name(Required)
Privacy(Required)