Can a DPO Trust AI for Legal Advice? We Tested Claude with a Legal Plugin

For Data Protection Officers in Life Sciences, the temptation is real: could an AI assistant with a specialized legal plugin handle regulatory queries or review information notices, such as Informed Consent Forms? At MyData-TRUST, we ran a structured evaluation of Claude, Anthropic’s AI assistant, combined with its official legal plugin, across 13 regulatory Data Protection focused queries and 10 document reviews spanning five jurisdictions on four continents, from Europe to Asia and Latin America. The results reveal both the promise and the limits of AI-assisted Data Protection work.

🔍 What Did We Test and How?

Our evaluation focused on two core tasks that Data Protection professionals in Life Sciences perform regularly: answering regulatory queries about local Data Protection law, and reviewing Clinical Trial documents such as ICFs and CTAs for compliance. The temptation to leverage AI in this context is particularly high, given the large volume of documents to review and the short turnaround timelines.

We tested Claude with its official legal plugin across five countries on three continents: France and Switzerland in Europe, China and South Korea in Asia, and Chile in Latin America. We deliberately chose to cover both GDPR-aligned and non-GDPR jurisdictions with fundamentally different regulatory frameworks. Each output was scored by our in-house country experts on a scale from 0 to 5, with mandatory justification for every score, where 5 represents expert-level quality requiring no correction and scores below 3 were deemed unacceptable.

We tested two configurations: the legal plugin alone (which anyone can install), and the plugin enhanced with proprietary skills developed internally at MyData-TRUST leveraging our long lasting sector specific expertise, to isolate what the technology delivers out of the box versus what expert-engineered processes add.

⚖️ Legal Plugin Alone: Polished Surface, Fragile Foundation

Each response was evaluated against six criteria, including whether cited legal sources are real and verifiable, whether the legal information is factually correct, and whether the analysis covers all relevant obligations.

When used without our custom skills, the legal plugin produced answers that looked professional and well-structured. 77% of responses scored acceptably on clarity, and 85% on apparent relevance. On the surface, the output appeared trustworthy to a non-expert reader.

But the criteria that matter most for legal advice told a different story. Only 38% of responses provided accurate, verifiable legal references. Nearly half cited hallucinated sources, returned broken links, or offered no references at all. Legal accuracy was equally concerning: 54% of responses contained important factual errors. Without Data Protection expertise to catch these issues, 100% of responses failed on every substantive criterion.

Document reviews showed a similar pattern. Completeness was the weakest point: the plugin detected only 30 to 50% of required compliance findings across all countries tested. False positives further eroded reliability, flagging issues that were not actual non-compliance.

The implication is clear: an inexperienced practitioner relying on this output would deliver incorrect legal advice or miss critical gaps in consent forms, with potential consequences ranging from study delays due to regulatory findings to risks for patient rights.

⚠️ The “Silent Killer” Problem

What makes AI agents particularly dangerous in a legal context is their failure mode. AI delivers wrong answers with the same confident tone as correct ones. During our evaluation, we identified several recurring patterns:

  • Source-blind confidence: the tool uses the same authoritative tone whether citing official law text or an unverified blog post.
  • Silent law version mixing: outdated and current legislation are blended without any flag indicating which version is being cited.
  • Hallucinations indistinguishable from correct answers: fabricated legal references look identical to real ones.
  • No request for clarification: ambiguous or complex regulatory concepts are silently interpreted rather than questioned.
  • Inverse confidence effect: the less certain the tool is, the longer and more polished its output becomes.

These patterns mean that the output most likely to mislead factually is also the output most likely to be trusted by people from its form.

📉 When AI Gets It Wrong: Real-World Consequences

These risks are not theoretical. Several high-profile incidents have shown what happens when professionals rely on AI-generated legal content without adequate human verification.

In the United States, courts have sanctioned attorneys for submitting briefs with AI-fabricated citations, with fines reaching $10,000 for a single brief containing 15 made-up references. In the compliance space, a startup was recently accused of providing customers with fabricated evidence of compliance processes that never took place. The common thread: AI output looked authoritative and thus projected the sense of credibility, and the professionals who used it lacked the knowledge or the processes to catch the errors.

For Data Protection in Life Sciences, the stakes are even higher. A flawed ICF review does not just create legal exposure; it can directly affect the rights and safety of Clinical Trial participants.

🔄 With Expert-Engineered Skills: A Different Story

When we added MyData-TRUST’s proprietary skills, which consist in structured processes encoding both technical and domain expertise, the results improved dramatically.

For regulatory queries, all responses passed the professional reviewer threshold, with an average score of 4.5 out of 5. Legal accuracy was flawless across all countries. Legal references improved significantly, though occasional gaps remained in more complex or less explored jurisdictions like South Korea.

For document reviews, every country scored 4 or above on all criteria. Completeness jumped to over 90% detection of required findings in all five countries.

Legal query evaluation: plugin alone (Claude Legal) vs. plugin with MyData-TRUST skills (MDT Custom), scored on a 0–5 scale across six criteria

These skills are not simple prompts. Each one encodes thousands of lines of structured instructions combining verification processes with deep regulatory knowledge, designed to address the exact failure modes we observed in the plugin alone.

🚫 Red Lines That Remain

Even with these improvements, red lines persist. The technology remains non-deterministic: the same input when repeated does not guarantee the same output. Human validation by a qualified reviewer remains mandatory. Confidentiality is another concern: sending sensitive documents like ICFs and CTAs potentially containing commercially confidential or proprietary information to external AI providers raises questions that go beyond technical encryption.

The expert gap, the distance between near-expert AI output and true expert judgment, is precisely where the highest value lies. A non-compliant ICF directly impacts patient rights, exposes the sponsor to regulatory action, and can delay a study.

📌 What This Means for DPOs

AI-powered legal tools are becoming more capable, but capability without control is risk. Specifically in the compliance environment where anything that is not properly documented and justified (including legal references) is almost as if it would simply not exist. Out of the box, the technology is not ready for unsupervised Data Protection work. With the right expert-engineered safeguards, AI becomes a powerful accelerator, but it does not replace the fully-skilled professional who reviews, validates, and takes responsibility for the advice.

At MyData-TRUST, our evaluation confirmed that combining AI with deep regulatory expertise produces results that meet professional standards while improving efficiency and scalability. We are now working to make these capabilities available as a service, with built-in safeguards including document redaction using named entity recognition to address confidentiality concerns. We are also scaling our skills to new jurisdictions, a process that requires integrating country-specific regulatory knowledge supervised by our local experts. Our goal: enabling Data Protection professionals in Life Sciences to benefit from AI-assisted reviews with the quality controls and oversight that this work demands.

Authors: Hubert Stoop and Anastassia Negrouk

References

Prev post
Next post
Powered by MyData-TRUST

Want to subscribe to our newsletter ?

Name(Required)
Privacy(Required)