Blog

Governance

We Asked a Leading ChatBot If a Hospital's Newest AI System Was Compliant. It Said Yes.

We ran one real compliance question two ways — through the newest general-purpose AI chatbot and through LegisGate. The gap between the answers is the risk every AI regulatory intelligence team is quietly carrying.

6 min read · June 5, 2026

Most AI governance teams now reach for a chatbot when they need to know what regulations apply to a new AI deployment. It is fast, it is fluent, and it answers in seconds. We wanted to know what that habit actually costs. So we ran one real compliance question two ways.

The scenario was a realistic one. An AI clinical documentation tool, used by emergency-department physicians to transcribe patient conversations, draft clinical notes, and surface diagnostic input that flows into the patient record - across roughly 50,000 patients a year, in both the United States and the EU. The question was the one every deployment starts with: what regulations apply here?

We put that question to a leading general-purpose AI chatbot. Then we put the identical scenario through LegisGate. Here is what came back.


One said "you're fine." The other found two ways to draw enforcement on day one.

The chatbot's verdict, in its own words, was that the deployment was "well-positioned for a compliant deployment." Confirm a couple of agreements, run an impact assessment, post a disclosure, and proceed.

LegisGate returned 16 binding regulatory findings for the same deployment. Two of them were critical blockers - issues that have to be cleared before the system can responsibly go live. Not "consider these." Stop signs.

The trust scorecard

Side-by-side · identical scenario, identical inputs

General-purpose AI chatbotLegisGate
Overall0 / 6 - Confident, unverifiable6 / 6 - Perfect score
What a research output has to get rightChatbotLegisGate
Cited only law that is currently in force
Surfaced every binding framework in scope
Identified the critical go-live blockers
Returned the same result on every run
Traced each conclusion to the statute
Produced a reproducible, defensible record
MetricChatbotLegisGate
Out-of-date law cited as current10
Critical go-live blockers caught02

Scored on the six dimensions that decide whether research can be trusted as a system of record. The chatbot named several frameworks correctly - its failure was reliability, not knowledge: it missed binding obligations, cited law no longer in force, and returned a prose all-clear with no findings, severity, or blockers. LegisGate produced a severity-tiered, statute-traced Final Designation Report: 16 findings, two of them critical.


The chatbot wasn't wrong about everything. That's exactly the problem.

This is the part that should give governance leaders pause. The chatbot got real things right. It knew the major privacy frameworks. It understood the broad strokes. If it had simply been wrong, it would be easy to dismiss.

Instead, folded into the same confident answer were three different kinds of failure: a state AI law it cited that had already been repealed and replaced, several binding frameworks it never mentioned at all, and a medical-device question it waved away instead of flagging for review. Correct, out of date, and missing - all delivered in one even, authoritative tone.

Nothing marked which lines to trust. A human had to verify every one to find out - which means the research saved no one any work, it just hid the work.

That is the trap. A chatbot is a strong first draft. It is not a record. The day a governance team treats one as the other, the risk doesn't disappear - it moves off the AI and onto the people relying on it.


Why even the best model does this

AI drifts. Point it at a question and it generates a fluent, plausible answer - but even pointed at a single source of truth, it drifts away from that source as it writes. It paraphrases and loses precision. It fills gaps with confident invention. It re-weights what matters by feel rather than by law. A better model drifts more eloquently; it does not stop drifting. And the more eloquent it gets, the more its drift goes unchecked, because it earns trust it cannot actually back.

You don't solve that with a better prompt. You solve it with architecture.

That is what LegisGate is. Deterministic logic running over a single, maintained source of truth. Hardened templates that produce the same structure every time. A patent-pending process built to contain the drift while still using the power of AI - so the model does what models are good at, inside guardrails that don't let it quietly rewrite the law. Same scenario in, same defensible answer out, every single run. Complete within scope. Traced to the statute. A record an organization can reproduce and stand behind - not a chat-window answer that changes the next time someone asks.


About the test - so you don't have to take our word for it

We did not stack the deck with an outdated or low-grade model. The chatbot in this test was the newest and most advanced general-purpose AI model publicly available at the time of testing - a current frontier system. We queried it the way a busy governance team actually does: a direct, real-world question, not a forensic, hours-long interrogation.

That is the point. The gap you see above is not a story about a weak model. It is what happens when any chatbot, however capable, is used as a system of record instead of a first draft. The fix was never a smarter model. It was a process the answer can't drift out of.

This article is for informational purposes only and does not constitute legal advice. AI regulatory intelligence and compliance requirements vary by organization, jurisdiction, and use case. Consult qualified legal counsel before making compliance determinations or relying on this content for any legal, regulatory, or business purpose.

← Back to Blog
Talk to usWe're here to help
We Asked a Leading ChatBot If a Hospital's Newest AI System Was Compliant. It Said Yes. | LegisGate™