The Verification Trap at Scale

What legal AI incidents reveal about four governance failures that law firms keep making.

I recently had a tooth surgically removed due to a tooth fracture (caused, in part, by my unrequited love for food). The surgeon and his team spent the first several minutes making sure I was the right guy by confirming my name and date of birth, verifying which tooth needed to be removed, and going through a detailed verification checklist. It was at a small dental clinic, but the Mayo Clinic would have had to go through the same detailed mandatory “Time Out” verification protocol for a major surgery; the team must count every sponge, clamp, and needle, and double-check the patient and procedure. This verification is one of the most important things standing between a successful procedure and an absolute legal nightmare, ensuring that the person on the table is actually Bob here for a heart transplant, and not Bill who came in to restock the vending machine and decided to take a quick nap on the operating table.

The legal-specific incidents in the AI Incident Database, along with Damien Charlotin’s hallucination database make clear that legal AI hallucinations in court filings follow the same pattern. Verification adds time; it’s tedious. And skipping it produces the same embarrassing (in the best case), or more likely, catastrophic, outcomes regardless of whether you bill $200 an hour or $2000.

Yet, the incidents are getting more pervasive.

In June 2023, Mata v. Avianca made two lawyers infamous for filing six cases ChatGPT invented. The sanction was $5,000. You’d think a public embarrassment that loud would have taught the profession something. It didn’t.

As of August 2026, Charlotin’s hallucination database catalogues 1847 cases worldwide (more than 1200 in the U.S. alone) where courts found lawyers had relied on fabricated AI-generated material. The rate is accelerating at roughly eight new entries per day as of mid-2026, up from five per day just months earlier. The financial penalties have also compounded from $5,000 in 2023 to over $100,000 in 2026.

And the firms getting caught aren’t getting smaller. They’re getting bigger.

Sullivan & Cromwell—one of the most prestigious law firms on the planet—filed an emergency bankruptcy motion in April 2026 containing approximately 40 AI-generated errors: fabricated citations, incorrect pin cites, and misquoted bankruptcy provisions. The firm’s own apology letter acknowledged that “comprehensive policies and training requirements governing the use of AI tools in legal work” existed and “were not followed.”

This should keep managing partners up at night: Sullivan & Cromwell is the firm that advises OpenAI on safe and ethical deployment of AI. If they can’t prevent this internally, the question isn’t whether your firm has the right policy. The question is whether the right policy is sufficient.

I don’t think it is. We will, no doubt, see many more cases like this, given the continued frenzy about deploying AI—especially AI agents—in legal practice. And I think the reason is structural, not cultural. This is the verification trap.

I’ve written before about verification challenges in legal AI: for judgment-heavy work, the cost of expert verification approaches the cost of doing the original task. A junior associate using AI to draft a brief might produce text in 20 minutes that would have taken four hours to write from scratch. But a senior attorney verifying every citation, quotation, and precedent characterization may still take three hours, collapsing the efficiency gain.

This is why the hallucination problem hasn’t improved despite near-universal awareness. Awareness tells you to verify; the economics tell you verification is expensive; time pressure tells you to cut it short; and the hallucination cases keep accumulating.

The verification trap doesn’t manifest in a single way. Looking across the incident database, four distinct failure patterns emerge, and the governance response has focused almost entirely on the wrong one.

Failure Type 1: “We didn’t know”

The earliest cases were straightforward ignorance. Steven Schwartz in Mata v. Avianca didn’t know ChatGPT could fabricate case citations. He asked it to confirm the cases were real, and it did. Then he filed.

A self-represented appellant in an Arizona probate appeal submitted a brief where six of eight citations were deficient or entirely fictitious. An attorney in Norway filed hallucinated citations to the country’s Supreme Court. In December 2023, Michael Cohen, the former Trump attorney, used Google Bard (RIP, Bard!) to find precedent as a client, passed the results to his attorney, and the attorney filed them without verification.

These cases share a common feature: the person using the AI didn’t understand the failure mode. They treated AI output like they’d treat a Westlaw search result—as something that wouldn’t be fabricated from whole cloth.

The legal profession has responded to this failure with awareness campaigns, CLE requirements, and court standing orders. But the problem is that Type 1 failures are already becoming a minority of cases. The lawyers getting sanctioned in 2025 and 2026 aren’t ignorant. They know AI hallucinations exist, but they just can’t make the verification economics work.

That is, the awareness problem was solved. But the verification trap remained.

Failure Type 2: “We had a policy”

Picture an associate working on an emergency motion at 2 a.m. The deadline is in six hours; the AI draft is on the screen. The governanace policy exists, and the associate has completed the training. But none of it matters because verification takes time the deadline doesn’t allow.

This is the failure mode your firm is most vulnerable to.

Sullivan & Cromwell had a policy. The firm maintained, in its own apology letter, “comprehensive policies and training requirements governing the use of AI tools in legal work” that were “designed to prevent exactly this situation.” But the policy wasn’t followed.

Sullivan & Cromwell isn’t alone here. K&L Gates and Ellis George LLP were sanctioned $31,000 after filing briefs with AI-hallucinated citations, and then re-filing a corrected brief that still contained errors. They knew the first brief was contaminated and attempted to fix it. And they still couldn’t catch everything because thorough verification of a complex brief takes time—potentially as long as writing the brief. Even under sanction-induced time pressure, they still cut verification short.

In August 2025, a senior King’s Counsel in Australia filed defense submissions in a murder trial (a murder trial!) containing hallucinated quotes and nonexistent precedent. The error delayed the trial by 24 hours, and required a public apology in open court, including the deep embarrassment to the defense team.

The pattern here is experienced, senior attorneys at sophisticated institutions, working within governance frameworks, producing hallucinated output anyway. The policies exist, but the verification economics make them impractical under real-world conditions. And the gap between policy-on-paper and practice-under-pressure is where the sanctions accumulate.

Failure Type 3: “We bought the right tool”

This is the failure that law firms are least prepared for, and that has me the most worried.

After Mata v. Avianca, the industry’s dominant response was: ChatGPT is a consumer toy. Use a proper legal research tool. Problem solved. Thomson Reuters claimed its products “avoid hallucinations by relying on the trusted content within Westlaw.” Casetext (before its acquisition by Thomson Reuters) explicitly stated that CoCounsel “does not make up facts, or ‘hallucinate.’”

In April 2026, the Sixth Circuit published its opinion in United States v. Farris. A defense attorney had used Westlaw’s CoCounsel to draft an appellate brief. The court’s first clue was the file name: “CoCounsel Skill Results.” The attorney hadn’t even renamed the file before filing it. Think about that for a second. An attorney filed a brief with a federal appellate court, and the file name was the default output from an AI tool. No one opened it in a word processor, formatted it, or did the bare minimum of renaming the file to something that wouldn’t immediately signal to the court that the entire document was machine-generated. The brief cited real cases, but the quotations attributed to those cases were fabricated.

Apart from the fact that you should take vendor promises with a grain of salt (I have nothing against vendors, but their incentive is to sell you their tool), this is a more dangerous failure mode than the Mata v. Avianca-style hallucination. When AI invents a case entirely, catching it is simple: you look it up and it doesn’t exist. But when AI cites a real case but fabricates the quotation, catching the error requires reading the actual opinion and comparing it word-for-word against the brief. That is harder and takes longer, and the kind of verification burden that law firms try to avoid by buying better tools.

Stanford’s RegLab 2025 evaluation of commercial legal AI tools found that Lexis+ AI hallucinated on more than 17% of queries and Westlaw’s AI-Assisted Research hallucinated on more than 33%. These numbers include not just fabricated citations but mischaracterized holdings, inapplicable authority, and incorrect legal propositions attributed to real sources.

A tool that’s wrong one time in five, used by thousands of time-pressed lawyers daily creates a certainty mirage; false sense of security in the absolute correctness of the tool, given the four times it got things right.

The structural point is that buying a better tool doesn’t solve the verification trap. It shifts the failure mode from obvious fabrication (easy to catch) to subtle mischaracterization (hard to catch). The verification cost doesn’t go down. In some cases, it increases.

Failure Type 4: “It wasn’t our AI”

In Withers v. City of Aberdeen, a contract dispute in Mississippi, something unprecedented happened. Attorneys on both sides of the case independently used generative AI to draft filings. Both sides produced hallucinated citations, and neither side caught the other’s errors. The court found that all four attorneys had violated Federal Rule of Civil Procedure 11—the rule that makes every attorney who signs a filing personally responsible for the accuracy of its contents. The judge cancelled the trial, disqualified all four attorneys, barred the two pro hac vice attorneys from practicing in the district for two years, imposed fines, and referred the matter to state bar authorities.

One of the sanctioned attorneys was subsequently sanctioned again in a separate bankruptcy case two months after the Withers show-cause hearing. For the same failure. The court found her apology “insincere” and concluded that she had acted in bad faith.

Withers exposes a governance gap that almost no firm has addressed. The standard AI governance playbook is entirely inward-facing: verify your output before filing. But Withers produced a contaminated record because both sides used unreliable tools and nobody—not the lawyers, local counsel, or the court—caught it until after the damage was done.

The same structural gap appears in the Michael Cohen case. Cohen used Google Bard to find cases and passed the results to his attorney, who filed them without verification. The supply chain was contaminated before it reached the lawyer’s desk.

And then there’s the Anthropic irony. In Concord Music Group v. Anthropic—a copyright case against the company that builds Claude—Anthropic’s defense team admitted that an expert declaration contained a fabricated citation generated when Claude was used to format citations. The AI company, defending its own AI product in federal court, was caught by its own AI’s hallucination. The court called it “a plain and simple AI hallucination,” struck the relevant paragraph, and noted it “undermines the overall credibility” of the expert’s entire declaration.

Sullivan & Cromwell, the firm that advises OpenAI on safe and ethical AI deployment, filed a brief with 40 AI-generated errors. And Anthropic, the company that builds Claude, got caught by its own model’s hallucination while defending that model in court. If the AI company’s governance advisor and the AI company itself both failed at verification, the problem isn’t institutional competence or a governance failure, but a structural feature of the technology itself.

The verification trap is not a governance failure. It’s a structural feature of the technology.

Failure Taxonomy

Why awareness doesn’t solve the verification trap

According to the recent Supio AI Adoption Report, which surveyed 207 U.S. personal injury attorneys, 99% of plaintiff attorneys said they won’t use AI content they cannot verify, and 96% are very or extremely concerned about untraceable AI output.

The lawyers getting sanctioned in 2026 aren’t the ones who never heard of Mata v. Avianca. They’re the ones who heard about it, adopted policies, bought better tools, and still couldn’t close the gap between what verification requires and what their deadlines allow.

This is because the verification trap is an economic problem. Telling lawyers to “verify all AI output” is the equivalent of telling surgeons to “not make mistakes.” The instruction is correct, but insufficient. What surgeons actually do is build verification into the process architecture: the sponge count, the surgical checklist, the time-out before incision. These are structural interventions that make the correct behavior the default behavior, even under time pressure.

The legal industry hasn’t built its equivalent yet. And until it does, the database will keep growing every day.

Three governance actions for this quarter

I’m not going to suggest that you “adopt an AI governance policy.” You probably already have one. The question is whether it accounts for the verification cost.

Audit your verification time, not your AI usage. Track how long attorneys actually spend verifying AI-generated citations and legal propositions, compared to the time saved by using the tool. If the ratio approaches 1:1, the tool isn’t saving time. It’s redistributing it, and creating liability in the process.

Build verification into process architecture, not checklists. A checklist item that says “verify all citations” only addresses Type 1 (awareness) failures. A workflow that routes every AI-assisted brief through independent citation verification by a different attorney is a structural intervention. The distinction matters because checklists fail under time pressure, while workflows survive it.

Extend your governance perimeter to inputs. After Withers, the assumption that opposing counsel’s citations are reliable is a liability. After the Cohen case, the assumption that client-provided legal research is reliable is a liability. Your verification protocol needs to cover material you didn’t generate.

The sanctions and fees have compounded from $5,000 to over $100,000, and are likely to grow even bigger. The firms are getting bigger, from solo practitioners to AmLaw 100. The case count is accelerating at a clip of eight new entries per day. And the AI company that builds one of the most prominent models in the market couldn’t prevent its own product from hallucinating in its own legal defense.

And this analysis only covers tools that sit still and wait for a prompt. In the last few weeks, AI agents from OpenAI and Anthropic “escaped” their testing environments and broke into real companies, and neither lab was monitoring the agents when it happened. The verification problem doesn’t get easier when the AI stops waiting for instructions and starts acting on its own. It gets worse in ways that the current governance conversation isn’t remotely prepared for. I’ve written before about why AI agents in legal practice make me nervous, and nothing in the past month has changed my mind.

The trajectory doesn’t flatten because managing partners send a firm-wide memo. It flattens when firms restructure their workflows to account for the costs of verification, or when malpractice insurers do it for them.

The sponge count isn’t optional because hospitals wrote a policy about it. It’s mandatory because the consequences of skipping it made the policy insufficient.

Legal AI verification is heading to the same place. The only question is how many firms join the database before they get there.