AI hallucination in legal.

AI legal research tools hallucinate 17-34% of the time. They fabricate case citations with correct formatting, invent holdings from real courts, and cite precedents that never existed. Multiple attorneys have been sanctioned by courts for submitting AI-generated briefs with fabricated authorities.

The citation problem.

Legal AI hallucination is uniquely insidious because of how legal citations work. A citation must include a case name, reporter volume, page number, court, and year. The model has seen thousands of real citations in its training data. It knows the pattern perfectly.

So when asked to support a legal argument, the model generates citations that follow every formatting convention perfectly. The case name sounds real. The reporter volume is plausible. The page number is within range. The court is correct for the jurisdiction.

The only problem is that the case does not exist. The model predicted what a supporting citation should look like. It did not check whether that citation refers to an actual decision.

Documented sanctions.

The most well-known case is Mata v. Avianca (2023), where a New York attorney submitted a brief containing six completely fabricated case citations generated by ChatGPT. The court sanctioned the attorney and his firm. But this was only the first public case. Since then, courts across multiple jurisdictions have encountered and sanctioned attorneys for AI-generated fabricated citations.

What makes these cases notable is that the attorneys often did not know the citations were fabricated. They trusted the AI's output because it looked correct. Some asked the AI to confirm the citations existed, and the AI confirmed - hallucinating the confirmation as well.

Courts have responded by requiring attorneys to disclose AI use and verify all citations independently. This is manual verification - the slowest and least reliable detection method. It also undermines the productivity benefit that motivated adopting AI in the first place.

The attorney asked the AI if the citations were real. The AI said yes. When the AI's only reference is its own training data, it cannot distinguish between real and fabricated - both are just patterns.

Purpose-built legal AI still hallucinates.

The industry response was to build specialised legal AI tools with RAG - connecting the model to actual legal databases like Westlaw and LexisNexis. These tools retrieve real cases and feed them to the model alongside the query.

This helps. But the hallucination rate on purpose-built legal AI tools is still 17-34%. The model still misrepresents holdings, combines facts from different cases, and generates analysis that sounds like it follows from the cited case but does not.

The reason: RAG improves the input but does not verify the output. The model reads the retrieved case. It then generates its analysis probabilistically. Nothing checks whether the analysis faithfully represents what the case actually held.

The professional responsibility problem.

Lawyers have an ethical obligation to ensure accuracy. Rule 3.3 of the Model Rules of Professional Conduct requires candour toward the tribunal - an attorney cannot cite authority they know to be false. The complication: when the AI fabricates a citation, the attorney may not know it is false because it looks exactly like a real citation.

Ignorance is not a defence. Courts have held that attorneys are responsible for verifying every citation, regardless of how it was generated. Using AI does not delegate the duty of competence.

This creates an impossible choice for firms: use AI and spend time manually verifying everything (reducing the productivity benefit), or do not use AI and fall behind competitors who do. The resolution is not better prompting or better models. It is a verification layer that checks every citation against the actual legal database before the brief is finalised.

What legal AI actually needs.

Citation verification against live databases. Before any AI-generated brief reaches an attorney, every citation should be checked against an authoritative legal database. Does this case exist? Does the holding match what the AI claims? Is the case still good law?

Quotation verification. When the AI quotes from a case, the quote should be checked against the actual text. Models frequently generate plausible paraphrases and present them as direct quotes.

Flagging of unverifiable claims. When the AI generates legal analysis that cannot be traced to a specific source, it should be flagged as unverified rather than presented as established law.

This is the same verification architecture that applies to every domain. The AI needs a connection to reality - in this case, live legal databases - and a validation step that checks its output against that reality before delivery.

Verify before you file.

120 verifications a day free. No card, no signup.