- The public database of court decisions involving AI hallucinated material listed 1,868 cases on 11 August 2026, with the United States accounting for 1,297 of them.
- Self represented litigants appear more often than lawyers in that database, 1,096 against 724. The typical bad filing does not come from a law firm.
- Checking that a case exists is not enough. The database records 782 entries involving a misrepresented source and 510 involving a false quotation, alongside 1,556 fabricated ones.
- Contract disputes are the most common legal field in the data, at 435 cases, ahead of administrative, civil rights and employment matters.
- Under the EU AI Act it is the administration of justice that is high risk, not your contract review tool. Getting that distinction right saves a lot of misplaced worry.
Sort AI legal work by task and by who carries the consequence, and it stops being one question and becomes six. Some of those tasks are safe to hand over with a light review. One of them has already cost lawyers money, cases and public admonishment in identifiable courts this year. The difference is not how clever the model is. It is whether a wrong answer is caught by the next person who reads it, or filed.
What follows sorts the work that way. It is not a tool review, there is no vendor comparison in it, and every figure carries the date it was read.
How bad is the citation problem, really?
Bigger than the anecdotes suggest and differently shaped. The AI Hallucination Cases database maintained by Damien Charlotin collects court and tribunal decisions in which a judge explicitly found, or implied, that a party relied on hallucinated material. Read on 11 August 2026, it listed 1,868 cases, with the site noting it was last updated on 8 August.
The jurisdictional split, taken from the same page on the same day: 1,297 in the United States, 206 in Canada, 97 in Australia, 61 in the United Kingdom, 55 in Israel, with dozens of other countries in single or low double digits. It is not a US only phenomenon, but it is heavily a US one, which reflects both filing volume and how publicly American courts write about it.
The finding almost nobody reports
Break the same database down by who submitted the material and the picture inverts. Pro se litigants, meaning people representing themselves, account for 1,096 entries. Lawyers account for 724. Judges appear 27 times and court appointed experts 13.
That ratio changes who this article is for. The profession has spent two years treating hallucinated citations as a discipline problem inside law firms. In the recorded decisions it is more often something that happens to a person with no lawyer, filing their own contract or employment dispute, using a general purpose chatbot as a substitute for representation. If you run a small business and have ever considered handling a claim yourself, you are in the larger group.
What kinds of error actually occur?
Four, and they need different checks. The database tags each case by the nature of the problem, and the counts as read on 11 August 2026 were: fabricated material 1,556, misrepresented sources 782, false quotations 510, and outdated advice 33. Cases can carry more than one tag, so these do not sum to the total.
The operational consequence is the one people miss. If your verification step is does this case exist, you catch the 1,556 and you miss most of the 782 and the 510, because in those the case is real, the reporter citation is right, and what the filing claims it says is invented. Norton Rose Fulbright's roundup of 2026 sanctions describes a Fifth Circuit matter in February 2026 involving sixteen fabricated quotations and five misrepresentations, with a 2,500 dollar sanction and the court observing that fuller acceptance of responsibility would probably have reduced it.
Break it down further and the subcategories matter too. Case law dominates at 1,679 entries, but statutory and regulatory material appears 213 times, exhibits and submissions 132 times, doctrinal writing 41 times. There are 17 entries for citing overturned case law and 16 for repealed law, which is the quietest failure of all: everything checks out except that it is no longer good law.
Which legal tasks are safe to delegate?
The ones where a wrong output is obvious to the person reviewing it, and unsafe wherever a wrong output is plausible and unchecked. Here is the sort, with the verification step that cannot be skipped for each.
| Task | Suitability | Who bears the error | The check that cannot be skipped |
|---|---|---|---|
| First pass document review and privilege screening | Good. Volume work with a human second pass built in | The firm, contained | Sample the negatives, not the positives. The risk is what it filed away |
| Contract redlining against a playbook | Good, where the playbook is written down | The client, at signature | Confirm every deviation flagged is a real deviation, and read the clauses it did not flag |
| Summarising a deposition or a discovery set | Reasonable, with the source alongside | The firm | Spot check every summarised fact against the transcript line it claims to come from |
| Drafting from a precedent you supply | Good. The model is editing, not inventing | The firm | Diff it against the precedent. Read what changed, not what stayed |
| Legal research and citation | Highest risk. This is the category that produces sanctions | The lawyer personally, plus the client | Open every authority in a primary source and confirm it says what the draft claims |
| Predicting an outcome | Not suitable as a basis for advice | The client, invisibly | There is no verification step. Treat output as a prompt for your own analysis |
Two rows deserve a note. Privilege screening looks like the safest thing on the list and carries a specific trap: a review tool's false negatives are silent. Nobody reviews the pile that was set aside as not responsive, so an error there surfaces only if the other side finds it. Sample that pile deliberately.
Outcome prediction is the row people argue with. The objection is that experienced lawyers predict outcomes constantly, so why not the model. The answer is that a lawyer's prediction carries a reasoning chain the client can interrogate and a professional who is accountable for it. A model's prediction carries a confident sentence. The output can be useful as a starting point for your own analysis, and it is not something to relay to a client as an assessment.
What does the verification duty actually require?
Reading the authority itself, in a primary source, every time, before it goes anywhere a court or a counterparty will see it. Courts in 2026 have not treated the technology as a novel excuse. The Norton Rose summary quotes the Fifth Circuit making the point directly, that generative AI may be new but the same sanctions rules apply and the existing rules are well equipped to handle these cases.
The 2026 outcomes recorded there give a sense of range. A Fourth Circuit matter in March drew public admonishment for three nonexistent citations. A Sixth Circuit matter the same month involved more than two dozen fake citations and record misrepresentations, drawing attorneys' fees, double costs, punitive sanctions of 15,000 dollars each and a disciplinary referral. In another Sixth Circuit case, counsel was removed from the matter and denied all Criminal Justice Act compensation. In April, an Eastern District of California order was discharged with no sanction after the attorney gave a candid explanation.
That last one is worth dwelling on, because it appears twice in the record. Candour changes the outcome. The pattern across the reported decisions is that the sanction tracks the response to being caught at least as much as the original error.
A verification rule that survives contact with a deadline: paste nothing you have not opened. Not the citation, not the quotation, not the parenthetical. Opening it in a primary source takes about ninety seconds per authority, and it is the only step in this article that has been shown to have consequences for skipping.
The confidentiality decision, made once
Confidentiality is a design question rather than a per matter judgement call, and treating it as the latter is how firms end up with an ad hoc policy nobody can state. There are three inputs and they are answerable in an afternoon.
Where does the text go? A model running on your own hardware, a business tier hosted service with contractual terms about training, and a consumer chatbot are three different answers, not three flavours of one. The question is not whether a vendor is trustworthy. It is what its contract obliges it to do and what its default retention is.
Does it train on your input? Ask in writing. Consumer tiers and business tiers frequently differ on exactly this, and the answer is a contractual fact rather than a technical one.
What is the client owed? Separate from both of the above. Whether a client should be told that their matter was processed by a third party service is a question about the retainer and the relationship, and it does not become moot because the vendor promised not to train.
The practical structure that follows is a two lane system. Non identifying work such as research questions, general drafting and jurisdiction agnostic clause language goes down the convenient lane. Anything containing client facts goes down the constrained lane, or gets redacted before it travels. The redaction step is where discipline actually breaks, since a matter summary with the names removed is often still identifying. We covered the general shape of that problem for any small organisation in our piece on access controls for a business without an IT department, and the legal version differs mainly in the size of the consequence.
The formal risk vocabulary is worth borrowing even if you never adopt the whole framework. The NIST AI Risk Management Framework, published on 26 January 2023 and explicitly intended for voluntary use, organises the work into govern, map, measure and manage. For a small firm the useful part is the first two: write down what you are using and for which task, before arguing about metrics.
Is a legal AI tool high risk under the EU AI Act?
Usually not, and the confusion is costing people time. The Act's high risk tier names AI solutions used in the administration of justice and democratic processes, giving the example of solutions to prepare court rulings, along with law enforcement uses that may interfere with fundamental rights such as evaluating the reliability of evidence.
Read that carefully. The high risk category is aimed at the machinery of justice, courts and enforcement, not at a private firm's contract review tool. A commercial tool used to redline a supply agreement sits in the minimal risk category alongside most software, with transparency obligations attaching to the underlying general purpose model rather than to you.
The dates matter more than the tier for most readers. Prohibited practices applied from 2 February 2025, general purpose AI governance rules from 2 August 2025, transparency rules during August 2026, high risk systems in sensitive areas from 2 December 2027. We laid the tiers and the timeline out in full in our breakdown of the EU AI Act tiers and dates, and the short version for a law firm is that your compliance surface is mostly professional conduct rules rather than the Act.
What should a small business owner take from this?
That the group most exposed in the recorded data is not lawyers. With self represented litigants at 1,096 entries against 724 for lawyers, and contract disputes the single largest legal field at 435 cases, the modal recorded incident looks a lot like a business owner handling their own contract dispute with help from a chatbot.
Three things follow. If you are drafting your own commercial documents with AI assistance, the drafting is the safe part and the citation of authority is not; a contract does not need case law in it. If you are in an actual dispute, anything you file that cites a case has to be opened and read first, and the ninety seconds is not optional. And if a model tells you a statute says something helpful, that is precisely the moment to check, because the 213 entries tagged to legal norms are people who did not.
There is a related habit worth carrying across from marketing into legal work: assume anything generated will be checked by someone with an interest in finding the flaw. That is the same reasoning we applied to public facing text in our piece on what AI content detection actually detects. A court is simply the most rigorous version of that reader.
Where does the time actually go?
The uncomfortable part of the arithmetic is that the verification step eats most of the saving on exactly the task where the tool feels most impressive. A research memo that takes an afternoon to write by hand can be drafted in minutes and then needs every authority in it opened and read, which is unavoidable and is not fast. What you save is the search, the structuring and the first draft. What you do not save is the reading.
That changes which tasks are worth automating first, and it is close to the inverse of what most firms try. Document review scales because the verification is sampling rather than exhaustive. Redlining against a written playbook scales because the check is bounded by the number of clauses. Research does not scale in the same way, because the check grows with every authority the draft adds, and a longer draft is a bigger bill of verification rather than a bigger saving.
There is a practical corollary for billing that firms are still working through. If a matter takes two hours instead of six, the client's reasonable expectation is that it is billed as two, and the professional conduct question of what you may charge for time you did not spend is separate from any question about the technology. It is worth settling before it arrives on an invoice rather than after.
What about the model being manipulated rather than mistaken?
It is a smaller problem than hallucination today and a structurally nastier one. If a tool summarises documents supplied by a counterparty, those documents are untrusted input, and text inside them can be written to influence what the summary says. That is the same failure mode as prompt injection in any other setting, and we went through how reliably it works in the results of a large scale prompt injection test.
Nothing in the sanctions record turns on this yet. The reason to hold it in mind is that document review and discovery summarisation, the two tasks rated most favourably in the table above, are exactly the ones that ingest material an opponent wrote. The mitigation is unremarkable: treat a summary of an adversary's document as a reading aid rather than as evidence of what the document contains.
The short version
Delegate the tasks where an error is visible to the next reader. Keep research and citation on the shortest leash, because that is the category with a documented and growing record of consequences. Verify by opening the authority rather than by asking whether it exists, since misrepresentation and false quotation together account for a large share of the recorded cases. Decide the confidentiality question once, in writing, rather than per matter.
And note who is actually getting caught. The database says self represented litigants, in contract disputes, more often than lawyers. That is a fact about access to legal help as much as about AI, and it is the part of this story most likely to reach somebody who has never read a bar association opinion in their life.