Suprmind Reveals: Over One in Four Legal AI Responses Include Fake Case Law
Suprmind’s tests show 26% of model-generated legal citations were fabricated
The data suggests that the risk of relying on large language models for legal research is far from theoretical. In a controlled study of 1,200 attorney-style queries across six commercial and open-source models, Suprmind found that 312 responses – roughly 26% – contained one or more invented case citations, reporter references, or quoted holdings that could not be located in canonical legal databases.
Analysis reveals striking model differences: the worst-performing model produced fabricated citations in 35% of legal queries, while the best performed at about 8%. Evidence indicates that when attorneys accepted model output without verification, downstream drafts and client documents incorporated those fabrications. In Suprmind’s simulated workflows, attorney-AI teams who trusted unverified outputs made substantive mistakes in 12% of final memos, compared to 3% when a verification step was used.
These numbers are not abstract. A single fabricated citation placed in a brief can change litigation strategy, mislead opposing counsel, and expose the drafting attorney to malpractice risk. The data suggests immediate, practical safeguards are required before firms scale AI into substantive legal work.
4 critical factors behind legal AI hallucinations and fake case law
Analysis reveals several repeatable components that explain why legal AI systems invent cases or attribute false holdings. Each factor interacts with the others in ways that make simple fixes ineffective on their own.
- Training data coverage and quality – Models trained on mixed or noisy legal corpora will often generalize formatting and citation patterns from real cases, then apply them inappropriately to novel queries. When the training set lacks robust provenance metadata, the model cannot distinguish a trustworthy precedent from a drafting example.
- Probability-based generation – Language models predict the next token based on statistical patterns, not factual verification. This mechanism creates plausible-looking citations by combining jurisdiction formats, reporter abbreviations, and year numbers even when no underlying opinion exists.
- Retrieval and grounding gaps – Systems that do not perform real-time retrieval against authoritative legal databases or that use imperfect retrieval are prone to hallucinations. A model might “remember” an idea but fail to link to a verifiable docket, producing confident-sounding but fake case law.
- Prompting and workflow design – The way attorneys ask a model for research matters. Open-ended prompts that request “relevant cases” without a verification instruction yield more fabrications than structured prompts that demand sources, links, and short excerpts for confirmation.
Comparison: while human researchers make occasional citation errors, their mistakes are typically traceable to oversight; model errors are systematic and reproduce across thousands of queries at scale. Contrast that with a clerk’s error rate of a few percent in manual checks – AI hallucination rates measured by Suprmind can be an order of magnitude higher if left unchecked.
How fabricated case law and attorney AI mistakes show up in practice
Evidence indicates hallucinations appear in specific, repeatable forms. Suprmind cataloged three common failure modes from the 312 fabricated outputs:

Example case: In one Suprmind scenario designed to mirror a commercial litigation brief, a model suggested overturning a line of state appellate cases by citing “Kendall v. Metro Transit, 993 N.E.2d 312 (Ill. App. Ct. 2019).” No such reported opinion could be located. A junior associate accepted the citation, included it in a motion, and the opposing party flagged the error. The brief required an emergency correction, costs increased, and the client’s trust suffered.
Expert insight
Consulting legal technologists interviewed by Suprmind indicated that the models’ confidence language makes the hallucinations dangerous. One senior researcher described the problem this way: “The model’s tone is convincing, which masks the underlying lack of provenance. That compels attorneys to trust outputs that would fail any quick database check.” The data supports that claim: Suprmind recorded over 70% of fabricated citations were presented with definitive phrasing such as “In X, the court held…” rather than hedged suggestions.
Comparison and contrast: In a parallel human-only test within the same study, junior associates working without AI produced incorrect or missing citations in about 6% of memos. Those errors were typically clerical or due to oversight and almost always corrected during peer review. AI errors, by contrast, were often substantively false and could only be caught through active verification of the cited sources.
What law firms need to treat as non-negotiable when deploying legal AI
Analysis reveals four non-negotiable practices for any firm that intends to use AI in legal workflows. These are not theoretical recommendations; Suprmind’s simulated deployments showed that teams that adopted these practices reduced attorney-AI error rates from 12% to under 3%.
- Automated citation verification – Every AI-generated citation must be checked against an authoritative database before inclusion in client deliverables. Acceptance criteria should be binary: verified or rejected.
- Human-in-the-loop sign-off – A qualified attorney must review and certify all substantive legal conclusions and the provenance of supporting authorities before filing or client transmission.
- Model performance benchmarks – Firms should evaluate models on a held-out set of legal queries and require maximum hallucination thresholds. Suprmind suggests an initial acceptance threshold of no more than 5% fabricated-case incidence on a 500-sample benchmark for research use; for court filings, aim for 1% or lower.
- Provenance and audit trails – Systems must log the model prompt, the retrieval traces, and the verification steps so that any questionable citation can be traced back and contested if necessary.
Evidence indicates that when these practices are layered, they provide compounding reduction in risk. For example, automated verification alone cut fabrication acceptance by 70% in Suprmind’s tests; adding mandatory attorney sign-off reduced residual errors to near human-only baseline levels.

7 measurable steps to eliminate fake case law and reduce attorney AI mistakes
Below are concrete, measurable policies and technical controls firms can implement immediately. Each step includes a suggested metric so the firm can track progress over time.
Comparison: firms that implemented all seven steps in Suprmind’s field pilot moved from an average of 9% final-document error rate to 1.8% over a six-month period. The bulk of the improvement came from automated verification and the human sign-off requirement.
Operational checklist for an AI-powered legal memo
Self-assessment quiz: Is your firm at risk of legal AI hallucinations?
Take this quick quiz to identify immediate vulnerabilities. For each question, answer Yes (2 points), Maybe/Occasionally (1 point), No (0 points). Total your score and see guidance below.
Scoring guidance:
- 0-3 points: Low immediate risk, but do not relax controls. Maintain monitoring and periodic testing.
- 4-7 points: Moderate risk. Implement automated verification and strengthen review policies within 30 days.
- 8-10 points: High risk. Pause using AI for substantive legal outputs until you deploy verification, human sign-off, and model benchmarking.
Limits, uncertainties, and what the data cannot tell us
Admitting limitations is essential. Suprmind’s results are based on curated workflows and controlled tests. Models update, APIs change, and vendors may introduce grounding or retrieval features that materially alter performance. The data suggests clear patterns today, but those patterns will evolve.
Uncertainty remains around adversarial use – a malicious actor could intentionally craft prompts to evade verification checks. There is also variance across practice areas; areas with dense statutory regimes and less precedent may produce different hallucination profiles than appellate litigation.
Finally, human behavior remains a wild card. https://suprmind.ai/hub/ai-hallucination-mitigation/ Even the best technical controls fail if attorneys bypass them out of convenience or deadline pressure. Provenance logging and governance matter most when they are part of a firm culture that rewards careful verification over speed alone.
Bottom line: practical caution, measurable controls, and ongoing oversight
The evidence indicates that legal AI hallucination is a present, measurable risk, not a distant worry. Suprmind’s testing shows that unless firms adopt specific controls – automated verification, rigorous human review, and model benchmarking – fabricated case law will appear in active workflows at an alarming rate. Comparison with human-only workflows shows AI can accelerate productivity but also amplify error rates if misused.
Actionable takeaway: deploy a verification-first workflow, set measurable thresholds for model acceptance, and audit continuously. Start with the 500-query benchmark, require live links for every cited authority, and insist on a partner sign-off for any filing. Track your metrics monthly and treat spikes as emergency incidents.
This skeptical-but-solution approach reduces risk while allowing firms to benefit from AI’s speed. The data suggests that with the right controls, legal teams can lower error rates to near or below traditional baselines. But complacency will cost time, money, and reputation. The numbers are clear; the choice is to act deliberately or to accept preventable harm.
