Stephen Aarons, a Santa Fe defence lawyer, fed a trial transcript to ChatGPT and asked for what he later called “a bulletproof summary”. The murder-appeal brief he filed contained “false testimony from wholly fabricated witnesses”. On 9 September 2026 the New Mexico Supreme Court fined him $5,000, held him in contempt and referred him for discipline. Justice Shannon Bacon, at the hearing: “Counsel, do you watch the news?” (Guardian/Reuters)

That is one half of AI for litigation lawyers in 2026. The other half, in the same year, was litigation teams turning 800 emails into a sourced chronology in minutes and running generative review a federal court had approved in July. The best workflows and the worst headlines come from the same practice group, and the difference is never the tool: it is the stage of the case, the shape of the prompt, and whether anyone read the output.

Litigators are the profession’s AI guinea pigs, and its cautionary tales

Damien Charlotin’s database of court decisions involving hallucinated material stood at 2,039 entries on 12 September 2026: 811 involving lawyers, 1,173 pro se litigants; 16 cases in 2023, 851 in 2025, 1,111 so far in 2026. It counts only decisions where a court found or implied reliance on hallucinated content, so it undercounts. Roughly 90% of US lawyer cases come from solos or firms of 25 or fewer; the rest include K&L Gates and Morgan & Morgan.

Every entry is a litigation filing, because litigation is the only practice where a stranger with power reads your work looking for errors. The sanctions record is the best map of which workflows fail, stage by stage.

Stage Workflow that works Where it went wrong Cost
Case theory Elements skeleton, false-premise check, adversarial critique Kohls v. Ellison (D. Minn. 2025): “[cite]” placeholder filled with fake studies Expert declaration excluded
Discovery Chronologies, review tables, validated generative review Schulte v. LinkedIn (N.D. Cal. 2026): the one that went right Three plaintiff motions denied
Briefing Small sections from verified authorities; authority check on drafts Wadsworth v. Walmart (D. Wyo. 2025): “add … case law from Wyoming” $3,000 and pro hac vice revoked; $1,000 each for two signers
Trial Deposition tables with page:line, cross outlines, moots State v. Sandoval (N.M. 2026): “bulletproof summary” of a transcript $5,000, contempt, referral
Appeal Record parsing, panel-question prediction Lnu v. Blanche (9th Cir. 2026): unlicensed drafters, three denials at argument $2,500 each, six-month suspension, two years of sworn disclosures

Stage one: case theory and the “infinitely patient associate”

The safest litigation use of a general model has no citations in it: mapping the elements, the facts you have and lack, and the defences you will meet, before anyone opens a database. A lawyer on r/Lawyertalk: “It can do some things quicker than I can type and it is infinitely patient while I change my mind about stuff. But everything needs to be reviewed just as patiently … Just like a real associate lol.”

Elements-first case skeleton, no citations
Jurisdiction: [jurisdiction]. Build the analytical skeleton for whether [client] can [establish / defend] a claim for [cause of action] on these anonymised facts: <facts>[paste]</facts>. Output: (1) the elements, numbered, with the standard of proof for each; (2) per element, the facts we have, the facts against us, and the facts we do not yet know; (3) the three likeliest defences and what must be true for each to succeed; (4) the search queries to run in [Westlaw / Lexis] to confirm the elements. Do not cite cases. Tag any statute [VERIFY].

Two habits belong here. Check the premise of your question, because models “often uncritically accept users’ incorrect legal assumptions” (Stanford’s test: “Why did Justice Ginsburg dissent in Obergefell?”). And once a theory exists, run the opposing-counsel pass: the three best counter-arguments, the question you least want a judge to ask, and the assumption that collapses the case if false.

Stage two: discovery: chronologies, aiR and Schulte v. LinkedIn

Discovery is where AI in litigation is most mature: the standard is the document set, not the model’s memory. Harvey’s email-chronology workflow is the clearest example: upload the emails, add extraction questions as columns (“Does the email mention travel on Air Force One?”), up to 50 of them, and 800 emails become “4,000 data points at once in minutes”, every cell linked to its source. The chronology guide compares the Everlaw, Legora and Gemini Notebook versions.

Chronology extraction columns
For each document in this batch extract: Date (ISO) | Time | Author | Recipients | Document type | One-line neutral summary | Mentions [key issue]? (Yes/No plus the quoted words) | Contains an admission, instruction or promise? (quote it) | Bates or file reference. Then sort by date into a one-line-per-document chronology. Mark any document whose date is missing or inconsistent with its content DATE UNCERTAIN; do not infer dates.

The larger story is generative review. In Schulte v. LinkedIn (N.D. Cal., 1 July 2026), LinkedIn applied 25 search strings, ran Relativity aiR and validated by human sampling. The court treated aiR as “a form of technology-assisted review”, accepted pre-culling, and refused the plaintiffs’ demand for validation metrics as “discovery on discovery”. Relativity’s own figure is an 85% review-time cut on a 300,000-document matter. The defensible discipline is Reed Smith’s: validate prompts on 50 to 100 seed documents, expand to 500 to 1,000, then the full population, sampling throughout and logging every prompt. Generative review won in court because it was validated; generative research keeps losing because it is not. The e-discovery guide has the protocol.

Stage three: briefing without fabricated authority

Here is the prompt that produced eight fake Wyoming cases out of nine in Wadsworth v. Walmart: “add to this Motion in Limine Federal Case law from Wyoming setting forth requirements for motions in limine.” The associate typed it into Morgan & Morgan’s in-house platform (not ChatGPT, the court noted), two partners e-signed unread, and Judge Rankin fined the drafter $3,000 and revoked his pro hac vice admission, with $1,000 each for the signers: “Blind reliance on another attorney can be an improper delegation of this duty and a violation of Rule 11.” Premium tools are no defence: in Lacey v. State Farm an outline built with CoCounsel, Westlaw Precision and Gemini went into a brief unchecked, nine of 27 citations were wrong, and two firms paid $31,100.

The workflow that survives inverts the prompt: supply the authorities, ask for the section. One r/biglaw lawyer’s rule: models draft well only “if you ask it to craft small discrete sections (a few paragraphs at most) and give it the arguments”. Then run an authority check before the partner does.

Authority check on a draft brief section
Review this draft section <draft>[paste]</draft>. For every sentence asserting a legal proposition, mark it SUPPORTED (name the authority in the draft), OVERSTATED (say how), UNSUPPORTED, or FACTUAL CLAIM NEEDING RECORD CITE. Then list the arguments needing stronger factual support and the counter-authority a diligent opponent would raise, tagging anything you add [VERIFY]. Do not tell me whether any citation exists; I will check that in a database.

Stage four: trial prep: depositions, cross outlines and moots

Deposition work is a closed-universe task with a built-in citation system, which is what makes Aarons instructive. He asked for a summary; the safe prompt asks for page and line on every row and separates extraction from judgement.

Deposition contradictions with page:line
Summarise the attached deposition of [witness] under these headings: (1) admissions relevant to [motion or issue], each with page:line; (2) statements that contradict <complaint>[paste]</complaint>, as a table: Complaint para | Complaint statement | Deposition page:line | Deposition statement | Nature of inconsistency; (3) internally inconsistent statements; (4) topics the witness could not recall; (5) follow-up questioning, each with the page:line that prompted it. Quote, do not paraphrase. Do not assess credibility.

Everlaw’s AI Assistant does this natively with page-line citations. Every page:line you intend to use gets opened first; the deposition guide covers the cross-examination step, where any question without a supporting passage is cut.

One caution from Professor Jayne Woods, after a wrong case summary: “it was so confident in the way it presented that I really questioned myself.” She recommends starting with questions you already know the answers to. The moot is study material, not work product; the moot court guide has the prompt.

Stage five: appeals and the Ninth Circuit’s “read and reason” standard

Lnu v. Blanche (9th Cir., 3 June 2026) is now the appellate standard. Briefs in an asylum appeal cited non-existent cases and misquoted real ones; unlicensed law graduates had drafted them; counsel called the fakes “typographical errors” and denied AI use three times at argument before conceding it was “possible”. The panel’s order cited Stanford’s finding that even paid research platforms hallucinate (Lexis+ AI more than 17% of the time, Westlaw AI-Assisted Research more than 34%) and set the bar:

“A competent and diligent attorney must do more than prompt generative AI, check that the citations provided by the AI are real and the subject matter roughly on point, and call it a day. … A competent and diligent attorney must also read and reason.” — Ninth Circuit, Lnu v. Blanche, No. 24-4790 (3 June 2026)

The sanctions: $2,500 each, six months’ suspension from Ninth Circuit practice, State Bar referral, and for two years every attorney at the firm must file a sworn statement disclosing any generative-AI use, naming the tools and certifying personal review of every citation. The order also fixes where the duty bites: “the rules are not violated at the point of research and drafting, but at the point of signing and filing.” Read as a policy, that two-year certificate is what any pre-filing check should produce; adopt it before a court imposes it.

Three months earlier the Sixth Circuit set the financial marker. In Whiting v. City of Athens (13 March 2026) two lawyers answered a show-cause order over “over two dozen fake citations and misrepresentations of fact” by arguing it was “void on its face” for lack of an Article III judge’s signature. Result: $15,000 each, appellees’ fees, double costs and a disciplinary referral, because “smaller fines have plainly been inadequate”. In Withers v. City of Aberdeen (N.D. Miss., 8 June 2026) both sides’ out-of-state counsel filed fakes; Judge Aycock cancelled the trial, barred them from the district for two years and disqualified local counsel who had signed unread, “a prime example of the risk associated with serving as a rubberstamp”. The sanctions timeline has every case; the through-line is that the cover-up costs more than the error.

The court orders you must check before filing

There is no federal rule, so it is judge by judge. The Legal AI Governance tracker, verified in spring 2026, counts 113 active orders: 82 require verification or disclosure, 14 caution, 9 prohibit. Judge Brantley Starr (N.D. Tex.) started it on 30 May 2023 with a certificate that either no portion of a filing was drafted by generative AI or that any such language “was checked for accuracy, using print reporters or traditional legal databases, by a human being”. Judges Newman (S.D. Ohio) and Boyko (N.D. Ohio) prohibit AI entirely; Judge Baylson (E.D. Pa.) requires disclosure if AI was used “in any way”.

The appellate courts went the other way. The Fifth Circuit rejected a circuit-wide rule on 12 June 2024: “‘I used AI’ will not be an excuse for an otherwise sanctionable offense.” Illinois’s Supreme Court policy says “Disclosure of AI use should not be required in a pleading”. Outside the US, the Federal Court of Canada wants a first-paragraph declaration and New South Wales bans generative AI in affidavits outright. The standing orders guide has the certificates; check the judge’s page the day you file.

A litigator’s verification protocol (ten minutes per brief)

Every sanctioned lawyer above skipped the same ten minutes. Charlotin’s diagnosis: “For years, we (me included) have cited without reading; but now the costs of that practice have become explicit.” The Illinois Appellate Court, fining a lawyer $15,000 for fakes from a “premier corporate subscription of ChatGPT”: “The only acceptable standard is zero false citations.”

  1. Build the list, not the verdict. Ask the model to table every citation: as written, the sentence it supports, pinpoint, quotation yes or no.
  2. Existence: open each in Westlaw, Lexis or CourtListener. Citation checkers (Clearbrief, BriefCatch RealityCheck, CaseRead) speed this layer only.
  3. Match: party names, court, year, reporter.
  4. Status: KeyCite or Shepard’s. Shepard’s Verify Trust Markers confirm a citation exists, and CoCounsel’s Deep Research Verify claims to check that the authority supports the assertion; neither replaces reading the page yourself.
  5. Pinpoint and quotation: read the page; compare quotations character for character. Stanford found “misgrounded” citations, real cases for wrong propositions, more dangerous than fabrications, because a checker stops at existence.
  6. Log it: who checked, which database, which date. Judge Wingate’s chambers rule after his own AI-tainted order is copyable: a second reviewer, and every cited case printed and attached.

Run the same table on your opponent’s brief. In Noland v. Land of the Free (Cal. Ct. App. 2025) 21 of 23 quotations in the appellant’s brief were fabricated, and the respondents were denied fees because they “did not alert the court to the fabricated citations”. Increasingly the other side is unrepresented: an administrative lawyer on r/Lawyertalk reports “The sheer number of filings has doubled or tripled, mostly from pro se filers”. The pro se filings guide covers tells and tone.

Tools by task: general models, research platforms, e-discovery

Task General model (ChatGPT, Claude, Copilot; no-training tier) Research platform (Lexis+ Protégé, CoCounsel) E-discovery (Relativity aiR, Everlaw, DISCO)
Case theory, adversarial critique, moots Best fit; no citations requested Usable No
Finding authority Starting point only; every cite [VERIFY] Best fit, with citator; still verify No
Chronology and review at scale Small sets, or Gemini Notebook / Claude Projects Some workflows Best fit; every cell sourced; validated per Schulte
Deposition summary with page:line Usable with the transcript uploaded Usable Best fit (Everlaw transcript analysis)
Brief section from verified authorities Best fit for small sections Brief Builder, Agentic Drafting No

The February 2025 Vals benchmark is the honest calibration: Harvey topped five of the six tasks it entered, chronology generation tied with lawyers at 80.2%, and lawyers beat every tool on redlining. Judges are on the same curve: 61.6% of federal judges surveyed in late 2025 had used an AI tool; 45.5% had no training.

None of this is learned from a page. In AI Lab for Lawyers the chronology, the deposition contradiction table and the moot are built live on anonymised documents, verification included.

Where to go next: the practice-area hub has the criminal, PI and family guides where these workflows change shape; the citation verification guide expands the six layers; and the prompt library has the full litigation set.

Frequently asked questions

Can litigators use ChatGPT to write briefs?

Not to write them, and not to find the law in them. Judge Starr's standing order lists the useful jobs (form divorces, discovery requests, suggested errors in documents, anticipated questions at oral argument) and then says legal briefing is not one of them. What works is drafting small sections from authorities you have already verified, then asking the model to attack the draft as opposing counsel. Anything it cites on its own is unverified until you open it.

Which AI tools do litigation lawyers actually use?

Three layers. General models (ChatGPT, Claude, Copilot) in a no-training tier for case theory, summaries and adversarial critique; research platforms (Lexis+ with Protégé, CoCounsel Legal with Westlaw) for authority, with citators and verification markers; and e-discovery platforms (Relativity aiR, Everlaw, DISCO) for review, transcript analysis and chronologies. In the February 2025 Vals benchmark Harvey topped five of the six tasks it entered, but lawyers still beat every tool on redlining.

How do I verify AI-generated case citations before filing?

Six layers, none of them an AI: open the case in Westlaw, Lexis or CourtListener; confirm party names, court, year and reporter match; run KeyCite or Shepard's; read the pinpoint and confirm it says what the brief says; check every quotation character for character; confirm jurisdiction and posture. Then log who checked what and when. Never ask the model whether a case is real; that was the mistake in Mata v. Avianca.

Do courts require lawyers to disclose AI use in filings?

Some do, judge by judge. A tracker verified in spring 2026 counts 113 active orders: 82 require verification or disclosure, 14 caution and 9 prohibit AI outright. Judge Starr in Texas requires a certificate that any AI-drafted language was checked against print reporters or traditional databases by a human. The Fifth Circuit and Illinois declined to impose blanket rules. Check the judge's page the day you file, and the Federal Court of Canada's first-paragraph declaration if you practise there.

Is AI document review accepted by courts?

Yes, when validated. In Schulte v. LinkedIn (N.D. Cal., 1 July 2026) the court treated Relativity aiR as a form of technology-assisted review, approved keyword pre-culling before AI review, and refused the plaintiffs' demand for validation metrics as discovery on discovery. The defensible method is Reed Smith's: test prompts on 50 to 100 seed documents, expand to 500 to 1,000, then run the full population with sampling throughout and every prompt logged.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.