The people who see AI contract drafting fail are rarely the people who did the drafting. They are the litigators who get the file later. In July 2026 Conor Maher and Joshua Bentley of Ellis Jones, an English disputes firm, listed what they were pulling out of AI-drafted share purchase agreements: warranties drafted “too widely by sellers and too narrowly by buyers, leading to sellers accidentally warranting something is true when they can’t possibly know”; payment terms “so unclear that they risk being unenforceable”; and “references to foreign/outdated law (this risks a party contracting out of the jurisdiction of England and Wales)”.

Their cost arithmetic is the sharpest line in the piece. Saving “a few hundred pounds” on the drafting leads to “a few thousand (possibly a few hundred thousand)” in litigation.

None of this means lawyers should not draft with a model. It means the usual prompt (one instruction, whole document, no precedent) produces the documents Ellis Jones is now litigating.

The fallout is already in court: what litigators are seeing

Two English firms published post-mortems in July 2026. Ellis Jones covered SPAs and shareholders’ agreements: conflated share classes, and bad-leaver provisions drafted so poorly that “an underperformer can’t be forcibly removed”. Natasha Thomas of Awdry Law found AI shareholders’ agreements “too generic”, with miscalibrated reserved matters, leaver and deadlock provisions, so that “the agreement may store up disputes rather than prevent them”. Her verdict: “Boilerplate language may sound convincing, and often does so with remarkable confidence, but it can miss the points that matter most.”

A year earlier, Darwin Gray had reported an insolvency claim citing case law that “did not exist”. Three failure patterns run through all three accounts.

Failure pattern 1: warranties and payment terms that cannot work

A warranty is a promise about a fact. A model drafting seller warranties from a generic template has no idea which facts this seller can stand behind, so it produces warranties of things that, in Ellis Jones’s words, “cannot legally or factually actually be warranted”. Buyer-side drafts show the reverse: warranties narrower than the deal needs.

Payment mechanics fail the same way: deferred consideration and earn-outs are arithmetic wrapped in conditions, and the model writes the conditions in confident prose while leaving the arithmetic ambiguous. Ellis Jones singles out “deferred consideration … or earn-out provisions” as the usual casualties.

Failure pattern 2: US law leaking into English documents

Darwin Gray’s warning is direct: “AI tools are also often geared towards a US audience, and often therefore generate US legal concepts and termination into its work”, making documents unsuitable for England and Wales. Sullivan & Cromwell’s January 2026 M&A memo gives the mechanism: models are biased towards the publicly filed deals in their training data. It takes no great leap to guess where most publicly filed deals come from.

The fix is not a negative instruction: “draft under English law, use the defined terms in our precedent, and flag any concept you would normally draft under US law” beats “do not use US concepts”. German and Austrian practitioners face the same leak; the note on our German Gutachten prompt recommends adding “Keine Bezugnahme auf US-Recht” the moment you see it happen.

Failure pattern 3: inconsistent definitions across a document set

A shareholders’ agreement sits beside articles, a subscription agreement and often an option plan. Awdry Law finds AI drafts with “inconsistencies, vague definitions, and clauses that do not work properly with the rest of the document set”. Inside one document the same thing happens: Alon Kapen of Farrell Fritz describes a definition of “Developed IP” that, three layers deep, quietly narrows an assignment to first-year improvements, a chain the tool never traced.

Why whole-agreement prompts fail (40 to 80 interlocking pages)

Zevra, a French M&A prompt library, states it plainly: “A full SPA spans 40-80 pages with interconnected clauses … Language models cannot hold all variables across such length and produce generic, sometimes contradictory clauses.” Its method is clause isolation with precise parameters and pseudonymised deal data.

Two mechanical reasons. Accuracy degrades as the context fills; the guide to prompting long contracts has the evidence. And a whole-document prompt forces the model to invent every choice you did not specify, filling the gaps with the commonest pattern in its training data. That is how a UK deal acquires a Delaware indemnity.

The viral counter-example proves the point. On 19 March 2026 a non-lawyer, Nav Toor, posted “12 prompts that replace $15,000 in legal bills”, each opening with a persona such as “You are a senior corporate attorney at Skadden Arps” and asking for “a complete, ready-to-sign NDA”. Artificial Lawyer wondered whether “Claude simply read ‘Wachtell’ and thought ‘OK, this means do the contract in the style of any large commercial law firm’”, and concluded that “a prompt alone is not usually a replacement for the input of a good lawyer – at least for now.” The firm name changes the tone. It does not add the facts.

The building-block method: MAC, earn-out, locked box, indemnity, one at a time

Reverse intake before any first draft
I need a first draft of a [share purchase agreement] for a [UK private company acquisition, GBP 8m, buyer side] under English law. Do not draft anything yet. Ask me, in one message, the eight to twelve questions whose answers would most change the draft: structure (locked box or completion accounts), deferred consideration and any earn-out, warranty and indemnity cover, restrictive covenants, my client's deal-breakers, and which precedent I will supply. Group them by importance. After I answer, tell me which provisions you would draft first and why.

Answer the questions, then draft in blocks, each with the term sheet, the precedent for that provision, and explicit rules for silence and conflict:

Draft one provision from the term sheet and precedent
<term_sheet>…</term_sheet>
<precedent>…</precedent>
Draft ONLY the [purchase price and locked-box provisions] of a share purchase agreement under English law, following the precedent's defined terms, numbering and style exactly. Where the term sheet is silent on a point the precedent addresses, use the precedent's position and mark it [CONFIRM: term sheet silent]. Where the term sheet conflicts with the precedent, follow the term sheet and mark [TERM SHEET OVERRIDES]. Flag any concept you would normally draft under US law. List every defined term you used that does not yet exist in the precedent. Do not cite any case or statute.

Repeat for the earn-out, the MAC definition, the indemnity package and the leaver provisions. Where the negotiating range matters, ask for it. The Legal Prompts’ MAC prompt uses a toggle, “VERSION 1 - PRO-BUYER … VERSION 2 - PRO-SELLER: Narrow MAC definition limited to financial performance only”, and its SPA indemnification prompt labels every clause buyer-favourable, market-standard or “AGGRESSIVE - MAY FACE PUSHBACK”. The model’s sense of “market” is an inference from public deals, not data on your sector; say so in the prompt.

Finish with the sweep that catches the Awdry Law failures: every defined term and where it is defined; every cross-reference pointing nowhere; every threshold that appears inconsistently; every clause that conflicts with the articles. Then check item by item. Claude for Word says it recognises “multi-level numbering, defined terms, cross-references, and standard contract structures”; Spellbook advertises “automatically updating defined terms across multiple documents simultaneously”. Neither replaces the read.

Few-shot drafting from your own precedents

The strongest drafting prompts describe no persona. They show the model two or three clauses in your house style and ask for a new one in the same voice. The Singapore Academy of Law’s prompt guide, written with Microsoft, does it in one instruction: “Draft an IP indemnity clause in favour of the licensor. I am a lawyer acting for a licensor in a negotiation on a software IP licensing agreement. Ensure that the clauses are in formal legal language using the active voice and concise. Base your drafts on these samples: [sample 1] and [sample 2].” Label each sample with its deal context; samples from different deal types confuse the model.

Few-shot clause drafting from house precedents
Draft an IP indemnity in favour of the licensor for a software licence agreement under [English] law. Base it on these samples of our house style, each labelled with its context: <sample_1 context="SaaS, licensor side, 2025">…</sample_1> <sample_2 context="on-premise, licensor side, 2024">…</sample_2> <sample_3 context="OEM, licensor side, 2026">…</sample_3>. Match their structure, defined-term conventions and sentence length. Then, in a separate note, list the three ways the new clause differs in substance from the samples and why each is needed for a [SaaS] licence. Flag [NEW] any concept that appears in none of the samples.

For teams drafting the same clauses weekly, the samples belong in an anonymised clause bank inside a Claude Project or custom GPT. That set-up, and the draft-attack-rewrite loop for testing a clause against a hostile reader, are in ChatGPT prompts for lawyers and adversarial prompting for legal drafts.

The three-model drafting test

On 16 June 2026 Davide DeMango published a head-to-head drafting test: one prompt for a freelancer-favourable web design services agreement, run through GPT-4o, Gemini 1.5 Pro and Claude Sonnet 3.5. The models are 2024-era; read it as a test of prompting habits, not a current ranking.

Model tested Indemnity produced Third-party IP carve-out Other finding
GPT-4o “broad and bilateral”, exposing the freelancer No
Gemini 1.5 Pro Vague No No liability cap
Claude Sonnet 3.5 “a mutual indemnification structure with a carve-out for third-party IP infringement claims” Yes

Only one model drafted the clause a competent lawyer would have drafted, from a prompt that did not ask for the carve-out. Whichever model you use in 2026 (see the ChatGPT versus Claude comparison), specify the cap, the carve-outs, the side and the governing law. A model that knows your clause is a convenience; a prompt that specifies it is a control.

Board minutes and resolutions: the “do not” list

Drafting resolutions and written consents from an approved precedent is a safe, boring, useful task; Mayer Brown lists it among seven practical M&A uses, with the caveat that it needs “the right precedents and understanding of the relevant governance structure”. Transcribing board deliberations is a different matter.

The rules: decide whether to record at all; keep consumer-tier tools away from deliberations (the privilege guide explains why Heppner matters); confine AI to resolutions, consents and registers drafted from precedent, with the articles supplied and every quorum or majority threshold marked [CONFIRM]; and never let it summarise the discussion itself. The in-house guide has a governance intake for that boardroom conversation.

An AI contract drafting checklist before you send anything

Check Why How
Side and governing law in the prompt Generic drafts default to US patterns First line of every prompt; ask the model to flag foreign concepts
Every warranty tested against what the client can know Sellers “accidentally warranting” unknowable facts Read each warranty as the seller; add knowledge and disclosure qualifiers
Defined terms and cross-references swept Definitions drift across a document set Sweep prompt, then a manual trace of each key term
Client data anonymised, tool tier appropriate Consumer tiers train by default; Heppner Business or Enterprise tier; pseudonymise parties and amounts

Then read it, all of it, as if a trainee had drafted it. Vals’ February 2025 benchmark found lawyers still beat every tool on redlining, 79.7% against Harvey’s 65.0%; drafting is redlining’s harder sibling.

This clause-by-clause method, run on participants’ own precedents, is the drafting module of AI Lab for Lawyers, because it is the only method that survives contact with a counterparty.

Where to go next: the playbook review workflow checks a draft against your own positions; every prompt here is in the prompt library with its verification step; and the other transactional workflows sit in the use-cases hub and how lawyers actually use AI.

Frequently asked questions

Can ChatGPT draft a contract?

It can draft clauses, and drafts them well when you supply the deal terms, the governing law, your side and a precedent to copy. It cannot reliably draft a whole agreement in one prompt: definitions drift, cross-references point nowhere, and boilerplate from the wrong jurisdiction creeps in. Use ChatGPT Business or Enterprise for client terms, draft one section at a time, and read every word before it leaves the building.

Why are AI-drafted contracts risky?

Because the errors are plausible. Ellis Jones, an English disputes firm, reports AI-drafted SPAs where sellers unknowingly warranted facts they could not know, deferred consideration so unclear it risked being unenforceable, and references to foreign law that risked taking the deal outside England and Wales. The document reads like a contract, which is exactly why nobody checks it until there is a dispute.

Should I ask the AI for a whole agreement or clauses?

Clauses. A full SPA runs 40 to 80 interlocking pages and, in the words of one M&A prompt library, models 'cannot hold all variables across such length and produce generic, sometimes contradictory clauses'. Give the model your precedent and the term sheet, ask for one provision at a time, mark where the term sheet is silent, then run a consistency sweep over defined terms and cross-references.

Is Claude better than ChatGPT for contract drafting?

In one published head-to-head from June 2026 using 2024-era models, only Claude produced a mutual indemnity with a carve-out for third-party IP claims; GPT-4o's indemnity was broad and bilateral and Gemini 1.5 Pro omitted a liability cap. Current models differ less than prompts do. Whichever you use, specify the cap, the carve-outs and the side, and supply your own precedent.

Can AI draft board minutes?

It can, and you probably should not let it. Skadden partners warn that AI transcripts capture far more detail than traditional minutes and miss the inflections that signal sarcasm or jokes, which can chill board discussion, and a CEO's ChatGPT conversations about avoiding a $250m earn-out were used as evidence in Delaware in 2026. Use AI for resolutions and consents from an approved precedent, not for transcribing deliberations.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.