Ask a model to “review this contract” and it reviews it for nobody. It does not know which side you act for, what you would accept, where you would walk away, or who must approve a concession. Every serious AI contract review tool on the market in September 2026 has arrived at the same fix: give the model a playbook first, then let it flag deviations.
That makes AI contract review two jobs. The second, running the review, takes minutes and is what vendors demonstrate. The first, writing down your positions, is the one most teams have never done. LegalOn’s survey of legal teams found that 95% have playbook gaps and 34% have no playbook at all; only 5% called theirs comprehensive. Vendor numbers, but they match what I see in practice.
Why every tool converged on the playbook
Spellbook, LegalOn, Law Insider, GC AI, Harvey, Legora and Anthropic’s Legal plugin all describe the same loop: upload the contract, compare it with encoded positions, flag what deviates, propose a redline. Harvey puts it best: “AI flags deviations. The lawyer decides which deviations matter for this counterparty, this matter, this risk profile.”
Mayer Brown’s M&A team put the precondition bluntly in September 2025: “This use case requires a strong foundation—specifically, a negotiation playbook with your standard and fallback positions, supported by sample provisions.” Its example is the kind of rule you have to write down: a standard confidentiality term of 24 months, but 18 for certain strategic bidders.
Anthropic’s open-source review skill refuses to work without one. Ask it to review before setup and it answers: “Run /commercial-legal:cold-start-interview first — I need to learn your playbook before I can review against it.” The Claude for Legal README explains why: “Skipping setup is the single most common reason a skill produces generic output.” It is true of every tool on this page.
Step 1: write the playbook (preferred, fallback, walk-away)
A playbook is a table, not a memo: one row per clause type, with the position you want, the one you will accept, the one you will not, model wording for each, and who must approve a deviation. GC AI’s NDA playbook structures nine clauses this way. Three of its rows show the level of detail that works:
| Clause | Preferred | Fallback | Walk-away |
|---|---|---|---|
| Term | “2-5 years commercial; indefinite for trade secrets” | “fixed 3-5 years across board” | “perpetual on all or expires before value loss” |
| Residuals | “none” | — | “broad clause covering remembered information” |
| Remedies | — | — | “liquidated/punitive damages or one-sided fee-shifting” |
Quoted cells are verbatim from GC AI; dashes mark cells not reproduced here. An NDA playbook is a morning’s work; a vendor MSA playbook is a week.
You do not start from a blank page. Five agreements you actually signed contain your real positions, including those you never admitted to:
I am building a negotiation playbook for [mutual NDAs] for a [SaaS company, 400 staff] that is usually the [receiving party]. Attached are five agreements we signed after negotiation: <signed_1>…</signed_1> to <signed_5>…</signed_5>.
For each clause type (definition of confidential information; purpose; exclusions; term; compelled disclosure; return or destroy; residuals; remedies; governing law; liability), extract the position we accepted in each agreement. Then propose Preferred | Fallback | Walk-away, with one sentence of reasoning, a short model clause for each, and the escalation trigger (who must approve a deviation).
Present as a table, then as a numbered playbook. Flag every clause type where our five agreements are inconsistent. Do not cite any statute or case.A partner or the GC signs off every row; the model proposes, it does not decide. Save the result as the knowledge file in a Claude Project or custom GPT (the Claude Projects and custom GPTs guide covers the set-up).
Step 2: choose the surface: chat, Project, Word add-in or platform
The same playbook can run in four places. A chat window (ChatGPT Business or Enterprise, Claude Team or Enterprise, licensed Copilot; never a consumer plan for client paper) suits a one-off review but forgets the playbook every session and returns prose, not a redline. A Claude Project or custom GPT holds the playbook as a knowledge file for repeat reviews of one contract type. A Word-native add-in (Claude for Word, Spellbook, Law Insider, Microsoft’s Legal Agent) produces tracked changes in the document you are editing but sees that one document in isolation. A legal platform (Harvey, Legora, LegalOn, CoCounsel) applies playbooks across hundreds of agreements, at a price and with seat minimums.
Whichever surface you pick, the first prompt is not “review this” but “what is this, and whose side am I on”. Anthropic’s open-source review skill reads the title and the exhibit titles before the body, on the rule that “a 40-page MSA with ‘confidential’ throughout is not an NDA”, then determines the user’s side, then reviews. Copy that order: ask for the agreement type, the party you represent, the governing law, and every exhibit, order form or URL incorporated by reference and whether it is attached, then stop. A review run from the wrong side inverts every recommendation, and unattached order forms are where the real terms hide.
Step 3: the review prompt and the output table
With the playbook loaded and the side confirmed, the core prompt asks for one row per playbook item, with the operative words quoted. That requirement matters most, for a reason the Better Call GPT section explains.
Review the attached [NDA] against our playbook <playbook>…</playbook> from the perspective of [the receiving party]. Read the entire agreement before flagging anything.
For every playbook item, output one row: Playbook item | Contract clause (number and quoted operative words) | Status: GREEN (meets preferred) / YELLOW (within fallback) / RED (beyond walk-away, or missing) | Business impact in one sentence | Proposed replacement wording | Escalation required (Yes/No per playbook).
Then list: any provision the playbook does not cover but that shifts risk to us; any internal inconsistency between clauses; and a three-line summary for the business owner. Finally, list what you could not assess because a schedule or exhibit is missing. Do not cite any case or statute.GREEN, YELLOW and RED is the vocabulary of Anthropic’s /review-contract command, and it maps onto how Spellbook, LegalOn and Law Insider present their flags. Use it in your own playbook so output is consistent across tools.
Step 4: calibrate on five contracts you know
Nobody should trust a review tool on the strength of a demo. Swiftwater’s advice to in-house teams is the test: “Run five contracts you know well through any tool you are seriously evaluating. Compare the AI’s output against what you would have caught yourself.” Claude Academy’s guide to NDA Projects says the same: calibrate on three to five previously reviewed NDAs before rollout.
Score it in four columns: issues the lawyer found, the tool found, both, neither. The “tool only” column is part real misses and part noise; the ratio tells you how far to trust the tool at volume.
What Better Call GPT found: determination versus location
The most-cited evidence is still Onit’s “Better Call GPT” study from January 2024: six models, ten anonymised procurement contracts under US and New Zealand law, reviewed against a playbook, with senior lawyers as ground truth and junior lawyers and legal process outsourcers (LPOs) as comparators. The models were 2024-era, which makes the result a floor, not a ceiling.
| Reviewer | Issue determination (F-score) | Issue location (F-score) | Cost per contract |
|---|---|---|---|
| Junior lawyer | 0.860 | 0.667 | $74.26 |
| LPO | 0.874 | 0.770 | $36.85 |
| GPT-4-1106 | 0.871 | 0.686 | — |
| GPT-4-32k | 0.820 | 0.740 | $1.24 |
| Claude 2.1 | 0.809 | 0.701 | $0.02 |
The paper’s Claude cost line is not broken down by version. Two findings matter more. The models matched or beat junior lawyers at determining that an issue exists but were worse than the outsourcers at locating it, which is why the prompts above insist on quoted words: a flag without a location is a flag you cannot check. And the documented errors were misreadings of specific wording: against a standard requiring immediate written notice, two models treated a clause that only required each party to “promptly notify” the other as meeting it. Every group, human and machine, showed “a preference for precision over recall”; if you want completeness, say so.
Where AI contract review fails: cross-references, side letters, four-condition renewals
Cross-reference blindness. Alon Kapen of Farrell Fritz describes an IP assignment that “looks broad until the definition of ‘Developed IP’, three layers deep, limits it to first-year improvements”. His illustrative diligence failure is worse: a tool flags 47 change-of-control provisions, three invented from neighbouring language, while two real triggers in side letters are missed. “The professional presentation masks the underlying uncertainty.”
Stacked conditions and asymmetry. Spellbook’s list of where ChatGPT fails includes evergreen clauses with four stacked non-renewal conditions (“written notice via certified mail 120 days prior to anniversary date”), asymmetric liability caps and carve-outs, undefined “reasonable administrative fees”, arbitration forum and rules, and the conflation of direct, indirect and consequential damages.
None of these are exotic; they are the clauses a tired associate checks last. Run them as a deliberate second pass:
Second pass on the attached agreement. Check only these known trouble spots and quote the operative words for each: (1) auto-renewal: every condition for a valid non-renewal notice (form, method, days before anniversary, addressee) as a checklist; (2) liability: are caps and carve-outs symmetrical? show each party's cap and exclusions side by side; (3) damages: any conflation of direct, indirect and consequential loss; (4) fees: undefined "reasonable" charges, tiers or index-linked increases; (5) disputes: forum, rules, seat, discovery limits, jury or class waiver; (6) defined terms: trace "[Confidential Information / Developed IP / Services]" through every layer of definition and state its effective scope; (7) side letters, order forms or URLs incorporated by reference. Report NOT PRESENT where a trouble spot does not arise.Item (6) you then check by hand: it is where the chain of definitions decides the deal, and where the evidence says models stop early. If the agreement is long, feed it in blocks; the guide to prompting long contracts explains why.
Word-native options compared: Claude for Word, Spellbook, Microsoft Legal Agent
For most transactional lawyers the review must end as tracked changes in the document. Three Word-native routes exist in September 2026.
| Tool | Status | What it does | Limits |
|---|---|---|---|
| Claude for Word | Beta since 11 April 2026; Claude Team and Enterprise plans only | Native tracked changes; recognises “multi-level numbering, defined terms, cross-references, and standard contract structures” | One document in isolation; no persistent chat history; no Enterprise audit logs yet; inputs and outputs deleted within 30 days |
| Spellbook | Established Word (and Google Docs) add-in | Tracked-change redlines; benchmarks against 2,300-plus contract types; agentic mode reconciles a 150-page SPA against a playbook | No published pricing (Lawyerist’s main complaint); an add-on, not a standalone |
| Microsoft Legal Agent in Word | Announced 30 April 2026 via the Frontier early-access programme; Word on Windows desktop | Clause-by-clause playbook review; tracked-change redlines produced through a “deterministic resolution layer” rather than by the model rewriting every line | Early access only; Artificial Lawyer’s market feedback in June 2026 was that it is “just not good enough yet” |
Anthropic’s published prompts for Claude for Word are short because the playbook does the work: “Flag provisions that deviate from standard market position, ranked by severity”; “Make the indemnification mutual and insert our standard fallback language”; “What did the counterparty change, and which revisions are dealbreakers?” Anthropic’s caveat sits beside them: “always verify that outputs match your specific requirements and your firm’s standard positions.”
Redlining is also the one task where lawyers still beat the tools head to head: Vals’ February 2025 report scored the lawyer baseline at 79.7% against Harvey’s 65.0%. A tool’s redline is a draft, not a result.
The 34% problem: most teams have no playbook at all
If a third of legal teams have no playbook, then for a third of legal teams the honest answer to “can we use AI for contract review” is “not yet, and the missing piece is not software”. The fastest route is Step 1: five signed agreements, one extraction prompt, one partner review, half a day. The result improves every tool you will ever buy, which is why Anthropic’s Mark Pike said of the company’s legal plugin: “Don’t use it out of the box… it’s at its best when you customize it with your own legal playbooks.”
That half-day is the exercise we run in session two of AI Lab for Lawyers: each participant builds a mini-playbook for one agreement type and reviews a live NDA against it, first in a chat window, then in a Claude Project, so the difference is visible on screen. In-house teams handling volume will find triage rules in the in-house counsel guide.
Where to go next: if the review shows you need to redraft rather than redline, read AI contract drafting: what goes wrong; the Harvey, Legora and CoCounsel comparison covers pricing signals and seat minimums; the contract prompts are in the prompt library; and the other workflows sit in the use-cases hub and how lawyers actually use AI.
Frequently asked questions
Can ChatGPT review a contract reliably?
Against a written playbook, in a no-training tier such as ChatGPT Business or Enterprise, it can flag most deviations reliably enough to be useful: the Better Call GPT study found GPT-4 matched junior lawyers on identifying issues. Without a playbook, side and jurisdiction, it produces a generic list. It still misses cross-referenced definitions, side letters and stacked renewal conditions, so a lawyer reads every flagged clause in the original.
What is a contract playbook?
A contract playbook is a written table of your negotiating positions for each clause type in a given agreement: the preferred position, an acceptable fallback, the walk-away point, model wording for each, and who must approve a deviation. GC AI's NDA playbook, for example, gives term as preferred two to five years, fallback a fixed three to five, walk-away perpetual on everything. Tools then flag the contract against those rows.
Which is better for contract review, Claude or Spellbook?
They are different surfaces. Claude for Word (Team and Enterprise plans, beta since April 2026) writes native tracked changes from your own playbook; Spellbook is a purpose-built Word add-in with clause benchmarks across 2,300-plus contract types, zero-data-retention agreements with OpenAI and Anthropic, and unpublished pricing. Run five contracts you know through both before choosing.
How accurate is AI contract review?
The most-cited benchmark, Onit's Better Call GPT study, put GPT-4's issue-determination F-score at 0.871 against 0.860 for junior lawyers and 0.874 for outsourcers, but issue location was weaker (0.686 to 0.740 against 0.770). Vals' 2025 report found lawyers still beat every tool on redlining, 79.7% to 65.0%. Accuracy is good enough for a first pass and not good enough to skip the read.
Does AI catch cross-referenced definitions?
Not reliably. Farrell Fritz describes an IP assignment that looks broad until the definition of 'Developed IP', three layers deep, limits it to first-year improvements, a chain the tool never followed. Claude for Word says it recognises defined terms and cross-references, and Spellbook's agentic mode updates defined terms across documents, but you should still ask the model to trace each key term explicitly and check the chain by hand.