The general AI vs legal-specific AI question has a number attached. The most commercially important figure in Clio’s 2025 Legal Trends Report is not the 79% of legal professionals who use AI; it is the share using a legal-specific tool: 40%, down from 58% in 2024, while nearly half now use ChatGPT, Gemini, Claude or Perplexity (Clio). Lawyers quietly moved the other way in the year the platforms raised billions.

On the r/legaltech thread where a large firm’s innovation committee shared its quotes, one commenter wrote: “If only all of us had known we could slap a label on ChatGPT and sell it to law firms for millions.” The question is usually asked as “which is more accurate?” That is the wrong question, and the benchmarks show why.

The ABA’s 2024 Legal Technology Survey found 52% of lawyers using or considering ChatGPT against 26% for CoCounsel, 24% for Lexis+ AI and 5.9% for Harvey. The State Bar of Texas 2026 survey has ChatGPT at 63% of AI users, Copilot at 44% and Westlaw/CoCounsel, the top legal-specific tool, at 30%.

Inside large firms the platforms are pilots more than plumbing: ILTA’s 2025 survey found 71% of Harvey users still piloting. “All the associates hated Harvey and weren’t too keen on CoCounsel. They loved Westlaw’s AI research, but go figure the partners went with Harvey because they think it’s magic” (r/legaltech).

What the benchmarks say: accuracy is close, authority differs

Test General model Legal-specific tools Lawyers
Vals legal research, Oct 2025 (210 questions, per LawNext’s write-up) ChatGPT 80% accuracy; 70% authoritativeness Counsel Stack 81%, Alexi 80%, Midpage 79%; 76% authoritativeness 71% accuracy
Vals Legal AI Report, Feb 2025 Not tested Harvey 94.8% document Q&A, 65.0% redlining 70.1% document Q&A, 79.7% redlining
Stanford RegLab, 2024 GPT-4 hallucinated “at least 58%” on 800,000+ general legal queries (Dahl et al.) Lexis+ AI more than 17%; Westlaw AI-AR more than 34% (Magesh et al., 202 queries) n/a

Two readings. First, “Both legal AI and generalist AI can produce highly accurate answers to legal research questions”, in the words of the Vals report (LawNext); the six-point gap was on whether the answer rested on the right authority, which is what gets you sanctioned. Second, lawyers still beat every tool on redlining in the only head-to-head of the big vendors (Vals); Thomson Reuters, LexisNexis and vLex declined the research test. See legal AI benchmarks explained.

What a platform actually adds: citations, matter context, audit trail, contract

Once accuracy is near a tie, a platform’s real product is four things a chat window lacks.

  • Verifiable citations. Lexis+ with Protégé’s Shepard’s Verify Trust Markers flag citations that cannot be verified as existing; CoCounsel’s Deep Research Verify checks whether the cited authority supports the assertion. Neither replaces reading the case.
  • Matter context. The innovation-committee lawyer’s verdict on Claude Cowork: “the real tier is $200, and it’s still not a legal tool”, because it is not matter-centric. Harvey’s Vault holds up to 100,000 documents per project with every cell linked to its source; Legora’s Tabular Review puts documents in rows and questions in columns.
  • Audit trail and admin. Who ran what, on which documents, under which permissions; Harvey’s governance guide lists reviewability and matter-level isolation as design dimensions.
  • The contract. Thomson Reuters says CoCounsel requests are “processed under the identity ‘Thomson Reuters’” with “systemic controls to turn off third parties’ abuse-monitoring solutions”. A business plan on a general model gives you no-training terms, not that.

Tricia Kinney, General Counsel of BlueLinx, puts the in-house view bluntly: “I am a huge fan of using legal specific AI tools as opposed to consumer specific AI tools. You want them training in the same context that we’re operating in” (quoted by GC AI).

The wrapper paradox: a platform adds a subprocessor

The confidentiality argument for platforms is weaker than the sales deck implies. An r/legaltech analysis compared Anthropic’s commercial terms and DPA with Harvey’s platform agreement, DPA and security addendum: “both prohibit training on your content, both encrypt at AES-256 at rest and TLS 1.2+ in transit, both commit to 48-hour breach notice, and both offer SOC 2 and audit rights. Harvey has two fair edges, a higher liability cap for a data breach and more operational detail in its security addendum… Harvey is a layer on top of foundation models, and its own terms list model providers as subprocessors… Going direct removes a party from the chain instead of adding one” (r/legaltech).

The replies sharpen it: the real exposure is abuse-monitoring escalation to human review, which the wrappers negotiate away. So a wrapper’s precise value on confidentiality is a negotiated human-review carve-out and a bigger breach cap, not “safer” in the abstract. The questions for any vendor are in the AI vendor due-diligence checklist.

‘Pot vs rice cooker’: the Shapiro-Waisberg debate

Zack Shapiro of two-lawyer Rains LLP made the case for the general model in February 2026: a well-configured general-purpose AI, with skills encoding the lawyer’s own judgement, beats vertical legal AI for a small firm. “The entire gap between ‘AI is a toy’ and ‘AI changed my practice’ lives in the quality of your instructions” (Shapiro on X).

Noah Waisberg, who founded Kira and now runs Zuva, answered with the debate’s best analogy: “Saying ‘foundation models mean legal AI is cooked’ is like saying ‘rice cookers are totally useless when you have a pot.’” And the warning for anyone above two lawyers: “as you invest more and more into your Claude prompts and workflows, what you’re building starts to look a lot like software… except without the QA, the versioning, the user feedback loops, or the ability to survive someone leaving the firm” (Zuva). BigLaw clients, he argues, pay for the last 20%.

Both are right about different firms: a pot is enough if you cook every day and know what you are doing; a rice cooker is for the kitchen where six people take turns. The two-person model is examined in AI-native law firms explained.

When ChatGPT or Claude is enough: the solo and small-firm case

Practitioners who get the most from general models describe the same tasks. “I do divorce, and AI is great for it, no need for a ton of legal research or cites, and it cranks out standard petition and motions easily.” “Where it really shines is finding me my needle in the haystack of a large file.” Both from r/Lawyertalk, both from lawyers who verify.

The conditions are strict: a business tier with no training by default (ChatGPT Business or Claude Team, $20 to $25 a seat); anonymised inputs; a written playbook, because LegalOn found 34% of legal teams have none; and a verification rule for anything with a citation. Under those conditions:

Playbook review with a general model (business tier, anonymised)
Working rules: apply [jurisdiction] law only; use only the materials below; tag any case or statute [VERIFY]; never invent facts; label inferences INFERENCE.
<playbook>[preferred / fallback / walk-away positions per clause type]</playbook>
<contract>[anonymised agreement]</contract>
I act for the [customer]. Read the entire contract before flagging anything. For every playbook item output one row: Item | Clause (number, operative words quoted) | GREEN (meets preferred) / YELLOW (within fallback) / RED (beyond walk-away or missing) | Business impact in one sentence | Proposed redline. Then list any clause the playbook does not cover that shifts risk to us, and anything you could not assess because a schedule is missing.

The 30-day small-firm set-up around a general model is in AI for solo and small law firms.

When a platform earns its seat: volume, data rooms, client pressure

Three triggers justify the price. Volume: GSK Stockmann reports up to 75% time saved on unstructured data rooms with Harvey, and Harvey’s benchmark team concedes what a chat window cannot fix: agents’ “bias is to search efficiently, not completely. Diligence requires reversing this intuition.” Protective orders: Morgan v. V2X (D. Colo., 30 March 2026) barred confidential discovery material from any AI platform unless the provider is contractually prohibited from training on inputs and from third-party disclosure, which “practically bars the use of most ‘low-to-no-cost’ AI tools”. Client pressure: “There is increasing client pressure to have some sort of ‘name brand’ (i.e. Harvey, Legora, Lexis, West) AI”, as one r/legaltech commenter put it; Litera’s 2026 survey found 51% of firms had a client directly influence an AI investment decision.

The unbundling trend: narrow, practice-specific tools

The choice is no longer binary. On the same thread: “my observation is that we appear to be moving away from the big platforms (Lexis AI, Legora, Harvey) and deciding much more on narrowly focused practice specific tools”, naming a stack of “half a dozen, practice focused tools”: Spellbook for corporate, Patlytics for IP, Relativity aiR for review, Gavel Exec for contracts. The economics support it: Relativity folded aiR for Review and aiR for Privilege into standard RelativityOne at no extra charge from early 2026, and Clio Work is published at $199 a user a month.

The caution is Robin AI, which failed a roughly $50 million raise in October 2025 and sold its managed-services arm two months later. Narrow tools sit closest to the model providers’ own roadmaps; check terms, exit and data claw-back before building a workflow on one. The platform head-to-head is in Harvey vs Legora vs CoCounsel.

Question If yes If no
Does a client’s OCG or a protective order require a named platform or specific terms? Platform, or a general model on terms that meet the order Next question
Do you review sets a chat window cannot hold: data rooms, productions, hundreds of contracts? Platform with tabular review, or e-discovery AI Next question
Will you file research done with the tool? Grounded platform with a citator, plus your own reading Next question
Have you standardised a workflow and written the playbook? A general model in a Project or Skill, or a narrow Word-native tool Write the playbook first
Fewer than 50 lawyers and none of the above? Business-tier general model, anonymisation, verification log Revisit in six months

Before any purchase, run the same test on both candidates, on “historic matters from 2022 to 2025 where you already know the outcome”, as Harvey’s own pilot advice puts it.

Five self-tests to run on any tool before you buy
Design five self-tests I can run on [tool] to probe legal hallucination under [jurisdiction] law: (1) a false-premise question about a dissent that was never written; (2) a fictitious judge or party; (3) an overruled precedent presented as current; (4) a jurisdiction trap, a [Texas] question that invites [California] law; (5) an "are these citations real?" trap where I supply one real and one invented citation.
For each: the exact prompt, the correct behaviour, the failure behaviour, and how I record the result. Do not run the tests; I will.
Vendor questionnaire for a platform or a wrapper
Prepare a vendor questionnaire for [tool] with a must-have / nice-to-have column: written no-training clause covering inputs, outputs, files and embeddings; retention default and what survives zero data retention; whether abuse monitoring and human review can be turned off, and who reads flagged content; subprocessors and foundation models with regions; SOC 2 Type II, ISO 27001 and 42001; admin controls (SSO, audit logs, retention, disabling feedback); deletion and claw-back at matter end; notification if a production order hits our data.

The full bake-off protocol is in how to evaluate legal AI tools.

Where to go next: the ranked overview is best AI tools for lawyers in the tools hub; prompts that make a general model behave like a configured one are in the prompt library. AI Lab for Lawyers rests on the premise behind this guide: most lawyers can get a large share of the value from tools they already own once they know how, so its four live sessions use browser tools only, with ChatGPT, Claude Cowork, NotebookLM and Harvey side by side.

Frequently asked questions

Is ChatGPT enough for a small law firm?

For most of the work, yes, provided it is ChatGPT Business or Enterprise (no training by default), documents are anonymised, and every authority is verified in a database. A divorce lawyer on r/Lawyertalk called it great for standard petitions and motions with no need for citations. It is not enough for legal research you intend to file, for data-room diligence at volume, or for clients whose outside-counsel guidelines demand a named legal platform.

What does Harvey do that ChatGPT cannot?

Three things a chat window lacks: matter-centric workspaces (Vault holds up to 100,000 documents per project, with every extracted cell linked to its source), playbook and precedent retrieval across the firm's own deals, and a commercial stack built for firms, including zero-data-retention terms with model providers, SOC 2 and ISO 42001 attestations and admin controls. It scored 94.8% on document Q&A in Vals' 2025 report. It does not remove the duty to read and verify.

Are legal AI platforms just wrappers?

Partly. An r/legaltech analysis of Anthropic's and Harvey's commercial terms found both prohibit training, both encrypt at the same standards, both commit to 48-hour breach notice, and Harvey's own terms list the model providers as subprocessors, so going direct removes a party from the chain instead of adding one. What a wrapper buys is negotiated abuse-monitoring carve-outs, a higher breach liability cap, workflow tooling and grounded legal content. Whether that is worth the seat price depends on volume.

Do clients expect firms to use a named legal AI?

Some do, and it is growing. A lawyer on the r/legaltech pricing thread reported increasing client pressure to have some sort of name-brand AI such as Harvey, Legora, Lexis or Westlaw, and Litera's 2026 survey found 51% of firms had a client directly influence an AI investment decision in the last year. But Thomson Reuters found that while two-thirds of corporate respondents want their outside firms to use AI, fewer than 20% actually mandate it.

When should a firm buy a legal AI platform?

When one of four things is true: you run document volumes a chat window cannot hold (data rooms, large productions); a client's guidelines require a named platform with contractual no-training and retention terms; your research must be grounded in a citator-backed corpus for filing; or you have standardised a workflow and want it to run the same way for every lawyer. Below that, test a business-tier general model on three past matters first.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.