In February 2025 the UK firm Hill Dickinson checked its network logs and found more than 32,000 hits on ChatGPT in a single week. It had no rollout; it had a rollout nobody had approved. It blocked general access, and the UK Information Commissioner’s Office replied that “the answer cannot be for organisations to outlaw the use of AI and drive staff to use it under the radar”.

That is the starting condition for every law firm AI implementation in 2026. You are not introducing AI; you are governing something most of your lawyers already use, often on personal accounts. This playbook sequences 90 days from that fact, using what firms that published numbers did and what sanctioned firms failed to do.

Start where you are: 34% shadow AI, 9% enforced policy

The baseline: the 8am 2026 Legal Industry Report (1,300+ US legal professionals): 69% personally use general-purpose generative AI for work, 54% have had no training and none is planned, 43% have no policy and none is planned, and 9% work under a written policy that is enforced. Its author, Nicole Black, puts firm-level rollout at roughly 41%; the gap is lawyers using AI “off the books”. Thomson Reuters’ 2026 Future of Professionals report (1,816 professionals, 62 countries): 34% use unsanctioned tools their organisation cannot monitor, and nearly 30% of mid-career professionals would leave within two years if AI fails to deliver.

The same report is the case for sequencing rather than buying: 66% of professionals in firms with an active AI strategy say AI meets or exceeds expectations, against 22% without one. The shadow AI guide is the diagnosis; the rest of this page is the treatment.

Days 1-15: governance, tier policy, tool inventory

Clifford Chance’s inside account of its rollout, given to Legal IT Insider, has the order right: principles and policy before tools. Its risk chief’s traffic-light matrix maps data classes to tool tiers, summarised as “If it’s green, knock yourself out”, and the firm never blocked generative AI. Three tasks for the first fortnight.

Name the owner. Among corporate legal departments, 85% now have a dedicated resource or committee for AI (CLOC 2026). A firm needs one partner who answers for AI, a risk or general-counsel voice, IT, and two practising lawyers who will use the tools.

Inventory what is in use. Nicole Black’s advice: “First, look at the software you’re already using… Talk to your employees. Don’t just decide.” Ask which tools, which tiers (personal ChatGPT Plus is a different product from ChatGPT Business or Enterprise), which tasks, and whether client material has gone into a consumer account. The ChatGPT confidentiality guide explains why the tier matters more than the vendor.

Publish the three tiers. Green: public information, general legal concepts, marketing, learning the tool, on any product. Amber: client material, anonymised, in a commercial no-training tier only, with a log of what went where. Red: privileged strategy, witness statements, protective-order material and anything to be filed, in an enterprise or legal platform with contractual retention terms, counsel-directed and documented. The full policy comes on day 45.

Days 16-30: choose two tools, run a bake-off

Do not choose a platform from a demo. Harvey’s advice: keep pilots “short and specific, 60 to 90 days, in a single practice area with objectives set in advance”, and “test the tool yourself on historic matters from 2022 to 2025 where you already know the outcome”. Axiom’s analysis of the 7% of in-house teams that scaled beyond pilots agrees: 8-12 week pilots on one use case, rules set first.

Pick two candidates that answer the tasks your inventory surfaced, usually one general-purpose tool in a commercial tier (ChatGPT Business or Enterprise, Claude Team or Enterprise, Microsoft Copilot with enterprise data protection) and one legal platform. Then run the bake-off the way Vals runs its benchmarks: identical instructions and documents to the tools and to a human baseline, on matters where the answer is known. Swiftwater’s version for in-house teams: “Run five contracts you know well through any tool you are seriously evaluating. Compare the AI’s output against what you would have caught yourself.”

Design the bake-off scoring sheet
We are evaluating [Tool A] and [Tool B] against a lawyer baseline on [contract review against our NDA playbook / chronology from an email set] using [five] closed matters from 2022-2025 where we know the right answer.
Design the scoring sheet: for each matter and tool, columns for issues correctly found, issues missed, false positives, fabricated or misgrounded authorities (a count, not yes/no), time to usable draft, time to verify, and a 1-5 rating from the reviewing lawyer with one sentence of reasoning.
Add a section for facts we must confirm from the vendor's own terms: training on inputs, retention, region, audit logs. Do not score anything yourself.

One warning: Vals found lawyers still beat every tool on redlining (79.7% against 65.0%), so a tool that loses to your associates on markups is normal, not disqualifying. And one question for the winner, from Oz Benamram’s survey of 130 large-firm AI leaders: “Are your lawyers choosing the tool because it’s better for the work, or because they are used to it?”

Days 31-45: policy and confidentiality triage

By day 45 the one-page policy exists and has been acknowledged. Its core is the tier table from day 15 plus the settings that make it real. On consumer tiers, training on conversations is on by default; the opt-outs sit in the settings (ChatGPT: Data Controls, “Improve the model for everyone”; Claude: Privacy, “Help improve Claude”), and none of them makes a consumer product acceptable for client data. Two courts have settled that: United States v. Heppner (S.D.N.Y.) held a defendant’s consumer-Claude exchanges neither privileged nor work product, and the UK Upper Tribunal held that putting client letters into ChatGPT “is to place this information on the internet in the public domain”.

Draft the one-page traffic-light AI policy
Draft a one-page generative-AI use policy for a [60-lawyer firm in [jurisdiction]]. Approved tools with tier, named: [ChatGPT Business, Claude Team, Microsoft Copilot with enterprise data protection, [legal platform]]. A data classification (public / internal / confidential / privileged) mapped to which tool may receive it. RED uses (client data in consumer tools, unverified fact-finding, automated decisions), YELLOW uses (research, review, first drafts, all with verification), GREEN uses (admin, marketing, scheduling). The verification rule for anything filed or sent to a client. Client consent. Billing for actual time only. Incident reporting within 24 hours. Training and acknowledgment. Quarterly review. Then a five-question onboarding quiz. Plain language.

Regulators expect more than a document. The SRA’s warning notice of 17 August 2026, issued after 42 reports of potential AI misuse, requires that client data “is not used to train AI models except where explicitly authorised”. The Divisional Court in Ayinde (June 2025) told firm leaders that “promulgating such guidance on its own is insufficient” and that the court will ask “whether those leadership responsibilities have been fulfilled”.

Days 46-60: training that works (not the 90-minute webinar)

Thomson Reuters describes the composite failure: a unanimous partnership vote, a firm-wide announcement, then “only a fraction of attorneys had logged into the AI tools more than once”; lawyers “complained that no one had trained them on the new tools beyond a single 90-minute webinar”; diagnosis, “a focus on procurement rather than capability building”. Law360’s survey shows the correlation: 80% of frequent AI users had training and 71% of non-users had none; roughly two-thirds of BigLaw lawyers were trained by their firm against 40% at midsize firms and under 15% at small ones.

The format that works is known. Paul Weiss’s first session was a PowerPoint on ethics, confidentiality and prompting; the firm found it “ineffective” and split it into a general-education session and a hands-on prompting workshop. Jackson Lewis’s rule: “Off-the-shelf training from a vendor is not going to be as effective as if it’s co-produced with internal folks.” K&L Gates runs a separate supervisory course for partners.

Design the two-session training programme
Design a two-session AI training programme for [our corporate group, 25 lawyers]. Session 1 (60 minutes): how the tools work, our three confidentiality tiers with the settings on screen, our policy, three sanctions cases, three live demonstrations of failure (a fabricated citation, a misread clause, a jurisdiction error). Session 2 (90 minutes, hands-on): each participant brings one real anonymised task; four exercises on our approved tools; a verification exercise on a brief with two planted fake citations I will supply; each participant writes one prompt-library entry. Include a timed challenge with a leaderboard, 15 minutes of pre-work and a weekly 15-minute clinic for six weeks. Output an agenda, materials list and attendance record.

For EU firms the record matters. Article 4 of the EU AI Act has required deployers to take measures on AI literacy since 2 February 2025; the European Commission’s Q&A says “simply relying on the AI systems’ instructions for use or asking the staff to read them might be ineffective” and that “organisations can keep an internal record of trainings”. The training guide compares the programmes; AI Lab for Lawyers is built as this block, four live two-hour sessions on participants’ own anonymised documents, available as private cohorts for firms.

Days 61-75: champions and use-case library

Training decays without people inside the practice groups who keep it alive. Barnes & Thornburg’s model is explicit: its “AI Practice Champions work across the firm to evaluate emerging tools, improve legal workflows, and help integrate AI into day-to-day client service”, embedded “directly inside practice teams”.

Pay the champions in the currency that counts: Ropes & Gray offers every lawyer up to 100 creditable “Innovation Hours” a year, and Ashurst found its lawyers “were very competitive”, so leaderboards and time-boxed trials drove engagement.

The champions’ first product is the use-case library: every prompt the firm approves, with an owner, a tier and a last-tested date, because, as one r/biglaw lawyer put it, “a model will update and I will need to tweak the prompt again”. Mark Pike, who leads Anthropic’s legal product: “Don’t use it out of the box… it’s at its best when you customize it with your own legal playbooks.” The site’s prompt library is starting stock; the firm’s entries should look like this.

Format a prompt as a governed library entry
Format the following prompt as a firm prompt-library entry: Title; Owner (name and practice group); Task it performs; Approved tools and tiers; Inputs required and what must be anonymised first; The prompt itself with [placeholders]; Expected output and format; Mandatory verification steps before the output is used; Known failure modes seen in testing; Last tested on (date, tool, model version, matter type); Version number. One page. Flag anything in the prompt that would breach our tier policy if run on a consumer tool.

Prompt to format:
[paste]

Days 76-90: metrics, billing rules, client communication

Measure before you scale, because almost nobody does: Thomson Reuters’ 2026 AI in Professional Services report found only 18% of firms track AI return on investment, and Paul Weiss told Bloomberg Law after 18 months with Harvey that verification “makes any efficiency gains difficult to measure”. Four numbers are enough at day 90: active use (weekly users over licensed users), time per task against the day-30 baseline, error rate from the verification logs, and sentiment from a short anonymous survey.

Then fix billing before clients ask. The ethics floor is uniform: ABA Formal Opinion 512 requires hourly billers to charge actual time, and North Carolina’s 2024 FEO 1 holds that a lawyer whose three-hour draft now takes one hour may bill one. The business problem is that efficiency on an hourly matter is, in a phrase quoted by the ABA, “an 80% discount they never asked for”, and 86% of solos have changed nothing about their pricing. The billable hour guide has the models; the day-90 task is to price one predictable matter type flat against what it now costs to deliver.

Finally, tell clients, because they are guessing. ACC/Everlaw found 59% of in-house teams do not know whether their outside firms use generative AI; a Seward & Kissel CFO’s warning: “If you can’t explain to your clients what you’re doing in AI space, your clients will assume that you are overpriced.”

Prepare the client conversation on AI and fees
Prepare talking points for a meeting with [client] about how we now use AI on their [matter types], in this sequence: what changed about the work; what changed for the client; what value that created; how pricing should reflect it. Include where AI saves time on their matters and where it does not, our verification commitment (every authority checked at source and logged), our billing rule (actual time only, no charge for learning tools, no reconstructed "equivalent time"), and two fee options (a fixed fee for [matter type]; a capped hourly arrangement). Lead with value, not concessions. One page.

What worked and what backfired: A&O to Hill Dickinson

Rollouts with numbers: A&O, Ashurst, Macfarlanes, Dentons, Gunderson, CMS

Firm What they did Published numbers Lesson
A&O (now A&O Shearman) Harvey beta from late 2022, announced February 2023 across 43 offices ~3,500 lawyers, ~40,000 queries in the beta; staff now save 2-3 hours a week; ContractMatrix cuts review time ~30% Volume first; measured value later and modest
Ashurst Three “Vox PopulAI” trials, November 2023 to March 2024, then firm-wide Harvey in June 2024 411 people, 23 offices; ~80% time saved on UK corporate filings, 59% on sector research, 45% on first-draft briefings; experts could not tell 50% of AI outputs from human work; “frequent hallucinations”; 88% “more prepared” Trial design with a blind study and a human baseline
Macfarlanes Harvey firm-wide from 2023; Amplify for clients’ in-house teams “80%+ of lawyers now using it regularly” Adoption is the metric that precedes value
Dentons Built fleetAI on GPT-4 in August 2023 Data “not used to train the model… erased after 30 days” Negotiate data terms before the rollout
Gunderson Dettmer Built ChatGD in 2023, then Perplexity Enterprise firm-wide “More than half of the partners logged in on the first day”; 80% of lawyers active; 35,000+ queries a month; NPS 68 (vendor case study) Partner visibility on day one drives the curve
CMS Harvey pilot from March 2024 with ~300 lawyers over 12 months, then 3,000+ users, then 7,000+ across 21 member firms 95% adoption in the initial rollout; 93% reported productivity gains, ~117.9 hours saved per lawyer per year (firm and vendor figures) A year-long pilot with numbers before the enterprise deal

The A&O beta remains the best-documented start. Legal IT Insider’s report from February 2023 carried David Wakeling’s line, “I have never seen anything like Harvey”, and an anonymous UK top-50 CTO’s rejoinder that lawyers are told at law school not to use Wikipedia, “only to go to one of the leading law firms and be given a research tool that hallucinates”. Both were right; the second is why days 31 to 60 exist.

Ashurst’s write-up is the one to copy for method: six hypotheses, controlled experiments co-designed with participants, and a blind study scored by a four-expert panel. One expert who spotted a hallucinated verdict then “dismissed the remaining output and applied critical scoring across the board”: the trust barrier in one sentence, and the reason the verification exercise sits in session two.

Bans that reversed: Mishcon, Hill Dickinson

The North Carolina Bar says it plainly: “Simply telling lawyers not to use AI is unrealistic”, and lawyers under pressure “may turn to free, consumer-grade tools (like the free version of ChatGPT) on personal devices”. A ban moves usage to the tier with the worst terms and removes the firm’s ability to see it. The enforceable restriction is narrow: no client data on consumer tiers, with a sanctioned tier that is actually available.

Mid-size firms: the laggard problem

The firms most exposed are not the small ones. Nicole Black’s observation from the 8am data is that solos pivot fast and large firms have dedicated teams, while “these mid-sized firms in the middle are often the slowest because they just don’t have the resources available to devote to tech adoption and change management”. Clio’s 2026 data has 86% of mid-sized firms adopting AI; Law360’s has 40% of midsize lawyers trained. High adoption, low training, no champions: the profile of the firm in the 90-minute-webinar story.

The fix is the same sequence at smaller scale: one owner, a tier policy, two tools tested on five closed matters, two training sessions per practice group, three champions and a spreadsheet of metrics. Tom Martin’s benchmark for smaller firms is 1-3% of revenue and a fixed monthly non-billable block. Solos have their own 30-day version.

Templates and checklists

Days Block Deliverables Evidence on file
1-15 Governance and inventory Named owner and group; tool and tier inventory; three-tier table Inventory spreadsheet; tier table circulated
16-30 Bake-off Two candidate tools; five closed matters; scoring sheet; baseline times Completed scoring sheets; vendor terms checklist
31-45 Policy One-page traffic-light policy; settings walkthrough; acknowledgment Signed acknowledgments; opt-out settings confirmed
46-60 Training Two sessions per group; verification exercise; weekly clinic Attendance records; planted-citation exercise results
61-75 Champions and library Named champions with creditable time; prompt library v1 Library entries with owner and last-tested date
76-90 Metrics, billing, clients Four metrics against baseline; billing rule; one flat-fee matter type; client talking points Dashboard; engagement-letter clause; meeting notes

The right-hand column is not administrative. Legal AI Governance’s list of the artefacts to keep for a malpractice renewal: a written AI policy, a vendor due-diligence file, informed-consent language, training records, a pre-filing verification log, a usage log and an incident procedure. The 90 days produce all seven; the legal AI statistics page has the figures for a sceptical partnership.

Where to go next: the firm AI policy template is the day-45 document in full, the bake-off guide is the day-30 method, and the firm implementation hub holds the rest. For the training block, AI Lab for Lawyers runs as a private cohort for firms, with participants to date from Norton Rose Fulbright, Paul Hastings, Taylor Wessing, Simmons & Simmons and Wolf Theiss among others.

Frequently asked questions

How should a law firm implement AI?

In sequence, not by procurement. Govern what people already use (tier policy, tool inventory), run a short bake-off of two tools on historic matters where the answer is known, write a one-page traffic-light policy, train hands-on in two sessions rather than one webinar, embed champions in practice groups with a shared prompt library, then measure and fix billing before you tell clients. Ninety days is enough for a firm of any size; the order is what matters.

What is the first step in a law firm AI rollout?

An honest inventory. Thomson Reuters found 34% of professionals use AI tools their organisation cannot monitor, and Hill Dickinson found 32,000 ChatGPT hits in one week when it looked. Before choosing anything, find out which tools, tiers and personal accounts are already in use and for what, then publish a three-tier confidentiality policy (green, amber, red) that tells people what may go where. Everything else builds on that.

Should law firms ban ChatGPT?

No; ban the consumer tier for client data and provide a sanctioned tier. Hill Dickinson blocked general access after 32,000 hits in a week and the UK Information Commissioner's Office replied that 'the answer cannot be for organisations to outlaw the use of AI and drive staff to use it under the radar'. Mishcon restricted ChatGPT in March 2023 and rolled Legora out firm-wide in July 2025. Clifford Chance never blocked at all: 'If it's green, knock yourself out.'

How long does an AI rollout take?

Ninety days to a governed baseline; longer to full deployment. Harvey recommends pilots of 60 to 90 days in a single practice area with objectives set in advance; Axiom's data on successful in-house teams points to 8-12 week structured pilots on one use case; the consultancy Swiftwater sequences a 90-day plan in three 30-day blocks. Ashurst's three trials ran from November 2023 to March 2024 before its firm-wide Harvey rollout in June 2024.

How do you measure AI adoption in a law firm?

Baseline first, then four numbers: active use (Macfarlanes reports 80%+ of lawyers using Harvey regularly, Gunderson 80% active on Perplexity), time per task against the baseline (Ashurst measured 80%, 59% and 45% savings on three task types), error rate from verification logs, and user sentiment. Do not skip the baseline: only 18% of firms track AI ROI, and Paul Weiss says verification makes efficiency gains hard to measure.

What did Allen & Overy learn from Harvey?

That volume comes before value. In the beta announced in February 2023, around 3,500 A&O lawyers asked Harvey about 40,000 queries, and David Wakeling said he had 'never seen anything like Harvey' in 15 years. The measured results came later and were modest: staff save 2-3 hours a week on average on summarisation, analysis and translation, and ContractMatrix cuts contract review time by about 30%. An anonymous UK CTO's warning at the time: 'a research tool that hallucinates'.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.