In March 2026 a federal judge in Colorado wrote the first two questions of your AI vendor due diligence for you. The amended protective order in Morgan v. V2X bars confidential information from AI platforms unless the provider is contractually prohibited from “(1) storing or using inputs to train or improve its model; and (2) disclosing inputs to third parties except where essential”, which the court accepted “practically bars the use of most ‘low-to-no-cost’ AI tools” (Akin’s summary).
Most firms stop there. That is how you end up with a vendor that truthfully says “we never train on your data” while keeping flagged prompts for years, routing them through a model provider you never contracted with, and charging five figures to return your files.
The 25 questions below carry the answer major vendors give as of September 2026 and the trap behind it. Per-tool detail is in the confidentiality hub.
Why the security questionnaire is not enough
A security questionnaire misses two things.
A wrapper adds a party. A widely read r/legaltech comparison of Harvey’s and Anthropic’s commercial terms found that “both prohibit training on your content” and “both commit to 48-hour breach notice”, and that Harvey’s own terms list model providers as subprocessors: “Going direct removes a party from the chain instead of adding one” (r/legaltech). A wrapper buys negotiated downstream terms, not “safer”.
The exit is where the money is. The NC Bar’s May 2026 survey digest, citing Clio, reports data-extraction fees of £12,888 in the UK and A$24,861 in Australia as an example of vendor lock-in (NC Bar).
Ask everything in writing, with a clause reference.
You are a data-protection lawyer acting for a law firm buying an AI tool. The vendor's terms and DPA are below.
Answer each numbered question in <questionnaire> in a table: Question | Vendor's answer (quote the operative words) | Clause | Gap.
Where the documents are silent, write "NOT ADDRESSED". Quote binding text only. End with the five gaps that most need a rider.
<terms>[paste]</terms>
<dpa>[paste]</dpa>
<questionnaire>[paste the 25 questions]</questionnaire>Block 1: training and retention (questions 1 to 5)
1. Do you use our inputs, outputs, files or embeddings to train or improve any model, and which clause says so? ABA Formal Opinion 512 requires informed consent before client information enters a “self-learning” tool. Use the Morgan words, “train or improve”.
2. What is the default retention, and who changes it? OpenAI deletes removed Enterprise conversations within 30 days; admins set retention. Anthropic Enterprise: “By default, data is retained indefinitely unless a custom retention period is set”.
3. Is zero data retention available, on which product, and what survives it? OpenAI: API only, “subject to prior approval”; CSAM-flagged images are kept “even if Zero Data Retention… is enabled” (OpenAI). Temporary Chat is not ZDR. Anthropic: eligible API use and Claude Code for Enterprise, “subject to Anthropic’s approval”, and it “still retains User Safety classifier results” (Anthropic). The chat product your lawyers use is usually outside ZDR.
4. What happens to flagged, rated or reported content? Anthropic keeps flagged inputs and outputs for two years and classifier scores for seven, and its privacy policy uses flagged or reported conversations for model improvement “Even if you opt-out”. Expect a named answer to “who can read it, and can you turn it off?”
5. Can our people land on the consumer edition? Perplexity: downgrade to Free and “AI data retention is enabled by default” again. Expect SSO-enforced workspaces.
Block 2: subprocessors and the wrapper chain (6 to 9)
6. Which foundation models do you call, and are their providers listed as subprocessors? Thomson Reuters built its next-generation CoCounsel Legal on Anthropic’s Claude Agent SDK. Expect a list with regions and change notice.
7. What have you obtained from them in writing? Harvey “contractually prohibits model providers from training on customer data” and “requires Zero Data Retention (ZDR) by model providers” (Harvey); Lexis+ with Protégé: “We never use customer data to train AI models”. Ask to see the flow-down.
8. Under whose identity do our requests reach the model provider? Thomson Reuters has the right answer: “processed under the identity ‘Thomson Reuters,’ never identifying the Thomson Reuters customer who made the request.”
9. What about connectors and MCP servers? OpenAI treats MCP servers as third parties with their own retention. Harvey now runs inside Microsoft Copilot, which can route to Anthropic models that “are currently excluded from the EU Data Boundary” (Microsoft).
Acting for a law firm, map every party that can touch our data when we use [vendor] for [use case].
From the subprocessor list and DPA below, build a table: Party | Role | Region | What it receives | Training prohibited? (quote) | Retention (quote) | Human review? (quote).
Where the documents do not support a cell, write "UNKNOWN".
<documents>[paste]</documents>Block 3: residency and inference location (10 to 12)
10. Where is our data stored at rest? OpenAI: EU residency for new ChatGPT Enterprise workspaces and in-region ZDR API projects. Anthropic: “us” is the only workspace geo. Harvey, Legora and Noxtua offer EU hosting (table below). Details in the EU data residency guide.
11. Where does inference run? OpenAI added in-region GPU inference in Europe on 16 January 2026. Anthropic’s own API offers only “global” or “us”; EU processing for Claude exists only via AWS Bedrock or Google Vertex.
12. What is carved out? Microsoft: “The EU Data Boundary doesn’t apply to web search queries. In addition, Anthropic models are currently excluded from the EU Data Boundary.” Ask for the footnote.
Block 4: certifications, from SOC 2 to ISO 42001 (13 to 15)
13. Which certifications, for which product, and can we read the report? Scope is the trap: Google Workspace lists SOC 1/2/3, ISO 42001 and FedRAMP High, yet Gemini Notebook “does not support ISO, SOC, or FedRAMP compliance”.
14. Do you hold ISO/IEC 42001? It certifies an AI management system, as ISO 27001 does information security. Anthropic, Google Workspace, Harvey and Legora list it; Noxtua lists it alongside BSI C5; OpenAI’s business-data page lists ISO 27001, 27017, 27018 and 27701 but not 42001. A governance signal, not a confidentiality guarantee.
15. For DACH firms: will you sign the professional-secrecy agreement? § 43e BRAO requires a contract at least in text form with instruction on the criminal consequences; ÖRAK allows a provider that will not accept § 40 Abs 3 RL-BA terms only “abstrakte und anonyme Fragen ohne Mandantenbezug”. JuriScout’s DACH comparison describes Noxtua as addressing § 203 StGB at architecture level and Harvey as not. See the BRAK, DAV and ÖRAK guide.
Block 5: incidents, holds and breach notification (16 to 19)
16. What is the contractual breach-notification period? Harvey’s and Anthropic’s commercial terms both commit to 48 hours, per the comparison above. Anything vaguer needs negotiating.
17. What is your incident history? Check, do not ask. OpenAI’s March 2023 bug exposed Plus subscribers’ chat titles and payment data, cited in the Italian Garante’s €15m fine of December 2024; Anthropic’s trust centre lists CVE-2026-22561 in the Claude for Windows installer.
18. What happens under a hold, subpoena or search warrant, and will you tell us? From 13 May to 26 September 2025 the New York Times preservation order forced OpenAI to keep every deleted Free, Plus, Pro and Team chat; only Enterprise, Edu and ZDR-API customers were excluded (the NYT order explained). Expect “we notify you unless legally prohibited” in the contract.
19. If the product runs agents, what can they access and do, and who is accountable? Harvey’s June 2026 agent guide names six governance dimensions: scope of access, authorised actions, reviewability, matter-level isolation, deployment governance and accountability. “Agents do not sign documents. Lawyers do.” (Harvey).
Block 6: model provenance and output rights (20 to 22)
20. Which model answers our prompts, and do you tell us when it changes? Anthropic’s “covered models”, from 9 June 2026, have prompts and outputs “retained for 30 days to support our safety work, on every platform where these models are offered”, which reaches wrappers too (Anthropic, covered models).
21. Who owns prompts, outputs and anything fine-tuned on our data? The Law Society of England and Wales: “establish rights over generative AI prompts, training data and outputs”. OpenAI says fine-tuned models “are for your use alone”.
22. How are sources shown? Lexis has Shepard’s Verify Trust Markers, CoCounsel Legal Deep Research Verify. Neither shifts the duty: an Illinois appellate court held in 2026 that paying for “‘premier’ or ‘corporate’ versions” does “not negate an attorney’s obligation to verify”.
Block 7: exit, export and deletion (23 to 25)
23. What does it cost to leave? The extraction fees above show what “we’ll help you migrate” can mean. Expect standard-format export within a stated period, with no fee on client files.
24. Can we delete at matter end? Morgan v. V2X also requires a contractual right to delete confidential information on request. Traps: Perplexity’s “cannot be deleted or removed”; OpenAI keeps files in custom GPTs and Projects until those are deleted.
25. What survives termination? Legal holds, flagged content, feedback and, on consumer plans, opted-in training data “in a de-identified format for up to 5 years”.
The answers major vendors give today
| Vendor (tier) | Trains by default | Retention / ZDR | Residency | Certifications listed |
|---|---|---|---|---|
| ChatGPT Enterprise / API | No | 30 days; admin-set on Enterprise; ZDR on eligible API endpoints | EU at rest for new workspaces; EU inference option | SOC 2 Type 2, ISO 27001/27017/27018/27701, CSA STAR |
| Claude Team / Enterprise / API | No | 30 days; Enterprise custom; ZDR on eligible API and Claude Code for Enterprise; covered models 30 days | Storage US only; EU via Bedrock or Vertex | SOC 2 Type II, ISO 27001, ISO 42001, HIPAA Type 1 |
| Harvey | Never | Requires ZDR by model providers | EU (Frankfurt) region | SOC 2 Type II, ISO 27001/27701/42001 |
| Legora | Not stated | BYOK | EU and US | ISO 42001, ISO 27001, SOC 2 Type 2 |
| Noxtua | Not stated | Not stated | Frankfurt; Deutsche Telekom’s AI Factory since March 2026 | BSI C5, ISO 42001, SOC 2 |
“Not stated” means the cited page is silent. The platforms are compared in Harvey vs Legora vs CoCounsel.
A scoring template for AI vendor due diligence
Score each block 0 to 3: 0 not addressed, 1 marketing only, 2 contract with gaps, 3 a clause you would show a regulator. Litigation and criminal teams weight block 5 at 20%; transactional and in-house teams weight blocks 3 and 7 at 15%. Automatic fails: training that cannot be switched off contractually; unnamed model providers; no breach-notification period; an extraction fee on client files.
Below 2 in block 1 or 7 goes back with a rider. Sheet, answers and rider form the vendor due-diligence file insurers now request: Aon’s Stan Sterna reports underwriters asking “Do you use AI? Do you police it? Do you have protocols in place?” It also feeds the bake-off guide and implementation playbook.
You are a [jurisdiction]-qualified technology lawyer acting for a law firm as customer.
From the gap list below, draft a rider with one clause per gap: no training on inputs, outputs, files or embeddings; retention and ZDR scope; abuse monitoring; subprocessor list, flow-down and change notice; storage and inference location; breach notice within [48] hours; notice of legal process unless prohibited by law; deletion at matter end; export in [format] within [30] days at no charge.
Use the vendor's defined terms, mark statutory references [VERIFY], and do not invent obligations the gap list does not support.
<gaps>[paste the "NOT ADDRESSED" rows]</gaps>Where to go next: is ChatGPT confidential for lawyers covers block 1’s settings, AI for in-house counsel covers the client’s side, and the prompts above are in the prompt library. In AI Lab for Lawyers we work through this checklist against one real vendor’s published terms, so that your next questionnaire asks about classifier scores rather than password rotation.
Frequently asked questions
What should a law firm ask an AI vendor before sharing client data?
Seven things, in writing with clause references: whether inputs, outputs, files and embeddings are used to train or improve any model; default retention and whether zero data retention is available and what survives it; which foundation models and subprocessors are involved and what terms flow down to them; where data is stored and where inference runs; certifications with their scope; breach notification, incident history and legal-hold procedure; and export format, fees and deletion at exit.
What is zero data retention?
An arrangement under which the provider does not store your prompts and outputs at all, rather than deleting them later. OpenAI offers ZDR on eligible API endpoints subject to prior approval; ChatGPT chat plans instead delete within 30 days. Anthropic offers it for eligible API use and Claude Code for Enterprise, still retains safety classifier results, and keeps prompts to its "covered models" for 30 days. A wrapper's ZDR claim concerns its model provider; ask what the wrapper itself keeps.
Does ISO 42001 matter for legal AI?
As a governance signal, yes; as a confidentiality guarantee, no. ISO/IEC 42001 certifies an AI management system, the way ISO 27001 certifies information security. Anthropic, Google Workspace, Harvey, Legora and Noxtua list it; OpenAI's business-data page lists ISO 27001, 27017, 27018 and 27701 but not 42001. Check the scope: Google Workspace holds it, yet Google says Gemini Notebook does not support ISO, SOC or FedRAMP compliance. A certified vendor can still retain flagged content.
Who are the subprocessors behind Harvey?
Harvey's public security page refers to model providers without naming them; it says it contractually prohibits model providers from training on customer data and requires zero data retention by model providers. A practitioner analysis of Harvey's platform agreement noted that its terms list model providers as subprocessors, and A&O Shearman's ContractMatrix is described as built on the Harvey API with OpenAI via Azure. Ask for the current list with regions, a change-notification right, and the flow-down terms for each name.
How do I check an AI vendor's breach history?
Look rather than ask. Read the vendor's trust centre (Anthropic's lists CVE-2026-22561 in the Claude for Windows installer), search regulator decisions (the Italian Garante fined OpenAI €15m in December 2024 partly for not notifying a March 2023 bug that exposed Plus subscribers' chat titles and payment data), and search the press for product incidents such as the roughly 4,500 shared ChatGPT conversations Google indexed in mid-2025. Then ask the vendor for its incident log and compare.