Published December 19, 2025 · Last updated September 2026

Harvey Subprocessors, SCIM and Confidentiality: The Security Review to Run Before You Buy Legal AI Software

What Harvey publishes about subprocessors, SCIM, SAML SSO, IP allow-listing and data residency, what stays behind its trust portal, and the exact wording to use when you ask.

Review a legal document right now

Upload a contract, brief, lease or exhibit and LegalSoul returns the issues, the risky clauses and the page cites in under a minute. Published pricing, no seat minimum, no quote process.

Short answer. Harvey states that it does not use inputs, outputs or uploaded documents to train underlying models, and it contractually requires zero data retention from the model providers behind it. It holds SOC 2 Type II and ISO 27001, and it documents SAML SSO, audit logs and IP allow-listing. The two answers it does not put on the open web are the subprocessor list, which sits behind an access request at its trust portal, and SCIM provisioning, which its public security material does not state either way.

That gap matters, because those are exactly the two things a client questionnaire or a malpractice carrier asks you about. Below is what Harvey publishes, what it does not, and the wording to use when you ask. Last updated September 2026.

Where can I find the Harvey subprocessor list?

Harvey publishes its subprocessor list through its trust portal at trust.harvey.ai, which runs on SafeBase and gates documents behind an access request. The list is not posted openly on the web, so budget for a request and, in most cases, an NDA before you can read it. Build that step into your diligence timeline rather than assuming a same-day download.

Two things are worth knowing before you request it. First, Harvey is hosted on Microsoft Azure, so the cloud layer is Microsoft regardless of which model providers appear on the list. Second, Harvey and LexisNexis announced a research alliance that puts authoritative legal content inside the Harvey workflow, so expect the list to reflect that relationship. Neither fact substitutes for the current signed list, which is the only version your client will accept.

Does Harvey support SCIM provisioning?

Harvey's public security material does not state SCIM support either way. It documents SAML SSO, audit logs and IP allow-listing, but SCIM is a separate protocol and SSO does not imply it. If automated deprovisioning is a requirement, ask the vendor in writing and get the answer in the order form or the security exhibit, not in an email from a sales engineer.

The practical question behind SCIM is simple: when an associate leaves on a Friday, what removes their access? With SCIM, your identity provider does it. Without SCIM, someone at your firm does it by hand, and that person needs to be named in your offboarding checklist. Either answer is workable. An unanswered one is not.

Does Harvey support SAML single sign-on?

Yes. Harvey documents SAML SSO as part of its enterprise controls, alongside audit logs, IP allow-listing, role-based access control and logical workspace separation. If your firm already runs Okta, Entra ID or another identity provider, that means Harvey access follows the same lifecycle as the rest of your stack rather than living in a separate password list.

Does Harvey support IP allow-listing?

Yes. IP allow-listing appears in Harvey's published enterprise controls next to SAML SSO and audit logs. In practice that lets a firm restrict access to its office ranges or its VPN egress addresses. It is worth confirming during procurement whether allow-listing is available on your specific contract tier, since enterprise control availability often varies by plan.

Where can my Harvey data be hosted, and can I keep it in the EU?

Harvey runs on Microsoft Azure and offers in-region processing for the EU and Switzerland, and for Australia, for customers with data localization requirements. If a client engagement letter or a regulator has told you where the data must sit, name the region in the contract rather than relying on a marketing page, and ask whether the requirement covers backups and support access as well as primary processing.

Can anyone at my institution see my work in Harvey?

Not by default. Harvey enforces role-based access control with logical workspace separation, and it syncs and enforces your firm's existing ethical wall policies, blocking restricted users from walled material. Harvey states it never creates, modifies or deletes those walls itself, which means the walls in Harvey are only as correct as the ones in your source system.

That last point is where firms get caught. If your conflicts system has a stale wall, Harvey will faithfully enforce the stale version. Reconcile the source before you roll out, not after a partner asks why a matter appeared in a search.

Does Harvey train its models on my client documents?

No. Harvey states it does not use inputs, outputs or uploaded documents to train the underlying models, and it contractually requires zero data retention from the model providers it uses. Firms that want a model tuned to their own material can request one explicitly, in which case it is built exclusively for that customer rather than folded into a shared model.

What certifications does Harvey hold?

Harvey holds SOC 2 Type II and ISO 27001, both renewed annually, plus ISO 27701 for privacy information management and ISO 42001 for AI management systems. Its recent audits were performed by Schellman, and it engages NCC Group and Bishop Fox for third-party penetration testing. It has also completed an Australian IRAP assessment.

Ask for the report, not the badge. A logo on a website tells you a certification existed at some point. The report tells you the scope, the audit window and the exceptions, and the scope is where the interesting information lives.

Harvey security controls at a glance

ControlStatus in Harvey's public materialWhat to do about it
Training on customer dataStated: no training on inputs, outputs or uploadsGet the clause in the contract, covering all three
Zero data retention by model providersStated as contractually requiredAsk which providers, and confirm it survives renewal
Subprocessor listGated behind an access request at trust.harvey.aiRequest early, and ask for advance notice of changes
SAML SSODocumentedConnect it to your identity provider on day one
SCIM provisioningNot stated publiclyConfirm in writing, or document a manual offboarding step
IP allow-listingDocumentedConfirm availability on your contract tier
Audit logsDocumentedConfirm retention period and export format
Ethical wallsSynced and enforced, never created or edited by HarveyReconcile your source system before rollout
CertificationsSOC 2 Type II, ISO 27001, ISO 27701, ISO 42001, IRAPRequest the current reports and read the scope
Data residencyAzure, with EU and Switzerland or Australia in-region optionsName the region in the contract if a client requires it
Penetration testingAnnual, by third parties including NCC Group and Bishop FoxAsk for the most recent summary letter
Published pricingNone. Quote onlyModel the cost before the diligence work, not after

The part firms underestimate: the diligence itself costs money

Everything above is answerable. What surprises smaller firms is how much partner time it takes to get the answers, because a gated trust portal, an NDA and a security exhibit are a procurement process whether or not you have a procurement department. For a firm of forty attorneys with a security team, that cost is noise. For a firm of six, it can exceed the first year of licence fees in billable time.

That is the real fork in the road. If your practice needs SSO, ethical wall enforcement, iManage integration and a SOC 2 report you can hand to a client, the diligence is worth running and Harvey answers those requirements. If what you actually need is a fast, accurate read on the agreement sitting on your desk, you are buying a platform to do one job, and there are cheaper ways to get there. We laid out both sides of that trade in the Harvey AI alternative comparison for small and midsize law firms, including where Harvey is plainly the stronger choice.

LegalSoul sits on the other side of that line on purpose. It does document review, it publishes its prices, and the first review runs without an account, so you can judge the output in the time it takes to read one page of a security questionnaire. What it does not have, today, is SAML SSO, a SOC 2 report or an ethical wall engine, and if those are on your requirement list you should buy accordingly.

The name on the tool matters less than how you deploy it. What counts is whether your AI copilot keeps privilege intact, blocks leaks, and holds up under a client audit. One sloppy setting on retention, training, or access makes for a very long week.

Here’s what follows: a plain-English checklist of what “safe” actually means for legal work, the ethics and privacy rules you need to hit, and the must-have controls (zero data retention, no training on your data, solid RBAC, encryption, and real audit logs you can ship to your SIEM). We’ll talk deployments and data residency, how to preserve privilege with a vendor, and the guardrails that prevent bad outcomes. You’ll also get a due-diligence list, a quick decision guide, common gotchas, and how LegalSoul is built to protect client information from day one.

Short answer and who this applies to

Short answer: yes, an AI copilot can be safe for privileged work, if you control the setup, the data path, and the contract. Safety isn’t about the logo. It’s about specific confidentiality requirements your legal team can enforce every day. If you’re juggling strict OCGs, regulated data, or cross-border matters, raise the bar to the same level you use for your DMS or eDiscovery vendor.

Why be picky? IBM’s 2024 Cost of a Data Breach report puts the average incident at $4.88M. The UK ICO has warned that generative AI can expose personal data through training or logs if settings aren’t nailed down. So: zero data retention on by default, no training on your prompts and outputs, and tight access controls are baseline now, not extras.

One practical note: “safe” also means your workflow fits real life. If associates can’t pick the right client/matter in two clicks, prompts get misfiled and audits get ugly. Bake process into the product, matter pickers, always-on DLP/PII redaction, and logs you can export without wrestling a CSV. That combo reduces leaks and speeds client approvals.

What “safe” means for law firms in 2025

Think in three parts: confidentiality (no unauthorized access, no surprise retention, no model training on your inputs), integrity (traceable, reviewable outputs with clear logs), and availability (reliable service, redundancy, tested recovery).

Professional duties still rule the day. ABA Model Rule 1.1 (tech competence) and Rule 1.6(c) (reasonable efforts to prevent disclosure) set the floor, and Formal Opinion 477R pushes risk-based security for protected client info. OCGs now regularly require SOC 2 Type II, ISO 27001, exportable audit logs, and DPA/SCC readiness. Many RFPs ask outright about customer-managed keys (CMK) and regional processing options.

Reality check: in 2024, several Am Law RFPs insisted on “no training on your data” by default and “audit logs with SIEM integration.” Vendors without immutable logs or per-matter segregation didn’t make the cut. Try this test: could you deliver a client audit pack in 48 hours, exports, DLP policy proof, and a subprocessor list? If not, keep tuning.

Regulatory and ethical framework to consider

Line up your use with ethics rules and privacy laws in the places your data lives and travels. Ethics: ABA Model Rules 1.1, 1.6, 5.3, plus state guidance on cloud/AI. Privacy: GDPR/UK GDPR (lawful basis, DPIAs, SCCs), CPRA (service provider terms, sensitive data), and HIPAA where relevant (BAA, minimum necessary). The EU AI Act starts phasing in 2025 to 2026 and leans hard on risk management, transparency, logging, and data governance, so build toward that now.

Regulators are clear on the risks. The UK ICO’s guidance flags training on personal data, data minimization, and security testing. The EDPB stresses transparency around model training and honoring data subject rights. If your prompts contain personal data, this all applies. For cross-border work, decide on US/EU/UK residency and transfer mechanisms up front.

Easy habit: treat each AI use like a mini DPIA. Write down purpose, data types, retention, subprocessors, and safeguards (e.g., DLP/PII redaction, zero-retention). Save it with the matter file. When clients or ethics counsel ask “what did you do,” you’ll have receipts.

Data handling standards you should require from any AI copilot

  • Zero data retention by default. Prompts and outputs shouldn’t live anywhere outside your tenant unless you flip an explicit switch.
  • No training on your data. Lock it in contract: your prompts/outputs aren’t used for foundation or fine-tunes unless you opt in.
  • Transparent data flow. Show diagrams, list subprocessors, and document regions in the security pack.
  • Pseudonymized logs with selective redaction. Audit what happened without exposing client secrets.
  • Clear deletion SLAs and encrypted backups. No mystery caches.

By late 2024, most enterprise providers offered “no training on your data” and “zero-retention” modes. Client questionnaires now probe with specifics (e.g., “If someone pastes PII, does any system retain it?”). Expect to prove it.

Two add-ons worth asking for: per-workspace or per-matter keys (KMS/CMK) tied to client retention schedules, and an admin setting that forces a quick classification step (client, matter, sensitivity) before sending. It’s a light touch that massively improves audit quality and prevents misroutes.

Security controls and certifications to demand

  • Encryption: TLS 1.2+ in transit, AES-256 at rest, ideally with customer-managed keys.
  • Identity and access: SSO/SAML, SCIM, granular RBAC, MFA, and least-privilege defaults.
  • Logging: immutable, tamper-evident activity logs with alerts and exports to your SIEM (Splunk, Sentinel, etc.).
  • Testing: independent pen tests (summary under NDA) and a secure SDLC.
  • Certs: SOC 2 Type II and ISO 27001; ISO 27701 is a strong bonus for privacy.

Why this matters: most OCGs now ask for SOC 2/ISO evidence and recent pen-test summaries. IBM’s 2024 report also found that companies using security AI and automation shortened breach lifecycles by roughly 100 days, which translates to real savings. Filters for prompt injection and anomaly alerting help here.

Don’t forget the admin layer: require dual-control for sensitive changes (like disabling DLP) and “break-glass” procedures. Log who changed what and when, then ship those events to the SIEM. Re-certify access regularly, stale accounts and broad roles are avoidable risk.

Deployment models and data residency options

  • Multi-tenant SaaS with strong isolation and zero retention: quickest to roll out; fine for many commercial matters.
  • Single-tenant or private VPC/VNet: tighter isolation, dedicated resources, and precise regional control.
  • On-prem or air-gapped: rare but useful for export-controlled or national security-adjacent work.

Data residency in the US, EU, or UK can be a dealbreaker. Some clients demand in-region processing and backups; others accept SCCs with documented safeguards. Scrutiny of transfers isn’t fading, so pick a footprint that avoids headaches instead of lawyering around them.

Common pattern: an EU-based diligence room for M&A, and a US region for litigation, both using customer-managed keys and zero-retention. Per-matter workspaces keep barriers clear.

One contract tweak that pays off: add “portability on termination.” You get full exports of audit logs, configs, and matter spaces in a machine-readable format. Less lock-in, more confidence.

Preserving privilege and client confidentiality

Privilege usually holds when a vendor acts as your agent and you take reasonable precautions. Bar opinions on cloud services back this up when you have confidentiality, access controls, and supervision in place. In 2025, “reasonable” means strong contracts, zero-retention by default, no training on your data, and clear internal policies.

Make it tangible:

  • Use DPAs and NDAs that bind subprocessors and forbid secondary use.
  • Segregate work by client/matter, and limit cross-matter search.
  • Require tagging with matter IDs and sensitivity labels for every interaction.

Example: a litigation team preps for depos inside a matter workspace. DLP rules block SSNs or health data from leaving firm boundaries. Audit logs show who touched what. If challenged later, you’re covered.

One habit helps a lot: keep a “privileged vendor” registry and note it in your guidance. When your logs and memos clearly show the vendor’s agency role, privilege fights get easier, and client audits go faster.

Governance, safety, and misuse prevention

Governance is the harness. Enforce DLP, PII detection/redaction, content filters, and defenses against prompt injection and data exfiltration. Keep a human in the loop for high-impact tasks, and add approval gates for automations that can touch client systems.

Build a “trust but verify” rhythm:

  • Require citations for legal conclusions.
  • Test model or prompt changes in staging with red-team scenarios.
  • Track quality with a simple metric like “corrections per 100 outputs.”

NIST’s AI Risk Management Framework (2023) pushes impact assessments, measurement, and ongoing monitoring. Regulators expect proof that your safety filters actually work, especially when personal data is in play.

Try “matter sensitivity budgets.” For hot matters, cap the number of automated steps without human review and require a second reviewer when outputs cite sources outside your record. It feels like second-partner review, adapted for AI.

Vendor due diligence and contracting checklist

Run procurement like you mean it:

  • Security pack: SOC 2 Type II, ISO 27001, pen-test summary, secure SDLC.
  • Data processing: DPA, subprocessor list with notice, SCCs if needed, and data flow diagrams.
  • Controls: SSO/SAML, SCIM, granular RBAC, CMK/KMS, DLP/PII redaction, audit logs with SIEM export.
  • Operations: SLAs, uptime targets, RTO/RPO, breach notification windows (e.g., 72 hours), incident playbooks.
  • Product safety: prompt-injection defenses, content filters, model governance details.

RFPs in 2024 to 2025 started to standardize: “no training on your data,” “zero-retention by default,” and “evidence of regional processing.” Clients won’t accept promises, they want proof.

Add a “configuration annex” to the contract listing the exact privacy/safety toggles you’ve enabled (retention, training, residency, logging). Any change requires written approval. It’s your snapshot in time for audits and renewals.

Operational rollout and training plan

Start small and intentional. Pick a few use cases (brief outlines, deposition prep, clause analysis). Agree on success metrics (time saved, quality, attorney satisfaction). Lock guardrails (zero retention, matter tagging, citations). Red-team with real prompts. Write it down.

Set up clean plumbing: per-matter workspaces, data barriers, least-privilege roles. Provision users via SCIM. Turn on DLP/PII redaction globally. Pipe logs to the SIEM on day one so security can actually see what’s happening.

Teach the basics fast: what not to paste, how to tag, how to ask for new workflows. A two-page “Prompting with Privilege” guide beats a marathon webinar. Reinforce inside the product with small nudges.

Nice quality-of-life move: save AI “session wrap-ups” to your DMS, final output, matter metadata, and a link to the audit trail. Work product and evidence, side by side.

Questions to ask before approving any AI copilot

  • Is zero-retention enforced by default across all workspaces and tenants?
  • Will you contractually guarantee no training on our prompts/outputs unless we opt in?
  • What deployment options are available (multi-tenant, single-tenant, private VPC, on-prem)?
  • Do you support regional processing and backups in the US, EU, and UK?
  • Which subprocessors touch our data, why, and how are they audited?
  • What identity controls are in place (SSO/SAML, SCIM, MFA, granular RBAC)?
  • What audit logs exist, are they immutable, and how do they integrate with our SIEM?
  • How do you defend against prompt injection and data exfiltration, and can we see red-team results under NDA?
  • What are your SLAs, breach notification timelines, and RTO/RPO?
  • Do you offer customer-managed keys and per-workspace encryption?

Ask for a full security pack, sample logs, and a live admin demo. Then run a 30-day pilot with retention/training locked in the contract and verify behavior in your own logs. Trust, verify, document.

How LegalSoul protects confidential client data

LegalSoul is built for legal confidentiality. You can choose multi-tenant with strong isolation, single-tenant, or private VPC/VNet, with processing regions in the US, EU, and UK to meet data residency needs. Zero-retention is on by default, and we never train on your data unless you opt in.

Access and security align to enterprise norms: SSO/SAML with MFA, SCIM provisioning, and fine-grained RBAC down to client, matter, and practice group. Customer-managed keys and per-workspace encryption let you match client retention schedules. Every action is logged and exportable to your SIEM.

Governance features include DLP and PII redaction, defenses against prompt injection, abuse monitoring, and optional human-review steps. Certifications: SOC 2 Type II and ISO 27001, with independent pen-test summaries under NDA. Our security pack covers data flows, subprocessors, and model governance.

Small touch, big payoff: a matter-aware UI that asks for client/matter and sensitivity before sending. Clean, defensible audit trails without slowing anyone down.

Risk assessment matrix and decision guide

Match sensitivity to setup:

  • High sensitivity (bet-the-company disputes, export-controlled, PHI): private VPC or single-tenant, regional processing, CMK, zero retention, no training on your data, strict RBAC, dual-control for admins, human-in-the-loop, and minimal external connectivity.
  • Medium sensitivity (complex commercial, employment): hardened multi-tenant, zero retention, DLP/PII redaction, per-matter workspaces, citations required, model change control.
  • Low sensitivity (internal research, marketing): standard controls and a firm “no client data” rule.

Go/no-go red flags:

  • No guaranteed zero-retention by default.
  • Training on your prompts/outputs without explicit opt-in.
  • Missing immutable audit logs or no SIEM export.
  • No clear subprocessor list or processing regions.

Flow it like this: quick DPIA for the use case, pick deployment by sensitivity and geography, run a 30-day pilot to verify controls, then lock a configuration annex. Keep the artifacts, questionnaires, pen-test summary, DPA, subprocessor list, sample logs, in the matter file so you’re audit-ready.

Common pitfalls to avoid

  • Using consumer accounts or public playgrounds with client data. Defaults there are not your friend.
  • Leaving default retention on or allowing silent model training.
  • Weak identity practices. Without SSO/SAML, SCIM, and good RBAC, you’ll collect stale accounts and broad access.
  • No auditability. If you can’t export logs to your SIEM on demand, you can’t answer hard questions.
  • Skipping red-team testing. Prompt injection and data exfiltration aren’t theoretical, test them.
  • Undertraining staff. Most mishaps are process errors: wrong matter, unnecessary PII, trusting outputs without checks.

Quick win: make the matter picker and sensitivity label mandatory, and keep DLP/PII redaction always on. Most accidents stop right at input.

Key Points

  • Safety is about controls, not brands: zero data retention by default, no training on your data, SSO/SAML/SCIM with granular RBAC, encryption, immutable audit logs with SIEM exports, and SOC 2 Type II/ISO 27001.
  • Ethics and privacy: meet ABA Model Rules 1.1/1.6/5.3 and GDPR/CPRA; keep privilege by treating the vendor as your agent, using DPAs/NDAs, regional processing, and per-matter segregation.
  • Deployment and governance: pick multi-tenant/private VPC/on-prem based on sensitivity and geography; enable DLP/PII redaction, prompt-injection defenses, human review, and ongoing red-teaming.
  • Buying and rollout: use a tight due-diligence checklist (subprocessors, CMK/KMS, SLAs, incident windows), run a 30-day pilot with locked settings, and capture a configuration annex; LegalSoul ships these controls ready to go.

Conclusion: Is an AI copilot safe for confidential matters in 2025?

Yes, when you control the setup. You need private or regional deployment, zero data retention, no training on your data, strong identity, encryption, and immutable logs. Layer on DLP/PII redaction and tested defenses against prompt injection and exfiltration, and document your choices to satisfy ethics rules and OCGs.

Next steps: run a scoped pilot with locked privacy settings, verify everything in your SIEM, and save a configuration annex you can reuse in audits. Want a platform built for legal work with these controls on by default? Try LegalSoul. You’ll get flexible deployments (including private VPC), zero-retention defaults, no model training on your data, enterprise-grade access/logging, and governance that matches how lawyers actually work.

Related reading before you sign

If you are still building the shortlist, the Harvey AI alternative comparison for small and midsize law firms puts the pricing, seat minimums and security controls side by side. If Harvey is not the only quote on your desk, the CoCounsel pricing breakdown does the same job for the Thomson Reuters side, including what is disclosed and what is only reported. For the research half of the stack, see the comparison of Westlaw Precision AI, Lexis+ AI and Bloomberg Law AI, and for drafting, the contract review tool comparison. Our own rates are on the pricing page, in public, with no quote process.

Unlock professional-grade AI solutions for your legal practice

Sign up