Does SOC 2 cover AI? Not with a dedicated Trust Services Criterion, no. SOC 2's five criteria, Security, Availability, Processing Integrity, Confidentiality, and Privacy, were finalized years before generative AI became a normal part of how software gets built. There's no AI checkbox waiting to be checked, no sixth criterion added for language models.
But an engineer who pastes three months of support tickets into ChatGPT to summarize a complaint pattern, then drops the output into a board deck, has just created exactly the kind of evidence gap a SOC 2 auditor now asks about. Nobody logged the prompt. Nobody reviewed what left the building. The system boundary diagram nobody updated doesn't show any of it happened.
The honest answer is that SOC 2 itself hasn't changed, but what auditors ask about has. Auditors are reading the same five criteria through an AI lens: are AI tools access-controlled the same way any other system is, is a hosted model provider treated as a vendor with real risk, is training data and prompt content handled under the same confidentiality rules as any other sensitive data. None of that requires a new criterion. All of it requires new evidence.
This is about what changes in your own audit scope and evidence, not which compliance platform to buy. For a vendor-by-vendor comparison, see the breakdown of the top SOC 2 compliance platforms for AI companies.
By the end of this, you'll know exactly which Trust Services Criteria actually touch AI, when a model provider becomes a subservice organization, why shadow AI is turning into the most common finding in SOC 2 audits, and what a SOC 2+ report even is.
Here's what I'll cover:
- Why SOC 2 has no dedicated AI criterion, and what that actually means in practice
- How the stakes and timing have shifted over the past two years
- A full Trust Services Criteria to AI-risk mapping
- When a foundation model provider becomes a subservice organization
- Shadow AI, and why it's already the most common AI-related finding
- What auditors are actually asking about AI systems in 2026
- The SOC 2+ / ISO 42001 stacking option almost nobody talks about
Does SOC 2 Cover AI With a Dedicated Criterion? Not Yet: Here's Why
SOC 2 is built around five Trust Services Criteria: Security, Availability, Processing Integrity, Confidentiality, and Privacy. Every control in a SOC 2 report maps back to one of these, and an auditor's entire job is testing whether the company's actual controls satisfy them.
None of the five were written with a language model in mind. Schellman, one of the audit firms that actually performs these examinations, states this plainly: SOC 2 was never built to test "fairness, bias, responsible and ethical use, and safety." That's an honest scope boundary, not a gap to apologize for.
ISO 42001, the AI management system standard, exists specifically to test those questions. SOC 2 does not, and treating it otherwise in a sales conversation or a report is the kind of overstatement that gets caught during due diligence.
So does SOC 2 cover AI? Not with its own line item, no. What's actually happening is narrower, and in practice more consequential: auditors are interpreting the same five criteria against a new category of evidence. Pentest Testing Corp frames it well: auditors aren't inventing new checkboxes, they're asking companies to demonstrate that criteria they've reported against for years still hold up when the system in question is a model instead of a database.
That's a real distinction from a platform-selection question. This article covers what changes in your own audit scope and evidence once AI enters the picture, not which compliance platform handles that best. The comparison of SOC 2 platforms built for AI companies covers that separately.
Does SOC 2 Cover AI Differently Than It Did Two Years Ago?
Two years ago, an AI startup could walk into a SOC 2 Type 2 audit scoped like any other SaaS company and mostly get away with it. That's getting harder. An auditor who doesn't ask the right AI-specific questions produces a report that won't hold up to real enterprise due diligence later. One who does ask, and finds undocumented gaps mid-engagement, turns a routine audit into a scramble.
Enterprise security questionnaires are catching up too. A buyer that asked only about SOC 2 status two years ago is increasingly likely to ask a follow-up question about which AI tools touch their data specifically, and a vague answer reads as badly as no answer at all.
Cost is the other place the stakes show up, and here the honest answer is a range with real sourcing gaps rather than a clean number. SOC2Auditors.org puts first-year Type 2 spend for an AI startup at $40,000 to $120,000, citing a source it names only as "Knowlee, 2026," with no further detail available to verify independently.
LastPass, working from a breakdown by auditor tier rather than AI-specific scope, puts a wider range of $15,000 to $450,000. Treat both as directional, not a quote: the real number for any given company depends far more on audit firm, scope, and how much AI-specific evidence already exists than either figure fully captures.
Mapping AI Risk to the Trust Services Criteria
Here's the version of "does SOC 2 cover AI" that actually matters in practice: not whether a sixth criterion exists, but how the five that do exist get tested once AI is part of the picture.
| Trust Services Criterion | Traditional test | AI-era test |
|---|---|---|
| Security (CC6.1, CC6.8) | Access controls, malware protection | Who can access the model API, fine-tuning environment, and vector store; controls against model theft |
| Availability (A1.2) | Uptime, backup and recovery | Contingency planning for an outage at your LLM provider, since you don't control OpenAI's or Anthropic's uptime |
| Processing Integrity (PI1.2, PI1.4) | Data processed completely and accurately | Model drift monitoring, output-reliability baselines, and tested rollback procedures for model updates |
| Confidentiality (C1.2) | Data classified and protected | How training data and prompt/inference logs are handled, including redacting PII before storage, not after |
| Privacy (P1.1) | Notice and consent for personal data | Whether customer data used in a prompt or fine-tuning run was disclosed as an AI use in the first place |
A soc 2 trust services criteria ai review, done properly, walks each of these five rows individually. It isn't one generic AI checkbox bolted onto the report, which is exactly why so few pages covering this topic map it this cleanly.
Two rows deserve a closer look, because they're the ones companies most often under-scope. A soc 2 processing integrity ai review tests whether model outputs match documented, expected behavior, not whether the model is accurate in some abstract sense. An auditor wants evidence of drift monitoring and a tested rollback path, not a claim that the model "works well."
A soc 2 confidentiality ai review is really a review of training data and prompt logs specifically, a different and often thinner control than the database encryption most companies already have covered.
The AI Controls SOC 2 Auditors Now Test For
Pentest Testing Corp lists the concrete evidence auditors are actually asking for in 2026, and it's a useful checklist regardless of which platform or process produces it:
- A model registry with version history and named approver sign-off
- Prompt and inference logs with PII redacted before storage, not retroactively
- Tested rollback procedures for model updates, not just a plan on paper
- Drift-monitoring output with defined baselines and alert thresholds
- Documented accountability for what an autonomous agent did and why
GoTeleport adds one more angle worth naming specifically: for AI agents that take action without a human in the loop, auditors increasingly treat "no human request" itself as a major accountability gap, not a neutral fact about how the system works.
Is Your LLM Subservice Organization Actually In Scope?
A subservice organization, in SOC 2 terms, is a third party whose controls are relevant to your own system boundary because it's performing part of the service you deliver. Payment processors and cloud hosts have played this role for years. The question now is whether a foundation model provider does too.
SOC2Auditors.org answers this plainly in its own FAQ: "Your LLM provider is a subprocessor, not part of your own SOC 2 system boundary. Your SOC 2 attests to your controls over how you use them."
The real test isn't whether you use OpenAI or Anthropic at all. It's whether the model is core to how the service actually gets delivered, or just a convenience feature layered on top. A support tool that occasionally summarizes tickets with a hosted model is a different risk profile from a product whose core function runs on that model's output.
The core question in any soc 2 subservice organization ai review is exactly this: is the provider integral to delivery, or incidental to it.
Pentest Testing Corp names the common mistake directly: "Treating a frontier LLM provider like any other SaaS tool in your vendor register is one of the more common gaps we see." A vendor register that lists a foundation model provider the same way it lists a scheduling tool misses the point of vendor risk assessment entirely.
The mistake isn't using a foundation model. It's putting it in the vendor register next to a scheduling app and calling the risk assessment done. — Upendra Varma, CTO at ComplyJet
LLM Vendor Risk Assessment and the AI Supply Chain Risk SOC 2 Auditors Are Now Flagging
Linford & Co. introduces a useful frame here: fourth-party risk, meaning a SaaS vendor's own dependency on a foundation model becomes something you're implicitly relying on too.
Two named 2026 events illustrate why this stopped being theoretical, according to Linford's own citations. The Department of Defense reportedly designated Anthropic a supply chain risk in March 2026, per reporting Linford attributes to Politico and CBS News. OpenAI reportedly shut down its Sora product roughly 15 months after launch that same month, per reporting Linford attributes to CNN.
Both are worth independently verifying before repeating as settled fact, but they illustrate a real, non-hypothetical point: foundation model providers can change or disappear on a timeline a company doesn't control.
Linford also offers a genuinely practical detection method for undocumented AI vendor use: compare documented AI token usage against actual billing. A gap of roughly 3x between what's logged and what's billed is a real signal that someone is using an AI tool the vendor risk process never caught. The same technique doubles as a shadow AI detection method, covered next.
One more honest note, sourced directly to Linford: "The AICPA hasn't yet provided formal guidance on how to handle [AI]." That's not a gap in this article, it's the actual state of the standard as of this writing, and it's worth knowing rather than assuming formal guidance already exists somewhere.
Does SOC 2 Cover AI Tools Like ChatGPT and Claude? Shadow AI SOC 2 Gaps, Explained
Using ChatGPT or Claude doesn't automatically break a SOC 2 report. Using them without anyone in the company knowing does.
Sonomos states the actual finding pattern plainly: "The most common AI-related finding in SOC 2 audits is not that the organization approved AI tools without proper controls, it is that employees are using AI tools that the organization has not reviewed or approved at all." That's shadow AI, and it's a genuinely different problem from the shadow SaaS auditors have dealt with for years.
Linford's shadow AI research explains why it's harder to catch: an unauthorized SaaS signup at least shows up somewhere, a new login, a new charge, an entry in an identity provider's logs. Shadow AI use often doesn't. An employee pastes data into a browser extension or logs into a personal ChatGPT account through OAuth, and none of it touches the systems a security team normally watches.
The instinct to just ban AI tools outright doesn't hold up either. A blanket ban pushes usage further underground rather than eliminating it, since the underlying work people are trying to do with the tool doesn't go away. The fix Sonomos describes, policy, inventory, contracts, DLP, evidence, works because it manages the risk instead of pretending the tool isn't there.
Consumer vs. Enterprise AI Tools: Why the Tier You Use Changes the Answer
This is the clearest differentiator in how we're covering this topic, and it's the direct answer to two of the most common questions people actually ask: is ChatGPT SOC 2 compliant, and is Claude SOC 2 compliant.
Neither question has a single answer, because it depends entirely on which tier is deployed. Consumer-tier ChatGPT, Claude, Gemini, and Copilot generally don't come with a Data Processing Agreement, and inputs can be used to train the underlying model. Enterprise tiers of the same tools typically do provide a DPA and a contractual no-training guarantee.
An auditor who understands this asks a more specific question than "is AI allowed here." They ask which tier is actually deployed, and whether that tool is covered by the vendor risk management program at all. A team running the free version of ChatGPT for internal notes and a team running ChatGPT Enterprise under a signed DPA are not the same risk, even if both would answer "yes, we use ChatGPT" to the same question.
What SOC 2 Auditors Are Actually Asking About AI in 2026
Pulling together real, attributed questions rather than a generic checklist, here's what auditors are actually asking, sourced to Sonomos and LastPass:
- Does the organization have a policy addressing employee use of generative AI tools?
- Are AI vendors included in the vendor risk management program, the same way any other vendor is?
- Who authorized the AI tool or agent, specifically?
- What data and systems can it access, and why does it need that access?
- Are its actions traceable after the fact, or does the trail go cold?
Sonomos frames the practical response as a five-step control stack: policy, then inventory, then contracts, then data loss prevention, then evidence. That order matters. A policy with no inventory behind it is aspirational. An inventory with no contract review behind it just documents risk without managing it.
The OWASP LLM Top 10 Angle: Why LLM02 and LLM06 Come Up in Audits
One genuinely useful crosswalk, adapted here from Pentest Testing Corp's use of it, connects OWASP's Top 10 for LLM Applications directly to SOC 2 criteria. Two categories show up in audit conversations specifically: LLM02, Sensitive Information Disclosure, maps onto the Confidentiality and Privacy criteria, and LLM06, Excessive Agency, maps onto the access-control and monitoring controls under Security (CC6.1, CC7.2).
Picture a support chatbot built on retrieval-augmented generation. A cleverly worded support ticket triggers the model to retrieve and surface a different customer's data in its response. That single failure demonstrates LLM02 and LLM06 at the same time: sensitive information disclosed, through a system given more retrieval agency than the situation actually called for.
It's a useful mental model for scoping AI controls even outside a formal OWASP review: ask what the system could disclose, and separately, what it could do without being asked.
ISO 42001 SOC 2 Stacking: How Far Do You Need to Go With AI Governance?
This is the one real differentiator almost nobody covering this topic operationalizes, including the most thorough pages written about it.
Schellman introduces a concept it calls SOC 2+: stacking ISO 42001's Annex A, 38 AI-governance controls, into Section 4 of a standard SOC 2 report. The idea is straightforward even if the execution isn't trivial: a company demonstrates AI governance maturity inside a report it's already producing, without pursuing full, separate ISO 42001 certification.
Think of it as a middle option. On one end, a company says nothing about AI governance in its SOC 2 at all. On the other, it pursues full ISO 42001 certification, a real undertaking with its own audit cycle and cost. SOC 2+ sits between the two: useful for a company whose AI use is real and worth demonstrating governance over, but not yet central enough to the business to justify a standalone certification.
A soc 2 plus ai governance approach along these lines doesn't replace ISO 42001 certification. It demonstrates real governance maturity inside a report the company is already producing, which is a meaningfully lower lift than starting a second certification from scratch.
Common Mistakes Companies Make Scoping AI Into a SOC 2 Audit
- Treating a frontier LLM provider like any other SaaS tool in the vendor register instead of evaluating whether it's actually a subservice organization.
- Assuming SOC 2 covers AI bias, fairness, or safety. It doesn't, and stating otherwise in a sales conversation or a report is a real overstatement risk, not a rounding error.
- Standing up AI-specific controls right before the audit window opens, without the 6 to 12 months of operating history Type II evidence actually requires. A control with no track record reads to an auditor the same as no control at all.
- Ignoring shadow AI because it doesn't show up in firewall logs the way an unauthorized SaaS signup does. The token-usage-versus-billing check covered earlier catches exactly this kind of gap.
- Not distinguishing consumer-tier from enterprise-tier AI tool usage when documenting the vendor risk management program, as if every version of the same tool carried the same risk.
- Conflating "we use AI" with "our LLM vendor is a subservice organization." The two are different questions with different evidence requirements, and answering one doesn't answer the other.
- Waiting for formal AICPA guidance before doing anything. None exists yet, and auditors are already testing against the existing five criteria in the meantime. Waiting is its own risk.
Where ComplyJet Fits In: Does SOC 2 Cover AI Vendor Risk in Its Workflow?
That's the practical version of everything above: a foundation model provider treated the same way any other vendor with real risk would be, evidence collected as controls actually run rather than reconstructed the week before an audit, and one flat, per-company price that doesn't change as a team's AI usage grows.
FAQs
Does SOC 2 Cover AI Explicitly, or Just Interpret Existing Criteria?
No dedicated AI criterion exists. Auditors apply the existing five Trust Services Criteria, Security, Availability, Processing Integrity, Confidentiality, and Privacy, to AI-specific evidence instead of testing against a separate AI standard.
Is ChatGPT SOC 2 Compliant?
It depends on the tier. Consumer ChatGPT generally lacks a Data Processing Agreement and a training-opt-out guarantee, while ChatGPT Enterprise typically provides both. Ask which one is actually deployed before answering the question either way.
Is Claude SOC 2 Compliant?
Same answer as ChatGPT: the consumer tier and the enterprise or API tier carry different DPA and no-training contract terms, and the answer depends entirely on which one a team is actually using.
Is My AI Vendor a Subservice Organization?
If the vendor's model is core to how you deliver your service, not just an internal productivity tool, treat it as a subservice organization in your vendor risk program. A product whose core feature runs on a hosted model is the clearest case. If it's incidental to delivery, like a tool used only for internal notes, standard vendor review still applies, just with less weight.
Does SOC 2 Cover AI Risk From Shadow AI Tools?
Indirectly, yes. Unmanaged AI tool use is commonly flagged as an access-control and vendor-risk gap under the existing criteria, even without an AI-specific rule to point to.
What Do SOC 2 Auditors Ask About AI Systems?
Whether a policy exists for employee AI use, whether AI vendors sit in the vendor risk management program, who authorized each tool, what data and systems it can access, and whether its actions are traceable after the fact.
Can You Audit an AI Black Box?
Auditors can't audit model internals directly, but they can audit the controls around it: access, monitoring, vendor contracts, and drift detection. That's what a SOC 2 actually tests, model or no model. The model's internal weights stay a black box either way; what changes is whether the controls surrounding it produce real, testable evidence.
Does SOC 2 Cover AI Bias or Fairness?
No. SOC 2 wasn't designed to test fairness, bias, or ethical AI use, and that's an explicit scope boundary rather than an oversight. ISO 42001 is the more relevant standard for that specific question.
Related Reading
- Top SOC 2 Compliance Platforms for AI Companies — for readers deciding which compliance platform to use, not what changes in their own audit scope
- Is SOC 2 a Certification? — foundational SOC 2 scope and terminology background
- SOC 2 Audit Process, Start to Finish — the end-to-end audit process this article's evidence requirements feed into
- SOC 1 vs SOC 2 — for readers unsure which report type applies to them at all
- What Is Compliance Automation? How It Works — the broader automation-tooling context behind the evidence collection this article describes


