Trimio Field Notes

Your AI Audit Trail Has a Hole in It: The Encrypted Reasoning Problem

June 2, 2026 6 min read securitygovernancecomplianceagentic-ai

Your AI agents are reasoning before they act. They're weighing options, evaluating tool calls, forming plans — and then executing against your production systems. When something goes wrong, your compliance team will ask: what was the model thinking? The answer, at every major AI provider today, is that you can't know. The thinking happened in an encrypted block that was consumed and discarded before the response arrived.

This is the encrypted reasoning problem. It's not a theoretical gap. It's a concrete hole in your AI audit trail — one that regulators are starting to notice, and that your security team almost certainly hasn't mapped yet.

0
Reasoning tokens visible via API
OpenAI reasoning models generate and bill reasoning tokens, but the content is fully discarded — you get the count, not the content.
Art. 22
GDPR Explainability Requirement
GDPR requires "meaningful information about the logic involved" in automated decisions that significantly affect individuals.
25,000
Min. tokens reserved for reasoning
OpenAI recommends reserving at least 25,000 tokens per reasoning call — a budget that can exceed your entire output token count.

The thinking happens in a black box

Essential
Both Anthropic and OpenAI reasoning models produce internal reasoning traces that are structurally inaccessible to the API caller. You pay for them, but you cannot read them — by design.

When you call a reasoning model today — Claude with Extended Thinking enabled, or OpenAI's GPT-5.5 — the model runs an internal deliberation step before generating its response. That deliberation is not invisible to the model: it shapes every downstream decision, tool call, and output the model produces. But it is invisible to you.

The two providers have implemented this differently, but the practical result is the same.

Anthropic (Extended Thinking): The API returns a thinking content block alongside the final text block. That thinking block contains a signature field — an opaque cryptographic blob that Anthropic uses for internal integrity verification. The actual reasoning content may be summarized or omitted entirely depending on the model and configuration. Newer models like Claude Opus 4.7 have deprecated manual extended thinking in favor of adaptive thinking, where the model decides autonomously how much to think — and on Claude Mythos, the display parameter defaults to "omitted", meaning you receive no reasoning content at all unless you explicitly request a summary. (Anthropic Extended Thinking docs)

OpenAI (Reasoning Models): GPT-5.5 and related reasoning models use internal reasoning tokens to "think" before responding. Per the official API documentation: "reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens." You can see the count of reasoning tokens consumed in the output_tokens_details.reasoning_tokens field. You cannot see the content — ever. (OpenAI Reasoning Models docs)

The architecture is intentional. Providers encrypt or discard internal reasoning to prevent prompt injection attacks that exploit visible chain-of-thought, to prevent users from fine-tuning on internal reasoning shortcuts, and to preserve model integrity across multi-turn conversations. The security rationale is sound. The compliance gap it creates is real.

You're billed for tokens you'll never read

Essential
Reasoning tokens are billed as output tokens at full rate. On GPT-5.5 at xhigh effort, the model may burn tens of thousands of reasoning tokens on a single call. You see the dollar amount. You don't see the deliberation that drove it.

This is the FinOps dimension of the problem, and it compounds the compliance gap.

OpenAI recommends reserving at least 25,000 tokens for reasoning and outputs when experimenting with reasoning models. At xhigh effort — recommended for "security and code review, enterprise productivity, and challenging coding workflows" — the model can generate far more than that. Every reasoning token is billed at the standard output token rate.

On GPT-5.5 at the current API rate of $15/M output tokens, 25,000 reasoning tokens per call is $0.375 in hidden deliberation cost — before the model has written a single visible character of response. Across an agentic workflow that makes 50 model calls per task, that's $18.75 in reasoning-token overhead per task execution, none of it recoverable in your audit log.

The billing and the auditability problems are connected: you're paying for a reasoning process you cannot observe, inspect, or reconstruct. For a routine developer task, that's a cost annoyance. For an automated decision touching a customer's account, a contract approval, or a security triage, it's a governance failure waiting for a regulator to notice.

What compliance frameworks actually require

Essential
GDPR Article 22, SOC 2 CC6.1, and ISO 27001 A.12.4 all establish requirements that AI's encrypted reasoning directly undermines — not because the regulations say "no encrypted reasoning," but because they require auditable decision logic.

Three frameworks your legal or compliance team is likely already working against:

GDPR Article 22 establishes the right to explanation for automated decisions. Article 22(3) requires that controllers "implement suitable measures to safeguard the data subject's rights and freedoms and legitimate interests, at least the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision." Critically, Article 22 requires "meaningful information about the logic involved" in automated decision-making. When the logic is encrypted and inaccessible, that requirement is structurally unmet — not through any explicit prohibition, but because you cannot produce the explanation a data subject can request.

SOC 2 CC6.1 (Common Criteria for Logical and Physical Access Controls) and the broader SOC 2 Trust Services Criteria require that entities identify and document significant system processes and controls. An AI agent that makes decisions through a reasoning process you cannot log is a control gap. Auditors increasingly treat AI-assisted decisions like any other automated control — they want to see the decision logic, the inputs, and the outputs.

ISO 27001 Annex A.12.4 (Logging and Monitoring) requires that event logs be produced, protected, and regularly reviewed. Tool calls made by a reasoning model, triggered by reasoning you can't read, are events. If you don't log the tool call triggers — and cannot reconstruct why the model decided to call a given tool — your logging posture has a structural gap.

None of these frameworks explicitly say "you must capture model reasoning tokens." What they require is auditable decision logic. Encrypted reasoning creates a zone where decisions are made and cannot be reconstructed. That's the compliance gap.

The proxy layer captures what the model reveals

Essential
You can't capture encrypted reasoning. But a proxy layer captures everything before and after the black box: the full prompt, every tool call the model made, every output produced. That observable record is your auditable decision trail.

The encrypted reasoning problem has a practical mitigation that doesn't require cracking provider encryption: shift your audit surface from the reasoning process to the observable behavior.

Here's what you actually need in an AI audit log:

None of these are inside the encrypted reasoning block. All of them are observable at the API boundary — which is exactly where an LLM proxy layer sits.

The reasoning trace tells you how the model thought. The proxy log tells you what the model did: what it was given, what actions it chose, and what it produced. For audit and compliance purposes, the behavioral record is both more reliable and more actionable than a reasoning trace would be. You can reproduce the context. You can replay the tool call sequence. You can trace the output back to a specific prompt and a specific authorization.

This is not a complete substitute for reasoning transparency — researchers and model developers legitimately want reasoning traces. But for enterprise AI governance, the observable behavioral log is the right audit surface. It's what a human reviewer would reconstruct anyway if they were investigating a decision.

Zero Data Retention makes the gap permanent

Essential
If your org has a Zero Data Retention arrangement with Anthropic, the reasoning block isn't just inaccessible to you — it doesn't exist anywhere after the response is returned. You must capture the observable record at the API boundary, or it's gone.

Anthropic's Extended Thinking feature is eligible for Zero Data Retention (ZDR). Under a ZDR arrangement, data sent through the feature "is not stored after the API response is returned." That applies to the thinking blocks.

This means that if your organization has a ZDR agreement — which is common in financial services and healthcare — there is no Anthropic-side copy of the reasoning trace to subpoena, reconstruct, or audit. The reasoning happened, influenced the output, and is now gone from every system in the world except for whatever you captured at the API boundary before the connection closed.

If you're not logging at the proxy layer, that means your audit trail for every ZDR-covered reasoning call is: inputs in, outputs out, nothing in between. For a routine summarization task, that's fine. For an agent approving a transaction, routing a support escalation, or generating a legal document, that's a governance gap you cannot patch retroactively.

The right time to wire the proxy logging layer is before the first production reasoning call — not after the first audit finding.

What your AI governance team should ask right now

Essential
Three concrete questions to close the encrypted reasoning audit gap before it becomes a compliance event.

Three diagnostic questions worth walking through this week:

  1. Which of our AI-assisted workflows use reasoning models? Identify every deployment using Claude with Extended Thinking, any OpenAI o-series or GPT-5.x model, or any model with adaptive thinking enabled. These are the workflows with a reasoning blind spot. If you don't have a complete inventory, you don't know the scope of the gap.
  2. What are we logging at the API boundary today? For each reasoning-model workflow, confirm: are we capturing the full prompt, the full tool call sequence, and the final output in a durable, queryable log — not just a provider dashboard? If the answer is "we can pull it from the provider console," that's not an audit log. Provider consoles have retention limits, access restrictions, and no query layer.
  3. Do we have a ZDR arrangement, and do we understand what it covers? ZDR is worth having for data minimization reasons. But it changes the audit posture. If you have ZDR, the proxy log is the only copy. Confirm that every ZDR-covered workflow is logging observable behavior before the ZDR kicks in — i.e., at the API boundary, not downstream.

The encrypted reasoning problem isn't going away — providers have strong incentives to keep reasoning opaque. But the governance gap it creates is solvable at the infrastructure layer, not the model layer. The proxy is your observation point for the decisions your AI makes. Build the logging there, and the encrypted black box stops being an audit liability.

Trimio is the LLM API gateway purpose-built for AI cost governance and observability. Every call — including tool invocations made by reasoning models — is logged, attributed, and queryable. See how it works.

Trimio
Stop guessing. Start governing.
trimio is the LLM API gateway purpose-built for AI cost governance — visibility, routing, caching, and budget enforcement in one layer.