Trimio Field Notes

The Cache Collapse Problem: How AI Proxies Silently Restart Your Sessions

May 24, 2026 7 min read cachingmulti-providersession-management

Turn 9 of a 20-turn agent session. Claude Sonnet, routed via least-cost routing to Fireworks/Qwen3 for the cheaper turns. Everything working as expected — until the model says:

"Let me check on that — I want to make sure I have the full context before proceeding."

It said it seven more times. The agent had effectively restarted. Not from a model failure. Not from a prompt engineering mistake. From a cache collapse — the proxy's compression layer had silently rewritten messages inside the active cache prefix, collapsing 88k → 23k cached read tokens (a 73% reduction), forcing the model to rebuild context from scratch on a session already 9 turns deep.

This is the kind of failure that looks like flaky AI behavior but is actually a deterministic infrastructure bug. Here's what happened, and how we fixed it across three PRs.

what a session anchor is and why it matters

Essential
A session anchor is the floor index below which the proxy will not rewrite, compress, or deduplicate messages. It preserves the contiguous cache prefix that the provider has already priced as cache-read tokens. Without it, the compression layer can destroy the cache hit that cost you money to build.

Providers like Anthropic use a cache prefix model: if the first N tokens of your prompt are identical to a previous request, they're billed at the cache-read rate (~10% of normal input cost). The longer your stable prefix, the more you save.

In a multi-turn agent session, the cache prefix grows naturally: system prompt + turn 1 + turn 2 + ... + turn N-1 is re-sent with every new turn, and all of it qualifies as cached after the first turn. A well-managed 20-turn session might have 80k+ tokens in cached reads by turn 10 — at Anthropic's rates, that's roughly $0.24 per turn saved vs. $2.40 in input cost. Across thousands of sessions, it's the difference between a profitable product and a loss leader.

The session anchor is the mechanism that protects this. Concretely, it's an index: do not touch messages below position N. The compression and deduplication pipeline can do whatever it wants above the anchor — but below it, the prefix is frozen.

the root cause: anchor gated on Anthropic-outbound

Essential
The session anchor was coupled to the Anthropic-outbound check. Routing to Fireworks/Qwen3 meant the anchor was never set — so compression ran freely inside the cache prefix, destroying 65k tokens of cached context in a single turn.

In Trimio's proxy, the session anchor was implemented in the same code path that injects Anthropic's cache_control markers. The logic looked roughly like this:

if pcIsAnthropic {
    injectCacheControl(messages)
    SetSessionAnchor(anchorIndex)
}

The problem: pcIsAnthropic was a single flag that conflated two distinct things:

  1. Anthropic-outbound: whether to inject cache_control markers in the wire format (only valid for Anthropic endpoints)
  2. Anthropic-shape: whether the request uses the Anthropic message format (which many providers support, including Fireworks/Qwen3 via their OpenAI-compatible API)

When a session was routed to Fireworks/Qwen3 via LCR, pcIsAnthropic was false — so no anchor was set. The compression layer, seeing no anchor floor, was free to rewrite any message in the session. On turn 9, it compressed a large system prompt segment that happened to fall inside what had been the cache prefix on previous turns. Result: 88k → 23k cached tokens, and a session restart loop.

the compression-rewrites-inside-cache-prefix failure mode

This failure mode is subtle because it requires two conditions to coincide:

  1. A provider route that doesn't set the session anchor
  2. A compression trigger (token budget pressure, dedup scoring above threshold) that fires on messages inside the historical cache prefix window

Condition 2 is more likely than you'd think. The compression layer looks at the full message list and picks the highest-value compression targets — and long, repeated system prompt blocks are prime candidates. Those blocks are also exactly what the cache is built on.

The result is a silent correctness failure: the request succeeds, the model responds, costs look normal — but the session's effective context has been quietly amputated. The 7-message restart loop we observed was the model trying to reconstruct context it no longer had access to via its cache.

the three-PR fix

Essential
Three PRs, cleanly separated: decouple the anchor from the Anthropic gate (PR #622), extend implicit anchor support to all provider shapes (PR #624), and enforce the anchor floor in the compression pipeline (PR #623). Order matters — #622 first, or #624 has nothing to set the anchor against.

PR #622 — decouple session anchor from Anthropic-outbound gate

Split pcIsAnthropic into two flags:

This is the foundational change. Anchor management now runs on message shape, not on provider identity. A Fireworks/Qwen3 request using the Anthropic message format gets anchor tracking — it just doesn't get cache_control injected (which would be invalid on that endpoint).

PR #624 — extend implicit anchor support to all providers

Added isImplicitCachedProvider() — a function that returns true for providers with implicit caching semantics (where the provider caches automatically without requiring cache_control markers in the request). Current list: Fireworks, Qwen, Gemini, OpenAI.

For these providers, SetImplicitAnchor() is called per turn, setting the anchor to the current message list length. The anchor doesn't need cache_control to be meaningful — it just needs to tell the compression layer "don't touch below this line."

PR #623 — compression pipeline respects CacheAnchorFloor

The compression and deduplication features now check the CacheAnchorFloor before operating on any message. Any message at or below the anchor index is off-limits. The floor is provider-aware: for explicit cache providers (Anthropic), it's derived from the cache_control markers. For implicit providers, it's derived from the implicit anchor set in PR #624.

The change is a single guard in the compression message selector:

if msgIndex <= session.CacheAnchorFloor {
    continue // never rewrite inside the anchor window
}

Simple. But it required PRs #622 and #624 to be meaningful — without a properly set anchor, the floor is 0 and the guard is a no-op.

why this matters at scale

The cache collapse problem isn't specific to Trimio's implementation. Any AI proxy that combines these three features will eventually hit this failure mode:

  1. Token compression or deduplication
  2. Multi-provider routing (including LCR to non-Anthropic providers)
  3. Long-running sessions with stable system prompts

If your proxy was built with Anthropic-first assumptions (as most are, given Anthropic's cache_control API surface), the session anchor is probably gated on Anthropic-outbound. That means Fireworks, Qwen, Gemini, and OpenAI routes are running without protection. The first time a high-value cached session gets routed to a cheaper provider — exactly what LCR is supposed to do — the compression layer eats the cache prefix.

The specific numbers from our incident: a 20-turn session, 9 turns in, 88k cached tokens collapsed to 23k, 7-message restart loop before we diagnosed it. In production, this shows up as degraded response quality on longer sessions — the kind of thing that gets filed as "the AI seems worse today" rather than traced to the underlying infrastructure behavior.

Trimio
Cache protection, at the proxy layer.
Trimio's session anchor system handles cache prefix protection across all providers — Anthropic, Fireworks, Gemini, OpenAI — with no code change required in your application. Your sessions stay coherent regardless of which provider LCR routes to.