Turn 9 of a 20-turn agent session. Claude Sonnet, routed via least-cost routing to Fireworks/Qwen3 for the cheaper turns. Everything working as expected — until the model says:
"Let me check on that — I want to make sure I have the full context before proceeding."
It said it seven more times. The agent had effectively restarted. Not from a model failure. Not from a prompt engineering mistake. From a cache collapse — the proxy's compression layer had silently rewritten messages inside the active cache prefix, collapsing 88k → 23k cached read tokens (a 73% reduction), forcing the model to rebuild context from scratch on a session already 9 turns deep.
This is the kind of failure that looks like flaky AI behavior but is actually a deterministic infrastructure bug. Here's what happened, and how we fixed it across three PRs.
Providers like Anthropic use a cache prefix model: if the first N tokens of your prompt are identical to a previous request, they're billed at the cache-read rate (~10% of normal input cost). The longer your stable prefix, the more you save.
In a multi-turn agent session, the cache prefix grows naturally: system prompt + turn 1 + turn 2 + ... + turn N-1 is re-sent with every new turn, and all of it qualifies as cached after the first turn. A well-managed 20-turn session might have 80k+ tokens in cached reads by turn 10 — at Anthropic's rates, that's roughly $0.24 per turn saved vs. $2.40 in input cost. Across thousands of sessions, it's the difference between a profitable product and a loss leader.
The session anchor is the mechanism that protects this. Concretely, it's an index: do not touch messages below position N. The compression and deduplication pipeline can do whatever it wants above the anchor — but below it, the prefix is frozen.
In Trimio's proxy, the session anchor was implemented in the same code path that injects Anthropic's cache_control markers. The logic looked roughly like this:
if pcIsAnthropic {
injectCacheControl(messages)
SetSessionAnchor(anchorIndex)
}
The problem: pcIsAnthropic was a single flag that conflated two distinct things:
cache_control markers in the wire format (only valid for Anthropic endpoints)When a session was routed to Fireworks/Qwen3 via LCR, pcIsAnthropic was false — so no anchor was set. The compression layer, seeing no anchor floor, was free to rewrite any message in the session. On turn 9, it compressed a large system prompt segment that happened to fall inside what had been the cache prefix on previous turns. Result: 88k → 23k cached tokens, and a session restart loop.
This failure mode is subtle because it requires two conditions to coincide:
Condition 2 is more likely than you'd think. The compression layer looks at the full message list and picks the highest-value compression targets — and long, repeated system prompt blocks are prime candidates. Those blocks are also exactly what the cache is built on.
The result is a silent correctness failure: the request succeeds, the model responds, costs look normal — but the session's effective context has been quietly amputated. The 7-message restart loop we observed was the model trying to reconstruct context it no longer had access to via its cache.
PR #622 — decouple session anchor from Anthropic-outbound gate
Split pcIsAnthropic into two flags:
pcIsAnthropicOutbound: true only when sending to an Anthropic endpoint. Controls cache_control injection into the wire format.pcIsAnthropicShape: true when the request uses Anthropic message format, regardless of destination. Controls fingerprinting and anchor store operations.This is the foundational change. Anchor management now runs on message shape, not on provider identity. A Fireworks/Qwen3 request using the Anthropic message format gets anchor tracking — it just doesn't get cache_control injected (which would be invalid on that endpoint).
PR #624 — extend implicit anchor support to all providers
Added isImplicitCachedProvider() — a function that returns true for providers with implicit caching semantics (where the provider caches automatically without requiring cache_control markers in the request). Current list: Fireworks, Qwen, Gemini, OpenAI.
For these providers, SetImplicitAnchor() is called per turn, setting the anchor to the current message list length. The anchor doesn't need cache_control to be meaningful — it just needs to tell the compression layer "don't touch below this line."
PR #623 — compression pipeline respects CacheAnchorFloor
The compression and deduplication features now check the CacheAnchorFloor before operating on any message. Any message at or below the anchor index is off-limits. The floor is provider-aware: for explicit cache providers (Anthropic), it's derived from the cache_control markers. For implicit providers, it's derived from the implicit anchor set in PR #624.
The change is a single guard in the compression message selector:
if msgIndex <= session.CacheAnchorFloor {
continue // never rewrite inside the anchor window
}
Simple. But it required PRs #622 and #624 to be meaningful — without a properly set anchor, the floor is 0 and the guard is a no-op.
The cache collapse problem isn't specific to Trimio's implementation. Any AI proxy that combines these three features will eventually hit this failure mode:
If your proxy was built with Anthropic-first assumptions (as most are, given Anthropic's cache_control API surface), the session anchor is probably gated on Anthropic-outbound. That means Fireworks, Qwen, Gemini, and OpenAI routes are running without protection. The first time a high-value cached session gets routed to a cheaper provider — exactly what LCR is supposed to do — the compression layer eats the cache prefix.
The specific numbers from our incident: a 20-turn session, 9 turns in, 88k cached tokens collapsed to 23k, 7-message restart loop before we diagnosed it. In production, this shows up as degraded response quality on longer sessions — the kind of thing that gets filed as "the AI seems worse today" rather than traced to the underlying infrastructure behavior.