Trimio Field Notes

Castform + Neon Beat GPT-5.6 Sol at 100× Less Cost. Here's What That Proves About Routing.

August 6, 2026 6 min read lcrretrievalroutingagenticfinops

Neon and Castform published a detailed case study this week that landed at #13 on HN with 336 points. The finding:

A small open-weights model, post-trained via reinforcement learning specifically for retrieval tasks, matches and beats GPT-5.6 Sol ($5/$30 per million tokens) on multi-hop search accuracy — at 100× lower cost per request.

The baseline: a typical multi-turn search with GPT-5.6 Sol costs approximately $0.03 end-to-end and takes more than 10 seconds. Castform's RL post-trained model handles the same retrieval task at a fraction of the cost, powered by Neon's Lakebase Search extensions.

This is a significant empirical result. And it illustrates a principle that goes beyond domain-specific model training: the right model for a task is not the biggest model available, it's the model best suited to that task's cost-quality tradeoff.

That principle is what Trimio's LCR engine implements at the proxy layer — automatically, without requiring you to train or maintain a domain-specific model.

Three kinds of agentic search — and what Castform chose

Essential
The HN thread on the Castform paper offered a clean taxonomy: (1) better retrieval infrastructure, (2) smarter harness with judges, (3) a model trained specifically for retrieval. Castform chose option 3 — RL post-training on domain data. Trimio implements option 2 at the proxy layer without requiring option 3.

The HN thread on the Castform announcement surfaced a useful taxonomy. One commenter described three approaches to reducing agentic search costs:

  1. Better retrieval infrastructure — Neon's Lakebase Search extensions, vector indexes, semantic chunking. Reduces the scope of what the model needs to process.
  2. Smarter harness with judges — routing calls to appropriate models, using lightweight classifier models to evaluate retrieval quality before calling expensive models for synthesis, decomposing multi-hop queries into sub-queries routed at appropriate cost tiers.
  3. A model trained specifically for retrieval — RL post-training on the specific retrieval task, producing a purpose-built model that outperforms generalist frontier models at orders-of-magnitude lower cost.

Castform chose option 3. It's a powerful approach — and it requires custom training infrastructure, domain-specific data collection, RL training runs, evaluation pipelines, and ongoing model maintenance. It's the right choice for companies that have retrieval as a core, high-volume workload and the engineering capacity to invest in it.

For everyone else, option 2 is available today, without a training pipeline.

What Trimio does at the proxy layer — right now, without custom training

Multi-hop retrieval is one of the highest-volume API call patterns in enterprise AI deployments. An agentic search loop makes multiple calls per query — decomposition, sub-retrieval, synthesis, re-ranking, final answer generation. Each of those calls has different cost-quality requirements.

Essential
Not every call in a retrieval loop needs GPT-5.6 Sol. Sub-retrieval calls and early-stage decomposition can route to DeepSeek V4 Flash at $0.14/$0.28 per million tokens. Only synthesis and final answer generation require frontier-tier quality. Trimio's LCR engine makes this distinction automatically.

Consider a five-call retrieval loop where the current default is GPT-5.6 Sol ($5.00/$30.00/MTok) across all five calls:

All calls to GPT-5.6 Sol
$0.03/query
Current baseline (Neon's measurement)
→
LCR-routed (sub-calls to cheap tier)
$0.004/query
Route 3 of 5 calls to DeepSeek V4 Flash

You don't have to reach 100× to make the routing argument work. Routing 60% of a retrieval loop's calls to a model that's 35× cheaper than GPT-5.6 Sol while maintaining frontier quality on synthesis calls captures 7-8× cost reduction with no change to output quality on the calls that matter.

And you don't need to train a custom model to get there. Trimio's LCR engine routes based on call metadata — system prompt characteristics, estimated context length, task type tags that developers add in seconds — without requiring a new training run or model deployment.

The compounding math: retrieval at scale

Enterprise AI deployments with agentic search workloads don't run one query. They run thousands of queries per day across dozens of knowledge domains. The compounding effect at scale:

The Castform approach achieves 100× reduction through purpose-built model training. Trimio's proxy-layer routing approach achieves 7-8× reduction through intelligent call routing. For most enterprise buyers, 7-8× is the right first target — achievable in a day, without a training pipeline.

The two approaches compound for teams that want to invest further: if you've done the Castform work and have a domain-trained model, Trimio routes to it alongside frontier models, giving you the purpose-built model for high-volume retrieval calls and frontier quality for the synthesis calls that need it.

Essential
Castform proves that the right model for a retrieval task is not the biggest model — it's the model optimized for that task. Trimio implements this principle at the proxy layer for every enterprise without requiring a training infrastructure investment. The principle is the same; the path to capture the savings is different.

Why frontier models are the wrong default for retrieval sub-calls

GPT-5.6 Sol at $5.00/$30.00/MTok is priced for tasks that require broad world knowledge synthesis, complex multi-step reasoning, and high-precision long-form generation. It is an excellent model for those tasks.

It is not an excellent cost choice for:

These tasks require specific, structured output — often a small JSON object, a ranked list, or a 2-3 sentence extraction — not general-purpose world-knowledge synthesis. A model that costs 35× less can produce identical output for these sub-tasks. Using GPT-5.6 Sol for relevance ranking because it's already in the context is the same logic error as using a backhoe to plant a flower bulb.

The Castform paper's core insight — that RL post-training on a specific task produces a model that outperforms general-purpose frontier models on that task at 100× lower cost — is exactly the same insight that motivates task-aware routing in Trimio's LCR engine. The mechanism is different (RL training vs. routing logic), but the underlying principle is identical: task specificity unlocks cost efficiency that general-purpose frontier models cannot match.

The routing table implication

When Castform's models become available via API (which is the natural progression for a company publishing research results like these), they should slot directly into routing tables as a retrieval-optimized tier. The routing logic is straightforward: detect retrieval sub-call patterns → route to Castform/purpose-built tier → route synthesis and final answer calls to frontier tier.

This is the retrieval routing tier that Trimio's LCR engine will support — the same way it supports DeepSeek V4 Flash for high-volume low-complexity calls today. The abstraction is: route each call to the cheapest model that meets the quality threshold for that specific call type.

The Castform result gives that abstraction a concrete empirical floor: 100× cost reduction on retrieval alone is real and published. The routing layer is how you capture it at production scale without training custom models for every call type in your pipeline.

Essential
100× cost reduction on retrieval is the ceiling of what domain-specific model training can achieve. Routing-layer cost reduction is the floor you can achieve today, with no training infrastructure, by putting the right model in front of the right call type. For most enterprise workloads, that floor is the right starting point.

The bottom line

The Castform paper is technically strong and the result is impressive. RL post-training on domain-specific retrieval data produces models that beat frontier generalist models at 100× lower cost. This is the academic and empirical validation of a principle the industry has been building toward for two years.

That principle, implemented at the proxy layer without custom training, is what Trimio's LCR engine does for every enterprise deployment. The first step is not a training pipeline — it's a URL change. One line of configuration routes every call through a proxy that applies cost-aware routing logic to every call in your pipeline, from the highest-volume retrieval sub-calls to the frontier synthesis calls that actually need GPT-5.6 Sol.

Trimio's LCR engine routes across retrieval tiers (DeepSeek V4 Flash, Luna) and frontier tiers (GPT-5.6 Sol, Claude Opus 5) automatically, based on call characteristics. No custom model training required. See how it works.

Trimio
Route the right call to the right model.
Trimio's LCR engine routes every call in your pipeline to the cheapest model that meets quality requirements for that specific call type — frontier models for synthesis, cheap tiers for retrieval sub-calls. No training required.