Trimio Field Notes

Open-Weight Models Just Made Routing 10× More Valuable

April 26, 2026 5 min read open-weightmodel-routingfinopsglm-5

Until early 2026, the routing argument for AI gateways had a footnote attached to it: "…assuming you're willing to use the open-weight models." The footnote was doing a lot of work. The cheaper alternatives were genuinely cheaper, but they were also visibly less capable. Most enterprise teams stayed on premium models because the quality gap was real.

In spring 2026, the footnote stopped being necessary. A wave of frontier-quality open-weight and open-API models hit production pricing. The "weak side" of routing now has genuine quality, not just cheap fallback. The savings math changed from "small" to "transformational."

The new entrants

Essential
glm-5 (Q94, MIT, $1.92/M output) is the headline — frontier-grade quality at workhorse pricing, 7.3x cheaper than GPT-5.2. Plus Kimi K2.5, MiniMax, MiMo, and OpenAI's own gpt-oss family.

Six weeks of new tracker entries on CostGoat:

ModelVendorQualityOutput ($/M)License
GLM-5Zhipu AI / Z.AI94$1.92MIT
Kimi K2.5Moonshot89$2.00Open API
MiniMax-M2.5MiniMax79$1.15Open API
MiMo-V2-FlashXiaomi77$0.29Open weights
gpt-oss-120bOpenAI62$0.18Open weights
gpt-oss-20bOpenAI45$0.14Open weights

The most consequential entry is GLM-5. A 744-billion-parameter mixture-of-experts model (40B active per token), MIT licensed, 200K+ token context, quality score 94 — within striking distance of frontier models from OpenAI and Anthropic. At $1.92/M output, it costs 7.3× less than GPT-5.2 ($14/M output, quality 96) for what most production benchmarks treat as functionally equivalent quality.

The second most consequential: Kimi K2.5 at $0.44/$2.00/M and quality 89, with a 262K context window. Production-grade for long-context summarization and analysis.

The third: OpenAI's own gpt-oss family. Yes, OpenAI now offers two open-weight models in its API: gpt-oss-120b at $0.18/M output and gpt-oss-20b at $0.14/M output. Quality is below frontier (62 and 45 respectively), but the price is near zero. For high-volume routine work — summarization, classification, simple extraction — these are the cheapest credible options OpenAI ships.

Why this changes routing math

Essential
The old "willing to use a quality-79 model" footnote is gone. Quality 94 is now in the cheap tier, so routing savings apply to 5-10x more workloads than they did a quarter ago.

Before this wave, the routing argument went something like this:

"If you're willing to use a quality-79 model for some workloads, you can save 10×. But quality-79 isn't great for production-facing features, so the savings only apply to back-office work."

After this wave, the routing argument is:

"Quality 94 is now available at $1.92/M output, a fraction of the cost of frontier models. Production-facing features can route to GLM-5 with quality essentially indistinguishable from frontier alternatives. The savings now apply to most production workloads, not just back-office."

That's the 10× more valuable claim. The routing savings now apply to a workload band that's 5-10× larger than it was a quarter ago, because quality-94 is now in the cheap-tier pricing band for the first time.

Same-quality routing examples

Essential
At 100M output tokens/month: GPT-5.2-Pro → GPT-5.2 saves $15,400/mo; GPT-5.5 → gpt-oss-120b is a 98.7% output cost reduction within OpenAI's own SKUs.

Here's what the math looks like at production volume. Assume 100 million output tokens per month — a reasonable volume for a mid-sized SaaS company that has rolled out an internal AI assistant to 1,000 employees.

Source modelTarget modelQuality deltaMonthly cost on sourceMonthly cost on targetSavings
GPT-5.2GLM-596 → 94$1,400$192$1,208/mo
GPT-5.2-ProGPT-5.296 → 96 (same)$16,800$1,400$15,400/mo
Claude Opus 4.6Kimi K2.5100 → 89$2,500$200$2,300/mo
GPT-5.5 standardgpt-oss-120b(within OpenAI)$3,000$18$2,982/mo

The last row is worth re-reading. Within OpenAI's own SKU lineup, routing from GPT-5.5 standard ($30/M output) to gpt-oss-120b ($0.18/M output) is a 98.7% output cost reduction. Quality drops from 100 to 62 — meaningful — but the trade-off is now an explicit one to evaluate per workload, not a fait accompli on every call.

Savings · per 100M output tokens / month
The same workload routed to a quality-equivalent open-weight target. Vertical axis is monthly $ saved; longer bars = more captured.
GPT-5.2 → GLM-5 96 → 94
$1,208
Opus 4.6 → Kimi K2.5 100 → 89
$2,300
GPT-5.5 → gpt-oss-120b within OpenAI
$2,982
GPT-5.2-Pro → GPT-5.2 96 → 96, same
$15,400
Source: CostGoat · workload model: 100M output tokens/month

The capital signal

Essential
AI absorbed 81% of Q1 2026 venture funding ($242B of $300B) — frontier and open-weight competitors are both massively capitalized. The trajectory is more frontier quality at the cheap end, not less.

The capital is there to keep this trend going. Q1 2026 venture funding (source):

The frontier labs are massively capitalized. So are their open-weight competitors. Z.AI, Moonshot, MiniMax, Xiaomi, and others are funded specifically to ship frontier-grade models at near-zero pricing. The trajectory is "more frontier-quality models at the cheap end of the market," not less.

What CFOs should do this quarter

Essential
Audit current routing (70-90% of traffic typically over-provisioned), pilot GLM-5 on a non-critical path, and put routing in the gateway — not application code that every team has to rebuild.

Three actions:

1. Audit your current routing pattern. What fraction of production traffic is hitting premium-tier models that could be served by GLM-5, Kimi K2.5, or another mid-tier alternative? In most organizations, this number is 70-90%. The savings opportunity is substantial.

2. Pilot a routed workflow on a non-critical path. Pick one internal AI feature — summarization, transcript processing, internal chatbot — and route it to GLM-5 or another quality-90+ alternative. Measure quality differential. In our experience and across public benchmarks, the quality difference for non-reasoning tasks is statistically insignificant.

3. Build the routing into your gateway, not your application code. Routing logic that lives in application code means every team rebuilds it. Routing logic that lives in the gateway is a config change, not an engineering project.

The bottom line

Essential
The "weak side" of LLM routing isn't weak anymore. GLM-5 at quality 94 for $1.92/M output is the most actionable single change available to most enterprise AI budgets in 2026. The math is done. The infrastructure to capture it is built.

The "weak side" of LLM routing isn't weak anymore. GLM-5 at quality 94 for $1.92/M output is the most actionable single change available to most enterprise AI budgets in 2026. The math is done. The infrastructure to capture it has been built. The only remaining question is whether your organization is paying 7× more than necessary because the routing footnote that used to be there is still in your team's mental model.

It isn't. Update the model.

Trimio is the LLM API gateway with built-in routing across 325+ models including GLM-5, Kimi K2.5, MiniMax, and the gpt-oss family. We make the savings automatic. See how it works.

Trimio
Stop guessing. Start governing.
trimio is the LLM API gateway purpose-built for AI cost governance — visibility, routing, caching, and budget enforcement in one layer.