Until early 2026, the routing argument for AI gateways had a footnote attached to it: "…assuming you're willing to use the open-weight models." The footnote was doing a lot of work. The cheaper alternatives were genuinely cheaper, but they were also visibly less capable. Most enterprise teams stayed on premium models because the quality gap was real.
In spring 2026, the footnote stopped being necessary. A wave of frontier-quality open-weight and open-API models hit production pricing. The "weak side" of routing now has genuine quality, not just cheap fallback. The savings math changed from "small" to "transformational."
glm-5 (Q94, MIT, $1.92/M output) is the headline — frontier-grade quality at workhorse pricing, 7.3x cheaper than GPT-5.2. Plus Kimi K2.5, MiniMax, MiMo, and OpenAI's own gpt-oss family.Six weeks of new tracker entries on CostGoat:
| Model | Vendor | Quality | Output ($/M) | License |
|---|---|---|---|---|
| GLM-5 | Zhipu AI / Z.AI | 94 | $1.92 | MIT |
| Kimi K2.5 | Moonshot | 89 | $2.00 | Open API |
| MiniMax-M2.5 | MiniMax | 79 | $1.15 | Open API |
| MiMo-V2-Flash | Xiaomi | 77 | $0.29 | Open weights |
| gpt-oss-120b | OpenAI | 62 | $0.18 | Open weights |
| gpt-oss-20b | OpenAI | 45 | $0.14 | Open weights |
The most consequential entry is GLM-5. A 744-billion-parameter mixture-of-experts model (40B active per token), MIT licensed, 200K+ token context, quality score 94 — within striking distance of frontier models from OpenAI and Anthropic. At $1.92/M output, it costs 7.3× less than GPT-5.2 ($14/M output, quality 96) for what most production benchmarks treat as functionally equivalent quality.
The second most consequential: Kimi K2.5 at $0.44/$2.00/M and quality 89, with a 262K context window. Production-grade for long-context summarization and analysis.
The third: OpenAI's own gpt-oss family. Yes, OpenAI now offers two open-weight models in its API: gpt-oss-120b at $0.18/M output and gpt-oss-20b at $0.14/M output. Quality is below frontier (62 and 45 respectively), but the price is near zero. For high-volume routine work — summarization, classification, simple extraction — these are the cheapest credible options OpenAI ships.
Before this wave, the routing argument went something like this:
"If you're willing to use a quality-79 model for some workloads, you can save 10×. But quality-79 isn't great for production-facing features, so the savings only apply to back-office work."
After this wave, the routing argument is:
"Quality 94 is now available at $1.92/M output, a fraction of the cost of frontier models. Production-facing features can route to GLM-5 with quality essentially indistinguishable from frontier alternatives. The savings now apply to most production workloads, not just back-office."
That's the 10× more valuable claim. The routing savings now apply to a workload band that's 5-10× larger than it was a quarter ago, because quality-94 is now in the cheap-tier pricing band for the first time.
Here's what the math looks like at production volume. Assume 100 million output tokens per month — a reasonable volume for a mid-sized SaaS company that has rolled out an internal AI assistant to 1,000 employees.
| Source model | Target model | Quality delta | Monthly cost on source | Monthly cost on target | Savings |
|---|---|---|---|---|---|
| GPT-5.2 | GLM-5 | 96 → 94 | $1,400 | $192 | $1,208/mo |
| GPT-5.2-Pro | GPT-5.2 | 96 → 96 (same) | $16,800 | $1,400 | $15,400/mo |
| Claude Opus 4.6 | Kimi K2.5 | 100 → 89 | $2,500 | $200 | $2,300/mo |
| GPT-5.5 standard | gpt-oss-120b | (within OpenAI) | $3,000 | $18 | $2,982/mo |
The last row is worth re-reading. Within OpenAI's own SKU lineup, routing from GPT-5.5 standard ($30/M output) to gpt-oss-120b ($0.18/M output) is a 98.7% output cost reduction. Quality drops from 100 to 62 — meaningful — but the trade-off is now an explicit one to evaluate per workload, not a fait accompli on every call.
The capital is there to keep this trend going. Q1 2026 venture funding (source):
The frontier labs are massively capitalized. So are their open-weight competitors. Z.AI, Moonshot, MiniMax, Xiaomi, and others are funded specifically to ship frontier-grade models at near-zero pricing. The trajectory is "more frontier-quality models at the cheap end of the market," not less.
Three actions:
1. Audit your current routing pattern. What fraction of production traffic is hitting premium-tier models that could be served by GLM-5, Kimi K2.5, or another mid-tier alternative? In most organizations, this number is 70-90%. The savings opportunity is substantial.
2. Pilot a routed workflow on a non-critical path. Pick one internal AI feature — summarization, transcript processing, internal chatbot — and route it to GLM-5 or another quality-90+ alternative. Measure quality differential. In our experience and across public benchmarks, the quality difference for non-reasoning tasks is statistically insignificant.
3. Build the routing into your gateway, not your application code. Routing logic that lives in application code means every team rebuilds it. Routing logic that lives in the gateway is a config change, not an engineering project.
The "weak side" of LLM routing isn't weak anymore. GLM-5 at quality 94 for $1.92/M output is the most actionable single change available to most enterprise AI budgets in 2026. The math is done. The infrastructure to capture it has been built. The only remaining question is whether your organization is paying 7× more than necessary because the routing footnote that used to be there is still in your team's mental model.
It isn't. Update the model.
Trimio is the LLM API gateway with built-in routing across 325+ models including GLM-5, Kimi K2.5, MiniMax, and the gpt-oss family. We make the savings automatic. See how it works.