Trimio Field Notes

GPT-5.6 Luna Cut 80%: The Routing Table Just Reshuffled

August 2, 2026 5 min read openaigpt-5.6lunapricinglcr

On July 30, 2026, OpenAI updated the GPT-5.6 pricing page. No blog post. No press release. Just a price change on a model that launched 34 days earlier:

For context: when GPT-5.6 launched on June 26, the 5× input-price spread between Luna ($1) and Sol ($5) was the largest single-provider cost gradient the LLM market had ever shipped. Thirty-four days later, that spread is 25×. And Luna is no longer a budget option — it's the second-cheapest frontier-tier model in the market.

Essential
Luna at $0.20/$1.20/MTok is now cheaper than every frontier model except DeepSeek V4 Flash ($0.14/$0.28). It is 79% cheaper on output than its launch price. Any cost estimate built on Luna at $1.00/MTok is now 80% too high.

The revised routing table

Essential
The 5-tier LCR routing table has a new floor. DeepSeek V4 Flash ($0.14) → GPT-5.6 Luna ($0.20) → Sonnet 5 intro / Terra ($2.00) → K3 / Gemini 3.6 Flash / Grok 4.5 ($1.50–3.00) → Opus 5 / Fable 5 / Sol ($5.00+). The gap between tier 1 and tier 3 is now 15× on input — down from 21× at launch.

Here's where the market stands as of August 2, 2026:

TierModelInput / Output (per MTok)Notes
1 — Ultra-cheapDeepSeek V4 Flash / MiMo V2.5$0.14 / $0.28Cheapest; agent-grade benchmarks
2 — Budget frontierGPT-5.6 Luna$0.20 / $1.2080% cut July 30; OpenAI-native API
3 — Mid-tierSonnet 5 intro / Terra$2.00 / $10–12Sonnet 5 intro expires Aug 31
4 — Near-frontierK3 / Gemini 3.6 Flash / Grok 4.5$1.50–3.00Specialized routing
5 — Quality ceilingOpus 5 / Fable 5 / Sol$5.00 / $25–30Max-effort tasks

The structural shift: Luna is now a tier-1 competitor, not a tier-3 model. At $0.20/MTok input, it is 43% more expensive than DeepSeek V4 Flash but offers OpenAI-native API compatibility and no export-control concerns. For enterprises that can't route to Chinese providers, Luna is now the cheapest frontier option by a wide margin.

What this means for your cost model

Essential
A workload moving 10M input tokens/month and 5M output tokens/month from Sonnet 5 ($2/$10) to Luna ($0.20/$1.20) saves $68,000/year. From Opus 5 ($5/$25), the savings are $165,000/year. Any LCR routing rule still weighted toward mid-tier models for routine traffic is leaving significant money on the table.

Consider a standard agentic coding workload: 10M input tokens/month, 5M output tokens/month. That's a mid-sized engineering team running Claude Code or an equivalent harness.

ModelMonthly CostAnnual Cost
Opus 5 ($5/$25)$175,000$2,100,000
Sonnet 5 ($2/$10)$70,000$840,000
GPT-5.6 Terra ($2/$12)$80,000$960,000
GPT-5.6 Luna ($0.20/$1.20)$8,000$96,000
DeepSeek V4 Flash ($0.14/$0.28)$2,800$33,600

The spread between Luna and Sonnet 5 for the same workload: $62,000/month. That's not a rounding error. It's an engineering hire.

The Terra squeeze

Terra at $2/$12 is now in an awkward position. It's 10× more expensive than Luna on input and 10× on output — but the quality gap between Terra (85.1% TerminalBench 2.1) and Luna (82.5%) is 2.6 percentage points. For workloads where that 2.6% matters, Terra is the right choice. For everything else — and "everything else" is most traffic — Luna at $0.20/$1.20 makes Terra economically irrational.

Essential
Terra's 2.6-point TerminalBench advantage over Luna costs 10× more per token. The routing question is: which 2.6% of your workload actually needs it? Route that to Terra. Route the rest to Luna. The LCR engine does this automatically.

This is the LCR thesis in its purest form: the same provider now offers a 10× cost gradient within a single model family. If your routing rule is "use Terra for everything," you're overpaying by 10× on every request that Luna could handle. The routing engine's job is to identify which requests those are — in real time, per request, based on complexity signals.

Sonnet 5 intro pricing: 29 days remaining

Anthropic's Sonnet 5 intro pricing ($2/$10, expiring August 31) was competitive with Terra at launch. After the Luna cut, it's not. Sonnet 5 at $2/$10 is 10× more expensive than Luna on input for a model in the same quality tier. When the intro expires and Sonnet 5 moves to $3/$15, the gap widens to 15×.

This doesn't mean "always route to Luna." Quality differences between Sonnet 5 and Luna on specific tasks (particularly long-context reasoning and tool-use chains) may justify the premium. It means the default routing assumption should be Luna unless quality signals indicate otherwise — not Sonnet 5 unless cost forces otherwise.

The sales impact

Essential
Any enterprise customer that received cost estimates based on Luna at $1.00/MTok is looking at 80% lower real costs. The ROI math just got dramatically stronger. Update every quote, every projection, every savings calculation — the old numbers are leaving money on the table.

If your sales team built ROI models on Luna's launch pricing, those models are now understating savings by 80%. A prospect who was shown $50,000/year in routing savings on a Luna-inclusive mix is actually looking at $90,000+. This is not a minor refresh — it's a fundamental change in the unit economics.

Three actions:

  1. Update every cost model that references GPT-5.6 Luna or Terra. The old prices are 80% and 20% too high respectively. Any model still using launch prices is giving prospects incorrect cost projections.
  2. Rerun LCR routing weights. Luna's new price point moves it from tier 3 to tier 2 in the routing table. Routing rules that excluded Luna for cost reasons should be re-evaluated — it may now be the optimal choice for workloads it was previously too expensive for (ironically, it was never too expensive; it's now even cheaper).
  3. Brief sales on the new positioning. "Trimio routes to the cheapest capable model" is a stronger pitch when the cheapest capable model just got 80% cheaper. The savings story writes itself.

The bottom line

OpenAI's Luna price cut is the most consequential pricing event since the GPT-5.6 launch itself. It compresses the cost gradient between ultra-cheap (DeepSeek V4 Flash) and mid-tier (Sonnet 5, Terra) from 21× to 15×. It makes Luna the default budget frontier model for enterprises that can't route to Chinese providers. And it makes every cost model built on launch pricing 80% too conservative.

If your routing engine isn't already weighting Luna at $0.20/$1.20, your routing table is stale. The market moved on July 30. The routing table should have moved with it.

Trimio's LCR engine automatically adjusts routing weights when provider prices change. See how it works.

Trimio
Stop guessing. Start governing.
trimio is the LLM API gateway purpose-built for AI cost governance — visibility, routing, caching, and budget enforcement in one layer.