Trimio Field Notes

DeepSeek Peak/Off-Peak Pricing: Time-of-Day Routing Is a New LCR Dimension

August 17, 2026 5 min read deepseekpricinglcrrouting

On August 13, DeepSeek announced something no other LLM provider has done at scale: time-based pricing tiers. Peak and off-peak rates, with off-peak at 50% of peak cost. The new pricing took effect at 16:00 UTC on August 16.

This isn't a promotional discount or a limited-time offer. It's a structural change to how DeepSeek prices API calls — one that introduces a new variable into the routing calculus that every LLM gateway needs to handle.

The pricing change

Essential
DeepSeek now charges different rates depending on when you call. Off-peak is 50% of peak pricing. For DeepSeek V4 Flash 0731 — already among the cheapest models in the market at $0.14/$0.28/MTok — off-peak pricing brings the effective rate to roughly $0.07/$0.14/MTok. That's cheaper than most cached calls on frontier models.

The specifics from DeepSeek's announcement:

DeepSeek V4 Flash 0731 at peak: $0.14/$0.28/MTok (input/output). At off-peak: approximately $0.07/$0.14/MTok. To put that in context: a frontier model like GPT-5.6 Sol costs $5/$30/MTok. DeepSeek Flash at off-peak is 71× cheaper on input and 214× cheaper on output than Sol — for tasks where Flash is capable enough.

The question isn't whether the off-peak rate is cheap. It's whether your routing infrastructure can take advantage of it.

Why time-of-day pricing changes the routing problem

Essential
Traditional LCR routing optimizes across models: route to the cheapest capable model for each request. Time-of-day pricing adds a second dimension: route to the cheapest model at the cheapest time. This only works for workloads that can tolerate scheduling delay — batch jobs, background analysis, non-interactive tasks.

Least Cost Routing to date has been a spatial problem: given a set of models with different prices and capabilities, route each request to the cheapest model that can handle it. The variables are: model price, model capability, request type.

DeepSeek's time-based pricing adds a temporal dimension. Now the routing decision is: given a set of models with different prices, capabilities, and time-based rate schedules, route each request to the cheapest model at the cheapest time — if the request can wait.

This splits workloads into two categories:

The second category is where the savings compound. An enterprise running 10,000 batch inference jobs per day through DeepSeek at peak rates could cut that cost in half by scheduling those jobs during off-peak windows. Not by switching models. Not by reducing usage. Just by shifting when the calls happen.

What this means for LCR engine design

Essential
A time-aware LCR rule needs three things: (1) the model's peak/off-peak schedule, (2) the request's latency tolerance, and (3) a scheduling mechanism that can queue latency-tolerant requests for off-peak execution. This is a new rule type — not a modification of existing cost-based routing.

The model_rates table that powers LCR routing needs a new structure. Today, a model entry has: provider, model name, input price, output price, context window, capability scores. With time-based pricing, it needs:

And the LCR rule engine needs a new parameter on incoming requests:

For requests with latency tolerance, the routing engine evaluates: is it currently off-peak for any time-tiered model? If yes, route to the cheapest model at its off-peak rate. If no, queue the request until the next off-peak window opens — or route immediately at the peak rate if the queue is full or the tolerance deadline is approaching.

This is a new routing pattern. It doesn't replace existing LCR logic — it adds a second layer for requests that can tolerate delay.

The DeepSeek-specific opportunity

Essential
DeepSeek is the first major provider to offer time-based pricing. If the pattern spreads to OpenAI, Anthropic, and Google — as competitive pressure mounts — time-of-day routing becomes a standard LCR feature, not a DeepSeek-specific optimization. The gateways that implement it first capture the savings first.

DeepSeek is the first, but they won't be the last. The economic pressure to fill off-peak GPU capacity is universal across providers. Every provider has idle inference capacity during low-demand hours. Time-based pricing is the mechanism to monetize that idle capacity — and once one provider offers it, the competitive pressure to match is real.

If OpenAI introduces off-peak pricing for GPT-5.6 Luna (already among the cheapest frontier-tier models at $0.20/$1.20/MTok), the off-peak rate could approach $0.10/$0.60/MTok — making it competitive with DeepSeek for quality-sensitive workloads. If Google follows for Gemini 3.7 Flash, the off-peak rate drops to ~$0.375/$1.875/MTok.

For enterprises, this means the savings from time-of-day routing will scale as more providers adopt the model. The gateway that supports time-aware LCR today is ready for the market as it evolves.

The bottom line

Essential
DeepSeek's peak/off-peak pricing is the first crack in the flat-rate-per-token model. Time-of-day routing is the LCR feature that captures the savings. The enterprises that implement it first will see 50% cost reduction on every latency-tolerant workload routed to DeepSeek — and the same architecture will capture savings from other providers as they follow.

DeepSeek's time-based pricing is a structural shift in how LLM API costs work. The flat per-token rate — the pricing model that every LLM gateway has been built around — is no longer the only model. Time-of-day pricing is here, and it will spread.

The routing infrastructure that supports time-aware LCR will capture savings that flat-rate gateways can't. For enterprises with significant batch or background inference workloads, the 50% off-peak discount on DeepSeek is a material cost reduction — available today, with no model quality tradeoff.

The question isn't whether time-of-day pricing will become standard. It's whether your routing layer is ready for it when it does.

Trimio's LCR engine supports model-specific pricing tiers and is being evaluated for time-of-day routing rules. See the routing architecture.

Trimio
Stop guessing. Start governing.
trimio is the LLM API gateway purpose-built for AI cost governance — visibility, routing, caching, and budget enforcement in one layer.