On August 13, DeepSeek announced something no other LLM provider has done at scale: time-based pricing tiers. Peak and off-peak rates, with off-peak at 50% of peak cost. The new pricing took effect at 16:00 UTC on August 16.
This isn't a promotional discount or a limited-time offer. It's a structural change to how DeepSeek prices API calls — one that introduces a new variable into the routing calculus that every LLM gateway needs to handle.
The specifics from DeepSeek's announcement:
DeepSeek V4 Flash 0731 at peak: $0.14/$0.28/MTok (input/output). At off-peak: approximately $0.07/$0.14/MTok. To put that in context: a frontier model like GPT-5.6 Sol costs $5/$30/MTok. DeepSeek Flash at off-peak is 71× cheaper on input and 214× cheaper on output than Sol — for tasks where Flash is capable enough.
The question isn't whether the off-peak rate is cheap. It's whether your routing infrastructure can take advantage of it.
Least Cost Routing to date has been a spatial problem: given a set of models with different prices and capabilities, route each request to the cheapest model that can handle it. The variables are: model price, model capability, request type.
DeepSeek's time-based pricing adds a temporal dimension. Now the routing decision is: given a set of models with different prices, capabilities, and time-based rate schedules, route each request to the cheapest model at the cheapest time — if the request can wait.
This splits workloads into two categories:
The second category is where the savings compound. An enterprise running 10,000 batch inference jobs per day through DeepSeek at peak rates could cut that cost in half by scheduling those jobs during off-peak windows. Not by switching models. Not by reducing usage. Just by shifting when the calls happen.
The model_rates table that powers LCR routing needs a new structure. Today, a model entry has: provider, model name, input price, output price, context window, capability scores. With time-based pricing, it needs:
And the LCR rule engine needs a new parameter on incoming requests:
For requests with latency tolerance, the routing engine evaluates: is it currently off-peak for any time-tiered model? If yes, route to the cheapest model at its off-peak rate. If no, queue the request until the next off-peak window opens — or route immediately at the peak rate if the queue is full or the tolerance deadline is approaching.
This is a new routing pattern. It doesn't replace existing LCR logic — it adds a second layer for requests that can tolerate delay.
DeepSeek is the first, but they won't be the last. The economic pressure to fill off-peak GPU capacity is universal across providers. Every provider has idle inference capacity during low-demand hours. Time-based pricing is the mechanism to monetize that idle capacity — and once one provider offers it, the competitive pressure to match is real.
If OpenAI introduces off-peak pricing for GPT-5.6 Luna (already among the cheapest frontier-tier models at $0.20/$1.20/MTok), the off-peak rate could approach $0.10/$0.60/MTok — making it competitive with DeepSeek for quality-sensitive workloads. If Google follows for Gemini 3.7 Flash, the off-peak rate drops to ~$0.375/$1.875/MTok.
For enterprises, this means the savings from time-of-day routing will scale as more providers adopt the model. The gateway that supports time-aware LCR today is ready for the market as it evolves.
DeepSeek's time-based pricing is a structural shift in how LLM API costs work. The flat per-token rate — the pricing model that every LLM gateway has been built around — is no longer the only model. Time-of-day pricing is here, and it will spread.
The routing infrastructure that supports time-aware LCR will capture savings that flat-rate gateways can't. For enterprises with significant batch or background inference workloads, the 50% off-peak discount on DeepSeek is a material cost reduction — available today, with no model quality tradeoff.
The question isn't whether time-of-day pricing will become standard. It's whether your routing layer is ready for it when it does.
Trimio's LCR engine supports model-specific pricing tiers and is being evaluated for time-of-day routing rules. See the routing architecture.