Trimio Field Notes

Inside the 720× Price Spread Hiding in OpenAI's Own Catalog

April 25, 2026 4 min read llm-pricingmodel-routingfinops

If you treat "GPT-5.5" as a single product with a single price, you are leaving money on the table. Possibly a lot of it.

OpenAI's current SKU lineup is not a price. It's a 720× spread.

The numbers

Essential
OpenAI's published catalog spans GPT-5.4-mini ($2/M output) → GPT-5.5 Pro ($180/M) at the same generation tier. Batch and Flex tiers cut another 50% off both ends. One vendor, one family, four orders of magnitude.

As of late April 2026, OpenAI's published catalog includes (per OpenAI and the cross-vendor tracker at Apidog):

ModelInput ($/M tokens)Output ($/M tokens)
GPT-5.4-mini$0.25$2.00
GPT-5.4 standard$2.50$15.00
GPT-5.5 standard$5.00$30.00
GPT-5.5 Pro$30.00$180.00
720×
Intra-OpenAI Spread
5.4-mini input ($0.25/M) → 5.5 Pro output ($180/M). One vendor, one family.
15×
Output Only
5.5 standard ($30) vs 5.4-mini ($2), same generation, output only.
50%
Batch Discount
Batch / Flex tiers cut another half off both sides of the spread.
Output price · OpenAI SKU lineup · linear scale
Same family, same vendor. The spread between mini and Pro is 90× on output alone.
GPT-5.4-mini workhorse
$2.00
GPT-5.4 standard
$15.00
GPT-5.5 standard "the latest"
$30.00
GPT-5.5 Pro
$180.00
Source: OpenAI, Apidog cross-vendor pricing · April 2026

A few comparisons worth letting sink in:

If your engineering team's default for everything is "send it to GPT-5.5," you are paying 6× more than necessary for any workload where 5.4 is acceptable, and 36× more than necessary for any workload where 5.4-mini is acceptable. And if you're routing through Pro, the spread expands further.

Why this happens

Essential
Frontier providers are no longer competing on a single point — they're competing on a price curve. The intra-vendor spread now exceeds the inter-vendor spread, so "vendor selection" is the wrong framing.

Frontier AI providers are no longer competing on a single point. They are competing on a price curve. OpenAI has explicitly versioned into four SKUs at the same generation tier — the same pattern Anthropic uses (Haiku/Sonnet/Opus) and Google uses (Flash/Pro/Ultra). The price differential within a single vendor's family now exceeds the differential between vendors.

For finance and procurement leaders, this means the old "vendor selection" framing is obsolete. The question is no longer which provider you use. It's which model, for which workload, under which pricing tier.

What "leaving money on the table" looks like in practice

Essential
5,000 engineers default-routed to gpt-5.5 for code formatting and commit messages — workflows that gpt-5.4-mini handles at 1/15th the cost with no detectable quality loss. The structural argument behind Uber's overrun.

A common pattern: a 5,000-engineer organization issues every developer access to GPT-5.5 because that's the latest. A subset of internal workflows — code formatting, lint suggestions, commit message drafting — get routed through it by default. Each workflow is fast, cheap-per-call, and quality is excellent.

At the end of the quarter, the bill arrives. The same workflows could have run on GPT-5.4-mini at 1/15th the output cost with no quality degradation a human could detect. The opportunity cost compounds across thousands of internal calls per day.

This is not a hypothetical: it's the structural argument behind Uber's publicly announced AI budget overrun in April 2026, where 70% of committed code is AI-generated and the realized cost-per-engineer ran 7-12× the published seat price.

What to do

Essential
Map workloads to tiers, default to mid-tier with escalation, use batch where latency tolerates it (50% off), and watch for new SKUs — the catalog grew four SKUs in two quarters.
  1. Map workloads to tiers. Not every internal use case needs the top SKU. A simple inventory — "summarize support tickets," "generate boilerplate," "draft technical specs" — quickly classifies into nano/mid/premium tiers.
  1. Route by workload, default to mid-tier, escalate when needed. A gateway that routes individual requests to the right SKU based on workload metadata captures the 6-15× spread without changing developer behavior.
  1. Use batch where latency tolerates it. OpenAI's batch tier is 50% off. If your nightly summarization job runs at 3am, it should be on batch.
  1. Watch for new SKUs as they ship. The "$180/M output Pro tier" was new in April. Within OpenAI alone, the catalog has grown by four SKUs in two quarters. The price spread is widening, not narrowing.

The bigger pattern

Essential
Treating "OpenAI" as a single procurement decision is leaving money the way treating "AWS compute" as one decision would. The question is no longer "should we use OpenAI?" — it's which OpenAI?

The intra-vendor price spread is now larger than the inter-vendor spread used to be. Treating "OpenAI" as a single procurement decision is leaving money in the same way that treating "AWS compute" as a single decision would have — you'd never do that with EC2 instance types, and AI models now demand the same discipline.

For a finance leader, the question on every internal AI workflow is no longer "should we use OpenAI?" It's "which OpenAI?"

Trimio is the LLM API gateway built for AI cost governance. We route each request to the right model for the workload, automatically — across vendors and across tiers within each vendor. See how it works.

Trimio
Stop guessing. Start governing.
trimio is the LLM API gateway purpose-built for AI cost governance — visibility, routing, caching, and budget enforcement in one layer.