Every AI gateway claims to route on cost. A few claim to route on quality. Almost none can answer the question a CFO will ask: "How do you know this score is real?"
That question is now answerable. As of this week, Trimio's Quality Prediction Engine (QPE) runs a complete, auditable chain from third-party benchmark data to per-request quality scores — and every step in that chain is documented, versioned, and defensible under audit.
The quality score is not a vendor assertion. It is a derivable number with a published methodology. Here's the full chain:
Previously, Trimio's model quality catalog was hand-curated YAML — a maintainer's nightmare and an auditor's red flag. As of PR #608 (May 22), a new MQVA (Multi-Quality Vector Aggregation) ingestor fetches live data directly from the Artificial Analysis API — an independent third-party benchmark aggregator — and writes provenance-stamped rows to the model catalog.
Key details:
GET /api/v2/data/llms/models with API key authsource_methodology_version = "aa_v2_2026-05-22"When the ingestor first ran against the live API, it surfaced two compatibility bugs: the API returns status: 200 as an integer (not a string), and legacy seed data used UUID keys instead of model slugs. Both were fixed within the hour (PR #610). The fact that these bugs surfaced — and were fixed — is itself a quality signal: live API ingestion is the right architecture because it breaks visibly when things change.
Even with good source data, a quality score is only useful if the projection weights meaningfully differentiate between model classes. Before PR #600, Trimio's dimension weights had cosine similarity of 1.0 across all 6 quality classes — meaning every class projected identically and the score carried no routing signal.
After PR #600: cosine similarity between any two class weight vectors is now ≤ 0.84. That's the threshold where the projection step actually discriminates between quality tiers. The practical result: 11 of 42 LCR routing cells flip to a different bucket when the differentiated weights are applied.
The quality prediction pipeline starts with a traffic classifier that assigns each request to one of 6 classes. That classifier needs to be accurate enough to trust its output. For three weeks, Trimio's classifier was stuck at macro 0.586 — below the 0.711 plan target.
PR #596 (May 21) shipped the consolidated classifier rework:
The classifier also gained a tool-calling labeler (Sonnet/Opus) with an enum-bound classify tool, eliminating the "prose-leak halts" that previously caused unpredictable classifier failures during training.
A classifier that achieves 0.716 macro accuracy on its own benchmarks is useful. A classifier that achieves 0.900 agreement with a human reviewer on an Opus reference set is defensible.
PR #602 documents the 30-row hand-review: 27 of 30 classifications matched human judgment (0.900). The three disagreements are transparently documented. This closes the PU-B exit criterion and establishes the human oracle that quality scores can be measured against.
Overnight Friday into Saturday, 9 PRs shipped the complete V1 Predictor stack into production proxy traffic:
lcr.EstimateTokensenable_s9_predictor toggle and PKD8 access controlThe score semantics are formally defined (D10):
Score = clamp((M_proj - p_low) / (p_high - p_low), 0, 1)"Score X means model M sits at X×100% of the curated anchor range for class C."
That is the complete quality score derivation in one line. It is auditable, reproducible, and explainable to a CFO.
Every API call through the proxy now produces four new fields in request_logs:
predicted_quality — the quality score for this specific requestprediction_confidence — how confident the classifier is in its class assignmentpredicted_quality_breakdown — the full 6-dimension vector (JSONB, access-controlled to org_admin+)predicted_quality_status — one of: scored, unscoreable_low, unscoreable_classifier_error, predictor_disabled, unscoreable_no_mqvaThe daily reliability bar (PR #606's U10) runs every morning at 07:00 UTC and emits predictor-bar-YYYY-MM-DD.{md,json} artifacts that track gate-drift over time. This is the same rigor that financial systems apply to model validation — because that's what it is now.
The quality score is not an academic exercise. It directly determines which model handles which request — and therefore how much each request costs.
Here's the current CostGoat pricing surface (May 23, 2026):
A 12× price delta at comparable quality. The quality score is the mechanism that identifies which requests can safely route from the expensive tier to the cheap one — without degrading user experience.
Every competitor in the AI gateway space routes on cost. Some route on latency. Trimio now routes on a continuous, ML-derived quality score grounded in live third-party benchmark data. That is a categorical differentiation — not a feature gap that closes with the next release.
The daily reliability cron (U10) checks four gates: R10, R12, R15, and R16. When all four pass, the quality score for that day is certified as reliable. When any gate fails, the score is flagged — and the routing engine knows to fall back to conservative behavior.
This is the same rigor that financial institutions apply to model validation. The difference is that most AI companies don't do model validation at all. They ship, hope, and apologize when the quality degrades.
Trimio's approach: every day, a structured report answers "is today's quality prediction trustworthy?" If the answer is no, the system degrades gracefully rather than silently producing bad scores.
The quality score chain is now complete: live benchmark ingestion → differentiated weights → human-validated accuracy → production deployment → per-request logging → daily reliability monitoring. Six steps, each auditable, each versioned, each defensible.
When a customer asks "how do you know this model will meet my quality bar?", the answer is no longer "trust us." It is: "Here's the score derivation. Here's the anchor list. Here's yesterday's reliability bar run."
That is the difference between a routing tool and a quality governance platform. The category is splitting. Trimio is on the right side of the split.
Trimio is the LLM API gateway with quality-aware routing — every call scored, every score auditable, every routing decision defensible under audit. See how it works.