Trimio Field Notes

What a Defensible AI Quality Score Actually Looks Like

May 23, 2026 8 min read quality-engineeringcost-governancemqva

Every AI gateway claims to route on cost. A few claim to route on quality. Almost none can answer the question a CFO will ask: "How do you know this score is real?"

That question is now answerable. As of this week, Trimio's Quality Prediction Engine (QPE) runs a complete, auditable chain from third-party benchmark data to per-request quality scores — and every step in that chain is documented, versioned, and defensible under audit.

The chain, end to end

The full quality score derivation
Third-party benchmark data (Artificial Analysis API) → 6-dimension normalization → differentiated class weights → anchor-normalized projection → per-request predicted_quality score in request_logs. Every step is versioned, documented, and auditable.

The quality score is not a vendor assertion. It is a derivable number with a published methodology. Here's the full chain:

Step 1: Live benchmark data ingestion (PR #608)

Previously, Trimio's model quality catalog was hand-curated YAML — a maintainer's nightmare and an auditor's red flag. As of PR #608 (May 22), a new MQVA (Multi-Quality Vector Aggregation) ingestor fetches live data directly from the Artificial Analysis API — an independent third-party benchmark aggregator — and writes provenance-stamped rows to the model catalog.

Key details:

When the ingestor first ran against the live API, it surfaced two compatibility bugs: the API returns status: 200 as an integer (not a string), and legacy seed data used UUID keys instead of model slugs. Both were fixed within the hour (PR #610). The fact that these bugs surfaced — and were fixed — is itself a quality signal: live API ingestion is the right architecture because it breaks visibly when things change.

Step 2: Differentiated projection weights (PR #600)

Even with good source data, a quality score is only useful if the projection weights meaningfully differentiate between model classes. Before PR #600, Trimio's dimension weights had cosine similarity of 1.0 across all 6 quality classes — meaning every class projected identically and the score carried no routing signal.

After PR #600: cosine similarity between any two class weight vectors is now ≤ 0.84. That's the threshold where the projection step actually discriminates between quality tiers. The practical result: 11 of 42 LCR routing cells flip to a different bucket when the differentiated weights are applied.

The routing signal
Before weight differentiation: all classes projected equally → quality score was meaningless. After differentiation: 11 of 42 LCR cells flip to a different routing bucket. The score now carries actual signal.

Step 3: Classifier accuracy gate (PR #596)

The quality prediction pipeline starts with a traffic classifier that assigns each request to one of 6 classes. That classifier needs to be accurate enough to trust its output. For three weeks, Trimio's classifier was stuck at macro 0.586 — below the 0.711 plan target.

PR #596 (May 21) shipped the consolidated classifier rework:

The classifier also gained a tool-calling labeler (Sonnet/Opus) with an enum-bound classify tool, eliminating the "prose-leak halts" that previously caused unpredictable classifier failures during training.

Step 4: Human spot-check oracle (PR #602)

A classifier that achieves 0.716 macro accuracy on its own benchmarks is useful. A classifier that achieves 0.900 agreement with a human reviewer on an Opus reference set is defensible.

PR #602 documents the 30-row hand-review: 27 of 30 classifications matched human judgment (0.900). The three disagreements are transparently documented. This closes the PU-B exit criterion and establishes the human oracle that quality scores can be measured against.

Step 5: The Predictor V1 in production (PRs #605–#607)

Overnight Friday into Saturday, 9 PRs shipped the complete V1 Predictor stack into production proxy traffic:

The score semantics are formally defined (D10):

Score = clamp((M_proj - p_low) / (p_high - p_low), 0, 1)

"Score X means model M sits at X×100% of the curated anchor range for class C."

That is the complete quality score derivation in one line. It is auditable, reproducible, and explainable to a CFO.

Step 6: Per-request quality logging (PR #606)

Every API call through the proxy now produces four new fields in request_logs:

The daily reliability bar (PR #606's U10) runs every morning at 07:00 UTC and emits predictor-bar-YYYY-MM-DD.{md,json} artifacts that track gate-drift over time. This is the same rigor that financial systems apply to model validation — because that's what it is now.

The CFO test
When a CFO asks "where does this quality score come from?", the answer is now: "Artificial Analysis benchmark data, normalized against a 300-row Opus reference set, projected through differentiated class weights, with 0.900 human agreement, re-evaluated daily." That is a defensible number. It is not a marketing claim.

Why this matters for routing decisions

The quality score is not an academic exercise. It directly determines which model handles which request — and therefore how much each request costs.

Here's the current CostGoat pricing surface (May 23, 2026):

GPT-5.5 (quality 100)
$30/M output
Value score: 3.3
→
Groq 4.3 (quality ~89)
$2.50/M output
Value score: ~35

A 12× price delta at comparable quality. The quality score is the mechanism that identifies which requests can safely route from the expensive tier to the cheap one — without degrading user experience.

Every competitor in the AI gateway space routes on cost. Some route on latency. Trimio now routes on a continuous, ML-derived quality score grounded in live third-party benchmark data. That is a categorical differentiation — not a feature gap that closes with the next release.

The reliability bar

The daily reliability cron (U10) checks four gates: R10, R12, R15, and R16. When all four pass, the quality score for that day is certified as reliable. When any gate fails, the score is flagged — and the routing engine knows to fall back to conservative behavior.

This is the same rigor that financial institutions apply to model validation. The difference is that most AI companies don't do model validation at all. They ship, hope, and apologize when the quality degrades.

Trimio's approach: every day, a structured report answers "is today's quality prediction trustworthy?" If the answer is no, the system degrades gracefully rather than silently producing bad scores.

The bottom line

The quality score chain is now complete: live benchmark ingestion → differentiated weights → human-validated accuracy → production deployment → per-request logging → daily reliability monitoring. Six steps, each auditable, each versioned, each defensible.

When a customer asks "how do you know this model will meet my quality bar?", the answer is no longer "trust us." It is: "Here's the score derivation. Here's the anchor list. Here's yesterday's reliability bar run."

That is the difference between a routing tool and a quality governance platform. The category is splitting. Trimio is on the right side of the split.

Trimio is the LLM API gateway with quality-aware routing — every call scored, every score auditable, every routing decision defensible under audit. See how it works.

Trimio
Quality you can prove.
Live benchmark data. Human-validated accuracy. Daily reliability reports. Not marketing claims — auditable numbers.