An essay by Sean Goedecke hit Hacker News this morning with 1,034 points and 438 comments. The thesis: LLMs don't democratize expertise — they amplify it.
The evidence is a conversation Terence Tao — one of the greatest living mathematicians — had with ChatGPT while exploring a potential counterexample to the Jacobian Conjecture. Tao's messages are short and precise. He routes the model into talking-to-mathematicians mode. When the output "looks weird," he redirects: "this looks more complex than I was hoping for." He makes the leaps; the model executes. The resulting exchange is, by Tao's own account, unusually productive.
Then Goedecke contrasts this with what happens when a novice tries the same model on the same type of task: they get stuck in a feature discussion vortex, never land the thing, and conclude the model wasn't helpful. Same model. Qualitatively different output. The difference is entirely the expertise of the prompter.
1,034 engineers upvoted this. They recognize it from their own work.
Here's what the essay doesn't address: if expertise amplification is real, your routing strategy is leaving money on the table — or quality on the floor — by treating all calls the same.
Goedecke identifies three things an expert does that a novice can't:
The HN thread added a useful counter-example: a hair stylist with zero coding background built a working Telegram bot and set up Arch Linux with Hyprland using Kimi K3 — because she was motivated and determined, which functions as a form of domain fluency. Motivation compresses the expertise gap by forcing the user to engage deeply enough to actually direct the model. But it's still the same mechanism: directed prompting from someone who knows enough to know when they're not getting the right answer.
A solo developer with Tao-level domain expertise can extract Tao-level output from a cheap model. But in an enterprise with 500 engineers deploying AI coding tools, the expertise distribution is wide. You have the senior engineers who direct the model precisely, and you have the juniors who don't know when the output is wrong and paste it anyway.
This creates a counterintuitive cost problem: your most expensive model calls are frequently being made by the people least equipped to extract value from them. The junior engineer who routes everything to Opus 5 because it "feels smarter" is getting worse ROI per dollar than the senior who extracts better output from Luna at one-fifth the price.
There's also a structural problem with agentic workflows. An agentic loop has different turn types:
Most enterprises route all three turn types identically. Opus 5 for everything, or GPT-5.6 Sol for everything, or whatever the team decided when they set up the integration. The cost bill accumulates uniformly across planning, execution, and evaluation — regardless of the fact that execution turns could be routed to V4 Flash at $0.14/MTok without meaningful quality degradation.
The Goedecke essay is empirical evidence for something Trimio was built around: not all turns in an AI workflow are equal, and routing should reflect that.
Specifically:
This is quality-aware routing in practice. The intelligence is in knowing which turn type is which — and routing accordingly, rather than defaulting to uniform expensive-model-for-everything or uniform cheap-model-for-everything.
The Goedecke essay identifies the cost of the novice failure mode: getting stuck in the feature discussion vortex, never landing the task. But there are two enterprise failure modes that mirror this:
Every call hits Opus 5 or Sol. The planning turns are excellent. The execution turns are 25× more expensive than they need to be. At enterprise scale — thousands of agentic workflow calls per day — the overpay on execution turns is material. A workflow with 20 turns where 15 are execution and 5 are planning is spending 75% of its token budget on turns that didn't need a frontier model.
The CFO saw the AI bill and ordered a cost reduction. Everything now routes to Luna or DeepSeek V4 Flash. The execution turns are now priced correctly, but the planning turns have degraded. The senior engineer who was previously able to direct Opus 5 precisely now gets a model that can't keep up with the domain sophistication of their prompts. Output quality drops. The "cheaper per turn" math looks good; the productivity math doesn't.
Neither failure mode is obvious from a dashboard that only shows aggregate cost. You need per-turn-type routing data to see it.
There's a second-order implication in the Goedecke essay that's easy to miss. He says: "LLMs make everyone a generalist but reward specialists." The specialist's advantage is structural — it doesn't disappear as models improve. It compounds through the model layer.
For enterprise AI infrastructure, this means:
1,034 engineers upvoted Goedecke's essay because they recognize the mechanism. They've seen a senior engineer get productive AI assistance on a hard problem that a junior engineer's identical prompt couldn't touch. They've experienced the difference between a model operating in expert-mode and a model explaining things to an amateur.
The essay frames this as a human skill gap. It is — but it's also an infrastructure opportunity. A routing layer that assigns the right model to the right turn type in an agentic workflow is doing for every engineer what Tao's expertise does for him: making sure the model's capability is matched to the task's requirement.
The alternative is letting every engineer route by instinct — which means your most expensive model calls go to the engineers least equipped to extract value from them, and your cheapest model calls go to the engineers most capable of producing good output from any model at all. That's the expertise amplification problem, running in the wrong direction.
Trimio's LCR engine routes agentic workflow turns to the model tier appropriate to the task — sending planning and direction calls to Opus 5, execution calls to Luna or V4 Flash, and surfacing per-engineer cost attribution so you can see the quality-cost tradeoff by team. See how it works.