The question is usually framed wrong.
People ask whether local LLMs can replace frontier models. Today, for serious coding and reasoning work, the answer is no. They are smaller, less consistent, and easier to break.
But replacement is the wrong bar.
The better question is: how many of your prompts never needed a frontier model in the first place?
If a local model can answer those prompts quickly and correctly, every one of them is a small cost removed from the system. The harder prompts can still go to a frontier model. The user should not have to decide manually. The router should.
That is the idea I wanted to test. So I built one.
The boundary keeps moving
There is a line, somewhere in your workload, where local models stop being good enough and frontier models become worth paying for. Above the line, you want the smartest, most expensive model you can get. Below the line, you are paying a frontier price for a problem a much smaller model could have solved.
Most people draw that line by picking one model and using it for everything. Frontier users overpay for trivial questions. Local-only users get burned by quality drops on harder work. Neither is honest about the fact that the line exists.
The line moves with hardware. It moves with the specific local model you choose. It moves as new local models come out. Two months from now it is somewhere else. Two years from now it is much further to the right.
If the line moves, the right architecture is not a single model. It is a router that finds the line and respects it.
What the benchmarks showed
I ran the same set of prompts through sixteen models: four local on my GPU, twelve frontier or open models through their APIs. Same coding tasks, same difficulty levels, same judging rubric.
A few results mattered more than the rest.
The right local model handles a meaningful slice of real work. On trivial and easy prompts, qwen2.5-coder:7b scored 5/5/5 on correctness, completeness, and clarity. That is the same median as gpt-5.4-nano and claude-haiku-4-5, two of the cheaper frontier tiers. On medium prompts it slipped to 4/4/5: still mostly correct, slightly less thorough. On hard prompts the gap opened up, mainly on completeness. The local model wrote clean code but missed edge cases and error handling that frontier models covered.
That pattern is exactly what the routing design needs. The prompts where the local model matches frontier output are the ones it should handle. The prompts where it loses thoroughness are the ones worth escalating. The benchmark did not just say “local is good enough sometimes.” It told the router where the boundary sits.
The wrong local model is worse than paying for frontier. qwen3:4b is a smaller model, and on paper it should be cheaper to run. In my benchmark run, it burned thousands of tokens on internal reasoning before producing a visible answer, and the answers were often weaker. Smaller does not mean cheaper if the model spends its budget thinking out loud.
Frontier models still earn their cost on harder prompts. Architecture questions, security review, multi-step reasoning, and subtle correctness bugs are where local models start to slip. The gap is real. It would be dishonest to pretend otherwise.
The takeaway is not that local models won. It is that the routing gradient exists. Some prompts belong local. Some belong frontier. Most LLM tools today force you to pick one tier and live with the trade-off. The specific model names in these results will age. The gradient won’t.
My results are not your results
I should be upfront about the hardware.
I am running this on an AMD Ryzen 9 9950X3D with 128 GB of RAM and a Radeon RX 7900 XTX with 24 GB of VRAM. That is high-end consumer hardware, not a five-year-old laptop. The local models I can run quickly, and the ones I can run at all, are partly a function of that GPU.
Your results will look different. On a smaller GPU or CPU-only setup, fewer prompts can stay local. Some models I treat as the default will be too slow or too memory-hungry for your machine. The cutoff line moves with the metal.
That does not change the architecture. It changes which prompts land on which side of the line. On weaker hardware, the router sends more prompts to frontier models. On stronger hardware, it keeps more local. The point is not a universal benchmark number. The point is a routing layer that adapts to the machine in front of it.
The router
I built a small FastAPI gateway. It runs locally on my machine at 127.0.0.1:8001 and exposes OpenAI-compatible endpoints. Anything that can talk to OpenAI can talk to it.
The behavior is simple. Every incoming prompt gets scored by a classifier. The score maps to one of three tiers: local, mid, or frontier. Local goes to Ollama on my GPU. Mid and frontier go to specific provider APIs. The user does not see the routing. They send a request, they get a response, the cheapest capable model handled it.
The classifier itself is not magic. It looks at the user’s actual message and scores it on signals: code blocks, architecture vocabulary, multi-step instructions, comparison verbs, simplicity hints, and so on. Each signal contributes points to a complexity score. Score thresholds map to tiers.
The interesting design choice was what the classifier looks at. Coding tools like Claude Code, Codex, and Letta wrap every short user message in large framework scaffolding: system prompts, tool definitions, permission context, prior tool results. A “what is 2 plus 2” question can arrive as a thirty-thousand-character payload. If the classifier scores the whole blob, every prompt looks complex. So the classifier extracts the last user-authored text block and scores only that. The framework noise is ignored.
That single decision is what made the router usable inside real tools instead of just curl. A router is not really routing prompts until it can separate the user’s intent from the tool’s scaffolding. Naive routers fail this step silently. They route everything to the most expensive tier because the prompt looks complex, when most of that complexity belongs to the framework, not the user.
Don’t optimize for cheapest. Optimize for “no worse than ideal.”
The point of routing is not to keep as much as possible local. The point is to keep as much as possible local without losing quality.
The way I measure that is a metric I call “no worse than ideal.” For each benchmark prompt, the dataset has a label for which tier should be enough. The classifier passes if it routes that prompt to a tier at least as capable as the ideal. Routing an easy prompt to a frontier model is fine, just a little wasteful. Routing a hard prompt to a local model is a failure.
On my benchmark, the classifier hit one hundred percent on “no worse than ideal” across twenty-five prompts. Exact-tier match was nineteen out of twenty-five. The remaining six were over-routed by one tier: five easy or trivial prompts sent to mid instead of local, and one medium prompt sent to frontier instead of mid. None were under-routed.
Twenty-five prompts is a small benchmark, and the result is directional rather than definitive. The interesting claim is not the exact percentage. It is the shape of the failures. When this classifier is wrong, it is wrong in the cheap direction. An over-routed prompt costs cents. An under-routed prompt costs trust. The router should bias toward escalation when it is unsure, and on this dataset, it does.
The messy part
Most write-ups stop at the curl test. The harder work was making the router behave inside actual tools.
The most revealing problem came from Codex. The Codex CLI uses the OpenAI Responses API, which is a different format from Chat Completions, so the gateway has to translate between them, including the full streaming event sequence. That part was a protocol quirk and a fixable amount of code. The harder problem was authentication. When Codex is logged in through a ChatGPT account, it refuses to honor a custom base URL at all. The request goes to OpenAI no matter what the config says. That is not a protocol limitation. It is a vendor decision. If you want routing, you have to authenticate Codex with an API key instead of the subscription you are already paying for.
Claude Code and Letta Code had their own quirks. Claude Code sends every request with a fixed model name regardless of difficulty, so the gateway has to know to ignore that name and apply the classifier instead. Letta Code’s local backend bundles tool definitions into the user message as content blocks, which inflates the apparent prompt size and confuses naive classifiers. Both were fixable with specific code: a list of model names to bypass, a parser that pulls only the user’s last text block out of a content-block array, a fallback that accepts placeholder API keys and uses the gateway’s own.
None of that is glamorous. All of it is necessary if routing is going to be invisible to the tool you actually use to write code. And the Codex auth restriction is worth noticing on its own, because it is the cleanest example of what the rest of this essay is building toward.
What this proves
The router works today. On my hardware, for my workload, a meaningful share of prompts stay local. The rest go to frontier models. I do not pick the model for each request. The router does. In routine use, quality has not meaningfully dropped, because the router escalates anything it is unsure about. Cost should be lower whenever routine prompts stay local, without forcing hard prompts onto weaker models.
That is the immediate result. The pattern’s value grows as local models improve, but the router itself does not have to change. A single-model architecture cannot say the same.
Why the repo is public
I am putting the code on GitHub. The repo is called LLMux. To be clear about what it is: this is reference code, not a product.
It runs on my machine. It is the evidence behind this article. If you want to read the code, fork it, take ideas from it, or run a version yourself, that is fine. But I am not promising backward compatibility, issue triage, or feature requests.
The reason to make it public is not adoption. It is honesty. The argument is more credible when the implementation is visible. And the pattern should not be trapped behind a company paywall.
Routing should not be a feature of the platform
Hosted platforms are starting to ship their own routing. That is good for users in the short term, because someone else does the work, the price drops, and the experience stays simple.
It is risky in the long term. When routing lives inside the platform, the platform decides the trade-off between cost, latency, quality, privacy, and margin. The user has no way to inspect the decision and no way to override it. Every prompt gets routed according to the vendor’s incentives, not the user’s. The Codex auth example earlier in this essay is the small version of the problem. The hosted-routing version is the big one.
A router you run yourself reverses that. You see the score. You see the tier. You see which model handled the request and what it cost. You can override a single prompt with a directive in the message. The trade-off is yours, not the platform’s.
That distinction is going to matter more as the local tier gets stronger. The day local models can handle a meaningful share of the work that today requires frontier APIs is the day platform-controlled routing becomes a real margin lever. The pattern should be in users’ hands before that day arrives, not after.
What comes next
The router today makes a decision and lives with it. It does not yet know when it was wrong.
The repo has the start of a feedback layer. Every request gets logged with its classifier score, routing decision, and response. The user can prefix a follow-up with !escalate to push the next attempt higher, and the gateway retroactively labels the previous request as under-routed. A separate process samples logged requests and asks a judge model whether the right tier was chosen.
None of that has turned into a learned router yet. The router has memory now. It has not learned anything from that memory. That is a separate article, if there is enough data to say something honest about it.
The bet
The bet is not that local models are frontier models. They are not, today, and there is no point pretending otherwise.
The bet is that the line between “local is enough” and “send it to a frontier model” keeps moving in one direction. New local models keep getting smaller, faster, and more capable. Consumer hardware keeps getting more memory and more compute. The local model anchoring this router did not exist a year and a half ago. The cutoff line has moved a long way in a short time, and it is not going to stop.
The router does not need to predict where the line goes. It just needs to respect the line wherever it currently sits. That is a pattern worth getting right early, because the longer you spend hand-picking models for every prompt, the more habituated you get to a trade-off that does not need to exist.
Don’t pick a model. Route the prompt.