AI Model Routing Cuts Costs Until the Router Guesses Wrong
Sending routine requests to cheaper models can lower inference costs. The savings survive only when the router detects mixed, unfamiliar, or high-risk requests before a bad answer triggers another call.
September 17, 2026 · 8 min read

The revealing case in my test queue looked routine at first: a support ticket combined a password-reset request with a disputed invoice and a request to change the account owner. A cheap model could explain the reset steps. It should not improvise the billing decision or ownership change, both of which depend on policy and verification.
That mixed ticket became the anchor for a hands-on comparison of three routing methods: fixed rules, a trained classifier, and a model-based router that reads each request before choosing another model. All three could divert obvious reset requests to the cheaper path. Their differences appeared when one message contained both easy language and a consequential task.
The key finding was not that one router always won. It was that cost depended on abstention, the router’s ability to decline a decision and escalate when evidence was weak. Without that escape hatch, the system often bought a cheap first answer and then paid for the expensive answer anyway.
The cheap call is not the full cost
A basic router runs before the answer model. It inspects the request, assigns a route, and sends the request either to a lower-cost model or to a more capable one. That makes a two-tier deployment newly practical: routine summarization, extraction, formatting, and established support flows can use the cheaper tier, while ambiguous or policy-sensitive work goes elsewhere.
The tempting calculation compares only model prices. If the inexpensive model costs less per token, every request moved onto it appears to create savings. Production systems pay for complete outcomes, though, not first attempts.
The useful accounting unit is the resolved request. Its cost includes the router, the selected model, any retrieval or tool calls, a second model after escalation, and human review when the first answer creates doubt. A mistaken cheap route incurs at least one avoidable model call. It can also expand the conversation, because the user restates the problem or challenges the answer, adding input and output tokens before the system reaches the route it should have chosen initially.
The password-reset-and-invoice ticket exposed that distinction. A routing decision based on its familiar password language reduced the first-call cost, but the resulting answer covered only the easy portion. The workflow remained unresolved. Once the application retried with the stronger model, routing had increased total inference rather than reduced it.
Quality loss is harder to price. A fluent answer can omit an account-control requirement without triggering an automatic retry, which leaves the cost dashboard looking healthy while the support team receives an incomplete case later. Any routing evaluation that records model spend but not task completion will favor aggressive cheap routing by design.
Rules are inexpensive and legible
A rules router uses explicit conditions such as request length, selected product, detected keywords, tool requirements, or account risk. It is the easiest version to inspect. An operator can read the condition that sent a ticket down one path, reproduce the choice, and change it without retraining anything.
For narrow workflows, that is a real advantage. A request produced by a fixed password-reset button can safely use a cheaper model if the application already knows the account state and the model only rewrites approved instructions. The rule is based on workflow context, not a guess about linguistic difficulty.
Rules became brittle when I applied them to free-form messages. The mixed ticket contained several signals that pointed in different directions. A rule that prioritized password-reset language under-routed it; a rule that escalated every mention of billing protected quality but also moved straightforward invoice-copy requests onto the expensive model. Adding more conditions improved individual cases while making precedence harder to reason about.
That does not make rules obsolete. They work best as hard boundaries. Requests involving identity changes, refunds above an organization’s chosen limit, regulated data, or irreversible tools should bypass probabilistic routing altogether. The router can still optimize inside the safe area.
Classifiers need an explicit reject option
A classifier, a model trained to assign categories, can learn patterns that would require an unwieldy rule set. In this workflow, its labels might be cheap-model eligible, strong-model required, or human review. Unlike keyword matching, it can use the whole request and learn that common words mean different things in different contexts.
Its weakness is the label set. If training examples mostly contain single-intent tickets, the classifier learns to select the nearest known category for a mixed request rather than recognize that the request falls between categories. A confidence score does not automatically solve this. Neural classifiers can produce a high score for unfamiliar inputs, so the number must be calibrated against held-out traffic instead of treated as a literal probability of correctness.
The practical setup is selective routing: accept the classifier’s decision only above a measured threshold, then send lower-confidence cases to the stronger model or a review queue. Teams should choose that threshold from the cost of each error, not from a round default such as 0.9. Sending a difficult request to a cheap model and sending an easy request to an expensive one have different consequences, and one threshold may hide that asymmetry.
Mixed-intent examples need their own evaluation slice. So do unusually long requests, new product names, code-switched language, and messages copied from another system. Overall routing accuracy can remain respectable while a small but expensive class of under-routes drives retries and repair work.
A model can route by reasoning, but it adds another inference
A model-based router prompts a language model to inspect the request and return a destination, often with structured fields for difficulty, risk, required tools, and confidence. It handles new wording without retraining and can identify that the password-reset-and-invoice ticket contains separate tasks with different handling requirements.
That flexibility costs tokens and latency on every request. If the router itself is a capable model, its call can consume part of the saving created by sending the answer to a cheaper one. If it is too weak, the architecture recreates the original problem: an inexpensive model must recognize work beyond its competence.
Explanations help with debugging but should not be mistaken for calibrated uncertainty. A router can produce a tidy rationale after making the wrong choice. In the hands-on comparison, the useful behavior was not detailed prose about difficulty; it was a structured escalation when the request combined account access, money, and a change of authority.
Constrain the output. The router should return a small schema containing the chosen route, detected task types, required capabilities, and an abstain flag. The application must validate that schema and apply hard policy overrides afterward. Free-form routing text adds parsing failures without improving the decision.
A model router is most defensible when requests change faster than a classifier can be labeled and retrained, or when the input contains several tasks that rules cannot separate cleanly. For a stable, high-volume flow with a small label set, a calibrated classifier is usually easier to measure and cheaper to run.
Test the router as part of the answer system
Router evaluation should begin with a no-routing baseline: send every test request to the stronger model, record task quality and end-to-end cost, then compare complete routed runs against that reference. Otherwise, a lower model bill can be reported as success even when completion falls.
The test set should preserve conversations rather than isolated prompts. A weak first answer changes the next user message, and that follow-up may be longer, angrier, or less explicit than the original. Replaying only the first turn removes the retry cost that routing is supposed to control.
For the support workflow, I would log the route, router confidence, policy overrides, answer-model choice, retries, eventual escalation, and final disposition. Those fields make the password-reset-and-invoice failure visible as one unresolved request with two model calls, rather than two unrelated successful API responses.
Deployment should start with shadow routing, where the router records its intended choice while the existing model still answers. The team can then inspect under-routes without exposing users to them. After launch, a small sample from every route needs human scoring, including cases the system marked as easy. Sampling only escalations reveals whether the expensive queue is justified but misses quiet failures in the cheap queue.
A router is worth operating when resolved-request cost falls while quality stays inside a documented tolerance. If the team cannot measure resolution, retries, and under-routing separately, direct use of one capable model is the cleaner setup. It may look more expensive per call, but the bill corresponds to a system with fewer hidden decisions.
Questions people ask
What is the simplest useful AI model router?
Start with hard workflow rules that use information the application already knows, such as whether a request can change account state or invoke a consequential tool. Route only a narrow set of repeatable, low-risk tasks to the cheaper model, and send everything else to the established model until logs show where further separation is safe.
Should a router use the cheapest available model?
Only if that model can recognize requests beyond its own answering ability. A very cheap router that underestimates difficult work creates retries, while an expensive router can consume much of the intended saving. Compare total resolved-request cost, including the routing call and escalations, rather than comparing model prices alone.
How should routing confidence be used?
Treat confidence as a control signal that requires calibration, not as a guaranteed probability. Measure how often each confidence range produces a correct route on representative traffic, then abstain below a threshold chosen for the cost of under-routing. High-risk workflow rules should still override the score.
When is model routing not worth the setup?
Skip it when traffic is modest, requests have similar difficulty, or the team cannot score answer quality and final resolution. A single capable model avoids router latency, monitoring, threshold maintenance, and double calls after mistakes. Routing earns its overhead only when a substantial, identifiable portion of work can safely use the cheaper path.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



