Skip to content

AI Industry & Models

Model Routing Cuts Cost Only When Wrong Routes Are Counted

Routers can send routine prompts to cheaper models and reserve stronger models for difficult cases. The savings disappear if quality misses and second calls stay outside the ledger.

Tobias LundIndustry & Models Writer

September 22, 2026 · 7 min read

A laptop showing a support-ticket routing log with cheap, strong and escalated model paths.
A laptop showing a support-ticket routing log with cheap, strong and escalated model paths.

Consider one support ticket asking whether an order qualifies for a refund. The application can send a routine request with a clear purchase date to a low-cost model, while a ticket involving a policy exception, contradictory account notes or missing order data goes to a stronger model.

That split is newly practical because routing has moved from custom application logic into managed services and open-source tooling. In December 2024, Amazon Bedrock introduced intelligent prompt routing, which its documentation says selects between models in the same family based on predicted response quality. AWS advertises cost reductions of up to 30% while maintaining quality, a vendor result rather than an independent guarantee. OpenRouter also documents an Auto Router that chooses a model from a supported pool, while the 2024 RouteLLM project showed how teams can train a router from preference data.

The mechanism is straightforward. A router, meaning a classifier that chooses which model receives a request, inspects the support ticket and predicts whether the cheaper model can answer it. That decision may cost far less than generating the answer, but it creates a new failure point: the router can classify a difficult refund case as easy, produce a fluent wrong answer and never trigger the expensive model.

A lower invoice proves little unless that mistake appears in the evaluation.

Label the decision the router should have made

Start with a fixed test set of real, redacted support tickets and send every ticket to both candidate models. Score each output against the same rubric: whether it identifies the applicable policy, uses the correct order facts, avoids unsupported claims and takes the permitted action. A human reviewer or a separately validated model judge can apply the rubric, though model-based grading needs periodic human checks because a judge can share the answer model’s blind spots.

The useful label is the cheapest model that passes the rubric. Evaluation teams sometimes call this an oracle route: the best routing decision known after all candidate outputs have been inspected. If the cheap model passes, its ticket receives a “cheap acceptable” label even when the stronger model writes a more polished reply. If only the stronger model passes, the correct label is “strong required.

” When neither passes, the router has no correct model to choose; the prompt, retrieval step or underlying workflow needs repair.

That produces four measurable outcomes. A correct cheap route saves money without crossing the quality threshold. A correct strong route protects quality. A false expensive route wastes money by sending an easy ticket to the stronger model.

A false cheap route is the dangerous one because the system returns an unacceptable answer while reporting a successful low-cost call.

The two errors should not carry equal weight. For the refund workflow, a false expensive route adds model spend and perhaps latency, while a false cheap route can misstate policy or authorize an action the support system should reject. Set separate limits for each error rather than collapsing both into one routing-accuracy percentage.

Keep the label beside the production decision. The support-ticket log should record a request identifier, the router’s chosen model, its confidence or score, the oracle route when available, the rubric result, input and output token counts, latency, escalation reason and final model used. Without the oracle label, the log shows what happened but cannot show whether the route was correct.

Observe the answer the router hid

A router prevents teams from seeing the main counterfactual: what the other model would have returned. If production sends the refund ticket only to the cheap model, a normal pass-or-fail review can detect a bad answer, but it cannot determine whether the stronger model would have fixed it. The reverse matters too. Sending a ticket to the costly model does not prove that the cheaper one would have failed.

During initial evaluation, run every test prompt through both models. After launch, use shadow evaluation, where a sampled request also runs against an alternative model without exposing that second answer to the user. The shadow call costs money, so the sample can shrink once error rates stabilize, but it should remain large enough to include the difficult ticket categories the router rarely sees.

Sampling only random traffic can miss those categories. Stratify the evaluation by documented prompt classes such as routine status checks, refund eligibility, policy exceptions and requests requiring account tools, then report false-cheap rates for each class. A router that looks acceptable across all traffic may still mishandle most policy exceptions if routine questions dominate volume.

Routing confidence also needs calibration. A score of 0.8 should correspond to a known likelihood that the selected cheap model will pass; otherwise, the number is merely useful for ranking. Plot pass rates by score band, choose the escalation threshold from observed results and rerun that calibration after changing either model, the system prompt, retrieval data or available tools.

Return to the refund ticket after each change. A stronger cheap model may turn yesterday’s “strong required” examples into safe low-cost routes, while a revised policy document can make old evaluation labels obsolete even though neither model changed.

Price the whole escalation path

The cheapest first call is not necessarily the cheapest completed ticket. Calculate cost from the full path:

`route cost = router cost + first model call + retrieval or tool charges + retry calls + escalation call`

For token-priced APIs, calculate each model call from its input tokens multiplied by the published input rate, plus output tokens multiplied by the output rate. Record the price sheet and date used because vendors change rates, discounts and caching terms. If a provider bills cached input, batch processing or tools differently, those charges belong in the same calculation.

There are two common escalation designs. A pre-generation router sends the refund ticket directly to one model, which avoids paying for a discarded cheap answer but depends heavily on classification quality. A generate-then-check design lets the cheap model answer first and escalates when a verifier rejects that answer; it may catch more failures, yet each rejected ticket pays for the cheap generation, the verification step and the stronger generation.

Some applications run two models in parallel and choose between their answers. That can reduce the latency of an escalation because both outputs arrive around the same time, but it pays for both on every routed request. For cost reduction, parallel generation is usually the wrong default unless the workflow values response time more than inference spend.

Price these branches separately. Report the share sent directly to the cheap model, the share sent directly to the strong model and the share that paid for more than one generation. Then calculate cost per accepted answer, not cost per API call. A failed cheap response followed by a successful strong response is one completed support ticket with two generation bills.

Latency needs the same treatment. Compare median and tail latency for completed requests, including retries and tool calls, rather than quoting the cheap model’s isolated response time. A routing classifier may add little delay, while a sequential escalation can make the hardest tickets wait through two full generations.

Keep one model as the control

The baseline should be deliberately boring: send every ticket to the stronger model using the same system prompt, retrieved documents, tools, output format and acceptance rubric. Measure its total cost, pass rate and completed-request latency on exactly the same ticket set used for the router.

This control answers the adoption question. A router is useful when it lowers cost per accepted answer while keeping false-cheap errors within the workflow’s limit. If it cuts raw model spend but lowers the acceptance rate, its apparent savings include work transferred to human agents, retries or customer corrections. Add those downstream events where they can be measured; otherwise, report them separately rather than assigning invented dollar values.

Do not compare a routed system against an older prompt or a weaker baseline configuration. The routing layer gets credit only for the difference created by model selection. Retrieval improvements, prompt rewrites and new tools should be tested independently or applied to both sides.

Run the comparison again whenever a model alias changes, a vendor retires a version or the prompt distribution moves. Managed routing reduces the engineering needed to choose a model, but it does not remove ownership of the decision. AWS’s documented router constraints, OpenRouter’s supported model pool and RouteLLM’s learned preferences all define which alternatives the router can consider; none can label the acceptable answer for a company’s refund policy.

The support-ticket log is the deciding artifact. If it can show the route taken, the route that should have been taken, every paid call and whether the final answer passed, the cost claim is testable. If it records only the selected model and its invoice, the savings are an assumption.

Questions people ask

How much production traffic should run through both models?

Run the full offline test set through both models before launch, then shadow a sampled share of production traffic. Choose the sample from traffic volume and the rarity of high-risk cases, and stratify it by ticket type so uncommon policy exceptions do not disappear inside a large pool of routine requests.

Can a model grade the router’s decisions?

Yes, a model judge can apply a fixed rubric at scale, but teams should compare its grades with human review before relying on it. Recheck disagreements and a continuing sample because model judges can favor verbose answers, miss policy details or rate outputs from related models inconsistently.

When is model routing not worth deploying?

Routing is usually not worth the added system when nearly every request needs the stronger model, model prices are close, or wrong answers carry more cost than the available evaluation can measure. In that case, one model produces a clearer bill, a shorter latency path and fewer hidden failure modes.

What metric should decide whether the router ships?

Use cost per accepted answer, paired with a separate ceiling for false-cheap routes. Compare both with the all-strong-model baseline on the same prompts, tools and rubric; a lower cost per call is insufficient when the routed workflow creates more rejected answers or paid escalations.

ShareFacebook
model evaluationai pricing and accessdeveloper toolingmodel routingllm evaluationinference costsmodel selection

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read