Model Routing Cuts AI Costs Until It Misreads One Refund
A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.
October 1, 2026 · 7 min read

The concrete case is a customer-support refund workflow. A ticket arrives with the customer’s message, order record, and a policy excerpt; the system must draft a reply and decide whether the case can proceed automatically or needs review.
Without routing, every ticket goes to the strongest approved model. That is easy to operate and expensive at volume. With routing, a small classifier or language model scores each request, sends apparently routine work to a lower-cost model, and reserves the stronger model for cases predicted to be difficult.
The dangerous ticket is short: “The replacement failed too. Please refund the original payment method.” Its length suggests an easy request. The second failed item makes it a policy exception, while “original payment method” may conflict with rules for replacements, store credit, or expired cards.
A router that treats prompt length as complexity can make the cheapest model responsible for the ticket that needs the most policy reasoning.
That failure belongs to the router only if the stronger model would have handled the same evidence correctly. This distinction is the center of the test.
The router makes a decision before the model does
A model router is a service that selects one model from an approved set for each request. It may use visible features such as token count and requested output format, a learned complexity score, or predictions of whether each candidate will produce an acceptable answer.
For the refund workflow, the router receives the customer message plus structured fields from the order system. It should choose between at least two candidates: a lower-cost model for routine cases and a stronger fallback for ambiguous policy work. More candidates can improve the cost curve, but each additional model creates another boundary that must be tested and monitored.
The execution order matters. First, the application retrieves the relevant policy text. Second, the router sees the complete prompt, including that text, rather than the customer’s message alone. Third, the selected model returns both a proposed action and a structured confidence signal.
Finally, application code checks hard rules before any refund action or customer reply proceeds.
Routing before retrieval is cheaper, but it hides the feature that makes the replacement ticket difficult. A two-sentence request can expand into conflicting policy clauses once the correct documents are attached. The router then needs either the retrieved context or a separate signal stating that the order has prior replacements, payment complications, or another exception flag.
Complexity is not prompt length
Length-based routing is attractive because token count is available before inference and costs almost nothing to calculate. It works for obvious extremes: a one-line classification request usually needs less reasoning than a long document comparison. The middle contains the expensive mistakes.
In the refund fixture, difficulty comes from interactions between fields. A request can be short while the order record shows an earlier replacement, a partial refund, or an account mismatch. Conversely, a long complaint may still reduce to a standard late-delivery response once the system checks the tracking state.
A practical complexity router should therefore score documented conditions rather than prose alone. For this workflow, useful signals include whether the requested remedy differs from the default policy and whether retrieved passages disagree. Those features describe the work the model must perform; character count merely describes the package it arrived in.
Latency targets create a different trap. Sending every urgent request to the fastest model can reduce time to first token, yet a fast wrong answer followed by escalation takes longer than one correct pass through the stronger model. Measure end-to-end completion time through the final accepted draft, including retries and review, rather than timing only the first response.
Predicted-quality routing is more direct. The router estimates which candidate will clear a defined acceptance threshold, then chooses the cheapest one expected to pass. That makes cost control newly practical at the request level, but only if “pass” means something testable: the refund action matches policy, cited order facts are present, and unsupported promises are absent.
Build the answer matrix before tuning the router
Start by running every evaluation ticket through every candidate model. This produces an answer matrix, a table showing which models passed each case under the same prompt, context, tool results, and grading rules.
The matrix separates model failure from routing failure. If all candidates mishandle the replacement refund, the prompt, evidence, grading rule, or model set is at fault. If the stronger model passes while the selected cheaper model fails, the router made an avoidable assignment. If both pass, choosing the cheaper candidate was correct even when their wording differs.
For each ticket, label the cheapest model that meets the acceptance threshold. That label becomes the oracle route, meaning the best known assignment after all candidate outputs have been graded. The production router is then compared with the oracle, not merely with the strongest model.
Two errors matter. Under-routing sends a case below the cheapest adequate model and risks a bad outcome. Over-routing sends it above that model and wastes money without improving the accepted result. Combining both into one accuracy score conceals the tradeoff, because a router that always selects the strongest model can look perfect while saving nothing.
Track routing regret as well, which is the added cost or quality loss relative to the oracle route. For a passed response, cost regret is the selected model’s request cost minus the oracle model’s cost. For a failed response, record quality regret separately rather than pretending that a refund-policy error can be converted cleanly into token spending.
The replacement ticket should remain visible as an individual row. Aggregate averages can hide a small class of exception cases, particularly when routine delivery and cancellation messages dominate the test set.
A fallback set needs different failure modes
A useful fallback set is not a ladder of models ranked only by size. Each candidate needs a role and an exit condition.
The first model handles tickets whose retrieved policy contains one clear rule and whose order record has no exception flags. The stronger model receives disagreements, missing evidence, and cases where the requested remedy conflicts with the default rule. A human queue remains the fallback when required facts are unavailable or policy language does not determine an action.
Retries should change something. Repeating the same prompt against the same model after a policy error adds cost without supplying new evidence. A valid escalation can switch models, retrieve a narrower policy section, or request human review. The log should state which condition caused that escalation.
For the replacement refund, the decisive log line is not merely `model=strong`. It is a routing record that preserves the selected model, predicted pass probability, relevant exception features, fallback trigger, and final disposition. That record lets an evaluator reconstruct whether the router saw the replacement history and ignored it, or never received it.
Keep hard controls outside the model. Even a strong model should not issue a refund when the application lacks a verified order identifier or when the amount exceeds the workflow’s approval boundary. Routing chooses who reasons over the case; it does not replace authorization.
Test the router in shadow mode first
Shadow mode means the router makes a selection without controlling the live response. The existing production model still answers, while the proposed router’s choice is logged and evaluated offline.
During this stage, send sampled tickets to the full candidate set and grade them with the same rubric. Sampling matters because running every live request through every model would erase the intended savings. Include all known exception categories, then add a changing sample of ordinary traffic so the test does not become a fixed benchmark that the router can memorize.
Before launch, set separate limits for under-routing and over-routing. The acceptable under-routing rate should be tighter for actions with customer or financial consequences than for low-risk drafting. Over-routing can usually tolerate a wider band because its immediate cost is money and latency, although persistent over-routing means the system is not earning its operational complexity.
After launch, audit by route boundary. Compare cases sent to the cheaper model with similar cases escalated to the stronger one, and review changes in retrieved context or order-state features. A single overall pass rate will not show whether a model degraded or the router started assigning it harder work.
Automatic routing is worth deploying when the answer matrix contains a substantial class of prompts that a cheaper candidate passes consistently and when the application can detect the exceptions. If the models fail on the same tickets, or the necessary routing features arrive only after the decision, a fixed strong model is the cleaner setup.
Questions people ask
How many models should a routing fallback set include?
Start with two models and a human or rules-based exit. One cheaper model creates the savings case, while one stronger model tests whether escalation repairs it. Add another model only when the answer matrix shows a distinct group of tickets that it handles at a better cost-quality point.
Should an
AI router use prompt length to estimate difficulty?
Prompt length can identify broad extremes, but it should not decide consequential cases by itself. The replacement-refund example is short yet difficult because the order history and policy interaction create the reasoning burden. Route on retrieved evidence and structured exception signals when those fields are available.
How do
I tell a routing error from a model error?
Run the same case through every approved candidate under identical conditions. If a stronger candidate passes and the selected model fails, count an avoidable routing error. If every candidate fails, investigate the prompt, retrieved evidence, grading rubric, and model set instead of blaming the router.
Does routing always reduce AI costs?
No. Classification calls, duplicated evaluation traffic, retries, and unnecessary escalation all add cost. Routing pays when a meaningful share of production requests can use a cheaper model without increasing failed outcomes, and when the savings exceed the router’s own inference and operating overhead.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



