Skip to content

AI Industry & Models

Model Routing Saves Money Only When You Can See Its Mistakes

A cheaper model can handle routine requests, but one bad routing decision may create retries and manual review. Logging the decision makes the real savings measurable.

Tobias LundIndustry & Models Writer

September 24, 2026 · 7 min read

A support dashboard showing one ticket's model route, escalation reason, retry count, and final resolution.
A support dashboard showing one ticket's model route, escalation reason, retry count, and final resolution.

Consider one support ticket, labeled R-1842 in a test fixture. A customer says an order arrived late, asks for a shipping refund, and mentions that a previous agent already promised account credit. The company has two language models available: a cheaper model for routine support and a more capable model for requests involving policy interpretation or conflicting account history.

A model router, which selects a model before the answer is generated, reads R-1842 and marks it routine. The cheaper model writes a fluent refusal because the standard refund window has closed. Nothing crashes. Latency looks normal, the response passes a basic format check, and the invoice shows the lower model cost.

The routing decision was still wrong. The earlier promise should have triggered escalation, and the customer now replies, another model handles the follow-up, and a human eventually reviews both messages. What looked like one inexpensive inference has become several inferences plus staff time.

That ticket is the useful test for a routing system. If an operator cannot reconstruct why R-1842 took the cheap path, which alternatives were available, and when the mistake became visible, the reported savings are incomplete.

Cheap routes hide a different failure mode

Direct model selection is easy to audit: an application calls one named model, records the response, and attributes the cost to that call. Routing inserts another decision before generation. The application may use rules, a small classifier, or a language model to estimate difficulty, risk, or required capability, then send the request down a cheaper or more expensive path.

The mechanism can reduce spend because many workloads contain repetitive requests that do not need the strongest available model. It also creates a classification problem. A false escalation costs more than necessary because an easy request reaches the expensive model. A missed escalation is harder to spot because the cheaper model may still produce polished text, valid JSON, or syntactically correct code.

R-1842 is a missed escalation. Its visible answer is plausible while its policy judgment is wrong, which means uptime dashboards and parser-error counts will stay green. Even an answer-quality monitor may miss the problem if it checks tone, length, and prohibited phrases rather than whether prior account commitments changed the correct outcome.

Routing therefore changes what teams must observe. Model output alone is insufficient. The decision that selected the model becomes part of the product behavior, and it needs the same durable record as the final answer.

The routing log needs an audit trail

A useful log record starts before the first model call. Give each request a stable identifier, preserve the router version, record the candidate models, and store the chosen route with a machine-readable reason. For R-1842, that might include `route=cheap`, `reason=routine_refund`, and a score representing the router's confidence that escalation was unnecessary.

The record also needs the evidence available at decision time. That does not require copying an entire private conversation into an analytics table. Teams can log structured signals such as whether prior-agent history was present, whether the request cited a promise, how many policy categories matched, and whether an attachment or tool result was missing. Sensitive text can remain in the controlled system that already holds the ticket, referenced through the request identifier.

Next, capture what happened after routing: the selected model, token usage, model-call status, validation result, retries, later escalation, human override, and final disposition. A request that reached the cheap model once and closed successfully has a different cost profile from one that triggered two retries before a reviewer rewrote the answer.

The override deserves its own field. If an agent changes R-1842 from “refund denied” to “credit honored,” the system should record the original route, the corrected route or action, and a standardized correction reason. Free-text notes are useful for investigation, but fixed labels make it possible to count how often the router misses prior commitments, ambiguous policy language, or requests spanning more than one intent.

Prompts and rules change, so the log must preserve versions rather than only current configuration. Without a version identifier, a team can discover that missed escalations rose during a week but cannot tie the change to a revised classifier prompt or threshold. A rollback then becomes guesswork.

Logging adds storage, engineering work, and governance obligations. Those costs are real. They also make a routed system newly practical to operate, because savings can be attributed to a particular decision policy rather than inferred from a lower aggregate model bill.

Escalation tests should target the boundary

A general model benchmark does not test the router. The router needs a labeled set of real request shapes where the expected decision is known: cheap path, expensive path, human review, or rejection before any model call. The label concerns routing, even when evaluators also score the eventual answer.

Start with cases close to the boundary. An ordinary shipping-delay question may clearly belong on the cheap path, while an explicit legal threat may clearly require another workflow. R-1842 matters because one phrase about a prior promise changes the route even though most of the ticket resembles a routine refund request.

Build such cases from documented production corrections and approved synthetic variations. For the refund fixture, remove the prior promise, paraphrase it, place it late in a long message, or include a contradictory account note. The expected route should change only when the policy-relevant evidence changes. This catches routers that rely on obvious keywords while ignoring context.

Run the suite whenever the router prompt, classifier, threshold, candidate model, or support policy changes. Track a confusion matrix, a table that separates correct routes from false escalations and missed escalations, but weight the cells according to their operational impact. Sending a routine FAQ to a premium model wastes inference spend. Sending a refund exception to the cheap model can generate customer contact and staff review, so equal error counts do not imply equal cost.

Production traffic should also run in shadow mode before a switch, meaning the new router makes decisions without controlling the live response. Compare its proposed route with the current system, then have reviewers inspect disagreements that carry high rework risk. Shadowing adds evaluation calls if the router itself uses a model, yet it avoids learning about systematic under-escalation from customers.

Keep a permanent holdout set that prompt authors do not tune against. If every failure from R-1842 becomes a visible training example, the router may learn that exact wording while continuing to miss less familiar statements of the same obligation. The holdout shows whether the routing rule generalized.

Savings have to survive the rework calculation

The first financial comparison is straightforward: estimate what the same traffic would cost if every eligible request used the more capable model, then compare it with router calls plus the mix of cheap and expensive model calls. Use provider invoices or internal token accounting rather than a theoretical list-price average when discounts, cached input, or batch processing affect the bill.

That calculation is gross savings. Net savings subtract the router's own inference and infrastructure cost, repeated calls, human review, support follow-ups, and any compensation linked to a bad automated decision. The accounting period must be long enough to catch delayed rework; R-1842 looks closed after the refusal and becomes expensive only when the customer replies.

A practical unit is cost per resolved request, where “resolved” has a documented business definition rather than meaning that the model returned text. For customer support, resolution might require no reopen within the team's normal service window and no corrective human action. For code generation, it might require passing the repository's tests without a later repair. Different workflows need different closure signals.

Quality belongs beside cost. Report the share of requests sent to each route, missed-escalation rate on the labeled set, production override rate, and cost per resolved request. A lower average model charge accompanied by more overrides is a transfer of cost from inference to operations, not proof that routing worked.

The decision threshold can then reflect the workflow. Raising escalation sends more uncertain cases to the capable model, increasing inference cost while reducing exposure to missed escalations. Lowering it does the reverse. There is no universal setting because the price of a wrong refund decision differs from the price of an imperfect internal summary.

Roll out where outcomes can be checked

Begin with a workflow that has a clear completion signal and an existing fallback. A support queue with recorded reopenings and human overrides is easier to evaluate than open-ended strategy writing, where disagreement may appear weeks later and never enter the routing log.

Set a release gate before traffic moves. The gate should specify an acceptable missed-escalation level on the holdout set, a maximum override rate in shadow traffic, and a net cost target measured per resolved request. These are internal operating thresholds, not universal benchmark numbers, and they should be stricter where errors trigger financial or policy commitments.

During rollout, keep a control slice on the previous routing policy or the single-model baseline. Compare similar requests over the same period, because changing ticket mix can make a router appear cheaper without improving any decision. If rework rises, the request-level record points back to the router version and reason code.

For R-1842, success is concrete. The log shows that prior-agent history was present, the escalation test expects the capable path, and the production router either follows that route or records an override that feeds the next evaluation cycle. The model bill matters only after that ticket stays resolved.

Questions people ask

Does model routing always require another AI model?

No. A router can use fixed rules, a conventional classifier, a language model, or a sequence of those methods. Rules are easier to inspect but can be brittle; model-based routers handle broader phrasing but add inference cost and need versioned prompts, confidence signals, and regression tests.

Which routing mistakes matter most?

Missed escalations usually deserve the closest review because the cheaper model can return a convincing answer while overlooking policy, context, or required tools. False escalations mainly consume extra inference capacity, although they can also add latency. The correct weighting depends on the documented cost of each outcome in the chosen workflow.

How much traffic should stay on the expensive model?

There is no fixed percentage. Choose the route from measured cost per resolved request and the tolerated missed-escalation rate, then adjust the threshold in shadow mode. A high-risk workflow may rationally escalate most uncertain requests even when that leaves less gross model-cost savings.

What is the minimum viable routing log?

Record a request identifier, timestamp, router version, available routes, selected route, reason or score, selected model, usage, validation result, retries, overrides, and final outcome. For R-1842, those fields reveal whether the router saw the prior promise and whether a later correction erased the apparent saving.

ShareFacebook
ai observabilitymodel evaluationai pricing and accessmodel routingai inference costsllm evaluationai observability

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read