Skip to content

AI Industry & Models

Shadow-Test a New AI Model Before It Reaches Customers

Mirror sampled production requests to a candidate model, hide its outputs, and compare each run. The method exposes task-specific regressions before a model swap reaches users.

Tobias LundIndustry & Models Writer

August 26, 2026 · 8 min read

A laptop showing paired production and candidate model traces for the same customer-support ticket.
A laptop showing paired production and candidate model traces for the same customer-support ticket.

Consider one customer-support ticket: a buyer asks for a refund after delivery. The production model receives the message, order status, a refund-policy excerpt, and definitions for tools that can inspect the order or create a refund. It checks the order, applies the policy, and drafts the reply.

A candidate model can look better in a benchmark yet mishandle that ticket. It might refuse a permitted refund, call the correct tool with the wrong order ID, or produce a stronger answer only after taking longer than the support interface allows. A shadow comparison catches those differences with real inputs while the current model remains in charge.

Shadow traffic is a copy of a production request sent to a second system whose output cannot affect the user. The candidate sees the same task, subject to privacy and access controls, but its answer goes to an evaluation log rather than the customer.

Mirror one ticket, not the whole system

Start at the point where the production application has finished assembling the model request. For the refund ticket, that means copying the system prompt, conversation, policy context, tool definitions, model parameters, and any structured order data visible to the current model. If the two models receive different evidence, the comparison measures the surrounding pipeline as well as the model.

Assign both runs one trace ID, which is an identifier linking every event generated by a request. Store the production output beside the candidate output, along with timestamps, model identifiers, prompt revision, tool-schema revision, token counts, errors, and the final disposition. The candidate response must never enter the user-facing response path, even when the production call fails.

Keep the first comparison narrow. Do not change the model, prompt, retrieval settings, and tool schema in one experiment, because a regression will then have four plausible causes. Hold the existing support workflow fixed and replace only the model endpoint. Prompt changes discovered during evaluation can become a second experiment.

Run the candidate asynchronously so that a slow or unavailable shadow endpoint cannot delay the live reply. This makes comparison newly practical on production traffic, but it does not make candidate latency irrelevant: measure the candidate from request dispatch through its final token and any simulated tool turns, using the same regional and network arrangement planned for deployment.

Sample traffic without hiding rare failures

Sending every production request doubles model inference for the mirrored path before evaluator costs and tool reads. A deterministic sample controls that bill: hash a stable request or account identifier, then mirror one in every N eligible requests. Hashing keeps repeat interactions together and prevents a single conversation from switching unpredictably between sampled and unsampled states.

Pure random sampling will underrepresent the cases most likely to block a rollout. For the support workflow, create separate slices for refund requests, ambiguous policy questions, missing order records, non-English messages, long conversations, and requests that caused the current model to call a tool. Sample each slice deliberately, while preserving its production frequency in the overall report or labeling any oversampling clearly.

The shadow destination is another place where customer data travels. Before copying a request, apply the same redaction rules used by the production path, verify retention settings, and confirm that the candidate endpoint is permitted to receive that data under the organization’s own controls. If a new provider cannot accept the ticket’s personal or contractual data, use redacted records or an approved environment rather than treating shadow status as a privacy exemption.

Log sampling decisions as well as sampled requests. Without the denominator, a team can count candidate failures but cannot calculate how often they occurred, which traffic slices produced them, or whether an apparent improvement came from a changing request mix.

Make candidate tool calls harmless

The refund ticket becomes dangerous at the tool boundary. Letting the candidate execute `create_refund` would turn an invisible comparison into a duplicate financial action, while blocking every tool would test a workflow different from production.

Separate proposals from execution. Allow the candidate to emit the tool name and arguments it would use, then intercept the call before it reaches any write-capable service. Record that proposed call and return a controlled result from a staging system, a snapshot, or a replay captured from the production run. Read-only tools can sometimes run against production, but only if their load, authorization, and data exposure are acceptable.

Replay has limits. If the production model looked up order 481 and the candidate asks for order 418, returning the result for 481 would conceal the candidate’s argument error. Match a replay only when the tool and normalized arguments agree; otherwise return a defined not-found or blocked result and score the divergence.

The comparison should capture the full tool sequence, including retries, malformed arguments, abandoned calls, and the text produced after a result arrives. For the same refund ticket, a candidate that first requests the order, then proposes an eligible refund, may be functionally correct even if its internal wording differs. A candidate that calls the refund tool before reading the order has changed the workflow in a way a text-only score will miss.

Compare paired runs on four axes

Evaluate each trace as a pair. Side-by-side scoring reduces noise from traffic mix because the production and candidate models handled the same ticket, policy excerpt, and available tools. Where human reviewers are used, randomize which output appears first and hide model names so that expectations about a newer or larger model do not become the result.

Quality should be tied to the support task rather than general fluency. Check whether the answer applies the supplied refund policy, uses tool results correctly, preserves required disclosures, and gives the customer an actionable next step. Automated evaluators can triage a large sample, but a human should adjudicate close cases and high-impact disagreements against the policy and tool record.

Treat refusals as their own outcome. The useful rate is not refusals divided by all tickets; it is unjustified refusals among requests the workflow should answer, alongside unsafe or disallowed compliance among requests it should decline. A candidate can improve average prose while quietly refusing more legitimate refund requests, and an aggregate quality score may blur that regression.

Measure latency at more than the median. Record time to first token, total response time, tool-wait time, and a tail percentile such as p95, meaning 95 percent of measured requests completed at or below that value. The shadow path does not slow current users, but its tail shows whether deployment would create timeouts, abandoned chats, or extra retries.

Tool behavior needs semantic comparison. Exact JSON equality is too strict when argument order or harmless formatting changes, yet tool name, normalized identifiers, required fields, call order, and write intent should match the workflow’s rules. Review disagreements from the refund slice first because an incorrect action there costs more than a differently worded greeting.

Track errors separately from refusals. Empty responses, invalid tool payloads, context-limit failures, provider timeouts, and evaluator failures have different remedies. Combining them into one failure rate can make a candidate look stable while an evaluation service, rather than the model, is dropping difficult cases.

Turn the evidence into a rollout gate

Write acceptance rules before opening the results. Set a minimum quality outcome, a maximum regression allowance for justified answers and refusals, a latency ceiling tied to the application timeout, and zero tolerance for prohibited write proposals. The exact limits belong to the product’s risk and cost constraints; choosing them after seeing the candidate invites selective interpretation.

Inspect both the weighted total and every critical slice. If the candidate wins overall but fails on long conversations or missing-order cases, keep the current model for those routes, adjust the prompt in a new experiment, or stop the replacement. A model router can preserve the incumbent for a weak slice, although that adds another decision system to monitor.

Shadow cost is approximately the sampled candidate inference cost plus evaluation calls, storage, and any duplicated read-only tool work. Token counts from the paired log let the team estimate the full-traffic bill using its contracted prices rather than assuming two models consume the same number of tokens. A candidate that writes longer answers or loops through tools can erase a lower per-token price.

Shadow testing still cannot reveal how users react to the candidate’s wording, whether they follow its instructions, or whether real tool execution creates an unexpected side effect. Once the candidate clears its gates, move to a canary release, which exposes a small, controlled share of users to the new model. Keep assignment stable by account, monitor the same slices, and retain an immediate route back to the incumbent.

For the refund ticket, the final promotion record should show one linked chain: both prompts, both outputs, proposed tool calls, policy-based quality labels, refusal classification, latency measurements, reviewer decisions, and the rollout gate they informed. That trace is more useful than a benchmark average because it shows exactly what the replacement would have done to a customer request already seen in production.

Questions people ask

How much production traffic should a shadow test copy?

Set the fraction from cost, traffic diversity, and the number of important slices rather than choosing a universal percentage. The sample is large enough when common cases stabilize and rare, high-impact workflows have enough paired examples for review. Log the unsampled denominator so results can be weighted back to the production mix.

Can the candidate model call production tools during shadow testing?

It can propose calls, but write-capable actions such as issuing refunds, sending messages, or changing records should be intercepted. Use recorded responses, staging services, or approved read-only calls. Preserve the proposed tool name, normalized arguments, sequence, and retries so tool behavior remains measurable without creating side effects.

How long should a shadow comparison run?

Run it through the traffic variations that matter to the application, including weekday patterns, policy changes, and less common request types. Calendar duration alone is a weak stopping rule. Stop when the predefined slices have sufficient paired coverage, reviewers have resolved major disagreements, and error and latency patterns are no longer moving materially.

Does a successful shadow test replace a canary rollout?

No. Shadowing measures hidden outputs against current production requests, but users cannot respond to those outputs and mutating tools remain blocked. A canary tests the missing interaction under limited exposure. Keep the incumbent available, route only a controlled share to the candidate, and attach rollback triggers to the same quality, refusal, latency, and tool metrics.

ShareFacebook
model evaluationdeveloper toolingai at workshadow testingmodel evaluationproduction aimodel rollout

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read