Skip to content

AI Industry & Models

Model API Deprecations Need a Rehearsal Before Shutdown

Changing the model name proves that an API call still runs. A five-gate rehearsal shows whether refusals, truncation, images, token counts, and failures still fit the application.

Tobias LundIndustry & Models Writer

September 5, 2026 · 8 min read

Laptop showing side-by-side model migration logs for a support ticket with an attached screenshot.
Laptop showing side-by-side model migration logs for a support ticket with an attached screenshot.

A model API deprecation often arrives with an apparent one-line fix: replace the retiring model identifier with the vendor’s recommended successor. That change can restore a successful HTTP response while quietly altering what the application does next.

Use one concrete job to expose the difference. Consider a support-draft service that receives a customer message and an optional screenshot, classifies the issue, and returns JSON containing `category`, `needs_human`, and `draft_reply`. The application rejects malformed JSON, routes sensitive cases to an employee, and shows the draft inside a text box with a fixed character limit.

A replacement model can accept that exact request and still break the workflow. It may refuse a borderline complaint that the old model answered, produce a longer draft that the interface cuts off, count the screenshot against limits differently, or return a retryable error under a different condition. None of those failures is visible in a basic connectivity test.

The migration therefore needs five gates: refusal behavior, output length, tokenization, multimodal input, and error handling. Passing them makes an endpoint swap newly practical. Failing one tells a small team where it needs an application change rather than another prompt tweak.

Freeze the contract before comparing models

Start with the behavior the application depends on, not the wording of the existing prompt. For the support-draft service, the contract includes valid JSON, three required fields, an allowed set of categories, a maximum draft length, and a rule that account-security cases set `needs_human` to true.

Record the request settings too. Temperature, maximum output tokens, system instructions, response-format controls, image placement, and tool definitions can all affect the result, while some successor models may ignore, rename, or reject parameters accepted by the retiring model. Vendor migration documentation should decide which settings are legal; the rehearsal decides whether the legal replacement remains useful.

Build a production-shaped test corpus from logged requests after removing personal or confidential information. Keep ordinary tickets, long threads, misspellings, hostile language, prompt-injection attempts, screenshots with small text, unsupported files, and messages near the application’s input limit. A corpus containing only clean examples measures prompt compliance under ideal conditions, not migration risk.

Each case needs expected properties rather than a single ideal answer. The billing category may be required, for example, while the exact wording of the reply can vary. That distinction prevents the team from treating harmless prose changes as regressions while missing a changed escalation decision.

Run the corpus through both models with captured request IDs, response bodies, token-usage fields, stop reasons, latency, and errors. A shadow run, meaning a call whose output is recorded but not shown to the user, lets the replacement process live-shaped traffic without controlling the support queue. It costs an additional model call per sampled request, so smaller teams can shadow a limited, privacy-reviewed slice rather than every ticket.

Gate 1 checks changed refusals

A refusal is the model declining to complete a request because of safety or policy controls. Replacement models can apply those controls differently even when the prompt and API route remain unchanged, and the surrounding application must recognize the form the refusal takes.

In the support-draft corpus, include legitimate but sensitive cases such as a customer reporting fraud, threatening self-harm, or pasting abusive language from another person. The useful comparison is not whether one model sounds more cautious. Check whether it still emits the required fields, whether `needs_human` changes, and whether a refusal appears as ordinary text, a dedicated response field, a particular stop reason, or no usable content.

Then exercise the fallback. If the parser receives no `draft_reply`, the service should route the ticket to a person rather than retrying the same blocked request until it exhausts its budget. A refusal-format change that crashes the parser means the migration is not drop-in, even if the replacement’s policy decision is acceptable.

Gate 2 measures output boundaries

The next pass measures length in two units: model tokens and application characters. A token is a chunk of text processed by the model, while the support interface still cares about visible characters and field sizes.

Compare the distribution of draft lengths rather than one average. Inspect the longest outputs, cases that reach the configured output cap, and responses whose stop metadata indicates truncation. If the replacement writes more before reaching the answer, a once-complete JSON object may now end halfway through `draft_reply`; increasing the token cap can repair completeness, but it also raises worst-case cost and may exceed the interface’s useful length.

Do not let the parser guess. When the response ends because it hit a limit, mark the run incomplete and use a defined fallback, such as one constrained retry followed by human routing. For the anchored support job, also render the result in the real text box. Valid JSON is little comfort if the final refund instruction sits beyond the portion an employee can see.

Gate 3 recalculates tokenization and limits

Tokenization is the conversion of text into the chunks a model reads and bills. A successor may use a different tokenizer, context limit, or accounting method, so the old local estimate can become inaccurate even though the input string has not changed.

Recount the longest support threads with the replacement model’s documented tokenizer or official counting method. Include system instructions, prior messages, schema text, tool definitions, and any vendor-specific overhead described in the documentation. Compare that estimate with the usage returned by successful API calls, then test requests just below and above the application’s intended ceiling.

This gate catches two separate problems. A request may exceed the replacement’s accepted context before generation begins, or it may leave too little room for the JSON reply after a long conversation consumes the budget. The fallback should be explicit: trim older messages, summarize approved portions, reject the attachment, or send the case to a person. Silent truncation is unsuitable because it can remove the customer’s first description of the problem while preserving later replies.

Token changes also affect rate limits and spending. Use the vendor’s current pricing and limit documentation to recalculate the support job from measured input and output usage; do not assume that a nominally cheaper successor lowers the bill if it produces longer drafts or requires more retries.

Gate 4 replays the screenshot path

Text-only success says nothing about a multimodal request, which combines formats such as text and images in one model call. The support service’s screenshot path should replay the actual transport used in production: hosted URL or encoded bytes, declared media type, message order, and accompanying instruction.

Use readable screenshots, tiny interface text, a corrupted file, an unsupported format, and an image near the vendor’s documented size boundary. Verify whether the successor accepts the same representation and whether image inputs use a different limit or token-accounting rule. Also check the application’s preprocessing, because resizing or recompressing an image to satisfy a new boundary can make a small error message unreadable.

Return to the contract after the image is accepted. The replacement still has to place its classification and reply inside the same JSON fields. If it reads screenshots well but requires a different request shape, altered preprocessing, or a second extraction call, the move may be worthwhile, but it is an application rewrite with a new latency and cost path.

Gate 5 forces failures on purpose

Happy-path comparisons miss the code most likely to run during a vendor incident. Force authentication failure, rate limiting, a timeout, malformed input, an oversized request, and a server-side error using vendor-supported test methods or a controlled wrapper around the client.

Record the HTTP status, SDK exception, response body, request identifier, and retry guidance. SDKs can rename exception classes, while APIs may distinguish temporary overload from invalid input through status codes or structured error types. The application should retry only failures documented as transient, use bounded backoff, and avoid retrying policy refusals or malformed requests that cannot improve on another attempt.

Idempotency matters if a retry can trigger an external action. The support-draft service only generates text, but a broader workflow might open a ticket or issue a credit after generation. Keep those side effects outside an uncertain model retry, or attach the application’s own operation key so one request cannot create duplicates.

Run this gate with logging and alerts enabled. The useful outcome is not merely that the client catches an exception; the on-call person must be able to distinguish bad credentials from exhausted limits without reading a raw response containing customer text.

Decide between a swap and a rewrite

Call the migration drop-in only if the same request schema remains supported, required outputs still validate, refusal and stop conditions reach the right fallback, multimodal inputs retain their path, and the existing retry policy classifies errors correctly. Cost and latency must also stay inside the team’s recorded operating budgets.

A rewrite starts when the team must change message structure, split image processing into another call, redesign parsing around free text, add a moderation or escalation stage, or alter product limits to accommodate longer responses. That label is useful for planning. It exposes engineering work that a replacement model name hides.

Before shutdown, send a small canary share of real support-draft traffic to the successor, with automatic rollback to the retiring model while it remains available. Compare validation failures, human escalations, truncations, retries, latency, and measured usage. Freeze unrelated prompt edits during this period, because changing the prompt and model together removes the baseline needed to explain a regression.

Archive the corpus, results, vendor migration notes, and final decision. The next deprecation then begins with an executable rehearsal rather than a remembered claim that the previous swap was easy.

Questions people ask

Is changing the model name ever enough?

Yes, if the replacement accepts the same request shape and passes all five behavioral gates within the application’s existing budgets. A successful API response alone is insufficient; the support-draft service must still validate JSON, route refusals, fit the text box, process screenshots, and recover from documented errors.

How large should a migration test set be?

There is no universal count. Cover each production path and each known boundary, then weight the corpus toward frequent requests and expensive failures. A small team gets more value from a compact set containing long threads, sensitive cases, broken images, and limit-adjacent inputs than from a large collection of routine greetings.

Should we run both models in production?

A time-limited shadow run is useful when privacy rules and budget permit it because the replacement sees production-shaped inputs without controlling user-visible output. Sample rather than duplicate every call if cost matters, strip data your policy does not permit for evaluation, and stop once the canary has enough evidence for the predefined gates.

What happens if the replacement fails one gate?

Classify the failure before extending the deadline pressure to users. A parameter rename may need a client patch, while changed refusals or multimodal transport can require new product logic. Keep the old model as rollback only until its documented shutdown, and route unsupported cases to the concrete fallback recorded in the contract.

ShareFacebook
model releasesdeveloper toolingmodel evaluationmodel api deprecationsapi migrationmodel testingdeveloper tooling

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read