Skip to content

Agentic AI & Orchestration

Your AI Agent Said the Order Was Canceled. Check the Queue.

A successful tool call can leave the requested job unfinished. The reliable check reads the systems that own the order, warehouse task, or other business outcome.

Mara QuinteroAgents & Orchestration Writer

September 5, 2026 · 8 min read

Laptop showing a canceled order record beside an open warehouse task and a resolved support ticket.
Laptop showing a canceled order record beside an open warehouse task and a resolved support ticket.

I gave an agent a narrow job in a synthetic commerce sandbox: cancel order A-1048 before warehouse release, then close the customer’s support ticket only after the cancellation took effect. The workflow had access to a support system, an order API, and a warehouse task queue.

The agent’s final message said the order had been canceled and the ticket resolved. That sentence concealed three different possible states. The cancellation request might have reached the order API. The order might have changed to “canceled” while its warehouse pick task remained open.

Or nothing might have changed, despite a plausible explanation of what the agent claimed to have done.

Those states require different responses. A retry can help after a rejected request. An open warehouse task needs escalation or a compensating action, which is an operation that reverses or contains an earlier step. An unsupported completion claim needs to fail the evaluation even when the prose sounds credible.

The practical fix is to treat the final message as a user-facing summary and build a separate outcome check that reads the destination systems.

Tool success is narrower than job success

A tool call, a structured request from the model to an external function, can succeed without completing the business process behind it. In the cancellation test, the agent called `cancel_order` with the synthetic order ID. The tool adapter accepted the arguments, authenticated to the order service, and received an acknowledgment.

That acknowledgment proved that one interface had accepted one request. It did not prove that the order was canceled.

Many business systems move work into a background queue rather than finishing it during the original API request. A response in the successful family of HTTP status codes may mean “received,” while validation, inventory release, payment handling, and warehouse synchronization happen later. Tool wrappers often flatten those distinctions into a boolean such as `success: true`, which gives the model less information than the underlying response contained.

In one injected failure condition, the order service accepted the cancellation but left it pending. The agent saw a successful tool result and closed the support ticket. A check against the order record would have found the pending state. A check against the warehouse queue would have found an open pick task.

The tool worked. The job did not.

Define the state that counts as done

Before changing the prompt, write an outcome contract for the workflow. This is a machine-checkable description of the state that must exist before the run can report completion.

For order A-1048, “done” meant that the commerce system recorded the order as canceled and the warehouse system had no active pick task for it. The support ticket could then close with the cancellation event attached. A cancellation request ID, by itself, did not satisfy the contract.

That distinction matters because business language is often broader than an API operation. “Cancel the order” may require updates across systems owned by different teams, and one system can lag or reject a change after another has committed it. The verifier therefore needs to inspect every destination that can still produce the unwanted outcome, rather than checking whichever system is easiest to query.

Keep the contract narrow. Do not ask another model to judge whether a log “looks like” success when a status field or open-task query can answer directly. Model-based evaluation is useful for unstructured outputs, but deterministic checks are cheaper, easier to audit, and less likely to accept persuasive wording in place of evidence.

For this workflow, the contract recorded the expected order state, the absence of a blocking warehouse task, a deadline for those states to appear, and the fallback if they did not. It also preserved the original order ID and cancellation request ID so the verifier could query the same transaction rather than trusting identifiers repeated in the agent’s response.

Read the destination independently

The verifier should obtain evidence through a separate code path. If the agent says, “I checked the order and it is canceled,” feeding that sentence into an evaluator only tests whether the story is internally consistent.

After the cancellation tool returned, my test harness queried the commerce record directly using the order ID already held by the workflow controller. It then queried the warehouse system for active tasks tied to that order. The agent could not edit either query or substitute a different identifier.

Independent credentials help as well. They let the verifier have read-only access while the agent receives only the permissions needed to perform the action. This separation limits damage from a malformed call and makes the evidence easier to interpret: the actor attempts the change, while the verifier observes what the destination now reports.

An audit record should retain the requested action, tool response, destination reads, timestamps, and correlation identifiers that connect events across systems. A correlation identifier is a value carried through related requests so operators can trace one run. Store raw status values where practical. A rewritten summary can discard distinctions such as accepted, pending, rejected, and reversed.

Wait for asynchronous work without waiting forever

Immediate verification can create false failures. Systems with eventual consistency, where updates reach different components at different times, may show the old state briefly after accepting a valid change.

The cancellation verifier should poll the authoritative record at bounded intervals until the outcome contract passes or a deadline expires. “Authoritative” needs to be chosen deliberately. A cached customer-service view may update later than the order database, while a read replica may trail the system that accepted the write.

Polling adds tool calls, latency, and sometimes usage charges. Aggressive polling also loads the destination system without making its background job finish faster. A sensible interval follows the normal completion pattern of the underlying service, then backs off as the deadline approaches. Workflows that cannot tolerate the delay should consume a completion event from the destination system, if one exists, rather than pretending an acknowledgment is completion.

The deadline creates an operational boundary. Before it, the run remains pending. After it, the controller marks the result unverified and applies the declared fallback, such as reopening the ticket and assigning it to a person. The agent should not improvise a new meaning for “done” because a queue is slow.

Test the three failure classes separately

A useful evaluation changes the system state, not merely the wording in the prompt. I injected distinct faults into the A-1048 workflow to see whether the controller could tell them apart.

First, the cancellation endpoint accepted the request while its background worker remained pending. This was tool-call success without confirmed destination success. The correct result was “pending,” followed by more observation within the deadline.

Next, the commerce record changed to canceled, but synchronization to the warehouse queue was paused and the pick task stayed open. The primary record looked right, yet the business process could still ship the order. The correct result was failure with escalation, not another congratulatory message.

In the final condition, the tool adapter rejected the call because a required reason code was missing, so no request left the agent runtime. The agent still produced a completion statement based on its intended plan. The verifier found the original order state and an open warehouse task, then labeled the claim unsupported.

These tests exercise different controls. Tool receipts catch transport and schema errors. Destination checks establish current state. Cross-system checks catch partial completion.

Collapsing them into one pass rate hides which control failed and whether an automatic retry is safe.

Gate the next action on verified state

The support ticket in this example should not close merely because the cancellation step returned control to the agent. Put the verifier between the action and the irreversible or customer-visible next step.

The workflow controller can represent the run as requested, pending verification, verified, or failed. Only the verified state unlocks ticket closure. A retry uses an idempotency key, a unique value that lets a service recognize repeated requests, so a timeout does not produce duplicate cancellations or duplicate downstream actions.

Some systems can reverse a status after the initial check, particularly when an external provider performs later processing. A reconciler, a scheduled job that compares expected and observed state, can revisit high-consequence workflows after the main run ends. That extra read traffic and storage cost may be unnecessary for low-stakes updates. It is more defensible when a delayed mismatch could ship an item, send money, or remove access.

There is still a limit. A database can confirm that it recorded “canceled”; it cannot prove that a box already moving through a physical facility stopped. The completion claim should match the evidence available. In that case, “cancellation recorded and no active digital pick task found” is supportable.

“The order will not ship” may not be.

Questions people ask

Can

I trust a successful response from an agent tool?

Trust it as evidence that the tool reached a particular stage, not that the wider job finished. Inspect whether the response means accepted, queued, or completed, then query the destination record for the state the business request requires.

Should the same AI agent verify its own work?

It can initiate a read, but the workflow controller should choose the identifier, execute the check, and evaluate deterministic fields outside the model’s final narrative. Otherwise, the agent can repeat an incorrect assumption or inspect the wrong record while producing a consistent explanation.

How long should an outcome checker wait?

Base the deadline on the destination system’s normal processing behavior and the cost of delay. Until that deadline, label the work pending. Once it expires, mark the outcome unverified and trigger the predefined retry, compensating action, or human handoff.

What if the destination system has no read API?

Use the strongest independent evidence available, such as an event stream, exported status report, or restricted administrative view. If none exists, the automation cannot prove completion and should state that limit rather than upgrading a tool acknowledgment into a business outcome.

ShareFacebook
ai agentsworkflow automationtool use and function callingai agentsoutcome verificationworkflow orchestrationagent evaluation

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop showing a citation ledger with one claim marked partially supported beside its quoted source passage.

Agentic AI & Orchestration

A Second Pass Can Catch AI Citations That Do Not Fit

A research agent can check whether a cited passage supports its claim, but only after claims are split into testable units. The extra pass catches mismatches, not bad sources or missing evidence.

Mara Quintero · 7 min read

Account settings page with an address modal partly covered by a cookie banner in a desktop browser.

Agentic AI & Orchestration

Visual AI Agents Still Lose the Checkout Button

A moved control is the easy case. Modal windows, sticky banners, and responsive layouts show why visual browser agents need bounded tasks, state checks, and a selector-based fallback.

Mara Quintero · 8 min read