Skip to content

Agentic AI & Orchestration

A Webhook Timed Out. The Agent Must Not Place the Order Twice

A timed-out tool call can leave an agent retrying an order that already exists. Recovery depends on treating the result as unknown until the remote system confirms what happened.

Mara QuinteroAgents & Orchestration Writer

October 3, 2026 · 8 min read

A laptop displaying a submitted toner order beside a log entry marked webhook timeout.
A laptop displaying a submitted toner order beside a log entry marked webhook timeout.

Consider an office operations agent replacing a toner cartridge. It has checked the approved vendor, confirmed the spending policy and reached step four: submit the order through a tool connector. That connector sends the request to an automation platform’s webhook endpoint, a URL that receives the tool call and starts the purchasing workflow.

The retailer accepts the order. Before the webhook handler can return the confirmation, the connection closes or the agent’s runtime stops waiting. The agent records a timeout.

Nothing in that timeout says the order failed. It says the caller did not receive a conclusive response.

That distinction is the center of reliable agent recovery. If the runtime translates every timeout into “try again,” the toner order may be placed twice. If it translates every timeout into “stop,” a genuinely failed order may never be placed. The correct state is unknown, and the orchestration layer has to resolve it before allowing another side effect.

Record the intended action before sending it

The recovery path starts before the webhook call. The runtime should create a durable operation record containing the local run ID, the intended vendor, the item and quantity, the approved amount, and a unique operation ID. Durable means the record survives a worker restart rather than living only in process memory.

That operation begins in a prepared state. Once the runtime sends the webhook request, it moves to submitted, but it should not move to confirmed until it has evidence from the remote system. A timeout leaves it submitted with an unknown outcome. This small state distinction prevents the language model from improvising around a network error.

The operation ID should also become the idempotency key, a unique token that tells a service to treat repeated requests as the same action. If the webhook handler or retailer supports idempotency, the runtime sends that key with the original toner order and stores it alongside the operation record.

A retry must reuse the same key. A fresh key represents a fresh action and can create another order.

The receiving service stores the key with the result of the first accepted request. When it sees the same key again, it returns the existing order or resumes the original operation instead of starting a second purchase. This gives the caller effectively-once behavior even though the network may deliver the request more than once.

The key also needs a stable payload. If the agent changes the toner quantity after a timeout but reuses the original key, the receiving service should reject the mismatch rather than guess which request is authoritative. Storing a fingerprint of the original request, a compact value derived from its contents, lets the webhook handler detect that change.

Check status before repeating the side effect

Return to the timed-out toner order. The runtime has an operation ID but no confirmation. Its next call should be a read, not another purchase.

The cleanest design gives the webhook handler a status endpoint keyed by the operation ID. The runtime asks whether the action is pending, succeeded or failed. If the retailer order exists, the handler returns the external order ID and the runtime marks the local operation confirmed. If processing is still underway, the runtime waits and checks later.

A definitive rejection, such as an invalid product code, can move the operation to failed without risking another order.

Status checks add latency, and repeated polling adds API traffic. That cost is usually smaller than reversing a duplicate purchase, but the runtime still needs limits. It should space checks farther apart over time, stop after a defined recovery window and send unresolved operations to review rather than polling forever.

A status request can time out too. That second timeout does not turn the purchase into a failure; it preserves the unknown state. The runtime should keep the same operation ID and avoid creating a new branch of the workflow merely because its attempt to inspect the first branch also lost a response.

Some services lack a dedicated status endpoint but allow lookup by a merchant reference, message ID or other caller-supplied identifier. The agent can search using the value stored before submission. Looking up an order by product name and approximate time is weaker because two legitimate orders can resemble each other, and fuzzy matching should not authorize an automatic retry.

Reconciliation catches what the request path missed

Idempotency handles repeated writes when every participant honors the same key. Status checks resolve many uncertain outcomes while the agent run is active. Neither covers every break in a chain that may include an agent runtime, an automation platform, a purchasing service and a retailer.

A reconciliation job compares local operations with remote records after the request path has had time to settle. Reconciliation means checking two systems’ records and resolving differences. For the toner order, the job finds local operations still marked submitted, queries the webhook handler or retailer, then records the remote order ID if one exists.

It can also catch the reverse mismatch: the local operation says confirmed, but a later remote state shows the order was canceled or rejected. That does not justify placing another order automatically. Cancellation may reflect a vendor decision, an approval change or a fraud control, so the workflow should return to the purchasing policy that governed the first attempt.

Reconciliation is where weak identifiers become expensive. If the local record contains only “buy toner,” a reviewer has little basis for matching it to a vendor order. If it preserves the operation ID, normalized request, submission attempts, webhook response and external reference, the same reviewer can determine whether an apparent duplicate is one retried request or two separately approved purchases.

The job does not need to run inside the model’s context. It belongs in ordinary application code with database constraints, authenticated API calls and an audit log. The model can explain an unresolved case to a person, but it should not decide that two nearly matching financial records are equivalent.

Put retries in the runtime, not the prompt

A prompt such as “avoid duplicate purchases” cannot enforce this behavior. The model may not see the earlier network attempt after a restart, and even a complete transcript cannot make a remote API deduplicate two requests.

The tool runtime should own the operation state and expose a narrower result to the agent: confirmed, failed or pending review. Database rules can prevent two active operations from claiming the same idempotency key, while the connector can refuse any payload that differs from the stored request. Those controls remain in force when the agent changes models, loses conversational context or resumes on another worker.

Retries still have a place. Read-only status calls are generally safer to repeat than order submissions, while a submission can be retried automatically only when the receiving path promises idempotent handling for the lifetime of the recovery window. The team must verify how long a provider retains keys; if the key expires while the local operation remains unresolved, a late retry may be treated as new.

For the toner workflow, the fallback is explicit. If neither the webhook handler nor the retailer can search by the stored operation ID, the runtime freezes the purchase, displays the submitted request and available logs, and asks a human to verify the vendor account before trying again. That takes longer. It is still the correct tradeoff for an irreversible or costly action.

The same pattern applies to messages, bookings and account changes, though the consequence of duplication differs. A repeated notification may be annoying; a repeated reservation may consume inventory. Each tool therefore needs its own retry policy rather than one agent-wide instruction to retry failed calls.

The minimum recovery contract

A tool is ready for autonomous use only if the orchestration layer can preserve intent across a lost response. In practical terms, the toner connector needs a durable local operation created before dispatch, one idempotency key reused across attempts, and a way to retrieve the remote outcome without repeating the purchase.

It also needs an unresolved state. Many workflow systems offer only success and failure because that is convenient for dashboards, but forcing a timeout into either bucket discards the information recovery depends on. “Unknown” or “submitted” should block downstream steps that assume confirmation, including sending a receipt, updating inventory or ordering the next supply item.

There is no general way to guarantee exactly-once execution across independent systems and an unreliable network. The practical target is effectively once: duplicate requests may occur, but stable identifiers, remote deduplication and reconciliation keep them from producing duplicate effects.

That setup costs storage, extra API reads and delayed completion when confirmation is missing. It also creates operational work because someone must own the queue of unresolved actions. For low-impact, reversible tools, that machinery may cost more than accepting an occasional duplicate. For purchases and external messages sent under an agent’s identity, blind retries are the more expensive design.

Questions people ask

Can a timed-out webhook still have completed the action?

Yes. The remote service may accept and commit the request before the response is lost, or the webhook handler may finish after the caller stops waiting. The timeout proves only that the agent did not receive a conclusive response, so the operation should remain unknown until checked.

How long should an agent wait before retrying?

There is no universal interval. The runtime should first check the operation’s status, then follow the remote service’s documented processing and idempotency behavior. If it cannot confirm that a repeated write will be deduplicated, elapsed time alone should not authorize another order.

What if the service does not support idempotency keys?

Store a caller-generated reference and search for the resulting order before retrying. If the service supports neither stable references nor reliable status lookup, automatic recovery is unsafe for costly actions; pause the workflow and have a person inspect the remote account or vendor records.

Should the

AI model decide whether the retry is safe?

No. Application code should enforce key reuse, payload matching, retry limits and operation states. The model can summarize logs or request approval, but it lacks the durable transaction record and database controls needed to prevent a second toner order after a restart.

ShareFacebook
ai agentsworkflow automationtool use and function callingai agentswebhooksidempotencyworkflow recovery

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop showing a citation ledger with one claim marked partially supported beside its quoted source passage.

Agentic AI & Orchestration

A Second Pass Can Catch AI Citations That Do Not Fit

A research agent can check whether a cited passage supports its claim, but only after claims are split into testable units. The extra pass catches mismatches, not bad sources or missing evidence.

Mara Quintero · 7 min read

Account settings page with an address modal partly covered by a cookie banner in a desktop browser.

Agentic AI & Orchestration

Visual AI Agents Still Lose the Checkout Button

A moved control is the easy case. Modal windows, sticky banners, and responsive layouts show why visual browser agents need bounded tasks, state checks, and a selector-based fallback.

Mara Quintero · 8 min read

Laptop showing a vendor support article beside an agent tool-call log with a blocked upload request.

Agentic AI & Orchestration

Test Browser Agents Before a Support Page Hijacks Them

A vendor support page is untrusted input, even when an agent needs it to finish a task. This walkthrough tests whether page text can trigger data leaks, unsafe tool calls, or account changes.

Mara Quintero · 8 min read