OpenAI’s Responses API Trades History Code for Retention Work
OpenAI can now retain a model run’s messages and tool context. That reduces orchestration code, but it makes deletion, storage duration and provider portability part of the API design.
September 25, 2026 · 8 min read

Take a support-ticket assistant with one bounded job: read a customer’s ticket, search the returns policy, draft a reply, then incorporate an agent’s correction. Under the familiar Chat Completions pattern, the application stores every message and tool result, builds the next request from that history, and sends the assembled context back to the model.
The Responses API changes who can hold that working record. An application may still manage everything itself, but it can instead give OpenAI a previous response ID or attach requests to a persistent Conversation. The first choice creates a response chain; the second creates a durable container for messages, tool calls and tool outputs.
That makes the support assistant easier to wire. It also means a deletion request can no longer be satisfied by clearing one transcript table unless the team knows which state OpenAI retained, for how long, and under which identifier.
The request can shrink without the context disappearing
Chat Completions is effectively stateless from the application’s perspective: each request includes the messages the model needs for that turn. The server processes the supplied context, while the client decides which old messages to retain, summarize or omit. Anthropic documents the same basic pattern for its Messages API, describing it as stateless and requiring applications to send conversation history with each request.
OpenAI’s Responses API accepts structured input and returns output items, which can include model messages, reasoning items and tool calls. To continue a run, the client can send `previous_response_id` with the new input rather than reconstructing the full item sequence in its own request.
In the ticket workflow, response one might contain the assistant’s policy-search call. The application executes a custom function, returns a `function_call_output` tied to the call ID, and receives the draft in response two. When the human agent says, “The item was opened, so remove the unopened-item language,” response three can refer to response two by ID rather than carrying the earlier ticket, function call and policy excerpt again.
The model still receives the relevant prior context. OpenAI’s documentation also says previous input tokens remain billable when `previous_response_id` is used, so a smaller HTTP payload should not be mistaken for a smaller model context bill. The gain is less client orchestration and fewer opportunities to serialize a tool result incorrectly, not free conversation memory or guaranteed lower generation latency.
There is another operational catch. OpenAI documents that instructions from a previous response are not automatically carried forward when `previous_response_id` is used. If the ticket assistant must never promise a refund before a human approves it, the application needs to supply that instruction again rather than assume the response chain preserved the policy layer.
A response chain and a Conversation are different commitments
A previous response ID suits a linear exchange. The application retains a pointer, OpenAI resolves the earlier response, and the next output extends the chain. That is enough for the support ticket while one worker handles one draft in one session.
The Conversations API is more durable. A conversation is a server-side object, meaning OpenAI stores its items so that later Responses calls can use them across sessions, devices or jobs. New inputs and outputs become part of that object, which makes a conversation useful when the ticket moves from an overnight drafting queue to a browser used by a human agent.
That convenience changes the data model. With client-managed history, the ticket ID is usually the primary record and the model transcript is a column, event stream or related table controlled by the application. With a Conversation, the application must also map the ticket to an OpenAI conversation ID, enforce access to that mapping and decide what happens when the ticket closes.
Do not treat the two server-side options as interchangeable. OpenAI’s data-control documentation lists stored Responses application state with a 30-day retention period by default, while Conversation objects and their items persist until deleted. A team that uses response chains for short drafting sessions has a different deletion burden from one that creates a Conversation for every customer account and leaves it open indefinitely.
The support assistant makes that difference visible. A response chain can expire as working state after the drafting window. A persistent Conversation may continue to hold the original complaint, the returns-policy excerpt and the human correction long after the final email has been copied into the help desk.
`store: false` restores control, with more client work
OpenAI lets developers set `store: false` on Responses calls when they do not want the response retained as retrievable application state. That choice removes the easiest server-side continuation path, because a later request cannot rely on OpenAI looking up an unstored response. The application goes back to carrying the necessary items itself.
This is closer to the Chat Completions operating model, though the Responses item schema still matters. Reasoning models may return reasoning items that need to be passed back alongside tool outputs so the model can continue correctly. For organizations using stateless calls, including those subject to Zero Data Retention controls, OpenAI documents encrypted reasoning content that the client can receive and resubmit without reading the underlying reasoning.
Storage settings also should not be confused with model training. OpenAI states that API inputs and outputs are not used to train its models by default, but training policy and retention are separate controls. Setting `store: false` addresses application-state storage; contractual data controls and abuse-monitoring retention still need to be checked against the organization’s account terms.
For the ticket assistant, the practical split is straightforward. If policy requires the company to remain the only durable holder of ticket text, keep a canonical event log, call Responses without stored state, and resend the needed messages, tool calls and outputs. If reducing orchestration is worth provider-held working state, use response IDs or Conversations and record every provider identifier needed for later retrieval and deletion.
Deletion has to follow the provider object graph
A customer erasure request exposes weak migrations quickly. Deleting the help-desk ticket does not automatically prove that a stored Response or Conversation associated with it has been deleted, while deleting one response does not replace an application-level policy for every related object.
OpenAI exposes deletion operations for stored responses and conversations. A production implementation therefore needs a reverse index from the company’s subject or ticket identifier to the relevant provider objects, rather than a lone conversation ID buried in a job log. The deletion worker should call the provider endpoint, record the result, and prevent a retry queue or cached payload from recreating the state afterward.
Portability requires a separate decision. A transcript made of user messages and assistant text can be normalized and replayed elsewhere. Provider-specific reasoning items, hosted-tool references and call identifiers may not transfer cleanly to another API, even when both vendors accept conversational messages.
The safest compromise is to keep a vendor-neutral record of the business event: ticket text, approved policy source, tool result, human edit and final reply. Provider response IDs can accelerate continuation, but they should not become the only record of why the assistant produced its answer. That record also gives the team a fallback when a stored response expires, a Conversation is deleted early, or the application moves to a stateless vendor.
Migrate the ticket, then run a deletion drill
A useful migration does not begin by replacing every Chat Completions call. Move the support-ticket workflow first, because it has a clear start, a human checkpoint and an obvious closing event.
Keep the existing application history during the trial. Send equivalent turns through Responses, compare the tool-call sequence and final draft, then inspect token accounting rather than assuming shorter request bodies reduced model charges. Repeat the safety instruction on each continuation that needs it.
Next, close a test ticket and execute the real retention path. Delete its local transcript, delete the associated OpenAI object, clear queued jobs, and verify that the application cannot reopen the conversation through an old identifier. Then export the vendor-neutral event record and replay the ticket without server-side state. If either operation depends on manually searching logs, the migration is not ready.
Stateful Responses become worthwhile when tool-heavy runs are failing at the joins between calls, or when several workers need to resume the same run. A four-turn drafting assistant with strict retention rules may be cheaper to operate conceptually with client-managed history, even if that means maintaining more code.
Questions people ask
Does the
Responses API reduce token costs by storing chat history?
Not by itself. OpenAI says earlier input tokens are still billed when a request uses `previous_response_id`, even though the client does not resend the full history in the HTTP body. Savings would need to come from better context selection or compaction, not from replacing messages with an ID.
How long does
OpenAI keep state from the Responses API?
OpenAI’s data-control documentation lists stored Response application state with a 30-day period by default. Conversation objects are a different storage choice and persist until deleted, so teams should not apply the response-retention assumption to a durable Conversation.
Can
I use the Responses API without OpenAI storing the response?
Yes. Set `store: false` and retain the context your application needs for later calls. That gives up server-side lookup through a stored response ID, and reasoning-model workflows may need encrypted reasoning content passed back with subsequent tool outputs.
What should
I store locally if I use server-side conversation state?
Keep the business record and a reverse index to provider objects. For the support workflow, that means the ticket, policy source, tool result, human correction, final reply, response or conversation IDs, and deletion status, so the run can be explained, removed or replayed without depending on one vendor’s internal format.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



