OpenAI Responses API After Assistants: Migration Patterns and Architecture Guide
The Assistants API shuts down on August 26. This guide covers concrete Responses API patterns for teams who have already migrated: thread management without server-side threads, tool orchestration, streaming, and cross-provider routing through an OpenAI-compatible gateway.
The Assistants API shuts down on August 26, 2026. If you are reading this, the deadline is either tomorrow or already past. The migration checklists and postmortem analyses exist elsewhere on this blog. This post assumes you have finished moving off Assistants and now need to build well on the Responses API.
What follows is a practical architecture guide: how conversation state works without server-side Threads, how tool orchestration replaces Run objects, how streaming differs, and where cross-provider routing through an OpenAI-compatible gateway fits in.
What the Responses API gives you that Assistants did not
The Assistants API managed state for you: Threads held messages, Runs polled for completion, and the server orchestrated tool calls across steps. That convenience came with lock-in. Thread and Run objects only existed on OpenAI's servers. You could not replay them through another provider, cache them locally, or inspect the full orchestration state in your own infrastructure.
The Responses API replaces that server-managed lifecycle with a stateless (or optionally stateful) request model. Each call to POST /v1/responses takes an input array and returns an output array. The model can call multiple tools within a single request. You own the state.
Key differences that affect architecture decisions:
- No server-side threads. Conversation state is either passed as
previous_response_id, managed through the Conversations API, or replayed manually ininput. - Agentic loop built in. The model can chain web search, file search, code interpreter, function calls, and MCP tool calls within one request without polling.
- Typed output items. Instead of
choices[0].message, you get anoutputarray with distinctreasoning,message,function_call, andweb_search_callitems. - Better cache utilization. OpenAI reports 40-80% improved cache hit rates compared to Chat Completions for equivalent workloads.
Conversation state without server-side Threads
Assistants stored conversation history in Thread objects. The Responses API offers three approaches to state, each with different trade-offs for portability and complexity.
Option 1: previous_response_id (simplest, OpenAI-only)
Pass the id from a previous response to chain turns:
first = client.responses.create(
model="gpt-5.6",
input="Explain the CAP theorem.",
)
second = client.responses.create(
model="gpt-5.6",
input="Now give me a concrete example.",
previous_response_id=first.id,
)
This is the closest analogue to Assistants Threads. OpenAI stores the context server-side and replays it automatically. The trade-off: this parameter is OpenAI-specific. DeepSeek's Responses API does not support previous_response_id (it operates as a stateless API). If you need cross-provider portability, use Option 2 or 3.
Option 2: Conversations API (new, OpenAI-only)
The Conversations API provides persistent, named conversations. Useful when multiple sessions need to reference the same conversation history. Like previous_response_id, this is OpenAI-specific.
Option 3: Manual state replay (portable)
Append the full output array from each response to your input array for the next request:
history = [{"role": "user", "content": "Explain the CAP theorem."}]
response = client.responses.create(
model="gpt-5.6",
input=history,
store=False,
)
# Replay all output items, including encrypted reasoning
history += response.output
history.append({"role": "user", "content": "Give me an example."})
next_response = client.responses.create(
model="gpt-5.6",
input=history,
store=False,
)
This works across any provider that supports the Responses API input format. DeepSeek accepts the same input item structure (messages, function_call, function_call_output, reasoning, web_search_call items). The cost is that you manage storage and replay yourself, and long conversations increase token usage per request.
Which to pick: Use manual replay if you route across providers or want full control. Use previous_response_id if you are committed to OpenAI and want minimal code. Do not mix approaches within the same conversation.
Tool orchestration: from Run polling to the agentic loop
The Assistants API required polling Run objects to check if the model wanted to call a tool, then submitting tool outputs and polling again. A multi-tool conversation could involve four or five round trips.
The Responses API eliminates this. When you configure tools in a request, the model can invoke multiple tools and incorporate their outputs within a single API call for built-in tools (web search, file search, code interpreter). For custom function calls, you still submit outputs manually, but the request/response cycle is simpler:
import json
response = client.responses.create(
model="gpt-5.6",
input="What is the weather in Tokyo and New York?",
tools=[{
"type": "function",
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
"strict": True,
}],
)
# The model may return multiple function_call items in output
tool_outputs = []
for item in response.output:
if item.type == "function_call":
# Execute your function
result = get_weather(json.loads(item.arguments)["city"])
tool_outputs.append({
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
})
# Submit tool outputs in a follow-up request
final = client.responses.create(
model="gpt-5.6",
input=[
*response.output, # replay the model's output
*tool_outputs, # your tool results
],
tools=[{...}], # same tool definitions
previous_response_id=response.id,
)
Built-in tools that replace Assistants features
| Assistants feature | Responses API replacement | Notes |
|---|---|---|
| Code Interpreter | {"type": "code_interpreter"} | Runs sandboxed Python, returns text/images |
| File Search (Retrieval) | {"type": "file_search", "vector_store_ids": [...]} | Same vector store infrastructure |
| Function calling | {"type": "function", ...} | Request shape differs from Chat Completions |
| Web browsing (beta) | {"type": "web_search"} | Server-side execution, citations in output |
Cross-provider tool support
DeepSeek's Responses API supports function and web_search tool types. Other built-in tools (file_search, code_interpreter, computer_use, mcp) are silently ignored. This means multi-provider routing for agentic workloads must account for tool availability per provider.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Streaming: from delta chunks to semantic events
Chat Completions streaming sends delta objects with incremental content. The Responses API uses semantic server-sent events (SSE). Each event has a type field that tells you exactly what is happening:
| Event type | What it means |
|---|---|
response.created | Request accepted, response generation started |
response.output_text.delta | Incremental text content (the equivalent of Chat Completions deltas) |
response.function_call_arguments.delta | Incremental function call argument JSON |
response.reasoning_text.delta | Chain-of-thought text (when reasoning is enabled) |
response.output_item.added / .done | An output item (message, function_call, etc.) starts/completes |
response.completed | Final event; carries the full response object with usage |
response.failed | Error during generation |
stream = client.responses.create(
model="gpt-5.6",
input="Summarize the latest AI research trends.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="")
elif event.type == "response.completed":
print(f"\n\nTokens: {event.response.usage.total_tokens}")
There is no data: [DONE] message. The stream ends with a response.completed, response.incomplete, or response.failed event.
DeepSeek's Responses API supports the same SSE event structure, making streaming behavior consistent when routing between OpenAI and DeepSeek through a compatible gateway.
Cross-provider routing considerations
The Responses API input/output format is becoming a cross-provider standard, but support varies:
| Provider | Responses API support | Key limitations |
|---|---|---|
| OpenAI | Full | Reference implementation |
| DeepSeek | Partial | No previous_response_id, no store, no background, no Conversations API. Tools limited to function and web_search. Unsupported parameters silently ignored. |
| DashScope (Qwen) | Via Chat Completions | DashScope uses OpenAI-compatible Chat Completions. Responses API format not natively supported, but Qwen models are accessible through Chat Completions routing. |
Routing architecture
When routing Responses API traffic through a gateway like TheRouter:
- Stateless requests route cleanly. If you use manual state replay (Option 3 above), the same request can go to any provider that supports the Responses API input format.
- Stateful parameters are provider-specific.
previous_response_id,store: true, and the Conversations API only work when requests reach OpenAI. A router cannot synthesize server-side state for a different provider. - Tool availability varies. Route requests with
web_searchorfunctiontools to providers that support them. Requests withfile_searchorcode_interpretermust go to OpenAI. - Fallback requires care. If OpenAI is unavailable and you fall back to DeepSeek, your application must handle the absence of built-in tool execution for tools DeepSeek does not support.
TheRouter routes OpenAI-compatible requests through configured providers. For Responses API traffic, this means requests using manual state replay and function-type tools route cleanly to any compatible endpoint. Requests that depend on OpenAI-specific features (previous_response_id, built-in tools beyond function and web_search) should be pinned to OpenAI.
Production checklist for post-Assistants architecture
Before shipping your Responses API integration to production, verify these items:
Error handling
- Handle
response.failedevents in streaming and checkstatusin synchronous responses - Implement retry logic for 429 (rate limit) and 500+ (server error) responses
- Set
max_output_tokensto prevent runaway generation costs
Token accounting
- Read
usage.input_tokensandusage.output_tokensfrom the response object - For reasoning models,
output_tokens_details.reasoning_tokensshows chain-of-thought token consumption - Cache hit tokens appear in
input_tokens_details.cached_tokens
State management
- If using manual replay, persist
outputarrays reliably (database, not just in-memory) - Encrypted reasoning items must be replayed as-is for reasoning models to maintain context
- Set
store: falsewhen you do not want OpenAI to retain responses
Fallback behavior
- Test each provider with your actual tool configuration to confirm which tools are supported
- Log and alert when a fallback provider silently ignores unsupported tools
- Consider separate routing rules for agentic (tool-heavy) vs. simple generation requests
Migration validation
- Verify that your response parsing handles both
output_textitems andfunction_callitems - Confirm structured output works with
text.formatinstead ofresponse_format - Run integration tests against both OpenAI and at least one alternative provider
FAQ
Does the Responses API cost more than Chat Completions?
For equivalent text generation without built-in tools, pricing is the same per token. OpenAI reports better cache utilization with Responses, which can reduce effective costs. Built-in tools like web search have their own per-search pricing.
Can I use the Responses API with reasoning models?
Yes. Starting with GPT-5.4, reasoning models have improved tool usage in the Responses API. Chat Completions does not support tool calling with reasoning_effort values other than none for GPT-5.4 and later.
What happens to my existing Chat Completions integrations?
Chat Completions remains supported. There is no announced deprecation. However, new features (built-in tools, Conversations API, background mode) are Responses-only.
Does DeepSeek's Responses API work identically to OpenAI's?
No. DeepSeek supports the core request/response format, function calling, and web search, but does not support stateful features (previous_response_id, store, Conversations API) or built-in tools like file_search and code_interpreter. Unsupported parameters are silently ignored.
Can I mix Chat Completions and Responses API calls in the same application?
Yes, but do not share state between them. Chat Completions uses messages/choices, Responses uses input/output. They are separate endpoints with different object shapes.
Further reading
- OpenAI: Migrate to the Responses API — official migration guide with code examples
- DeepSeek: Using the Responses API — DeepSeek's compatibility details and limitations
- OpenAI: Conversation State — the three state management approaches in detail
- Assistants to Responses API Migration Routing Guide — our step-by-step migration checklist
- Assistants API Shutdown Postmortem — architecture lessons from the deprecation
- Assistants API Sunset: Final Migration Checklist — last-minute migration steps
- Assistants API Alternatives Comparison — Responses vs. Chat Completions vs. third-party alternatives