← All articles

OpenAI Responses API After Assistants: Migration Patterns and Architecture Guide

The Assistants API shuts down on August 26. This guide covers concrete Responses API patterns for teams who have already migrated: thread management without server-side threads, tool orchestration, streaming, and cross-provider routing through an OpenAI-compatible gateway.

· TheRouter

The Assistants API shuts down on August 26, 2026. If you are reading this, the deadline is either tomorrow or already past. The migration checklists and postmortem analyses exist elsewhere on this blog. This post assumes you have finished moving off Assistants and now need to build well on the Responses API.

What follows is a practical architecture guide: how conversation state works without server-side Threads, how tool orchestration replaces Run objects, how streaming differs, and where cross-provider routing through an OpenAI-compatible gateway fits in.

What the Responses API gives you that Assistants did not

The Assistants API managed state for you: Threads held messages, Runs polled for completion, and the server orchestrated tool calls across steps. That convenience came with lock-in. Thread and Run objects only existed on OpenAI's servers. You could not replay them through another provider, cache them locally, or inspect the full orchestration state in your own infrastructure.

The Responses API replaces that server-managed lifecycle with a stateless (or optionally stateful) request model. Each call to POST /v1/responses takes an input array and returns an output array. The model can call multiple tools within a single request. You own the state.

Key differences that affect architecture decisions:

  • No server-side threads. Conversation state is either passed as previous_response_id, managed through the Conversations API, or replayed manually in input.
  • Agentic loop built in. The model can chain web search, file search, code interpreter, function calls, and MCP tool calls within one request without polling.
  • Typed output items. Instead of choices[0].message, you get an output array with distinct reasoning, message, function_call, and web_search_call items.
  • Better cache utilization. OpenAI reports 40-80% improved cache hit rates compared to Chat Completions for equivalent workloads.

Conversation state without server-side Threads

Assistants stored conversation history in Thread objects. The Responses API offers three approaches to state, each with different trade-offs for portability and complexity.

Option 1: previous_response_id (simplest, OpenAI-only)

Pass the id from a previous response to chain turns:

first = client.responses.create(
    model="gpt-5.6",
    input="Explain the CAP theorem.",
)

second = client.responses.create(
    model="gpt-5.6",
    input="Now give me a concrete example.",
    previous_response_id=first.id,
)

This is the closest analogue to Assistants Threads. OpenAI stores the context server-side and replays it automatically. The trade-off: this parameter is OpenAI-specific. DeepSeek's Responses API does not support previous_response_id (it operates as a stateless API). If you need cross-provider portability, use Option 2 or 3.

Option 2: Conversations API (new, OpenAI-only)

The Conversations API provides persistent, named conversations. Useful when multiple sessions need to reference the same conversation history. Like previous_response_id, this is OpenAI-specific.

Option 3: Manual state replay (portable)

Append the full output array from each response to your input array for the next request:

history = [{"role": "user", "content": "Explain the CAP theorem."}]

response = client.responses.create(
    model="gpt-5.6",
    input=history,
    store=False,
)

# Replay all output items, including encrypted reasoning
history += response.output
history.append({"role": "user", "content": "Give me an example."})

next_response = client.responses.create(
    model="gpt-5.6",
    input=history,
    store=False,
)

This works across any provider that supports the Responses API input format. DeepSeek accepts the same input item structure (messages, function_call, function_call_output, reasoning, web_search_call items). The cost is that you manage storage and replay yourself, and long conversations increase token usage per request.

Which to pick: Use manual replay if you route across providers or want full control. Use previous_response_id if you are committed to OpenAI and want minimal code. Do not mix approaches within the same conversation.

Tool orchestration: from Run polling to the agentic loop

The Assistants API required polling Run objects to check if the model wanted to call a tool, then submitting tool outputs and polling again. A multi-tool conversation could involve four or five round trips.

The Responses API eliminates this. When you configure tools in a request, the model can invoke multiple tools and incorporate their outputs within a single API call for built-in tools (web search, file search, code interpreter). For custom function calls, you still submit outputs manually, but the request/response cycle is simpler:

import json

response = client.responses.create(
    model="gpt-5.6",
    input="What is the weather in Tokyo and New York?",
    tools=[{
        "type": "function",
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
        "strict": True,
    }],
)

# The model may return multiple function_call items in output
tool_outputs = []
for item in response.output:
    if item.type == "function_call":
        # Execute your function
        result = get_weather(json.loads(item.arguments)["city"])
        tool_outputs.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(result),
        })

# Submit tool outputs in a follow-up request
final = client.responses.create(
    model="gpt-5.6",
    input=[
        *response.output,      # replay the model's output
        *tool_outputs,          # your tool results
    ],
    tools=[{...}],  # same tool definitions
    previous_response_id=response.id,
)

Built-in tools that replace Assistants features

Assistants featureResponses API replacementNotes
Code Interpreter{"type": "code_interpreter"}Runs sandboxed Python, returns text/images
File Search (Retrieval){"type": "file_search", "vector_store_ids": [...]}Same vector store infrastructure
Function calling{"type": "function", ...}Request shape differs from Chat Completions
Web browsing (beta){"type": "web_search"}Server-side execution, citations in output

Cross-provider tool support

DeepSeek's Responses API supports function and web_search tool types. Other built-in tools (file_search, code_interpreter, computer_use, mcp) are silently ignored. This means multi-provider routing for agentic workloads must account for tool availability per provider.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Streaming: from delta chunks to semantic events

Chat Completions streaming sends delta objects with incremental content. The Responses API uses semantic server-sent events (SSE). Each event has a type field that tells you exactly what is happening:

Event typeWhat it means
response.createdRequest accepted, response generation started
response.output_text.deltaIncremental text content (the equivalent of Chat Completions deltas)
response.function_call_arguments.deltaIncremental function call argument JSON
response.reasoning_text.deltaChain-of-thought text (when reasoning is enabled)
response.output_item.added / .doneAn output item (message, function_call, etc.) starts/completes
response.completedFinal event; carries the full response object with usage
response.failedError during generation
stream = client.responses.create(
    model="gpt-5.6",
    input="Summarize the latest AI research trends.",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="")
    elif event.type == "response.completed":
        print(f"\n\nTokens: {event.response.usage.total_tokens}")

There is no data: [DONE] message. The stream ends with a response.completed, response.incomplete, or response.failed event.

DeepSeek's Responses API supports the same SSE event structure, making streaming behavior consistent when routing between OpenAI and DeepSeek through a compatible gateway.

Cross-provider routing considerations

The Responses API input/output format is becoming a cross-provider standard, but support varies:

ProviderResponses API supportKey limitations
OpenAIFullReference implementation
DeepSeekPartialNo previous_response_id, no store, no background, no Conversations API. Tools limited to function and web_search. Unsupported parameters silently ignored.
DashScope (Qwen)Via Chat CompletionsDashScope uses OpenAI-compatible Chat Completions. Responses API format not natively supported, but Qwen models are accessible through Chat Completions routing.

Routing architecture

When routing Responses API traffic through a gateway like TheRouter:

  1. Stateless requests route cleanly. If you use manual state replay (Option 3 above), the same request can go to any provider that supports the Responses API input format.
  2. Stateful parameters are provider-specific. previous_response_id, store: true, and the Conversations API only work when requests reach OpenAI. A router cannot synthesize server-side state for a different provider.
  3. Tool availability varies. Route requests with web_search or function tools to providers that support them. Requests with file_search or code_interpreter must go to OpenAI.
  4. Fallback requires care. If OpenAI is unavailable and you fall back to DeepSeek, your application must handle the absence of built-in tool execution for tools DeepSeek does not support.

TheRouter routes OpenAI-compatible requests through configured providers. For Responses API traffic, this means requests using manual state replay and function-type tools route cleanly to any compatible endpoint. Requests that depend on OpenAI-specific features (previous_response_id, built-in tools beyond function and web_search) should be pinned to OpenAI.

Production checklist for post-Assistants architecture

Before shipping your Responses API integration to production, verify these items:

Error handling

  • Handle response.failed events in streaming and check status in synchronous responses
  • Implement retry logic for 429 (rate limit) and 500+ (server error) responses
  • Set max_output_tokens to prevent runaway generation costs

Token accounting

  • Read usage.input_tokens and usage.output_tokens from the response object
  • For reasoning models, output_tokens_details.reasoning_tokens shows chain-of-thought token consumption
  • Cache hit tokens appear in input_tokens_details.cached_tokens

State management

  • If using manual replay, persist output arrays reliably (database, not just in-memory)
  • Encrypted reasoning items must be replayed as-is for reasoning models to maintain context
  • Set store: false when you do not want OpenAI to retain responses

Fallback behavior

  • Test each provider with your actual tool configuration to confirm which tools are supported
  • Log and alert when a fallback provider silently ignores unsupported tools
  • Consider separate routing rules for agentic (tool-heavy) vs. simple generation requests

Migration validation

  • Verify that your response parsing handles both output_text items and function_call items
  • Confirm structured output works with text.format instead of response_format
  • Run integration tests against both OpenAI and at least one alternative provider

FAQ

Does the Responses API cost more than Chat Completions?

For equivalent text generation without built-in tools, pricing is the same per token. OpenAI reports better cache utilization with Responses, which can reduce effective costs. Built-in tools like web search have their own per-search pricing.

Can I use the Responses API with reasoning models?

Yes. Starting with GPT-5.4, reasoning models have improved tool usage in the Responses API. Chat Completions does not support tool calling with reasoning_effort values other than none for GPT-5.4 and later.

What happens to my existing Chat Completions integrations?

Chat Completions remains supported. There is no announced deprecation. However, new features (built-in tools, Conversations API, background mode) are Responses-only.

Does DeepSeek's Responses API work identically to OpenAI's?

No. DeepSeek supports the core request/response format, function calling, and web search, but does not support stateful features (previous_response_id, store, Conversations API) or built-in tools like file_search and code_interpreter. Unsupported parameters are silently ignored.

Can I mix Chat Completions and Responses API calls in the same application?

Yes, but do not share state between them. Chat Completions uses messages/choices, Responses uses input/output. They are separate endpoints with different object shapes.

Further reading

Models covered in this article

Help & contact