← All articles

GPT-6 Astra API Integration Guide: Model Slug, Service Tiers, Responses API, and Multi-Provider Routing

Everything you need to integrate GPT-6 Astra into production — the model slug, five reasoning effort levels, the full pricing table with Fast mode and long-context surcharges, breaking changes from GPT-5.6 Sol, and how to route Astra through multiple providers with a single base_url change.

· TheRouter

GPT-6 Astra is OpenAI's flagship reasoning model, released September 3, 2026. The model ID is gpt-6-astra. It carries a 1,050,000-token context window, outputs up to 128,000 tokens, and supports five reasoning effort levels from low to max. Standard API pricing is $10 per million input tokens and $50 per million output tokens — 2.5 times GPT-5.6 Sol's promotional rate. Astra is available through the OpenAI API, Microsoft Azure Foundry, AWS Bedrock, and third-party gateways like OpenRouter.

We wrote this guide because Astra is not a drop-in replacement for Sol. OpenAI removed temperature, top_p, and logprobs from the API surface. The none and minimal reasoning effort levels are gone. Tool calling now requires the Responses API. If you are running GPT-5.6 Sol in production, several things will break when you swap the model string. This guide covers the complete integration path — from your first API call to multi-provider routing through TheRouter.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Getting Started in 3 Minutes

Step 1: Get your API key. If you already have an OpenAI API key, it works with Astra. If not, sign up at platform.openai.com, create a project, and generate a key.

Step 2: Install the SDK. The OpenAI Python SDK v1.x works out of the box:

pip install --upgrade openai

Step 3: Make your first call. The Responses API is the primary surface for Astra:

from openai import OpenAI

client = OpenAI()  # uses OPENAI_API_KEY env var

response = client.responses.create(
    model="gpt-6-astra",
    reasoning={"effort": "medium"},
    input=[
        {
            "role": "developer",
            "content": "You are a senior API engineer. Be concise."
        },
        {
            "role": "user",
            "content": "Explain the difference between Chat Completions and the Responses API in two sentences."
        },
    ],
)

print(response.output_text)

Chat Completions still works for plain text generation. Swap the endpoint and message shape, and remove any temperature or top_p fields. The moment you need function calling or built-in tools (web search, code interpreter, file search), move to the Responses API.

Step 4 (optional): Route through TheRouter. Change one line to route Astra through TheRouter with automatic fallback:

client = OpenAI(
    api_key="YOUR_THEROUTER_KEY",
    base_url="https://api.therouter.ai/v1",
)

response = client.chat.completions.create(
    model="openai/gpt-6-astra",
    messages=[{"role": "user", "content": "Hello from TheRouter"}],
)

GPT-6 Astra at a Glance

ItemValue
Model IDgpt-6-astra
Context window1,050,000 tokens
Max output128,000 tokens
Knowledge cutoffApril 30, 2026
ModalitiesInput: text, image. Output: text
EndpointsChat Completions, Responses, Batch
Not supportedRealtime, Assistants, fine-tuning
Reasoning effortlow, medium, high, xhigh, max
Rate limits (Tier 5)15,000 RPM / 40,000,000 TPM
Available onOpenAI API, Azure Foundry, AWS Bedrock, OpenRouter

Enterprise workspaces get Astra disabled by default. An administrator must enable it before API calls succeed.

Pricing: Standard, Batch, Fast Mode, and Long Context

All prices per million tokens. Source: OpenAI pricing page (retrieved September 9, 2026).

TierInputCached InputCache WriteOutput
Standard$10.00$1.00$12.50$50.00
Standard, long context (over 272K input)$20.00$2.00$25.00$75.00
Batch / Flex$5.00$0.50—$25.00
Batch, long context$10.00$1.00—$37.50
Fast mode$20.00$2.00—$100.00
Fast mode, long context$40.00$4.00—$150.00

Fast mode delivers up to 2x the processing speed at 2x the price. It is unavailable for EU data residency endpoints.

How Astra Compares to GPT-5.6 Sol

ModelInputCached InputOutput
GPT-6 Astra$10.00$1.00$50.00
GPT-5.6 Sol (promo through Nov 21, 2026)$4.00$0.40$20.00
GPT-5.6 Terra$2.00$0.20$12.00
GPT-5.6 Luna$0.20$0.02$1.20

Astra costs 2.5x Sol's promotional rate on both input and output. Sol's promotion runs at least through November 21, 2026.

Worked Cost Example

An agent run that reads a 200,000-token codebase once, reuses it as a cached prefix across 30 calls, and generates 60,000 tokens of output total:

  • Astra: $2.00 first read + $5.80 (29 cached reads at $0.20 each) + $3.00 output = ~$10.80
  • Sol (promo): $0.80 first read + $2.32 (29 cached reads at $0.08 each) + $1.20 output = ~$4.32

The ratio holds at 2.5x. What you need to measure is whether Astra finishes in fewer calls and fewer output tokens for your workloads.

Reasoning Effort: Five Levels, No "None"

GPT-5.6 Sol exposed six effort levels from none to max. Astra exposes five:

LevelTypical use
lowSimple extraction, classification, formatting
mediumStandard chat, summarization, code review
highMulti-step analysis, complex code generation
xhighResearch-grade reasoning, multi-hop problems
maxBenchmark-grade performance, long-horizon agent tasks

OpenAI removed none and minimal. If your Sol integration uses either, switch to low and evaluate. Everything else keeps its current setting.

This matters for routing economics. Astra at low is still a reasoning model — there is no classical completion mode. For mechanical workloads (classification, extraction, simple formatting), routing to GPT-5.6 Terra or Luna remains the cost-efficient option.

# Astra at medium effort
response = client.responses.create(
    model="gpt-6-astra",
    reasoning={"effort": "medium"},
    input="Summarize this changelog in three bullet points.",
)

At max effort, expect time-to-first-token measured in minutes, not seconds. Budget for latency before you budget for tokens.

Breaking Changes from GPT-5.6 Sol

If you are migrating from Sol, review these changes before switching the model string:

  1. temperature and top_p removed. Passing either parameter returns an error. Remove them from your request payload.

  2. logprobs removed. Not available on Astra. If you rely on logprobs for confidence scoring, keep a Sol or Terra fallback path.

  3. none and minimal reasoning effort removed. Use low instead.

  4. prompt_cache_retention renamed. The field is now prompt_cache_options.ttl. Update your cache configuration.

  5. Tool calling requires the Responses API. Chat Completions still handles plain text, but function calling and built-in tools (web search, code interpreter, file search, computer use) require the /v1/responses endpoint.

  6. Enterprise opt-in required. Astra is disabled by default in Enterprise workspaces. An admin must enable it before API calls succeed.

  7. Safety monitoring can interrupt work. Astra ships with runtime safety filters that can stop a task it flags as suspicious. OpenAI acknowledges the filters currently produce false positives on innocuous work. Worth knowing before you wire Astra into an unattended pipeline.

# Before (Sol)
response = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[{"role": "user", "content": "Hello"}],
    temperature=0.7,       # REMOVE for Astra
    top_p=0.9,             # REMOVE for Astra
)

# After (Astra via Responses API)
response = client.responses.create(
    model="gpt-6-astra",
    reasoning={"effort": "medium"},
    input=[{"role": "user", "content": "Hello"}],
)

Multi-Provider Routing

GPT-6 Astra is available through three provider endpoints plus third-party gateways. Each has a different model identifier:

ProviderEndpointModel ID
OpenAIhttps://api.openai.com/v1/gpt-6-astra
Azure Foundryhttps://<resource>.openai.azure.com/Deployment-specific
AWS BedrockRegional Bedrock endpointopenai.gpt-6-astra-v1:0 (check Bedrock docs)
OpenRouterhttps://openrouter.ai/api/v1/openai/gpt-6-astra
TheRouterhttps://api.therouter.ai/v1/openai/gpt-6-astra

With TheRouter, you configure Astra as a primary model and set fallbacks. If OpenAI's endpoint returns a 5xx or hits a rate limit, TheRouter automatically routes the request to your configured fallback — another GPT model, Claude, or any OpenAI-compatible provider.

from openai import OpenAI

# Route through TheRouter with automatic fallback
client = OpenAI(
    api_key="YOUR_THEROUTER_KEY",
    base_url="https://api.therouter.ai/v1",
)

# If openai/gpt-6-astra fails, TheRouter falls back
# to your configured secondary provider
response = client.chat.completions.create(
    model="openai/gpt-6-astra",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain prompt caching in one paragraph."},
    ],
)

For details on configuring model fallbacks, see the TheRouter routing guide.

Common Errors and Fixes

ErrorCauseFix
temperature is not supportedPassing temperature to AstraRemove temperature from your request
top_p is not supportedPassing top_p to AstraRemove top_p from your request
logprobs is not supportedRequesting logprobsRemove logprobs; use a Sol/Terra fallback if needed
Invalid reasoning effortUsing none or minimalSwitch to low
model_not_found (Enterprise)Astra not enabled in workspaceHave an admin enable the model
rate_limit_exceededTier limit hitCheck your rate limit tier; upgrade or add retry logic
context_length_exceededInput exceeds 1,050,000 tokensTrim input or split into multiple calls

Production Checklist

Before deploying GPT-6 Astra to production:

  • Remove unsupported parameters. Strip temperature, top_p, logprobs from all Astra requests.
  • Update reasoning effort. Replace none/minimal with low. Verify latency at your chosen effort level.
  • Migrate tool calls to Responses API. Function calling requires /v1/responses, not Chat Completions.
  • Update cache configuration. Rename prompt_cache_retention to prompt_cache_options.ttl.
  • Budget for long-context surcharges. Prompts over 272K input tokens trigger 2x pricing. Monitor input token counts.
  • Set up fallback routing. Configure a fallback model (Sol, Terra, or a cross-provider alternative) for rate limits and outages.
  • Test safety filters. Run Astra through your automated pipelines and verify it does not halt unexpectedly on legitimate tasks.
  • Enable in Enterprise workspace. Have an admin opt in the model before your code goes live.
  • Evaluate cost vs. Sol. Run your representative workloads on both models. Compare total cost per task, not just per-token price.

TheRouter Integration Note

GPT-6 Astra is not yet listed in TheRouter's models page as of this writing. When it is added, you will be able to route to it with the standard openai/gpt-6-astra model identifier. Until then, you can configure a custom model route pointing to OpenAI's endpoint directly.

TheRouter's routing and fallback capabilities are especially relevant for Astra because of the 2.5x cost increase over Sol. A routing configuration that sends simple tasks to Terra or Luna and reserves Astra for complex reasoning can cut your overall spend significantly while keeping frontier-grade intelligence available for the tasks that need it.

Sources

Help & contact