GPT-6 Astra API Integration Guide: Model Slug, Service Tiers, Responses API, and Multi-Provider Routing
Everything you need to integrate GPT-6 Astra into production — the model slug, five reasoning effort levels, the full pricing table with Fast mode and long-context surcharges, breaking changes from GPT-5.6 Sol, and how to route Astra through multiple providers with a single base_url change.
GPT-6 Astra is OpenAI's flagship reasoning model, released September 3, 2026. The model ID is gpt-6-astra. It carries a 1,050,000-token context window, outputs up to 128,000 tokens, and supports five reasoning effort levels from low to max. Standard API pricing is $10 per million input tokens and $50 per million output tokens — 2.5 times GPT-5.6 Sol's promotional rate. Astra is available through the OpenAI API, Microsoft Azure Foundry, AWS Bedrock, and third-party gateways like OpenRouter.
We wrote this guide because Astra is not a drop-in replacement for Sol. OpenAI removed temperature, top_p, and logprobs from the API surface. The none and minimal reasoning effort levels are gone. Tool calling now requires the Responses API. If you are running GPT-5.6 Sol in production, several things will break when you swap the model string. This guide covers the complete integration path — from your first API call to multi-provider routing through TheRouter.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Getting Started in 3 Minutes
Step 1: Get your API key. If you already have an OpenAI API key, it works with Astra. If not, sign up at platform.openai.com, create a project, and generate a key.
Step 2: Install the SDK. The OpenAI Python SDK v1.x works out of the box:
pip install --upgrade openai
Step 3: Make your first call. The Responses API is the primary surface for Astra:
from openai import OpenAI
client = OpenAI() # uses OPENAI_API_KEY env var
response = client.responses.create(
model="gpt-6-astra",
reasoning={"effort": "medium"},
input=[
{
"role": "developer",
"content": "You are a senior API engineer. Be concise."
},
{
"role": "user",
"content": "Explain the difference between Chat Completions and the Responses API in two sentences."
},
],
)
print(response.output_text)
Chat Completions still works for plain text generation. Swap the endpoint and message shape, and remove any temperature or top_p fields. The moment you need function calling or built-in tools (web search, code interpreter, file search), move to the Responses API.
Step 4 (optional): Route through TheRouter. Change one line to route Astra through TheRouter with automatic fallback:
client = OpenAI(
api_key="YOUR_THEROUTER_KEY",
base_url="https://api.therouter.ai/v1",
)
response = client.chat.completions.create(
model="openai/gpt-6-astra",
messages=[{"role": "user", "content": "Hello from TheRouter"}],
)
GPT-6 Astra at a Glance
| Item | Value |
|---|---|
| Model ID | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Modalities | Input: text, image. Output: text |
| Endpoints | Chat Completions, Responses, Batch |
| Not supported | Realtime, Assistants, fine-tuning |
| Reasoning effort | low, medium, high, xhigh, max |
| Rate limits (Tier 5) | 15,000 RPM / 40,000,000 TPM |
| Available on | OpenAI API, Azure Foundry, AWS Bedrock, OpenRouter |
Enterprise workspaces get Astra disabled by default. An administrator must enable it before API calls succeed.
Pricing: Standard, Batch, Fast Mode, and Long Context
All prices per million tokens. Source: OpenAI pricing page (retrieved September 9, 2026).
| Tier | Input | Cached Input | Cache Write | Output |
|---|---|---|---|---|
| Standard | $10.00 | $1.00 | $12.50 | $50.00 |
| Standard, long context (over 272K input) | $20.00 | $2.00 | $25.00 | $75.00 |
| Batch / Flex | $5.00 | $0.50 | — | $25.00 |
| Batch, long context | $10.00 | $1.00 | — | $37.50 |
| Fast mode | $20.00 | $2.00 | — | $100.00 |
| Fast mode, long context | $40.00 | $4.00 | — | $150.00 |
Fast mode delivers up to 2x the processing speed at 2x the price. It is unavailable for EU data residency endpoints.
How Astra Compares to GPT-5.6 Sol
| Model | Input | Cached Input | Output |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol (promo through Nov 21, 2026) | $4.00 | $0.40 | $20.00 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 |
Astra costs 2.5x Sol's promotional rate on both input and output. Sol's promotion runs at least through November 21, 2026.
Worked Cost Example
An agent run that reads a 200,000-token codebase once, reuses it as a cached prefix across 30 calls, and generates 60,000 tokens of output total:
- Astra: $2.00 first read + $5.80 (29 cached reads at $0.20 each) + $3.00 output = ~$10.80
- Sol (promo): $0.80 first read + $2.32 (29 cached reads at $0.08 each) + $1.20 output = ~$4.32
The ratio holds at 2.5x. What you need to measure is whether Astra finishes in fewer calls and fewer output tokens for your workloads.
Reasoning Effort: Five Levels, No "None"
GPT-5.6 Sol exposed six effort levels from none to max. Astra exposes five:
| Level | Typical use |
|---|---|
low | Simple extraction, classification, formatting |
medium | Standard chat, summarization, code review |
high | Multi-step analysis, complex code generation |
xhigh | Research-grade reasoning, multi-hop problems |
max | Benchmark-grade performance, long-horizon agent tasks |
OpenAI removed none and minimal. If your Sol integration uses either, switch to low and evaluate. Everything else keeps its current setting.
This matters for routing economics. Astra at low is still a reasoning model — there is no classical completion mode. For mechanical workloads (classification, extraction, simple formatting), routing to GPT-5.6 Terra or Luna remains the cost-efficient option.
# Astra at medium effort
response = client.responses.create(
model="gpt-6-astra",
reasoning={"effort": "medium"},
input="Summarize this changelog in three bullet points.",
)
At max effort, expect time-to-first-token measured in minutes, not seconds. Budget for latency before you budget for tokens.
Breaking Changes from GPT-5.6 Sol
If you are migrating from Sol, review these changes before switching the model string:
-
temperatureandtop_premoved. Passing either parameter returns an error. Remove them from your request payload. -
logprobsremoved. Not available on Astra. If you rely on logprobs for confidence scoring, keep a Sol or Terra fallback path. -
noneandminimalreasoning effort removed. Uselowinstead. -
prompt_cache_retentionrenamed. The field is nowprompt_cache_options.ttl. Update your cache configuration. -
Tool calling requires the Responses API. Chat Completions still handles plain text, but function calling and built-in tools (web search, code interpreter, file search, computer use) require the
/v1/responsesendpoint. -
Enterprise opt-in required. Astra is disabled by default in Enterprise workspaces. An admin must enable it before API calls succeed.
-
Safety monitoring can interrupt work. Astra ships with runtime safety filters that can stop a task it flags as suspicious. OpenAI acknowledges the filters currently produce false positives on innocuous work. Worth knowing before you wire Astra into an unattended pipeline.
# Before (Sol)
response = client.chat.completions.create(
model="gpt-5.6-sol",
messages=[{"role": "user", "content": "Hello"}],
temperature=0.7, # REMOVE for Astra
top_p=0.9, # REMOVE for Astra
)
# After (Astra via Responses API)
response = client.responses.create(
model="gpt-6-astra",
reasoning={"effort": "medium"},
input=[{"role": "user", "content": "Hello"}],
)
Multi-Provider Routing
GPT-6 Astra is available through three provider endpoints plus third-party gateways. Each has a different model identifier:
| Provider | Endpoint | Model ID |
|---|---|---|
| OpenAI | https://api.openai.com/v1/ | gpt-6-astra |
| Azure Foundry | https://<resource>.openai.azure.com/ | Deployment-specific |
| AWS Bedrock | Regional Bedrock endpoint | openai.gpt-6-astra-v1:0 (check Bedrock docs) |
| OpenRouter | https://openrouter.ai/api/v1/ | openai/gpt-6-astra |
| TheRouter | https://api.therouter.ai/v1/ | openai/gpt-6-astra |
With TheRouter, you configure Astra as a primary model and set fallbacks. If OpenAI's endpoint returns a 5xx or hits a rate limit, TheRouter automatically routes the request to your configured fallback — another GPT model, Claude, or any OpenAI-compatible provider.
from openai import OpenAI
# Route through TheRouter with automatic fallback
client = OpenAI(
api_key="YOUR_THEROUTER_KEY",
base_url="https://api.therouter.ai/v1",
)
# If openai/gpt-6-astra fails, TheRouter falls back
# to your configured secondary provider
response = client.chat.completions.create(
model="openai/gpt-6-astra",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain prompt caching in one paragraph."},
],
)
For details on configuring model fallbacks, see the TheRouter routing guide.
Common Errors and Fixes
| Error | Cause | Fix |
|---|---|---|
temperature is not supported | Passing temperature to Astra | Remove temperature from your request |
top_p is not supported | Passing top_p to Astra | Remove top_p from your request |
logprobs is not supported | Requesting logprobs | Remove logprobs; use a Sol/Terra fallback if needed |
Invalid reasoning effort | Using none or minimal | Switch to low |
model_not_found (Enterprise) | Astra not enabled in workspace | Have an admin enable the model |
rate_limit_exceeded | Tier limit hit | Check your rate limit tier; upgrade or add retry logic |
context_length_exceeded | Input exceeds 1,050,000 tokens | Trim input or split into multiple calls |
Production Checklist
Before deploying GPT-6 Astra to production:
- Remove unsupported parameters. Strip
temperature,top_p,logprobsfrom all Astra requests. - Update reasoning effort. Replace
none/minimalwithlow. Verify latency at your chosen effort level. - Migrate tool calls to Responses API. Function calling requires
/v1/responses, not Chat Completions. - Update cache configuration. Rename
prompt_cache_retentiontoprompt_cache_options.ttl. - Budget for long-context surcharges. Prompts over 272K input tokens trigger 2x pricing. Monitor input token counts.
- Set up fallback routing. Configure a fallback model (Sol, Terra, or a cross-provider alternative) for rate limits and outages.
- Test safety filters. Run Astra through your automated pipelines and verify it does not halt unexpectedly on legitimate tasks.
- Enable in Enterprise workspace. Have an admin opt in the model before your code goes live.
- Evaluate cost vs. Sol. Run your representative workloads on both models. Compare total cost per task, not just per-token price.
TheRouter Integration Note
GPT-6 Astra is not yet listed in TheRouter's models page as of this writing. When it is added, you will be able to route to it with the standard openai/gpt-6-astra model identifier. Until then, you can configure a custom model route pointing to OpenAI's endpoint directly.
TheRouter's routing and fallback capabilities are especially relevant for Astra because of the 2.5x cost increase over Sol. A routing configuration that sends simple tasks to Terra or Luna and reserves Astra for complex reasoning can cut your overall spend significantly while keeping frontier-grade intelligence available for the tasks that need it.
Sources
- OpenAI: GPT-6 Astra announcement — retrieved September 9, 2026
- OpenAI API pricing — retrieved September 9, 2026
- Apidog: GPT-6 Astra API guide — retrieved September 9, 2026
- CloudZero: GPT-6 Astra pricing breakdown — retrieved September 9, 2026
- Azure: GPT-6 Astra on Azure Foundry — retrieved September 9, 2026
- Amazon: GPT-6 Astra on Bedrock — retrieved September 9, 2026