GPT-6 Sol and GPT-6 Luna API Guide: Pricing, Integration, and Migration from GPT-5.6
GPT-6 Sol ($2/$10 per 1M tokens) and GPT-6 Luna ($0.10/$0.50) launched September 22, 2026 — a 50%+ price cut from GPT-5.6. We cover model IDs, context windows, reasoning effort, SDK integration, migration from GPT-5.6, three-tier GPT-6 routing, cross-provider fallback with Claude Sonnet 5, and cost savings math for mixed workloads.
GPT-6 Sol and GPT-6 Luna are the mid-tier and value-tier models in OpenAI's GPT-6 family. They launched on September 22, 2026, at 18:00 UTC — 19 days after GPT-6 Astra. Sol costs $2 per million input tokens and $10 output. Luna costs $0.10 input and $0.50 output. Both represent a 50%+ price cut from their GPT-5.6 predecessors, both support a 1,050,000-token context window with reasoning effort from none to max, and both are available immediately in the API under model IDs gpt-6-sol and gpt-6-luna.
If you are running GPT-5.6 Sol or Luna in production, the migration is a model-name swap with one behavioral difference worth testing: GPT-6 models default to medium reasoning effort and will ask clarifying questions more often than GPT-5.6 did. Everything below — pricing tables, SDK code, caching changes, routing configurations — is sourced from OpenAI's pricing page and model documentation as of September 22, 2026.
GPT-6 Sol and Luna at a Glance
| Spec | GPT-6 Sol | GPT-6 Luna | GPT-6 Astra |
|---|---|---|---|
| Model ID | gpt-6-sol | gpt-6-luna | gpt-6-astra |
| Input / 1M tokens | $2.00 | $0.10 | $10.00 |
| Cached input / 1M tokens | $0.20 | $0.01 | $1.00 |
| Output / 1M tokens | $10.00 | $0.50 | $50.00 |
| Long context (>272K) input | $4.00 | $0.20 | $20.00 |
| Long context output | $15.00 | $0.75 | $75.00 |
| Context window | 1,050,000 | 1,050,000 | 1,050,000 |
| Max input tokens | 922,000 | 922,000 | 922,000 |
| Max output tokens | 128,000 | 128,000 | 128,000 |
| Reasoning effort | none–max (default medium) | none–max (default medium) | low–max (no none) |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 | March 31, 2026 |
| Batch pricing | 50% off standard | 50% off standard | 50% off standard |
Sources: OpenAI Pricing (retrieved 2026-09-22), Using GPT-6 (retrieved 2026-09-22).
Pricing Deep-Dive: What Changed from GPT-5.6
The headline is simple: 50% off across the board, measured against GPT-5.6's promotional rates.
| Model | Input | Output | Cut vs GPT-5.6 |
|---|---|---|---|
| GPT-6 Sol | $2.00 | $10.00 | 50% input, 50% output vs GPT-5.6 Sol ($4/$20) |
| GPT-6 Luna | $0.10 | $0.50 | 50% input, 58% output vs GPT-5.6 Luna ($0.20/$1.20) |
| GPT-6 Astra | $10.00 | $50.00 | Flagship tier, no predecessor comparison |
GPT-5.6 Sol's $4/$20 pricing was already a promotional rate (original launch was higher). GPT-6 Sol at $2/$10 is the new permanent price — not promotional. GPT-5.6 Sol's promotional pricing remains available through at least November 21, 2026, for users who have not yet migrated.
Batch pricing halves the standard rate again: Sol drops to $1.00/$5.00, Luna to $0.05/$0.25. For high-volume offline workloads — classification, extraction, summarization — Luna batch at $0.05 input is effectively free compared to any frontier model.
Long-context pricing kicks in above 272K input tokens: 2x input price, 1.5x output price. Sol long-context input is $4.00 (matching GPT-5.6 Sol's standard rate), and Luna long-context input is $0.20 (matching GPT-5.6 Luna's standard rate). If your prompts routinely exceed 272K tokens, the effective per-token cost is the same as GPT-5.6 at standard pricing — but you get GPT-6 quality.
Model IDs and API Access
Both models are available through the standard OpenAI Responses API and Chat Completions API.
from openai import OpenAI
client = OpenAI() # uses OPENAI_API_KEY env var
# GPT-6 Sol — strong reasoning for coding, analysis, multi-step tasks
response = client.responses.create(
model="gpt-6-sol",
input="Explain how to implement exponential backoff with jitter for API rate limiting.",
)
print(response.output_text)
# GPT-6 Luna — cost-efficient for extraction, classification, summarization
response = client.responses.create(
model="gpt-6-luna",
input="Extract the company name, revenue, and YoY growth from this earnings report.",
)
print(response.output_text)
// Node.js / TypeScript
import OpenAI from "openai";
const client = new OpenAI();
const response = await client.responses.create({
model: "gpt-6-sol",
input: "Write a Python function that implements a circuit breaker pattern.",
});
console.log(response.output_text);
Availability: API access is immediate for all API tiers. ChatGPT Work and Codex are available on Plus, Pro, Business, Enterprise, and Edu plans. Free and Go users get Luna in the desktop app. The models are not yet available in Chat (the web interface).
Reasoning Effort: Tuning Cost and Quality
GPT-6 Sol and Luna both support the full none to max reasoning effort range (Astra does not support none). The default is medium.
# Low-effort call for simple extraction — faster, cheaper
response = client.responses.create(
model="gpt-6-luna",
input="Classify this support ticket as billing, technical, or general.",
reasoning={"effort": "low"},
)
# High-effort call for complex analysis
response = client.responses.create(
model="gpt-6-sol",
input="Review this pull request for security vulnerabilities and suggest fixes.",
reasoning={"effort": "high"},
)
Setting effort to none disables extended thinking entirely — the model responds in a single pass. This is useful for latency-critical tasks where you want deterministic, fast responses without reasoning overhead.
Mid-conversation effort changes are new in GPT-6: you can increase effort for a difficult follow-up question without rewriting your prompt prefix, and the cache stays intact. Add a configuration_update input item to switch effort mid-conversation.
Caching: What Improved in GPT-6
GPT-6 introduces three caching improvements that directly affect cost:
- Higher default hit rates. OpenAI states that GPT-6 caching yields higher hit rates by default compared to GPT-5.6.
- Reasoning/tool changes no longer break the cache. In GPT-5.6, changing
reasoning_effortor tool definitions could invalidate your cached prefix. GPT-6 preserves the cache across these changes. - Explicit cache breakpoints. You can now tell the model where a cached prefix ends, giving you fine-grained control over what gets cached vs. what varies per request.
Cached input pricing is 10% of the standard input price on both Sol and Luna — $0.20 per million tokens on Sol and $0.01 on Luna.
When to Use Sol vs Luna vs Astra
The GPT-6 family maps cleanly to workload complexity:
GPT-6 Astra ($10/$50) — Frontier reasoning. Multi-day software projects, complex agentic workflows, tasks where getting the answer wrong costs more than the token spend. Use when quality is the only constraint.
GPT-6 Sol ($2/$10) — The everyday workhorse. Strong coding, analysis, tool calling, and multi-step reasoning. Sol scores within 1.1 percentage points of Claude Fable 5's highest score on DeepSWE v1.1, at a fraction of the cost. Use for production workloads that need reliable reasoning without frontier pricing.
GPT-6 Luna ($0.10/$0.50) — High-volume, cost-sensitive work. Classification, extraction, summarization, structured output generation, and simple tool calls. Luna scores 66.6% on DeepSWE v1.1 at 93% lower cost per task than Claude Opus 5. Use when you need to process thousands of requests per hour without worrying about cost.
Decision shortcut: If the task would take a senior engineer more than 30 minutes to do manually, start with Sol. If it is a task a junior engineer could do in 5 minutes, start with Luna. If it involves coordinating across multiple tools, browsers, or codebases for hours, use Astra.
Migrating from GPT-5.6
The migration is straightforward: change the model name. The API surface is identical.
- model="gpt-5.6-sol"
+ model="gpt-6-sol"
- model="gpt-5.6-luna"
+ model="gpt-6-luna"
Behavioral differences to test:
-
GPT-6 asks more questions. Astra's "initiative and follow-through" behavior — asking for clarification instead of assuming — carries over to Sol and Luna. If your prompts rely on the model making assumptions, you may need to add explicit instructions like "infer the user's intent and carry the task to completion without asking for clarification."
-
Default reasoning effort is
medium. GPT-5.6 also defaulted to medium, so no change here. But GPT-6 at medium effort is more capable than GPT-5.6 at medium effort — you may find that tasks that previously neededhigheffort now work fine atmedium, saving tokens. -
Knowledge cutoff moved forward. Sol's cutoff is April 20, 2026 (vs February 16 for GPT-5.6 Sol). Luna's cutoff is May 18, 2026. If your prompts reference events between February and May 2026, the model will now have that context built in.
-
Caching is more resilient. If you were careful about not changing reasoning effort or tools mid-conversation to preserve cache hits, you can relax that constraint with GPT-6.
Rollback path: GPT-5.6 Sol and Luna remain available in the API. GPT-5.6 Sol's promotional pricing ($4/$20) runs through at least November 21, 2026. You can run both models in parallel during migration — route a percentage of traffic to GPT-6, compare results, and cut over when confident.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Three-Tier GPT-6 Routing
For teams running mixed workloads, configure a three-tier GPT-6 routing stack that matches task complexity to model cost:
# TheRouter configuration example
routes:
- name: gpt6-astra-frontier
model: openai/gpt-6-astra
match:
tags: [frontier, multi-tool, computer-use]
priority: 1
- name: gpt6-sol-workhorse
model: openai/gpt-6-sol
match:
tags: [coding, analysis, tool-calling]
priority: 2
fallback: gpt6-luna-value
- name: gpt6-luna-value
model: openai/gpt-6-luna
match:
tags: [extraction, classification, summarization]
priority: 3
This configuration routes frontier tasks to Astra, everyday development work to Sol, and high-volume processing to Luna. Sol falls back to Luna if the Sol endpoint is unavailable or rate-limited — the quality drop is modest for most tasks, and the cost drop is 20x.
Cross-Provider Routing: Sol vs Claude Sonnet 5
GPT-6 Sol at $2/$10 matches Claude Sonnet 5 at exactly $2/$10. At identical pricing, the routing decision comes down to workload fit:
| Factor | GPT-6 Sol | Claude Sonnet 5 |
|---|---|---|
| Price (input/output per 1M) | $2 / $10 | $2 / $10 |
| Context window | 1,050,000 | 1,000,000 |
| Max output | 128,000 | 64,000 |
| Reasoning effort control | none–max | low–max |
| Prompt caching | Yes, improved in GPT-6 | Yes |
| Batch API | Yes, 50% off | Yes, 50% off |
| Tool calling | Yes, async tool calling new in GPT-6 | Yes |
| Strengths | Coding benchmarks, async tools, longer output | Agentic workflows, extended thinking, creative writing |
At price parity, we configure cross-provider routing as a reliability play: if OpenAI returns a 5xx or hits rate limits, fall back to Anthropic (or vice versa) without any cost penalty.
Cost Savings: Before and After
Consider a team running 100M input tokens and 20M output tokens per day on GPT-5.6 Sol:
| GPT-5.6 Sol | GPT-6 Sol | Savings | |
|---|---|---|---|
| Input cost/day | $400 | $200 | $200/day |
| Output cost/day | $400 | $200 | $200/day |
| Monthly total | $24,000 | $12,000 | $12,000/month |
Switch the Luna workloads separately — a team processing 500M input tokens per day on GPT-5.6 Luna:
| GPT-5.6 Luna | GPT-6 Luna | Savings | |
|---|---|---|---|
| Input cost/day | $100 | $50 | $50/day |
| Output cost/day | $120 (at $1.20) | $50 (at $0.50) | $70/day |
| Monthly total | $6,600 | $3,000 | $3,600/month |
Batch savings stack: If those Luna workloads can tolerate async processing, batch pricing drops the cost another 50% — to $1,500/month total.
Production Checklist
Before switching production traffic to GPT-6 Sol or Luna:
- Update model IDs in configuration:
gpt-5.6-soltogpt-6-sol,gpt-5.6-lunatogpt-6-luna - Test with your actual prompts — GPT-6 may ask clarifying questions where GPT-5.6 assumed
- Add
"infer the user's intent and carry the task to completion"to system prompts if autonomous behavior is required - Verify long-context pricing impact if prompts exceed 272K tokens (2x input, 1.5x output)
- Check caching behavior — reasoning effort changes no longer break the cache
- Set up fallback routing: Sol to Luna, or Sol to Claude Sonnet 5 for cross-provider resilience
- Monitor output quality for 24–48 hours before cutting over 100% of traffic
- Keep GPT-5.6 model IDs as rollback targets — they remain available through at least November 2026
FAQ
What are the model IDs for GPT-6 Sol and Luna?
Use gpt-6-sol and gpt-6-luna in the model parameter of any OpenAI API request.
Is GPT-6 Sol better than GPT-5.6 Sol? Yes. GPT-6 Sol scores higher on AutomationBench, DeepSWE, and Agents' Last Exam while costing 50% less. The knowledge cutoff is also more recent (April 2026 vs February 2026).
What is the cheapest OpenAI model? GPT-6 Luna at $0.10/$0.50 per million tokens is the cheapest model in the GPT-6 family and the cheapest current-generation OpenAI model. Batch pricing drops this to $0.05/$0.25.
Can I use GPT-6 Sol and Luna with the OpenAI Python SDK?
Yes. Install openai >= 1.0.0 and set model="gpt-6-sol" or model="gpt-6-luna". No SDK changes are needed.
When will GPT-5.6 Sol and Luna be deprecated? OpenAI has not announced deprecation dates. GPT-5.6 Sol's promotional pricing runs through at least November 21, 2026. Both models remain available in the API.
Does GPT-6 Luna support reasoning effort?
Yes. Luna supports the full none to max range, unlike Astra which does not support none. The default is medium.
Is GPT-6 Sol the same price as Claude Sonnet 5? Yes. Both are $2/$10 per million tokens. At identical pricing, the choice comes down to workload fit and provider reliability.
What is the context window for GPT-6 Sol and Luna? 1,050,000 tokens total, with 922,000 max input and 128,000 max output. Prompts over 272K input tokens incur long-context pricing (2x input, 1.5x output).
Can I route GPT-6 Sol through TheRouter?
Yes. TheRouter routes OpenAI-compatible requests through configured providers. Configure openai/gpt-6-sol or openai/gpt-6-luna as model targets with fallback routing to other providers.
What is new in the GPT-6 API compared to GPT-5.6? Async tool calling (continue reasoning while tools execute), mid-turn steering (send corrections while the model works), configuration updates mid-conversation (change reasoning effort without breaking cache), and improved default cache hit rates.