← All articles

OpenAI Ultrafast (Cerebras): GPT-5.6 Sol at 14x Speed — Service Tier Routing Guide

OpenAI's Ultrafast mode runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second — 14 times faster than Standard processing. We break down how service tiers work, what Ultrafast means for API routing decisions, and how operators can prepare their stacks for a latency-optimized inference path.

· TheRouter

On August 13, 2026, OpenAI previewed Ultrafast — a new service tier that runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second, roughly 14 times the speed of Standard processing. This is not a new model. It is the same GPT-5.6 Sol, run on different silicon, with a different latency profile.

For API operators, Ultrafast changes the routing calculus. Until now, getting frontier intelligence and low latency from the same model meant choosing Fast mode (2.5x Standard speed at 2x the price). Ultrafast pushes the speed ceiling much further, opening up use cases where even Fast mode was too slow.

This guide covers everything published so far about Ultrafast, how it fits into OpenAI's service tier system alongside Standard, Fast, Batch, and Flex modes, and what operators should plan for when access expands.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Sources: OpenAI — Previewing Ultrafast, retrieved 2026-08-19; Cerebras — Accelerating GPT-5.6 Sol Ultrafast, retrieved 2026-08-19; OpenAI API Pricing, retrieved 2026-08-19; OpenAI — Fast Mode Guide, retrieved 2026-08-19; MLQ — OpenAI previews GPT-5.6 Sol Ultrafast, retrieved 2026-08-19.

Getting Started in 3 Minutes

Ultrafast is currently in limited preview. You cannot opt in through the API today the way you can with Fast mode. Here is what you can do right now:

  1. Sign up for access at openai.com/form/ultrafast/.
  2. Route GPT-5.6 Sol requests through Fast mode as the current best latency option available to all API users, using service_tier: "fast" in your requests.
  3. Structure your routing logic so that switching from Fast to Ultrafast requires changing one parameter value, not rewriting your integration.

When Ultrafast becomes generally available, the expected integration will follow the existing service_tier pattern. Based on OpenAI's current tier architecture, a request would look something like this:

from openai import OpenAI

client = OpenAI()

# Fast mode — available now
response = client.responses.create(
    model="gpt-5.6-sol",
    input="Analyze this incident log and identify root cause.",
    service_tier="fast",
)

# Ultrafast — when access is granted
# response = client.responses.create(
#     model="gpt-5.6-sol",
#     input="Analyze this incident log and identify root cause.",
#     service_tier="ultrafast",  # expected parameter value
# )

Note: The service_tier: "ultrafast" parameter value has not been publicly confirmed. The example above follows the pattern established by "fast" and "flex". Check the OpenAI service tier documentation for the confirmed value when your account gains access.

OpenAI Service Tiers Explained

OpenAI now offers multiple processing modes for the same model, each trading off latency, cost, and availability guarantees differently. Understanding these tiers is essential for routing decisions.

TierSpeed vs StandardPricing (Sol, per 1M tokens)AvailabilityBest For
Standard1x (baseline)$5.00 in / $30.00 outGA, all accountsGeneral production traffic
BatchAsync (up to 24h)$2.50 in / $15.00 out (50% off)GAOffline processing, ETL
FlexVariable, may queue$2.50 in / $15.00 out (50% off)GACost-sensitive, tolerant of variance
FastUp to 2.5x$10.00 in / $60.00 out (2x Standard)GAUser-facing, latency-critical
UltrafastUp to 14x (~750 tok/s)Not publishedLimited previewReal-time agent loops, incident response

Pricing from OpenAI API Pricing, retrieved August 19, 2026. Ultrafast pricing has not been disclosed.

How service_tier works in the API

Every request to the Responses API or Chat Completions API accepts a service_tier parameter:

  • "auto" or omitted — routes to Standard
  • "fast" (or legacy "priority") — routes to Fast mode
  • "flex" — routes to Flex processing
  • "default" — returned in the response when Fast mode is downgraded to Standard due to ramp rate limits

The response object includes a service_tier field confirming which tier actually processed the request. This is important for monitoring: if you request "fast" but the response says "default", the ramp rate limiter kicked in and you were served at Standard speed with Standard pricing.

import OpenAI from "openai";

const openai = new OpenAI();

const response = await openai.responses.create({
  model: "gpt-5.6-sol",
  input: "Summarize this quarterly report.",
  service_tier: "fast",
});

// Check which tier actually processed the request
console.log(`Processed by: ${response.service_tier}`);
// "priority" = Fast mode was used
// "default" = downgraded to Standard

You can also set the default tier at the project level in OpenAI's dashboard under Settings > General > Project Service Tier, so all requests from that project default to a specific tier without per-request configuration.

What Makes Ultrafast Different

The hardware: Cerebras Wafer-Scale Engine

Standard and Fast mode run GPT-5.6 Sol on GPU clusters. Ultrafast runs the same model on Cerebras Wafer-Scale Engine chips, which take a fundamentally different approach to inference.

GPU inference on large models is bottlenecked by memory bandwidth — model weights must be repeatedly transferred between on-chip memory and off-chip storage during token generation. Cerebras packs 44 GB of SRAM on each wafer-sized chip, keeping weights on-chip and eliminating that data movement bottleneck. Tokens flow through model layers pipelined across wafers without the memory transfer overhead that limits GPU throughput.

The result: GPT-5.6 Sol generates up to 750 output tokens per second on Ultrafast, compared to roughly 50-55 tok/s on Standard processing.

Speed claims in context

OpenAI and Cerebras report specific speed comparisons:

Model / ModeOutput Tokens/secSource
GPT-5.6 Sol UltrafastUp to 750OpenAI
GPT-5.6 Sol Standard~50-55 (implied by 14x claim)Derived
Claude Fable 5~140 (implied by 5x comparison)Cerebras
Claude Opus 4.8 Fast~150 (implied by 5x comparison)Cerebras

Caveats: OpenAI has not published the prompt size, output length, concurrency, warm-up conditions, or time-to-first-token measurements behind the 750 tok/s figure. The "14x" claim describes output generation throughput, not necessarily total request latency including time-to-first-token. Independent benchmarks have not been published yet.

Cerebras also tested GPT-5.6 Sol Ultrafast on Humanity's Last Exam (HLE) — 2,500 PhD-level questions. Ultrafast completed the full benchmark in 11 hours 11 minutes. Claude Fable 5 needed 78 hours 27 minutes for comparable accuracy. That is a 7x wall-clock speedup on a real workload.

When Ultrafast Makes Sense

Not every workload benefits from 750 tok/s. The speed premium (pricing not published) means Ultrafast is overkill for batch classification or overnight ETL jobs. The sweet spot is workloads where:

Latency is on the critical path of a human interaction:

  • Incident response: analyzing logs and traces while the outage is still unfolding
  • Real-time voice/support: resolving complex issues without conversational lag
  • Commerce: answering product questions before the customer loses interest

Agent loops iterate faster than a human can context-switch:

  • Research workflows that used to require overnight batch runs can become same-day interactive sessions
  • Agentic coding where the model is executing multi-step plans and waiting for model output is the bottleneck

Competitive time pressure exists:

  • Financial signal analysis where minutes matter
  • Security incident response where detection-to-containment time is the key metric

When to stick with Standard or Fast

ScenarioRecommended TierWhy
Background document processingBatch or FlexCost matters more than speed
General chat applicationStandard50+ tok/s is fine for conversational UX
User-facing product with latency SLAsFast2.5x Standard speed at known pricing
Real-time agent loop, incident responseUltrafast (when available)Speed directly affects outcome quality
High-volume classification/extractionLuna + StandardLuna at $0.20/$1.20 per 1M is 25x cheaper than Sol

Common Errors and Fixes

Since Ultrafast is in limited preview, most operators will be working with Fast mode today. Here are the common issues:

Ramp rate limit downgrades

Symptom: You request service_tier: "fast" but the response returns service_tier: "default".

Cause: If your traffic exceeds 1M TPM and ramps by more than 50% within 15 minutes, Fast mode downgrades requests to Standard speed at Standard pricing.

Fix:

  • Ramp traffic gradually when switching models or tiers
  • Use feature flags to shift traffic over hours, not instantly
  • Avoid running ETL or batch jobs in Fast mode — use Batch tier instead

Model not supported in Fast mode

Symptom: InvalidRequestError when using service_tier: "fast" with certain models.

Fix: Fast mode supports GPT-5.6 Sol, Terra, Luna, GPT-5.5, GPT-5.4 and earlier GPT-4 series models. Fine-tuned models and embeddings are not supported. Check the pricing page for the current list.

Confusing priority vs fast in response

Symptom: You send service_tier: "fast" but the response says service_tier: "priority".

This is expected. For GPT-5.6 and earlier models, the response returns "priority" regardless of whether you sent "priority" or "fast". Both values route to the same processing path.

Production Checklist

Before adding Ultrafast (or Fast mode) to your production stack:

  • Monitor service_tier in responses — log the response field, not just your request parameter, to track actual tier usage
  • Set up tier-based cost alerts — Fast mode is 2x Standard pricing; Ultrafast pricing has not been disclosed but is expected to be higher
  • Implement graceful fallback — if Ultrafast/Fast is unavailable or rate-limited, fall back to Standard without failing the request
  • Separate latency-critical from batch traffic — route interactive requests through Fast/Ultrafast, batch work through Batch/Flex
  • Test ramp behavior — gradually increase Fast mode traffic to avoid triggering the 50%/15-minute ramp rate limit
  • Review data residency — Fast mode is compatible with data residency, ZDR, and BAA; confirm Ultrafast compatibility when access is granted
  • Track per-tier metrics — measure TTFT, total latency, and cost separately for each tier to validate the premium is worth the speed

TheRouter Integration Note

TheRouter routes OpenAI-compatible requests to configured providers. GPT-5.6 Sol is available through TheRouter's OpenAI provider integration, and the service_tier parameter passes through to OpenAI's API.

When using TheRouter to route GPT-5.6 Sol requests:

  • Standard, Fast, Batch, Flex — the service_tier parameter is forwarded to OpenAI. Set it in your request and TheRouter preserves it.
  • Ultrafast — when OpenAI expands Ultrafast access, the service_tier value will pass through the same way. No TheRouter configuration change is expected.
  • Fallback routing — you can configure model fallbacks so that if a GPT-5.6 Sol request fails, it falls back to Terra or another provider. The service_tier parameter applies only to the OpenAI path; fallback providers will use their own processing modes.

For details on routing GPT-5.6 models through TheRouter, see the OpenAI provider page and the GPT-5.6 Sol model page.

What We Know vs What We Do Not Know

KnownUnknown
Ultrafast runs GPT-5.6 Sol on Cerebras WSEUltrafast pricing (per-token premium)
Up to 750 output tok/s, 14x StandardExact service_tier parameter value
Limited preview since Aug 13, 2026GA timeline
Compatible with existing OpenAI API formatSupported regions and data residency
Early customers include Jane Street, Podium, Basis, RogoRate limits specific to Ultrafast
Cerebras partnership covers 750 MW through 2028Whether Ultrafast will extend to Terra/Luna
No quality compromise vs Standard SolTime-to-first-token measurements

This guide will be updated as OpenAI publishes Ultrafast pricing, confirms the API parameter, and expands availability.

FAQ

Is Ultrafast a new model?

No. Ultrafast runs the same GPT-5.6 Sol model as Standard and Fast processing. The difference is the inference hardware (Cerebras Wafer-Scale Engine instead of GPUs) and the resulting throughput.

Can I use Ultrafast today?

Only if you are in the limited preview group. Sign up for access notifications. In the meantime, Fast mode (service_tier: "fast") is the best available latency option for all API users.

How does Ultrafast compare to Fast mode?

Fast mode delivers up to 2.5x Standard speed at 2x the per-token price. Ultrafast delivers up to 14x Standard speed. The pricing premium for Ultrafast has not been published.

Will Ultrafast work with other GPT-5.6 models (Terra, Luna)?

OpenAI has only announced Ultrafast for GPT-5.6 Sol. There is no information on whether Terra or Luna will be supported.

Does Ultrafast affect model quality?

OpenAI and Cerebras both state there is no quality compromise. Cerebras tested GPT-5.6 Sol Ultrafast on Humanity's Last Exam and reported comparable accuracy to Standard processing, with a 7x wall-clock speedup over Claude Fable 5.

How does this relate to the OpenAI-Cerebras partnership?

OpenAI announced a partnership with Cerebras in January 2026, covering 750 megawatts of inference capacity through 2028, with an option for an additional 1.25 GW by 2030. Ultrafast is the latest product built on that infrastructure.

Models covered in this article

Help & contact