Back to Models

gpt-realtime-whisper

openaiopenai/gpt-realtime-whisper

Whisper-class automatic speech recognition over the realtime channel. Audio in, transcribed text out.

GPT-Realtime-Whisper is OpenAI's streaming speech-to-text model, released May 7, 2026 alongside GPT-Realtime-2 and GPT-Realtime-Translate. It converts live audio into text in real time, streaming incremental transcript deltas as the speaker talks β€” without waiting for utterances to finish. Designed for low-latency applications, it offers tunable delay settings from 'minimal' to 'xhigh', letting developers trade latency for accuracy depending on the use case.

TheRouter serves this model over one ordinary HTTP request β€” POST /v1/audio/transcriptions, multipart, the same contract as Whisper β€” and does not expose the realtime streaming channel at all. You upload audio, you get one JSON body back. There are no incremental deltas, no session to open, no WebSocket or WebRTC client to write, and no delay/turn-detection knobs. Every /v1/realtime path returns 501 realtime_websocket_not_supported, before authentication, for every model β€” so if you see that error, your API key is not the cause.

Best for
  • β€’ Near-live captions for conferences, webinars, broadcasts β€” on TheRouter this means segmenting speaker audio client-side and posting each segment, so subtitles lag by roughly one segment plus a round trip (seconds), and words can be lost or repeated at segment boundaries
  • β€’ Compliance and quality monitoring β€” capture conversations as text for regulatory archiving, quality assurance, or analytics pipelines
  • β€’ Multilingual transcription pipelines β€” pair with GPT-Realtime-Translate to capture source-language transcripts alongside translated audio, giving users both
  • β€’ Voice-controlled interfaces where text visibility matters β€” show users what the system heard alongside the voice agent's response
Reach for something else if
  • β€’ Conversational voice agents β€” for assistants that reason, call tools, and generate spoken responses, use GPT-Realtime-2 instead
  • β€’ Pre-recorded audio transcription β€” for files or bounded audio requests, use gpt-4o-transcribe ($0.006/min, higher accuracy) or gpt-4o-mini-transcribe ($0.003/min)
  • β€’ Live translation β€” GPT-Realtime-Whisper only transcribes, it does not translate. Use GPT-Realtime-Translate for speech-to-speech translation

How TheRouter serves this differently from the vendor

As the vendor operates it

A streaming realtime channel: a WebSocket session that emits incremental transcript deltas while the speaker is still talking, with tunable delay from 'minimal' to 'xhigh' and turn detection.

On TheRouter

One ordinary HTTP request: POST /v1/audio/transcriptions, multipart, one JSON body back. No incremental deltas, no session to open, no WebSocket or WebRTC client, no delay or turn-detection knobs. Every /v1/realtime path returns 501 realtime_websocket_not_supported, before authentication, for every model.

Context Length
--
Max Output
--
Audio Priceper audio minute
$0.0184/ minute of audio

Modalities

audio→text

Capabilities

STT

Pricing Breakdown

TypeRate
Audio$0.0184 / minute of audio

request is price per minute of audio, the only unit OpenAI publishes for this model.

Supported Parameters

filelanguagepromptresponse_formattemperature

Specifications

Release dateMay 7, 2026developers.openai.com β†—verified
Knowledge cutoffSeptember 30, 2024developers.openai.com β†—verified
Context window16,000 tokensdevelopers.openai.com β†—verified
Max output tokens2,000 tokensdevelopers.openai.com β†—verified
Billing unitPer minute of audio processed, derived from the duration field of each response β€” not per token and not per request. Ten 6-second requests and one 60-second request cost the same. OpenAI's own list price is $0.017/minute.developers.openai.com β†—verified
API endpointUpstream (OpenAI direct, NOT reachable through TheRouter): /v1/realtime/transcription_sessions, session type 'transcription', over WebRTC or WebSocket. On TheRouter the entry point is POST /v1/audio/transcriptions β€” every /v1/realtime path returns 501 realtime_websocket_not_supported before authentication runs, so an API key is never the cause. See the Realtime Transcription guide.developers.openai.com β†—verified
Latency tuningNOT AVAILABLE ON THEROUTER. audio.input.transcription.delay (minimal / low / medium / high / xhigh) is a realtime-session parameter and there is no realtime session here β€” sending it has no effect. The accepted parameters are exactly: file, language, prompt, response_format, temperature. On TheRouter, latency is governed by your client-side chunk length plus one HTTP round trip per chunk (measured 1.7-6.9 s per chunk), not by a delay setting.developers.openai.com β†—verified
Supported featuresOn TheRouter β€” Streaming: NO. The response is a single JSON body returned after the whole submitted audio is transcribed; there are no incremental deltas, because the streaming transport this model uses upstream is not exposed here. Near-realtime output is achieved by segmenting audio client-side and issuing one request per segment. Prompt steering: YES β€” the prompt parameter is accepted (it is unavailable on OpenAI's own GA realtime session). Function calling: no. Structured outputs: no. Fine-tuning: no.developers.openai.com β†—verified
Audio input formatOn TheRouter: a container file uploaded as a multipart file part β€” WAV, MP3, FLAC, M4A, MPEG, MPGA, OGG, or WebM, up to 25 MB. RAW PCM IS NOT ACCEPTED: the upstream realtime session takes 24 kHz mono PCM16 over its stream, but that channel is not exposed here, so PCM must be wrapped in a container (e.g. a WAV header) before upload. Each segment must decode standalone β€” a raw byte-slice of a longer file is not valid audio. The file part must carry a filename with a real extension; the format is inferred from it.developers.openai.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
Azure WER improvement
Azure Foundry documentation states GPT-Realtime-Whisper achieves approximately 50% lower Word Error Rate compared to previous gpt-4o realtime transcription. This is a relative claim from the provider, not an independently benchmarked absolute WER score.
~50% lower WER than gpt-4o realtime% reductionlearn.microsoft.com β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/audio/transcriptions   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -F "model=openai/gpt-realtime-whisper"   -F "file=@speech.mp3"

Transcription (TheRouter)

On TheRouter, GPT-Realtime-Whisper is served on POST /v1/audio/transcriptions β€” the same REST endpoint as Whisper β€” not on a chat completions or WebSocket path. TheRouter bridges the provider's realtime channel behind this unchanged public contract; request-level behavior is identical to any other transcription model. Supported params: file, language, prompt, response_format, temperature. Billed per minute of audio, not per token and not per request. There is no TheRouter chat/completions route for this model β€” calling it that way returns a 503. For the full protocol-level contract, near-realtime chunking pattern, and Java/Go/Python reference implementations, see the Realtime Transcription guide: https://therouter.ai/guides/multimodal/realtime-transcription

cURL
curl https://api.therouter.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -F file="@speech.wav" \
  -F model="openai/gpt-realtime-whisper" \
  -F language="en" \
  -F prompt="Technical terms: TheRouter, API" \
  -F response_format="json" \
  -F temperature="0"

More from openai

Similar models

Cross-provider sibling models

News & changes

2026-05-07

OpenAI launches GPT-Realtime-Whisper alongside GPT-Realtime-2 and GPT-Realtime-Translate

On May 7, 2026, OpenAI released three new realtime audio models: GPT-Realtime-2 (128K-context voice agent), GPT-Realtime-Translate (streaming translation at $0.034/min), and GPT-Realtime-Whisper (streaming speech-to-text at $0.017/min). GPT-Realtime-Whisper is designed for live transcription with tunable latency β€” the 'minimal' setting provides near-instant transcript deltas, while 'xhigh' favours accuracy over speed. The model connects via a dedicated transcription sessions endpoint with WebRTC (browsers) or WebSocket (server media pipelines). It is also available through Azure Foundry with approximately 50% lower WER than previous realtime transcription models.

re-authored by TheRouterdevelopers.openai.com β†—

Frequently asked

This model is documented as WebSocket-only. Is there a wss:// endpoint on TheRouter?

Yes, and it is partially live β€” measured on production 2026-07-29. POST /v1/realtime/sessions with this model and a valid key returns 200 with a session_token and expires_at (401 without a key, not 501). But connecting to wss://api.therouter.ai/v1/realtime with that token currently answers 503 upstream_unavailable β€” 'This model is not configured for realtime billing' β€” so a token can be obtained and not yet used. The conversational realtime family (openai/gpt-realtime, openai/gpt-realtime-2) is still gated and answers 404 model_not_found at session creation. Unimplemented sub-paths (/v1/realtime/transcription_sessions, /calls, /client_secrets) answer 501 rather than a bare 404, so a client falling back between them is not misled into blaming its own URL or key. What to build on today: POST /v1/audio/transcriptions, which works end to end and is unaffected by how the socket rollout resolves. Streaming on TheRouter also exists in the other direction β€” /v1/audio/speech (TTS) streams over SSE or chunked HTTP, never WebSocket.

How is GPT-Realtime-Whisper different from gpt-4o-transcribe?

Upstream, the distinction is streaming versus batch. On TheRouter that distinction disappears: both models are served over the same POST /v1/audio/transcriptions request, both return one complete transcript, and neither streams. What actually differs here is price and accuracy β€” and the trade-off runs opposite to what the names suggest. GPT-Realtime-Whisper's list price is $0.017/min against gpt-4o-transcribe's $0.006/min, so it is roughly 2.8x MORE expensive, while gpt-4o-transcribe typically achieves higher absolute accuracy on pre-recorded audio. Unless you specifically need this model, gpt-4o-transcribe is the better default on TheRouter on both counts. (Figures are OpenAI list prices; the charged rates are on each model's pricing card.)

Can I use GPT-Realtime-Whisper through the standard OpenAI chat completions SDK?

Not natively, and not on TheRouter either. Natively, OpenAI's GPT-Realtime-Whisper uses the dedicated /v1/realtime/transcription_sessions endpoint (WebRTC or WebSocket) β€” you manage a persistent connection, stream audio via input_audio_buffer.append events, and handle transcript deltas as they arrive. On TheRouter, this model is served only on POST /v1/audio/transcriptions β€” the standard REST transcription contract, request in and full transcript out. There is no TheRouter chat/completions route for this model; calling it that way returns a 503.

What does the 'delay' setting do and which one should I use?

The delay setting controls the latency/accuracy tradeoff for streaming transcription. Lower values (minimal, low) emit transcript deltas faster but may have higher word error rate because the model has less audio context. Higher values (high, xhigh) wait for more context before emitting text, improving accuracy at the cost of display delay. OpenAI recommends starting with 'low' for live captions and 'medium' for balanced use cases. Test with your actual audio β€” microphones, background noise, accents, and domain vocabulary all affect the best setting.

Can I steer the model's transcription with custom vocabulary or prompts?

Prompt steering is NOT supported for GPT-Realtime-Whisper in GA Realtime sessions. Open the session with a language hint instead (the 'language' parameter). Where prompt steering may become available in future snapshots, OpenAI recommends using short keyword lists (e.g., 'Keywords: metoprolol, atorvastatin, A1C') rather than full natural-language instructions. For production, treat keyword steering as an aid, not a guarantee β€” continue to manually evaluate names, numbers, dates, and domain terms.

How does GPT-Realtime-Whisper compare to traditional Whisper for latency?

Traditional Whisper (whisper-1) processes audio in chunks and must wait for the full utterance before transcribing β€” it is not natively streaming. GPT-Realtime-Whisper emits incremental transcript deltas continuously as audio arrives, with tunable delay from 'minimal' (near-instant partial text) to 'xhigh' (more context, better accuracy). For live captioning and streaming applications, the latency improvement is dramatic: users see text appearing mid-sentence rather than waiting seconds for the full turn. For pre-recorded files where latency doesn't matter, gpt-4o-transcribe or gpt-4o-mini-transcribe offer higher accuracy at lower cost.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datedevelopers.openai.com β†—2026-05-30verified
Knowledge cutoffdevelopers.openai.com β†—2026-05-30verified
Context windowdevelopers.openai.com β†—2026-05-30verified
Max output tokensdevelopers.openai.com β†—2026-05-30verified
Billing unitdevelopers.openai.com β†—2026-05-30verified
API endpointdevelopers.openai.com β†—2026-05-30verified
Latency tuningdevelopers.openai.com β†—2026-05-30verified
Supported featuresdevelopers.openai.com β†—2026-05-30verified
Audio input formatdevelopers.openai.com β†—2026-05-30verified
Azure WER improvementlearn.microsoft.com β†—2026-05-30single source
OpenAI launches GPT-Realtime-Whisper alongside GPT-Realtime-2 and GPT-Realtime-Translatedevelopers.openai.com β†—2026-05-30verified
This model is documented as WebSocket-only. Is there a wss:// endpoint on TheRouter?therouter.ai β†—2026-07-29to verify
How is GPT-Realtime-Whisper different from gpt-4o-transcribe?developers.openai.com β†—2026-05-30to verify
Can I use GPT-Realtime-Whisper through the standard OpenAI chat completions SDK?developers.openai.com β†—2026-05-30to verify
What does the 'delay' setting do and which one should I use?developers.openai.com β†—2026-05-30to verify
Can I steer the model's transcription with custom vocabulary or prompts?developers.openai.com β†—2026-05-30to verify
How does GPT-Realtime-Whisper compare to traditional Whisper for latency?developers.openai.com β†—2026-05-30to verify
Help & contact