Back to Models

gpt-realtime-whisper

openaiopenai/gpt-realtime-whisper

How TheRouter serves this differently from the vendor

As the vendor operates it

A streaming realtime channel: a WebSocket session that emits incremental transcript deltas while the speaker is still talking, with tunable delay from 'minimal' to 'xhigh' and turn detection.

On TheRouter

One ordinary HTTP request: POST /v1/audio/transcriptions, multipart, one JSON body back. No incremental deltas, no session to open, no WebSocket or WebRTC client, no delay or turn-detection knobs. Every /v1/realtime path returns 501 realtime_websocket_not_supported, before authentication, for every model.

API guide

Transcription (TheRouter)

On TheRouter, GPT-Realtime-Whisper is served on POST /v1/audio/transcriptions β€” the same REST endpoint as Whisper β€” not on a chat completions or WebSocket path. TheRouter bridges the provider's realtime channel behind this unchanged public contract; request-level behavior is identical to any other transcription model. Supported params: file, language, prompt, response_format, temperature. Billed per minute of audio, not per token and not per request. There is no TheRouter chat/completions route for this model β€” calling it that way returns a 503. For the full protocol-level contract, near-realtime chunking pattern, and Java/Go/Python reference implementations, see the Realtime Transcription guide: https://therouter.ai/guides/multimodal/realtime-transcription

cURL
curl https://api.therouter.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -F file="@speech.wav" \
  -F model="openai/gpt-realtime-whisper" \
  -F language="en" \
  -F prompt="Technical terms: TheRouter, API" \
  -F response_format="json" \
  -F temperature="0"

Near-realtime transcription by client-side chunking

There is no streaming transcription endpoint on TheRouter, so near-realtime output is produced by segmenting audio on the client and issuing one ordinary request per segment. Both snippets below were executed against production on 2026-07-29 and produced the output shown in the guide. Two things they demonstrate are the ones integrations get wrong: each segment must be a standalone decodable container (a raw byte slice of a longer file is rejected with 400), and the WAV header is NOT a fixed 44 bytes β€” real recordings carry LIST/FLLR chunks before data, so the RIFF chunks must actually be walked. Expect word loss or duplication at segment boundaries and 1.7-6.9 s per segment; for accuracy-critical work submit the whole recording in one request.

cURL
# Segment locally, then post each segment to the same endpoint.
# Each piece must decode on its own β€” split with a tool that writes a real
# container per piece (ffmpeg here), never with `split`/`dd` on the bytes.
ffmpeg -i speech.wav -f segment -segment_time 3 -c copy chunk-%03d.wav

for f in chunk-*.wav; do
  curl -sS https://api.therouter.ai/v1/audio/transcriptions \
    -H "Authorization: Bearer $THEROUTER_API_KEY" \
    -F file="@$f" \
    -F model="openai/gpt-realtime-whisper" \
    -F language="en" | python3 -c 'import json,sys; print(json.load(sys.stdin)["text"])'
done
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datedevelopers.openai.com β†—2026-05-30verified
Knowledge cutoffdevelopers.openai.com β†—2026-05-30verified
Context windowdevelopers.openai.com β†—2026-05-30verified
Max output tokensdevelopers.openai.com β†—2026-05-30verified
Billing unitdevelopers.openai.com β†—2026-05-30verified
API endpointdevelopers.openai.com β†—2026-05-30verified
Latency tuningdevelopers.openai.com β†—2026-05-30verified
Supported featuresdevelopers.openai.com β†—2026-05-30verified
Audio input formatdevelopers.openai.com β†—2026-05-30verified
Azure WER improvementlearn.microsoft.com β†—2026-05-30single source
OpenAI launches GPT-Realtime-Whisper alongside GPT-Realtime-2 and GPT-Realtime-Translatedevelopers.openai.com β†—2026-05-30verified
This model is documented as WebSocket-only. Is there a wss:// endpoint on TheRouter?therouter.ai β†—2026-07-29to verify
How is GPT-Realtime-Whisper different from gpt-4o-transcribe?developers.openai.com β†—2026-05-30to verify
Can I use GPT-Realtime-Whisper through the standard OpenAI chat completions SDK?developers.openai.com β†—2026-05-30to verify
What does the 'delay' setting do and which one should I use?developers.openai.com β†—2026-05-30to verify
Can I steer the model's transcription with custom vocabulary or prompts?developers.openai.com β†—2026-05-30to verify
How does GPT-Realtime-Whisper compare to traditional Whisper for latency?developers.openai.com β†—2026-05-30to verify
Help & contact