OpenAI launches GPT-Realtime-Whisper alongside GPT-Realtime-2 and GPT-Realtime-Translate
On May 7, 2026, OpenAI released three new realtime audio models: GPT-Realtime-2 (128K-context voice agent), GPT-Realtime-Translate (streaming translation at $0.034/min), and GPT-Realtime-Whisper (streaming speech-to-text at $0.017/min). GPT-Realtime-Whisper is designed for live transcription with tunable latency β the 'minimal' setting provides near-instant transcript deltas, while 'xhigh' favours accuracy over speed. The model connects via a dedicated transcription sessions endpoint with WebRTC (browsers) or WebSocket (server media pipelines). It is also available through Azure Foundry with approximately 50% lower WER than previous realtime transcription models.