Back to Models

GPT Audio 1.5 is OpenAI's generally-available audio-native model β€” the first GPT Audio model to shed the '-preview' suffix and serve as the recommended production replacement for the earlier gpt-4o-audio-preview and gpt-4o-mini-audio-preview variants. Released in February 2026 alongside the realtime sibling gpt-realtime-1.5, it is accessible through the standard Chat Completions REST API at api.therouter.ai/v1 with the same modalities and audio parameter pattern. The model accepts audio input and produces both a text transcript and a base64-encoded audio output in a single round-trip, eliminating the need for separate Whisper transcription and TTS pipelines.

From a TheRouter routing perspective, gpt-audio-1.5 is a significant milestone: it is the first fully GA OpenAI audio model, meaning production stability guarantees, no surprise deprecation of the same year, and dramatically lower pricing compared to the gpt-4o-audio-preview preview era ($40/80 vs $100/200 per million audio tokens β€” a 60% reduction). The model also improves upon gpt-audio (the first GA snapshot from January 2026) with enhanced instruction following, better multilingual voice quality, and improved tool-calling reliability. For TheRouter users, it replaces both gpt-4o-audio-preview (deprecated May 2026) and gpt-audio (original GA snapshot) as the default audio-in/audio-out Chat Completions model.

Best for
  • β€’ Production voice-in/voice-out applications β€” voice notes, audio messaging, podcast-style content generation, and interactive voice response (IVR) systems, now with GA stability guarantees
  • β€’ Audio summarisation with improved instruction following β€” long audio recordings transcribed, summarised, and structured (speaker labels, action items) in one API call
  • β€’ Multilingual voice applications β€” improved non-English accuracy makes it suitable for Japanese, Indic, and European language voice workflows
  • β€’ Tool-calling from voice β€” customers speaking complaints or requests can trigger structured tools (ticket creation, escalation, knowledge base search) in a single request with improved reliability over the preview model
Reach for something else if
  • β€’ Real-time conversational AI requiring <200 ms audio-to-audio latency β€” use gpt-realtime-1.5 on TheRouter (Realtime API over WebSocket) for persistent low-latency connections
  • β€’ Pure text-only workloads β€” standard GPT-4o or GPT-4o Mini text models are cheaper ($4/$16 vs $2.50/$10 per million text tokens for GPT-4o Mini) without audio modality overhead
  • β€’ High-volume streaming transcription β€” use gpt-4o-transcribe or gpt-4o-mini-transcribe on TheRouter for specialised speech-to-text workloads at a fraction of the cost

How TheRouter serves this differently from the vendor

As the vendor operates it

OpenAI serves GPT Audio 1.5 through the Chat Completions API for audio-in/audio-out requests with the same modalities and audio parameters documented for the vendor endpoint.

On TheRouter

TheRouter exposes the same Chat Completions contract at https://api.therouter.ai/v1: change only the base URL, API key, and model id to openai/gpt-audio-1.5. No realtime WebSocket behavior is implied by this page.

Context Length
--
Max Output
--
Input Priceper 1M tokens
$4.32/ 1M tokens
Output Priceper 1M tokens
$17.28/ 1M tokens

Modalities

textaudioimage→textaudio

Capabilities

Vision

Pricing Breakdown

TypeRate
Input$4.32 / 1M tokens
Output$17.28 / 1M tokens
Audio input$34.56 / 1M tokens
Audio output$69.12 / 1M tokens
Image input$5.40 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatmodalitiesaudio

Specifications

Release date2026-02-23 (alongside gpt-realtime-1.5)learn.microsoft.com β†—verified
StatusGeneral availability (GA) β€” replaces gpt-4o-audio-preview and gpt-4o-mini-audio-preview; those models deprecated 2026-05-07developers.openai.com β†—verified
Training cutoffNot publicly disclosedunknown
Output modalityAudio + text (dual output via modalities parameter); audio in wav, mp3, flac, opus, pcm16, aaclearn.microsoft.com β†—verified
Supported voicesAlloy, Ash, Ballad, Coral, Echo, Sage, Shimmer, Verse, Marin, Cedar (10 voices)learn.microsoft.com β†—verified
Max audio file size20 MB per input audio filelearn.microsoft.com β†—verified
Tool / function callingSupported β€” improved reliability over gpt-4o-audio-preview per Feb 2026 releasetechcommunity.microsoft.com β†—verified
Structured output (JSON)Not publicly disclosed for this audio modelunknown
LicenseOpenAI Terms of Service (proprietary, API-only)verified

Benchmarks

BenchmarkDistributionScoreSource
Big Bench Audio
OpenAI and Azure materials describe GPT Audio 1.5 as improving instruction following, multilingual audio quality, and tool calling, but they do not publish a standalone Big Bench Audio score for this Chat Completions model.
β€”Not publicly disclosedβ€”
Multilingual WER reduction
Azure Foundry's launch post says GPT Audio 1.5 improves multilingual audio quality, including Japanese and Indic-language use cases, but does not publish exact WER or accuracy figures.
β€”Not publicly disclosedβ€”
Tool calling accuracy
Azure Foundry's launch post names better instruction following and more reliable tool calling as GPT Audio 1.5 improvements, but does not disclose a pass rate or benchmark score.
β€”Not publicly disclosedβ€”

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "openai/gpt-audio-1.5",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Audio-in / audio-out chat

Audio chat via TheRouter's OpenAI-compatible endpoint β€” include the modalities and audio parameters to get a spoken response back.

cURL
# Encode audio to base64 first:
# AUDIO_B64=$(base64 -i input.mp3)

curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-audio-1.5",
    "modalities": ["text", "audio"],
    "audio": {"voice": "alloy", "format": "wav"},
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "What is in this audio recording?"},
          {"type": "input_audio", "input_audio": {"data": "'"$AUDIO_B64"'", "format": "mp3"}}
        ]
      }
    ]
  }'

More from openai

Similar models

Cross-provider sibling models

News & changes

2026-02-23

OpenAI releases gpt-audio-1.5 and gpt-realtime-1.5

OpenAI released gpt-audio-1.5 to the Chat Completions API and gpt-realtime-1.5 to the Realtime API on February 23, 2026. Both models deliver improved instruction following, multilingual performance, and tool-calling reliability. gpt-audio-1.5 becomes the first fully GA GPT Audio model, replacing the gpt-4o-audio-preview variants with significantly lower pricing ($40/$80 vs $100/$200 per million audio tokens). The gpt-4o-audio-preview and gpt-4o-mini-audio-preview models were deprecated effective May 7, 2026.

re-authored by TheRouterlearn.microsoft.com β†—

Frequently asked

How is gpt-audio-1.5 different from gpt-4o-audio-preview?

gpt-audio-1.5 is the generally-available successor to gpt-4o-audio-preview. Key differences: (1) GA status with production stability guarantees instead of 'preview'; (2) significantly lower audio token pricing ($40/$80 vs $100/$200 per million tokens); (3) improved instruction following, multilingual accuracy, and tool-calling reliability. The API parameters are identical β€” you only need to change the model name. The gpt-4o-audio-preview was deprecated on May 7, 2026.

re-authored by TheRouterdevelopers.openai.com β†—
Can I use gpt-audio-1.5 on TheRouter today?

Yes. gpt-audio-1.5 is fully available on TheRouter through the standard OpenAI-compatible /v1/chat/completions endpoint at api.therouter.ai/v1. Just set the model name to 'openai/gpt-audio-1.5' in your existing code. The API accepts the same modalities and audio parameters as gpt-4o-audio-preview β€” no migration changes required beyond the model ID.

What audio formats does gpt-audio-1.5 support?

GPT Audio 1.5 supports input audio in wav, mp3, flac, opus, pcm16, and aac formats. Audio input is provided as base64-encoded data in the input_audio content part. The maximum input audio file size is 20 MB. Output audio is returned as base64-encoded data in the specified format (default wav). The model outputs both a text transcript and the audio payload when modalities includes both 'text' and 'audio'.

re-authored by TheRouterlearn.microsoft.com β†—
How do I migrate from gpt-4o-audio-preview to gpt-audio-1.5?

Migration is straightforward: change the model name from 'openai/gpt-4o-audio-preview' to 'openai/gpt-audio-1.5' in your API requests. No other parameter changes are needed β€” the modalities, audio, and tool-calling APIs are identical. The new model is cheaper ($40/$80 vs $100/$200 per million audio tokens for text+audio, and $4/$16 vs $2.50/$10 for text-only tokens). Gpt-4o-audio-preview was deprecated on May 7, 2026, so migration is strongly recommended.

re-authored by TheRouterdevelopers.openai.com β†—
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datelearn.microsoft.com β†—2026-08-18verified
Statusdevelopers.openai.com β†—2026-08-18verified
Training cutoffβ€”β€”unknown
Output modalitylearn.microsoft.com β†—2026-08-18verified
Supported voiceslearn.microsoft.com β†—2026-08-18verified
Max audio file sizelearn.microsoft.com β†—2026-08-18verified
Tool / function callingtechcommunity.microsoft.com β†—2026-08-18verified
Structured output (JSON)β€”β€”unknown
Licenseβ€”β€”verified
Big Bench Audiotechcommunity.microsoft.com β†—2026-08-18unknown
Multilingual WER reductiontechcommunity.microsoft.com β†—2026-08-18unknown
Tool calling accuracytechcommunity.microsoft.com β†—2026-08-18unknown
OpenAI releases gpt-audio-1.5 and gpt-realtime-1.5learn.microsoft.com β†—2026-08-18verified
How is gpt-audio-1.5 different from gpt-4o-audio-preview?developers.openai.com β†—2026-08-18to verify
What audio formats does gpt-audio-1.5 support?learn.microsoft.com β†—2026-08-18to verify
How do I migrate from gpt-4o-audio-preview to gpt-audio-1.5?developers.openai.com β†—2026-08-18to verify
Help & contact