Grok Voice Agent Builder API Routing: Per-Minute Billing Changes Your Voice Operator Policy
xAI's Voice Agent Builder beta (July 1, 2026) brings Grok Voice into production at $0.05/min. Per-minute billing, 100 concurrent sessions, and telephony-included pricing introduce new operator routing decisions.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

xAI launched the Grok Voice Agent Builder in beta on July 1, 2026. The product is a no-code platform for configuring and deploying production voice agents on Grok Voice — but the operator implication is not the no-code interface. It is the billing model: $0.05 per minute of audio, voices included, telephony included, no separate platform fee. That is the first time a major AI model provider has brought a per-minute voice API into general production, and it requires every AI gateway operator to update how they think about voice workload routing.
What xAI shipped
The Voice Agent Builder announcement describes the product in terms aimed at builders who want to deploy voice agents without writing WebSocket integration code. But buried in the billing section are four constraints that define the operator policy surface:
- Billing unit: $0.05 per minute of audio — time-based, not token-based.
- Concurrency limit: 100 concurrent sessions per team.
- Session cap: 30-minute maximum session length.
- Telephony: Integrated at no additional cost; outbound and inbound phone calls included in the per-minute rate.
These are not UI details. They are API constraints that affect how a multi-provider gateway should handle voice workloads routed to Grok.
Why Grok Voice Agent Builder API routing is different from token routing
Every text or vision routing decision in a gateway today is denominated in tokens: input tokens, output tokens, cached tokens, reasoning tokens. Cost attribution is straightforward — tokens consumed × price per million.
Grok Voice Agent Builder API routing breaks that model in two ways.
First, cost is time-not-count. A 10-minute voice session costs $0.50 regardless of how many words were spoken, how many turns occurred, or how much reasoning the model performed. This means gateway cost attribution must track wall-clock session time per provider, not just token deltas. Teams that aggregate costs across text and voice routes in a single ledger will need a new accounting dimension.
Second, concurrency is capped at 100 sessions per team. Most text routing policies treat concurrency as a rate-limit concern (requests per second, tokens per minute) that degrades gracefully — a slow request still runs. Voice sessions are stateful connections. If 100 sessions are already running, the 101st request is not a delay; it is a hard rejection. Operators running voice workloads across multiple providers need a concurrency-aware router that tracks live session counts per provider and re-routes before rejection rather than after.
The 30-minute session cap and routing timeout policy
The 30-minute maximum session length introduces a deterministic termination event that text routing never has. A gateway routing a multi-turn coding session to Grok 4.3 can let the session run indefinitely. A voice session through Grok Voice will be terminated at exactly 30 minutes.
For operators, this means every Grok Voice session routed through the gateway must carry an expiry-time assertion: if elapsed_time >= 28:00, either complete the current turn gracefully or initiate a handoff to a new session before the provider terminates the connection. Routing policies that do not track session duration will experience unexpected disconnects after deployment.
The 30-minute cap is also a cost-ceiling signal. Each Grok Voice session cannot cost more than $1.50 at the current rate. For use cases like customer support where calls average 5–12 minutes, the per-session cost will be $0.25–$0.60 — low enough that Grok Voice competes directly with dedicated voice providers on price.
Telephony-included pricing changes the provider comparison
Grok Voice includes inbound and outbound telephony at no additional cost. Most comparable voice AI platforms charge separately for telephony infrastructure — a connection fee, a per-minute rate on the carrier side, and a separate per-minute rate on the AI model side. That dual-billing structure means the effective cost of a phone-based voice session is the sum of two rates.
Grok Voice collapses both into $0.05/min. For operators routing voice sessions that involve actual phone calls, this changes the comparison table significantly. The routing decision is no longer "AI model cost + telephony cost per provider" — it is a single dimension.
Operators should update their provider cost models to account for bundled telephony when evaluating Grok Voice against other voice routing options.
What to configure in your routing policy
To integrate Grok Voice into a production routing policy:
-
Add a
voiceworkload class to your gateway routing configuration. Grok Voice should not share rate-limit buckets with text models — the billing units and concurrency semantics are fundamentally different. -
Track live session count against the 100-session limit. Route new voice sessions to fallback providers when the Grok Voice session pool reaches ~85–90% utilization.
-
Tag each voice session with a start timestamp and route expiry. Trigger a graceful session completion or re-route sequence at 28 minutes to avoid provider-initiated disconnects.
-
Update cost attribution to denominate Grok Voice costs in minutes, not tokens. Aggregate them separately before computing per-request or per-day spend across providers.
-
Test telephony integration if your voice workloads include phone calls. Grok Voice's bundled telephony means your application can initiate calls without a separate carrier integration — validate that your routing layer passes the correct call metadata.
TheRouter's async media routing docs and the existing MiMo TTS voice agent routing analysis cover per-session cost attribution patterns that apply directly to Grok Voice session budgeting.
What comes next
The Voice Agent Builder launched as a beta. xAI's history with Grok Build shows that beta product APIs tend to stabilize quickly — the Voice Agent Builder's configuration surface (personality, knowledge, tools) will likely expand. Watch for changes to the 100-session concurrency limit and the 30-minute session cap, both of which are likely to relax as xAI adds capacity.
The $0.05/min rate is marked "currently" in the announcement — a standard qualifier for beta pricing. Operators building production voice routing policies on Grok Voice should build rate-agnostic cost models that can absorb a price change without requiring a routing policy rewrite.

GPT-Live voice API routing: full-duplex voice makes delegation policy the new control point
GPT-Live voice API routing is the next operator decision as OpenAI brings full-duplex voice, background model delegation, and realtime safeguards toward developers.

GPT-Realtime-2.1 and 2.1-Mini: The Voice Routing Decision Every AI Operator Needs to Make Now
OpenAI released GPT-Realtime-2.1 and GPT-Realtime-2.1-mini on July 6. The two-tier structure, configurable reasoning effort, and a new audio pricing baseline change how operators should route voice agent traffic.

xAI Grok 4.3 model routing: the flagship alias now needs policy
xAI Grok 4.3 model routing turns the new flagship alias, 1M context, cached-input pricing, and regional availability into an AI gateway policy decision.