xAI Grok 4.3 model routing: the flagship alias now needs policy

xAI Grok 4.3 model routing turns the new flagship alias, 1M context, cached-input pricing, and regional availability into an AI gateway policy decision.

TheRouter Newsroomvia xAI Docs
xAI Grok 4.3 model routing diagram showing alias, region, context, and cost policy lanes

xAI Grok 4.3 model routing is now a live policy question for AI gateway teams. xAI's official model page lists grok-4.3 as its most advanced flagship model, with text and image input, a 1,000,000-token context window, function calling, structured outputs, reasoning, cached-input pricing, and aliases that can move traffic to the latest Grok release. That is enough surface area that operators should not treat the upgrade as a simple model-name swap.

What changed in xAI Grok 4.3 model routing

The official Grok 4.3 model page describes Grok 4.3 as xAI's flagship model for non-hallucination rate, agentic tool calling, and instruction following. At a glance, it supports text and image input to text output, exposes the grok-4.3 model name, and maps grok-4.3-latest plus grok-latest as aliases.

The pricing page now makes the routing math more explicit. xAI lists Grok 4.3 at $1.25 per 1M input tokens, $0.20 per 1M cached input tokens, and $2.50 per 1M output tokens. The same model page lists 37 requests per second and 10,000,000 tokens per minute, with availability in us-east-1, eu-west-1, and us-west-2.

The broader xAI models page also says Grok 4.3 is the default choice for chat and coding, while dedicated APIs handle images, videos, and voice. It adds an important operational caveat: Grok has no access to realtime events unless server-side Web Search or X Search tools are enabled. In practice, Grok 4.3 is both the default high-end route and a route that still needs tool policy.

Why xAI Grok 4.3 model routing matters for operators

Alias behavior is the first routing risk. grok-4.3 is a named model, while grok-4.3-latest and grok-latest are moving targets. That is convenient for fast adoption, but it is not the same as deterministic production routing. A gateway should let teams choose whether a workload wants a pinned model, a family-latest model, or the provider-wide latest alias.

The 1M context window also changes budget policy. Large context is useful for long codebases, research bundles, and support transcripts, but it can hide waste when an application sends the same large prefix repeatedly. The cached-input price gives teams a reason to track cacheable prefixes separately from fresh input. If a router only records total input tokens, it cannot explain why one Grok lane is cheap and another lane is expensive.

Regional availability matters as well. The model page lists three regions, so production routing should not collapse Grok into a single global endpoint. A latency-sensitive agent in Europe, a U.S. batch job, and a compliance-bound workspace may all want different region rules. The provider profile should store region, alias policy, cache policy, and tool policy as separate fields.

The AI gateway policy for xAI Grok 4.3 model routing

A clean xAI Grok 4.3 model routing policy starts with three route names: grok-4.3-pinned, grok-4.3-family-latest, and grok-latest-experimental. The pinned route is for production workloads that need repeatability. The family-latest route is for teams that accept model refreshes within the Grok 4.3 family. The provider-latest route should be opt-in and monitored like a canary.

Cost attribution should separate input, cached input, and output. That matters because a 1M-context model can look expensive until cached prefixes are measured, or look cheap until output-heavy agent loops are counted. Teams should alert on cache-hit-rate drops, not just on total token spend.

Tool policy should be explicit. If the task requires current information, route with Web Search or X Search enabled and label it as a grounded lane. If the task is code review, summarization, or static analysis, keep realtime tools off and avoid paying latency for a capability the workload does not need. The TheRouter model catalog can help teams compare provider options, while the broader AI gateway documentation is the right place to model policy, fallback, and evidence.

Fallback should preserve semantics. Do not silently fail over from a grounded Grok route to an ungrounded model, or from a pinned route to a latest alias, unless the calling application opted in. The safer pattern is to return a controlled provider error when the requested route cannot preserve context size, tool availability, region, and alias policy.

What TheRouter users should watch

Start by auditing every xAI route name in application code and gateway configuration. Replace generic grok-latest usage in production with a named lane that says whether it is pinned, family-latest, or experimental. Then attach region and realtime-tool requirements to each lane.

Next, test the cache economics. Run the same long-prefix workload through Grok 4.3 with stable prefixes, measure cached-input share, and record latency plus total cost per successful task. If cache-hit rate is low, the problem may be prompt construction rather than model price.

Finally, add alias drift checks to release review. When a provider exposes convenient latest aliases, the router should record which underlying model served traffic and when the alias changed. xAI Grok 4.3 model routing is a useful reminder that a flagship model launch is not just a benchmark event. It is a new contract among aliasing, regions, tools, cache behavior, and fallback.

Help & contact