Google Cloud API Gateway model routing: Spec-first routing and its critical host constraint

Google Cloud API Gateway model routing is now in Public Preview: a serverless OpenAPI-spec layer that routes Gemini, Claude, and OpenAI requests — but only if all your backends live on Vertex AI.

TheRouter Newsroomvia Google Developers Blog
Google Cloud API Gateway model routing paths converging through a gateway with diverging model endpoints

Google Cloud API Gateway model routing entered Public Preview on August 4, giving teams a serverless way to expose one OpenAI-compatible endpoint while routing requests to Gemini, Claude, or OpenAI OSS-GPT models behind Vertex AI. The important detail is not just that Google added a router. It is where the router lives: inside an OpenAPI 3.x specification, under Google's x-google-api-management extension, with model selection expressed as configuration rather than gateway code.

For platform teams already standardizing on Vertex AI, this is a clean migration path. For teams that mix Vertex with DashScope, direct Anthropic API access, or other provider hosts, the preview has a hard boundary that should be caught before anyone promises a single universal router.

Google Cloud API Gateway model routing moves routing into OpenAPI

The preview lets operators define model routers directly in an OpenAPI document. A router has a defaultModel and a rules[] array, and each model points to a named backend. In the published example, virtual model names such as claude-opus-4-7 map to Vertex AI backend paths such as an Anthropic publisher endpoint under aiplatform.googleapis.com.

That design changes the handoff between application and platform teams. Before, clients often hardcoded a full Vertex AI path like https://aiplatform.googleapis.com/v1/projects/... for each model. After migration, a team can expose a stable endpoint such as /v1/chat/gemini-claude, let clients send a familiar OpenAI-style body with "model": "claude-opus-4-7", and keep the backend path mapping in the API spec.

Google also keeps client authentication at the Gateway boundary. Clients authenticate to API Gateway, not to each model provider. Rate limiting and token tracking sit at that managed layer too. That is attractive if your team wants fewer proxy processes to patch, scale, and observe.

Google Cloud API Gateway model routing has a shared-host boundary

The constraint is the line operators need to remember. Google says all backends referenced by a single router must share the same host, for example aiplatform.googleapis.com. Routing selects a different model and path on that shared Agent Platform host. It does not route across different hosts.

The practical consequence is direct. You cannot put a Vertex-hosted Gemini backend, Anthropic's direct API, and DashScope in one Google API Gateway router. If a team needs that provider mix, it needs separate gateways, or it needs a code-level proxy in front of Google API Gateway to make the cross-host decision.

The parameter where this shows up is specific enough to audit in code review: x-google-api-management.ai.models.routing.routers.<name>.defaultModel.backend. If that backend points into a router whose other models do not share the same host, the design is outside the preview's current shape.

Spec-first routing versus code-level proxy routing

Google's approach is a strong example of spec-first AI gateway design. The routing policy lives beside the HTTP contract. Platform teams can review it as configuration, apply normal API Gateway controls, and avoid running another open-source proxy tier.

That is not the same tradeoff as a code-level router such as LiteLLM or a custom TheRouter-style gateway. Code-level routing can make decisions across provider hosts, combine Vertex and direct APIs, add custom fallback logic, and normalize unusual response behavior. It also means your team owns the service, deployment, rate-limit behavior, and failure modes.

Azure API Management's Unified Model API is closer to Google's spec-first camp, but its cross-provider story is broader. Google's preview is more opinionated: you get a managed, serverless router if your backend universe already sits on Vertex AI.

That opinion can be useful. It reduces operational surface area. It also makes the architecture less portable. The right choice depends on whether your routing problem is mainly model selection inside Vertex, or provider selection across clouds and vendors.

What TheRouter operators should change now

If your team is already moving model calls through Vertex AI, this preview is worth testing behind a non-critical route. Start with one low-risk chat endpoint, map two Vertex-hosted models, and check how token tracking, latency, and error reporting behave compared with your existing gateway. The broad TheRouter documentation is a useful checklist for thinking about provider identity, model names, and client compatibility before changing production traffic.

If your architecture spans multiple provider hosts, treat Google's shared-host constraint as a design input rather than a footnote. Keep cross-provider policy in a router layer that can actually see every provider. Use API Gateway as the managed edge for Vertex-hosted traffic, or as one hop inside a larger routing system.

The closest migration pattern is simple:

  • Move clients from hardcoded Vertex model URLs to one stable Gateway path.
  • Keep virtual model names stable in client payloads.
  • Put Vertex backend paths in the OpenAPI spec, not in application code.
  • Keep cross-host routing outside this preview until Google expands the host model.

That last point is where this announcement matters most. Google Cloud API Gateway model routing is a useful managed primitive, but it is not a universal multi-provider router. For teams that recently followed Google’s broader Vertex AI Agent Platform migration, it fits the same direction: move more agent and model traffic into managed Google infrastructure, then decide deliberately where open routing still belongs.

Help & contact