Google Cloud API Gateway model routing: Spec-first routing and its critical host constraint
Google Cloud API Gateway model routing is now in Public Preview: a serverless OpenAPI-spec layer that routes Gemini, Claude, and OpenAI requests — but only if all your backends live on Vertex AI.

Google Cloud API Gateway model routing entered Public Preview on August 4, giving teams a serverless way to expose one OpenAI-compatible endpoint while routing requests to Gemini, Claude, or OpenAI OSS-GPT models behind Vertex AI. The important detail is not just that Google added a router. It is where the router lives: inside an OpenAPI 3.x specification, under Google's x-google-api-management extension, with model selection expressed as configuration rather than gateway code.
For platform teams already standardizing on Vertex AI, this is a clean migration path. For teams that mix Vertex with DashScope, direct Anthropic API access, or other provider hosts, the preview has a hard boundary that should be caught before anyone promises a single universal router.
Google Cloud API Gateway model routing moves routing into OpenAPI
The preview lets operators define model routers directly in an OpenAPI document. A router has a defaultModel and a rules[] array, and each model points to a named backend. In the published example, virtual model names such as claude-opus-4-7 map to Vertex AI backend paths such as an Anthropic publisher endpoint under aiplatform.googleapis.com.
That design changes the handoff between application and platform teams. Before, clients often hardcoded a full Vertex AI path like https://aiplatform.googleapis.com/v1/projects/... for each model. After migration, a team can expose a stable endpoint such as /v1/chat/gemini-claude, let clients send a familiar OpenAI-style body with "model": "claude-opus-4-7", and keep the backend path mapping in the API spec.
Google also keeps client authentication at the Gateway boundary. Clients authenticate to API Gateway, not to each model provider. Rate limiting and token tracking sit at that managed layer too. That is attractive if your team wants fewer proxy processes to patch, scale, and observe.
Google Cloud API Gateway model routing has a shared-host boundary
The constraint is the line operators need to remember. Google says all backends referenced by a single router must share the same host, for example aiplatform.googleapis.com. Routing selects a different model and path on that shared Agent Platform host. It does not route across different hosts.
The practical consequence is direct. You cannot put a Vertex-hosted Gemini backend, Anthropic's direct API, and DashScope in one Google API Gateway router. If a team needs that provider mix, it needs separate gateways, or it needs a code-level proxy in front of Google API Gateway to make the cross-host decision.
The parameter where this shows up is specific enough to audit in code review: x-google-api-management.ai.models.routing.routers.<name>.defaultModel.backend. If that backend points into a router whose other models do not share the same host, the design is outside the preview's current shape.
Spec-first routing versus code-level proxy routing
Google's approach is a strong example of spec-first AI gateway design. The routing policy lives beside the HTTP contract. Platform teams can review it as configuration, apply normal API Gateway controls, and avoid running another open-source proxy tier.
That is not the same tradeoff as a code-level router such as LiteLLM or a custom TheRouter-style gateway. Code-level routing can make decisions across provider hosts, combine Vertex and direct APIs, add custom fallback logic, and normalize unusual response behavior. It also means your team owns the service, deployment, rate-limit behavior, and failure modes.
Azure API Management's Unified Model API is closer to Google's spec-first camp, but its cross-provider story is broader. Google's preview is more opinionated: you get a managed, serverless router if your backend universe already sits on Vertex AI.
That opinion can be useful. It reduces operational surface area. It also makes the architecture less portable. The right choice depends on whether your routing problem is mainly model selection inside Vertex, or provider selection across clouds and vendors.
What TheRouter operators should change now
If your team is already moving model calls through Vertex AI, this preview is worth testing behind a non-critical route. Start with one low-risk chat endpoint, map two Vertex-hosted models, and check how token tracking, latency, and error reporting behave compared with your existing gateway. The broad TheRouter documentation is a useful checklist for thinking about provider identity, model names, and client compatibility before changing production traffic.
If your architecture spans multiple provider hosts, treat Google's shared-host constraint as a design input rather than a footnote. Keep cross-provider policy in a router layer that can actually see every provider. Use API Gateway as the managed edge for Vertex-hosted traffic, or as one hop inside a larger routing system.
The closest migration pattern is simple:
- Move clients from hardcoded Vertex model URLs to one stable Gateway path.
- Keep virtual model names stable in client payloads.
- Put Vertex backend paths in the OpenAPI spec, not in application code.
- Keep cross-host routing outside this preview until Google expands the host model.
That last point is where this announcement matters most. Google Cloud API Gateway model routing is a useful managed primitive, but it is not a universal multi-provider router. For teams that recently followed Google’s broader Vertex AI Agent Platform migration, it fits the same direction: move more agent and model traffic into managed Google infrastructure, then decide deliberately where open routing still belongs.

Vertex AI Extensions Shut Down November 26: Three Migration Paths for Agent Platform
Google deprecated Vertex AI Extensions on May 26, 2026, with a hard shutdown on November 26. Teams routing agent workflows through Google must migrate their Code Interpreter, Search, and Custom Extensions to the Gemini Enterprise Agent Platform before the deadline.

Is Vertex AI Deprecated? Imagen and Veo Endpoint Shutdown List for June 30, 2026
Vertex AI is not deprecated, but Google is retiring every Imagen 2.x/3.x/4.x and Veo 2.x/3.0 GA endpoint on June 30, 2026. Use this shutdown list to migrate image and video routing to gemini-2.5-flash-image and veo-3.1 endpoints.

DeepSeek V4 Pro Survives Its Own Deadline: What the Reversal Means for Your Routing Policy
DeepSeek announced on September 10 that V4 Pro would be retired today at 04:00 UTC. Instead, they reversed course in response to user demand, keeping V4 Pro live at unchanged pricing. Here is what the two-model landscape now looks like and which workloads belong on which.