Gemini Non-Global Endpoint Pricing Goes Live July 1: The 10% Regional Premium Every Routing Team Must Account For

Starting July 1, 2026, Google's Gemini 3 and later models charge a 10% premium on non-global (regional) endpoints. Teams routing to eu-west1, asia-northeast1, or other non-global locations must audit their cost models today.

TheRouter Newsroomvia Google Cloud
Editorial illustration of a globe split between regional and global cloud routing paths with cost indicators

Starting tomorrow, July 1, 2026, teams that route AI traffic to regional Gemini endpoints will pay more — and the cost difference is not a rounding error.

Google Cloud's Agent Platform pricing page carries a footnote that most teams have overlooked: "For non-global endpoints, pricing will go into effect for the Generally available Gemini 3 and later families of models on July 1, 2026. Before July 1, 2026, Global endpoint pricing applies to Non-global endpoints."

That single footnote has a concrete impact: a 10% premium on every token for Gemini 3.5 Flash and Gemini 3.1 Flash-Lite when requests go through regional (non-global) endpoints.

What happened

Google introduced a global/non-global pricing split when it launched the Gemini Enterprise Agent Platform. Until now, the policy included a grace period — non-global endpoints were billed at global rates even when you targeted regional locations like eu-west1, us-east4, or asia-northeast1. That grace period ends at midnight July 1, 2026 UTC.

After the cutover, the pricing for affected models is:

Gemini 3.5 Flash — Standard tier:

  • Global: $1.50 input / $9.00 output per million tokens
  • Non-global: $1.65 input / $9.90 output per million tokens (+10%)

Gemini 3.1 Flash-Lite — Standard tier:

  • Global: $0.25 input / $1.50 output per million tokens
  • Non-global: $0.275 input / $1.65 output per million tokens (+10%)

The same 10% premium applies across Priority and Flex/Batch tiers. Gemini 3.1 Pro Preview and Gemini 3 Flash Preview are not affected — those models carry a single price across all endpoints.

Why it matters for AI engineering teams

The 10% difference sounds modest until you account for the traffic volumes that make regional endpoints necessary in the first place. Teams that pin to eu-west1 or asia-northeast1 for data-residency, latency, or compliance reasons typically run high-volume workloads — exactly the use cases where 10% cost delta compounds significantly.

There are three scenarios where this hits hardest:

Data-residency compliance. EU-based teams using regional endpoints to comply with GDPR data-processing requirements cannot simply switch to the global endpoint without legal review. They were absorbing global pricing for free until today.

Latency-pinned routing policies. Teams that hard-coded regional endpoints to reduce round-trip latency may discover that the route now costs 10% more. The latency benefit may still justify the cost — but the tradeoff needs to be reassessed and documented.

Multi-provider routing with undifferentiated fallback. If your routing layer treats gemini-3.5-flash at eu-west1 identically to the global endpoint in terms of cost modeling, your budget projections from last week are wrong as of tomorrow.

The router/operator angle

This pricing change is a classic example of why cost modeling must track endpoint-level metadata, not just model names. A routing policy that assigns a cost estimate to gemini-3.5-flash without distinguishing the endpoint used will undercount actual spend by 10% for all non-global traffic.

Operators should take three steps before July 1:

1. Audit your active endpoint configurations. Review every Gemini API call in your routing layer or proxy. Identify whether the base URL includes a region prefix (e.g., us-east4-aiplatform.googleapis.com or an equivalent regional endpoint). Any request going to a non-global URL is subject to the new pricing.

2. Update cost models and billing alerts. If you use per-model spend alerts, recalibrate the thresholds for any configuration that mixes global and non-global routing. A 10% drift in a high-volume fallback path can blow monthly budget caps before the alert fires.

3. Evaluate whether to stay regional or reroute through global. For workloads where data residency is not a hard requirement, routing through the global endpoint restores the original pricing. This is a legitimate cost optimization, but confirm with your legal/compliance team that global routing is permissible for the data types in question.

One additional consideration: if you operate a multi-region fallback where the primary route is a non-global endpoint and the fallback is a global endpoint (or vice versa), those two paths now have different effective costs. Make sure your router's cost-based fallback logic reflects that asymmetry.

What to watch

Google has not announced whether additional Gemini models will join the non-global pricing tier in future. The current set is Gemini 3.5 Flash and Gemini 3.1 Flash-Lite — both widely used as efficient workhorses in production routing stacks. If you use either model in a regional endpoint configuration, the cost change is effective tomorrow.

For teams using the global endpoint exclusively: no action needed. The footnote explicitly states global endpoint pricing is unchanged.

Routing teams should treat this as a reminder that cloud provider pricing pages deserve the same change-tracking attention as API changelogs. A footnote that activates a 10% cost increase is functionally equivalent to a billing event — and it deserves to be caught in advance, not on your next invoice.

Help & contact