Gemini Non-Global Endpoint Pricing Goes Live July 1: The 10% Regional Premium Every Routing Team Must Account For
Starting July 1, 2026, Google's Gemini 3 and later models charge a 10% premium on non-global (regional) endpoints. Teams routing to eu-west1, asia-northeast1, or other non-global locations must audit their cost models today.

Starting tomorrow, July 1, 2026, teams that route AI traffic to regional Gemini endpoints will pay more — and the cost difference is not a rounding error.
Google Cloud's Agent Platform pricing page carries a footnote that most teams have overlooked: "For non-global endpoints, pricing will go into effect for the Generally available Gemini 3 and later families of models on July 1, 2026. Before July 1, 2026, Global endpoint pricing applies to Non-global endpoints."
That single footnote has a concrete impact: a 10% premium on every token for Gemini 3.5 Flash and Gemini 3.1 Flash-Lite when requests go through regional (non-global) endpoints.
What happened
Google introduced a global/non-global pricing split when it launched the Gemini Enterprise Agent Platform. Until now, the policy included a grace period — non-global endpoints were billed at global rates even when you targeted regional locations like eu-west1, us-east4, or asia-northeast1. That grace period ends at midnight July 1, 2026 UTC.
After the cutover, the pricing for affected models is:
Gemini 3.5 Flash — Standard tier:
- Global: $1.50 input / $9.00 output per million tokens
- Non-global: $1.65 input / $9.90 output per million tokens (+10%)
Gemini 3.1 Flash-Lite — Standard tier:
- Global: $0.25 input / $1.50 output per million tokens
- Non-global: $0.275 input / $1.65 output per million tokens (+10%)
The same 10% premium applies across Priority and Flex/Batch tiers. Gemini 3.1 Pro Preview and Gemini 3 Flash Preview are not affected — those models carry a single price across all endpoints.
Why it matters for AI engineering teams
The 10% difference sounds modest until you account for the traffic volumes that make regional endpoints necessary in the first place. Teams that pin to eu-west1 or asia-northeast1 for data-residency, latency, or compliance reasons typically run high-volume workloads — exactly the use cases where 10% cost delta compounds significantly.
There are three scenarios where this hits hardest:
Data-residency compliance. EU-based teams using regional endpoints to comply with GDPR data-processing requirements cannot simply switch to the global endpoint without legal review. They were absorbing global pricing for free until today.
Latency-pinned routing policies. Teams that hard-coded regional endpoints to reduce round-trip latency may discover that the route now costs 10% more. The latency benefit may still justify the cost — but the tradeoff needs to be reassessed and documented.
Multi-provider routing with undifferentiated fallback. If your routing layer treats gemini-3.5-flash at eu-west1 identically to the global endpoint in terms of cost modeling, your budget projections from last week are wrong as of tomorrow.
The router/operator angle
This pricing change is a classic example of why cost modeling must track endpoint-level metadata, not just model names. A routing policy that assigns a cost estimate to gemini-3.5-flash without distinguishing the endpoint used will undercount actual spend by 10% for all non-global traffic.
Operators should take three steps before July 1:
1. Audit your active endpoint configurations. Review every Gemini API call in your routing layer or proxy. Identify whether the base URL includes a region prefix (e.g., us-east4-aiplatform.googleapis.com or an equivalent regional endpoint). Any request going to a non-global URL is subject to the new pricing.
2. Update cost models and billing alerts. If you use per-model spend alerts, recalibrate the thresholds for any configuration that mixes global and non-global routing. A 10% drift in a high-volume fallback path can blow monthly budget caps before the alert fires.
3. Evaluate whether to stay regional or reroute through global. For workloads where data residency is not a hard requirement, routing through the global endpoint restores the original pricing. This is a legitimate cost optimization, but confirm with your legal/compliance team that global routing is permissible for the data types in question.
One additional consideration: if you operate a multi-region fallback where the primary route is a non-global endpoint and the fallback is a global endpoint (or vice versa), those two paths now have different effective costs. Make sure your router's cost-based fallback logic reflects that asymmetry.
What to watch
Google has not announced whether additional Gemini models will join the non-global pricing tier in future. The current set is Gemini 3.5 Flash and Gemini 3.1 Flash-Lite — both widely used as efficient workhorses in production routing stacks. If you use either model in a regional endpoint configuration, the cost change is effective tomorrow.
For teams using the global endpoint exclusively: no action needed. The footnote explicitly states global endpoint pricing is unchanged.
Routing teams should treat this as a reminder that cloud provider pricing pages deserve the same change-tracking attention as API changelogs. A footnote that activates a 10% cost increase is functionally equivalent to a billing event — and it deserves to be caught in advance, not on your next invoice.

Nano Banana 2 Lite Is Your New Default Gemini Image Endpoint — Here's the Routing Decision Framework
Google's Nano Banana 2 Lite (gemini-3.1-flash-lite-image) landed June 30 at $0.034/1K images and 4-second latency. If you are still routing to gemini-2.5-flash-image, you are on a legacy model. Here is the three-tier routing framework every image pipeline team needs.

Gemini 2.0 Flash Shutdown & EOL June 2026: What Replaced It and What to Do Now
Gemini 2.0 Flash and Flash-Lite hit end-of-life June 1, 2026 — the shutdown is done. Here is what model IDs were removed, which replacements (gemini-2.5-flash-lite, gemini-3.1-flash-lite) actually work as substitutes, the real cost delta, and the fastest way to audit…

Gemini 3.5 Flash pricing and routing: what changed at Google I/O 2026
Gemini 3.5 Flash launched at $1.50/$9 per million tokens — up to 6x the cost of earlier Flash models — and the new Interactions API changes how routing teams should handle Gemini Flash fallbacks and session continuity.