Gemini Provisioned Throughput Now Queues 7 Orders: What the Multiple Pending Orders GA Means for Routing Teams

Google made multiple pending Provisioned Throughput orders generally available on July 1 — you can now queue up to seven orders per model and region simultaneously, removing the one-at-a-time bottleneck that forced sequential 10-day activation waits.

Published via Google

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Seven routing paths consolidating into a single throughput node, representing Gemini's new ability to queue multiple provisioned capacity orders

On July 1, 2026, Google made multiple pending Provisioned Throughput orders generally available on the Gemini Enterprise Agent Platform. You can now submit up to seven pending orders for the same model and region simultaneously, without waiting for the previous order to activate. For AI engineering teams that route significant traffic through Gemini models on GCP, this removes the planning bottleneck that forced sequential 10-business-day activation windows.

What happened

Google's Gemini Enterprise Agent Platform Provisioned Throughput lets teams reserve dedicated inference capacity rather than competing for shared resources. Before July 1, placing orders was a serial process: you submitted an order, waited for it to activate (typically within 10 business days), then placed the next one. If you needed to ramp capacity for a planned traffic spike, you had to plan months ahead and manually initiate each step.

The July 1 GA release changes this. You can now queue up to seven standard Provisioned Throughput orders for the same Google model and region at the same time. Orders are processed based on commercially reasonable efforts and activate in sequence, but you no longer have to babysit the process: place your ramp orders up front, let the queue work.

The feature is available through the Google Cloud Console at the Gemini Enterprise Agent Platform Provisioned Throughput management page, and is governed by the aiplatform.provisionedThroughputs.create and .update permissions under roles/aiplatform.provisionedThroughputAdmin.

Why it matters for AI engineering teams

Provisioned Throughput (PT) is the primary tool for avoiding 429 resource-exhausted errors when running Gemini models at scale. Teams that have already moved from pay-as-you-go to PT know the activation delay is the main friction point: you can't guarantee capacity tomorrow with today's order. Multiple pending orders change the math.

Capacity ramp without manual follow-up. If you're scaling from 10k to 30k QPM over three months, you can now submit three sequential orders today — 10k, 20k, and 30k GSUs — each set to activate after the previous one. Planned scaling no longer requires calendar reminders and manual re-engagement with GCP procurement.

Overlap ordering for model migrations. PT orders are model-specific within a publisher. You can switch the assignment of an active PT order from Gemini 3.1 Flash to Gemini 3.5 Flash, but not from Google to Anthropic. Multiple pending orders make it practical to stage a model migration with overlapping PT reservations: keep the old model's PT active while the new model's PT warms up in queue.

Predictability for billing teams. PT is a commitment — you can't cancel mid-term. Multiple pending orders give billing and FinOps teams a single-approval workflow: one procurement cycle can authorize the full ramp schedule, rather than requiring re-approval at each activation step.

The router/operator angle

For routing teams, the relevant change is in how you architect your Gemini capacity plan:

Before July 1: Treat Gemini PT as a one-at-a-time commitment. Route traffic to pay-as-you-go if you need headroom, accept 429 risk or over-provision a single large order.

After July 1: Queue a ramp schedule. For teams using Gemini as one target in a multi-provider routing strategy, this means you can lock in GCP capacity ahead of demand and use PT as a reliable tier rather than a fallback.

Overages above your active PT level are automatically billed as pay-as-you-go. This means your routing fallback path from PT → pay-as-you-go is unchanged; multiple pending orders just let you shrink that fallback window as traffic grows.

Key things to check in your routing policy:

  • Confirm active PT quota covers your base traffic: A pending order is not active capacity. Only the current activated order counts toward your throughput guarantee.
  • Single Zone PT requires direct GCP rep engagement: Multiple pending orders are for standard PT only. Single Zone Provisioned Throughput remains a manual sales process.
  • Model changes require your current term to expire or a manual order modification: You can change the model assignment of a PT order but it takes effect on the next renewal or via a change request, not immediately.

What TheRouter users should watch or try

If you're routing traffic to Gemini 3.1 Pro, Gemini 3.5 Flash, or any Google-published model through your API gateway, the PT queue change is relevant for your capacity tier. Teams already using PT can now lock in their capacity ramp without sequential planning cycles. Teams on pay-as-you-go who have been deferring PT due to the activation-window friction now have a lower coordination cost to start the transition.

See the Gemini Enterprise Agent Platform Semantic Governance Policies story for the June 29 operator governance release that shipped alongside the PT updates, and the Gemini Model Armor Agent Gateway GA story for the content security layer that applies to all traffic passing through GCP Agent Gateway.

Help & contact