Gemini 3.5 Flash pricing and routing: what changed at Google I/O 2026

Gemini 3.5 Flash launched at $1.50/$9 per million tokens — up to 6x the cost of earlier Flash models — and the new Interactions API changes how routing teams should handle Gemini Flash fallbacks and session continuity.

Published via Google AI Blog

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract editorial illustration representing multi-tier routing decisions and cost-aware AI provider policy

For teams that have built routing policies around "Gemini Flash = cheap fallback," Google I/O 2026 just invalidated the assumption.

On May 19, Google launched Gemini 3.5 Flash — its newest flagship fast model — directly into general availability with a model ID of gemini-3.5-flash. No preview period. But the headline detail for routing teams is price: $1.50 per million input tokens and $9.00 per million output tokens. That is 3× the cost of the previous Gemini 3 Flash Preview ($0.50/$3) and 6× the cost of Gemini 3.1 Flash-Lite ($0.25/$1.50).

Google also introduced the Interactions API in public beta — a new server-side session management layer that mirrors the stateful patterns of OpenAI's Responses API. Whether your current proxy can route these calls transparently matters more than the benchmark numbers.

What happened

Google I/O 2026 (May 19–20) was the largest Gemini model cycle in terms of developer-facing changes. The centerpiece release for API teams:

  • Gemini 3.5 Flash (gemini-3.5-flash): Generally available from day one. 1,048,576 token context window, 65,536 max output tokens. Supports thinking mode. Google is using it internally across Gemini Search, Antigravity agent platform, Google AI Studio, and Gemini Enterprise.
  • Gemini Interactions API (public beta): A new API surface with built-in support for multi-step tool use, server-side conversation history, and orchestration primitives. It is explicitly compared to OpenAI Responses in design.
  • Gemini 3.5 Pro: Promised "next month" (June 2026). No pricing announced yet, but the pattern suggests it will land above Gemini 3.1 Pro's current $2/$12.
  • Context caching and batch discounts: Available on 3.5 Flash. 50% batch discount confirmed. Context caching rates not yet fully published at time of writing.

For comparison, Google's current Flash-tier lineup now spans a 6× price range:

ModelInput ($/1M)Output ($/1M)
Gemini 3.1 Flash-Lite$0.25$1.50
Gemini 3 Flash Preview$0.50$3.00
Gemini 3.5 Flash$1.50$9.00
Gemini 3.1 Pro Preview$2.00$12.00

Gemini 3.5 Flash now sits closer to the Pro tier in cost than to its Flash predecessors.

Why it matters for AI engineering teams

The naming convention "Flash" has trained routing policies at many teams to treat Gemini Flash models as the cheap-but-capable tier and Pro as the high-cost reasoning tier. That mental model breaks with 3.5 Flash.

Cost modeling needs a refresh. Teams routing high-volume, moderate-complexity tasks to any model named gemini-*-flash need to audit which model ID is actually in their config. A hard-coded gemini-flash or gemini-3-flash reference now routes to a fundamentally different price point if aliased, and a direct gemini-3.5-flash routing was not an option two weeks ago.

The Interactions API adds a proxy-passthrough risk. The new beta API surface manages session history server-side (like OpenAI Responses). If your routing layer passes raw request bodies through to the Gemini endpoint without understanding the Interactions API schema, you may silently fail or lose session continuity. Standard /v1/chat/completions calls still work — the Interactions API is additive, not a replacement — but any team planning to adopt it needs to verify that their router supports the new endpoint patterns.

Gemini 3.5 Pro is coming within weeks. A new top-tier Gemini model above Gemini 3.1 Pro means a third price/capability anchor to calibrate against within the Google provider. Teams building multi-model routing strategies in June 2026 will be working with an incomplete picture unless they plan for 3.5 Pro to arrive.

Context caching changes effective cost. For long-context use cases (agent memory, large codebases, long documents), context caching and batch discounts on 3.5 Flash can materially reduce the $9/M output price. The effective cost for a team routing code review tasks with a large codebase prefix may be significantly lower than the headline rate — but only if the router or application explicitly enables caching.

The router/operator angle

The practical change for routing teams is that the Google Gemini tier ladder now requires a three-level policy rather than two:

  1. Ultra-cheap tier: gemini-3.1-flash-lite for high-volume classification, simple summarization, routing pre-checks.
  2. Mid tier (the old "Flash"): gemini-3-flash-preview remains as a cost-effective middle option.
  3. Premium Flash: gemini-3.5-flash — the new default across Google products. Reserve for coding, agentic tasks, long-horizon reasoning, and multimodal work where quality gap over 3-Flash-Preview justifies 3× cost.

The implication for fallback chains: if you have a primary provider failure and fall back to Gemini Flash, your fallback cost has likely just tripled compared to a policy written in early 2026. A cost-aware fallback configuration needs to distinguish between gemini-3-flash-preview and gemini-3.5-flash explicitly.

The Interactions API also introduces a proxy architecture question. The new API uses server-side history management, which means session continuity is bound to the Google endpoint, not the client. A routing layer that redirects requests mid-session to a different provider (e.g., as a failover) will break Interactions API sessions. Teams using stateless /generateContent or /chat/completions calls are not affected — but this is worth tracking as the Interactions API moves toward general availability.

Cost attribution: Running Gemini 3.5 Flash at scale with thinking enabled produces significantly higher token counts per request than non-thinking Flash models. The benchmark data from Artificial Analysis shows running their standard suite against 3.5 Flash (high/thinking mode) cost $1,551 — more expensive than Gemini 3.1 Pro Preview at $892. Usage-based billing attribution in dashboards needs to account for reasoning token inflation.

What TheRouter users should watch or try

  • Audit your Gemini model IDs now. If any routing rule or model config references gemini-flash without a version lock, verify the model alias resolution before Google potentially changes the default. Explicit model IDs (gemini-3-flash-preview, gemini-3.5-flash) are safer for cost predictability.
  • Update fallback cost thresholds. If Gemini Flash was configured as a low-cost fallback destination, the new pricing means your cost ceiling on fallback has changed. Review routing policies that use Gemini Flash as the "cheap secondary" option.
  • Watch the Interactions API GA timeline. The new beta API is not required for existing /chat/completions-based integrations, but if you are building long-running agent workflows on Gemini, plan for the migration path now rather than after GA.
  • Evaluate Gemini 3.1 Flash-Lite for simple routing tasks. At $0.25/$1.50, Flash-Lite remains a compelling option for pre-routing classification, intent detection, and cheap summarization steps in agent pipelines. Adding it as an explicit routing target for simple workloads preserves the cost profile that "Flash" used to imply.
  • Prepare for Gemini 3.5 Pro. Within weeks, a new Google flagship Pro model will land. Keep one routing slot open and plan an evaluation run when it releases.

The core lesson from Google I/O 2026 is not that Gemini got smarter (it did) — it's that Google has repriced its inference economy, and routing policies written before May 2026 need to be updated with that in mind.

Models covered in this article

Help & contact