Gemini 2.0 Flash Shutdown & EOL June 2026: What Replaced It and What to Do Now
Gemini 2.0 Flash and Flash-Lite hit end-of-life June 1, 2026 — the shutdown is done. Here is what model IDs were removed, which replacements (gemini-2.5-flash-lite, gemini-3.1-flash-lite) actually work as substitutes, the real cost delta, and the fastest way to audit…
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

If your routing policy uses gemini-2.0-flash or gemini-2.0-flash-lite as a cost-optimized target — either as a primary or fallback — you have a hard deadline: June 1, 2026. Google's official Gemini API pricing page now displays a prominent warning that both models are deprecated and will be shut down in 7 days. Any call to these endpoints after June 1 will fail.
This is not a soft deprecation. The models have been in restricted access since March 6, 2026 (new customers could not add them). June 1 is the full shutdown. Affected teams need to audit model IDs in configuration now, not next sprint.
What is being shut down
The following model IDs become unavailable on June 1, 2026:
gemini-2.0-flash— The general-availability Flash model. Pricing was $0.10/$0.40 per million input/output tokens.gemini-2.0-flash-lite— The lite variant. Pricing was $0.08/$0.30 per million tokens.gemini-2.0-flash-001andgemini-2.0-flash-lite-001— The pinned versions. Already restricted since March 6 to existing customers only.
Note: gemini-3.1-flash-lite-preview was separately shut down today, May 25. If you used that preview model, it is already unavailable.
The migration cost reality
The word "migration" suggests a simple model ID swap. The cost reality is more complicated. Gemini 2.0 Flash was priced at one of the lowest points in the Gemini lineup. Every available replacement costs more:
| Replacement model | Input | Output | vs. Gemini 2.0 Flash |
|---|---|---|---|
gemini-2.5-flash-lite | $0.10/M | $0.40/M | Same price, but prior-gen (not Gemini 3) |
gemini-3.1-flash-lite | $0.25/M | $1.50/M | ~2.5–3.75x more |
gemini-3-flash-preview | $0.50/M | $3.00/M | 5x more |
gemini-3.5-flash | $1.50/M | $9.00/M | 15x more |
gemini-3.1-pro-preview | $2.00/M | $12.00/M | 20–30x more |
The closest price-parity option — gemini-2.5-flash-lite at $0.10/$0.40 — is a prior-gen model that Google will also eventually deprecate. The lowest current-gen (Gemini 3) replacement is gemini-3.1-flash-lite at $0.25/$1.50. For high-volume workloads where Gemini 2.0 Flash was carrying the cost load, that 2.5–3.75x jump is significant.
The cost offset depends heavily on whether you can enable context caching. Google Cloud's Gemini Enterprise Agent Platform now enables implicit caching by default, providing a 90% discount on cached input tokens. For workloads where large system prompts or documents repeat across requests — code review, document QA, classification with shared instructions — context caching can more than compensate for the per-token price increase on the underlying model.
Why it matters for AI engineering teams
The model ID is likely scattered across multiple configs. Gemini 2.0 Flash has been available since early 2025 and has been the go-to low-cost Gemini option for over a year. It likely appears in routing configs, fallback chains, cost-tier assignments, rate-limit bypass rules, and direct SDK calls across multiple services. A single audit of your main config is not enough — you need to search all repositories, environment variables, and infrastructure-as-code files for gemini-2.0-flash and gemini-2.0-flash-lite.
The fallback chain is the hidden risk. Primary model configs are usually well-documented. Fallback chains are not. If your routing layer falls back to gemini-2.0-flash when a primary model is unavailable or overloaded, June 1 turns your fallback into an outage path rather than a reliability mechanism. The deprecation does not just affect primary traffic — it removes a tier from every fallback chain that references it.
The Gemini 3 tier pricing jump changes workload economics. For teams that built production workloads specifically around the $0.10/$0.40 Gemini 2.0 Flash pricing point, migrating to gemini-3.1-flash-lite at $0.25/$1.50 changes the unit economics. A 1 billion token/month workload goes from roughly $300 (input-heavy, 2.0 Flash) to roughly $750 ($0.25/M input). That is a budget line item that needs finance visibility, not just an engineering config change.
The Gemini CLI deprecation compounds the deadline. Google also announced that the existing Gemini CLI is deprecated with a hard shutdown on June 18, 2026, encouraging migration to the new Antigravity CLI. Teams using Gemini CLI for scripting, local development, or CI pipelines face two consecutive deadlines (June 1 for model IDs, June 18 for the CLI) within three weeks.
The router/operator migration playbook
Treat this as a two-phase operation:
Phase 1: Audit (do this today)
- Search all config files, environment variables, SDK initializations, and routing rules for
gemini-2.0-flashandgemini-2.0-flash-lite. Include pinned versions (-001,-exp, any dated snapshots). - Identify which services use these models as primary vs. fallback. Log both.
- For each usage, determine the workload type: latency-sensitive real-time, cost-sensitive batch, code generation, classification, long-context document work, or multimodal.
Phase 2: Replace with the right model, not just any model
The replacement choice is not one-size-fits-all:
- Cost-sensitive, simple tasks, batch workloads: Start with
gemini-2.5-flash-lite(same price, prior-gen) and evaluate before committing togemini-3.1-flash-lite. If quality is sufficient, this is the lowest-friction migration. - Cost-sensitive, needs Gemini 3 quality:
gemini-3.1-flash-lite($0.25/$1.50). Enable context caching if you have repeated large inputs — the 90% cached-token discount will narrow the effective cost gap significantly. - High-volume classification or extraction with shared prompts:
gemini-3.1-flash-litewith implicit caching. Effective cached-input rate is $0.025/M, which makes it cheaper than Gemini 2.0 Flash at scale if your cache hit rate is reasonable. - Code generation or complex reasoning was using 2.0 Flash as a low-cost path: Migrate to
gemini-3-flash-preview($0.50/$3) orgemini-3.5-flash($1.50/$9) depending on quality needs. These are not equivalent replacements for all prompts — test before routing production traffic.
Context caching as a cost lever:
If you switch from Gemini 2.0 Flash to gemini-3.1-flash-lite and your workload has a large system prompt (>10K tokens) that repeats across requests, caching that prompt segment can bring your effective input cost to $0.025/M — far below Gemini 2.0 Flash's $0.10/M for those cached tokens. The 3.1-flash-lite migration may end up cheaper than today if your architecture supports caching.
Fallback chain update:
After replacing primary model IDs, update fallback chains explicitly. Do not leave gemini-2.0-flash or gemini-2.0-flash-lite in any fallback position. A decommissioned model in a fallback chain will fail silently until the fallback is triggered under load — at which point the failure surfaces as an outage, not a configuration warning.
What TheRouter users should watch or try
- Audit your provider configs for Gemini 2.0 Flash references now. If you route to the Gemini API through TheRouter, search your provider and routing rule configs for
gemini-2.0-flash. Replace with an explicit current model ID rather than relying on any automatic aliasing — the June 1 deadline is hard. - Test cost impact before switching fallbacks. The 2.5x–15x cost range across migration options means a blind swap to the "nearest equivalent" could meaningfully change your monthly bill. Test the target model against your actual prompt distribution before switching production routing.
- Enable context caching if you haven't. If your Gemini workloads use large, repeated system prompts or documents, implicit caching (90% cached-token discount) can absorb much of the Gemini 3.x price increase. This is available via the Gemini Enterprise Agent Platform on Google Cloud.
- Set a calendar alert for the Gemini CLI deadline too. June 18 is the Gemini CLI shutdown. If you have any scripts, CI pipelines, or local tooling that uses the old CLI, that is a second deadline requiring attention within 3 weeks.

Gemini Non-Global Endpoint Pricing Goes Live July 1: The 10% Regional Premium Every Routing Team Must Account For
Starting July 1, 2026, Google's Gemini 3 and later models charge a 10% premium on non-global (regional) endpoints. Teams routing to eu-west1, asia-northeast1, or other non-global locations must audit their cost models today.

Gemini 3.5 Flash pricing and routing: what changed at Google I/O 2026
Gemini 3.5 Flash launched at $1.50/$9 per million tokens — up to 6x the cost of earlier Flash models — and the new Interactions API changes how routing teams should handle Gemini Flash fallbacks and session continuity.

Nano Banana 2 Lite Is Your New Default Gemini Image Endpoint — Here's the Routing Decision Framework
Google's Nano Banana 2 Lite (gemini-3.1-flash-lite-image) landed June 30 at $0.034/1K images and 4-second latency. If you are still routing to gemini-2.5-flash-image, you are on a legacy model. Here is the three-tier routing framework every image pipeline team needs.