Nano Banana 2 Lite Is Your New Default Gemini Image Endpoint — Here's the Routing Decision Framework

Google's Nano Banana 2 Lite (gemini-3.1-flash-lite-image) landed June 30 at $0.034/1K images and 4-second latency. If you are still routing to gemini-2.5-flash-image, you are on a legacy model. Here is the three-tier routing framework every image pipeline team needs.

Published via Google DeepMind

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract editorial diagram showing three-tier image generation routing paths — speed, balance, and quality — converging at a unified AI gateway checkpoint

Google's June 30 launch of Nano Banana 2 Lite (gemini-3.1-flash-lite-image) is not just a new model release — it is the signal that the Nano Banana family is now complete enough to treat as a structured routing tier, not a single default endpoint. Paired with the simultaneous API release of Gemini Omni Flash for video, Google has effectively published a reference architecture for multimedia pipelines. If your team's image generation routing still points to gemini-2.5-flash-image (the original Nano Banana), you are on a legacy model that Google itself is redirecting away from.

What happened

On June 30, 2026, Google DeepMind released two models to developers simultaneously via the Gemini API and Google AI Studio:

  • Nano Banana 2 Lite (gemini-3.1-flash-lite-image): the fastest, cheapest entry in the Nano Banana family. $0.034 per 1,000 images at 4-second text-to-image latency. Google's recommended replacement for gemini-2.5-flash-image (the original Nano Banana).
  • Gemini Omni Flash (gemini-omni-flash-preview): a video generation and conversational editing model at $0.10 per second of output, covered separately in a prior analysis.

The Nano Banana family now has three tiers:

ModelAPI IDLatencyCost/1K imagesBest for
Nano Banana 2 Litegemini-3.1-flash-lite-image~4s$0.034Speed-sensitive, high-volume, draft/ideation
Nano Banana 2gemini-3.1-flash-imagefaster than Prolower than ProBalanced quality and cost
Nano Banana Progemini-3.0-pro-imageslowerhigherComplex, professional, accuracy-critical

The legacy model — original Nano Banana (gemini-2.5-flash-image) — is no longer recommended by Google. Per the official blog: "We recommend upgrading to Nano Banana 2 Lite for better quality, faster speeds and lower costs."

Why it matters for AI engineering teams

Three routing decisions just changed for any team calling the Gemini image API.

1. Model ID version discipline. If you hard-coded gemini-2.5-flash-image as your default image endpoint, you are now behind two generations. Unlike text models where a latest alias sometimes tracks the current best model, Google's image model IDs encode the generation and tier explicitly — gemini-3.1-flash-lite-image vs gemini-3.1-flash-image vs gemini-3.0-pro-image. You need a model routing config layer, not a hard-coded string, to move efficiently when new tiers land.

2. Tier selection is now a routing policy decision. Before Nano Banana 2 Lite, most teams defaulted to whichever single Gemini image model they onboarded with. Now, the same request type — generate an image from a text prompt — has three cost/latency/quality points that should be chosen based on workload context:

  • Draft or ideation workflows where users iterate quickly: Nano Banana 2 Lite at $0.034/1K, 4-second latency.
  • Production content requiring high character consistency or legible in-image text: Nano Banana 2 (the mid-tier).
  • Professional, accuracy-critical outputs where budget is secondary: Nano Banana Pro.

Routing based on a quality_tier request parameter — rather than a single static endpoint — means you can honor per-use-case SLOs without renegotiating provider contracts.

3. Cost accounting now requires tier tracking. At $0.034/1K for Lite vs a higher rate for Pro, the cost difference across a moderate production workload can be substantial. Billing reconciliation needs to track which tier was used per job, not just whether the call was to "the Gemini image API."

The router/operator angle

The most operationally significant aspect of this release is the image-to-video pipeline composition Google is now explicitly recommending: generate images at scale with Nano Banana 2 Lite, then pass the result as a reference frame into Gemini Omni Flash for video generation or animated editing.

This pipeline pattern — a fast, cheap image model feeding into a stateful video session — introduces async job coordination that most teams using synchronous text APIs have not had to solve. The image generation call is itself fast (4 seconds), but the video editing that follows uses the Interactions API, which is a session-based, multi-turn protocol. Stitching these together requires:

  • Job lifecycle tracking: which image output feeds which video session ID.
  • Fallback routing: if Omni Flash is rate-limited or in preview degradation, you need an independent fallback policy for the video leg that does not re-generate the image.
  • Cost attribution: the two legs bill at different rates ($0.034/1K images vs $0.10/sec video) and must be allocated to the right team/project in your billing ledger.

For teams using TheRouter's async media pipeline support, these image generation jobs can be submitted via the standard /v1/jobs/:id lifecycle, with the model parameter set to google/gemini-3.1-flash-lite-image. Tier routing — choosing Lite, mid-tier, or Pro based on request metadata — can be expressed as a routing rule without changing application code.

A practical routing rule set:

{
  "routes": [
    { "condition": "req.quality == 'draft'", "model": "google/gemini-3.1-flash-lite-image" },
    { "condition": "req.quality == 'balanced'", "model": "google/gemini-3.1-flash-image" },
    { "condition": "req.quality == 'pro'",     "model": "google/gemini-3.0-pro-image" }
  ]
}

This also positions your stack correctly for when the Nano Banana 2 Lite preview tag is dropped and a versioned stable ID is published — you update one routing config, not dozens of application call sites.

What TheRouter users should watch or try

  • Migrate away from gemini-2.5-flash-image if you have not already. Google considers Nano Banana 2 Lite the official replacement. Confirm your routing config is pointing to gemini-3.1-flash-lite-image before the legacy model ages out of the Gemini API changelog.
  • Set up tier-based routing rules for image generation if your workloads span draft, balanced, and quality use cases. This prevents both over-spending (routing all traffic to Pro) and under-serving (routing all production content to Lite).
  • Review the image-to-video composition pattern if your team is evaluating Gemini Omni Flash. The upstream image generation leg (Nano Banana 2 Lite) is cheap and fast enough to run ahead of each video session start, but it requires job coordination that breaks cleanly only if your gateway supports linked async job tracking.
  • Track preview model IDs actively. gemini-omni-flash-preview and gemini-3.1-flash-lite-image are both in API preview. Google has previously moved Gemini image preview models to deprecated status within 30-60 days after GA siblings launch. Set a routing config review cadence, not a manual reminder.
  • Check the Gemini API image generation docs for the full model comparison table and updated pricing when preview status changes.
Help & contact