Gemini Omni Flash API Routing: Google's Conversational Video Model Is Now an Operator Decision
Google's Gemini Omni Flash lands in public API preview with model ID gemini-omni-flash-preview, ~$0.10/sec pricing, and a new Interactions API that makes multi-turn video editing a routing policy problem.

When Google released Gemini 3.5 Flash and Gemini Enterprise Agent Platform earlier this quarter, the multimodal routing story still had a conspicuous gap: video generation remained a specialist toolchain with no native Google API path that matched the text-routing experience. That changes with Gemini Omni Flash, which hit public API preview on June 30, 2026.
The operative shift for AI engineering teams is not just that Google now has a video generation model. It is that the model uses the Interactions API — a session-scoped, multi-turn protocol where each video edit builds on the prior turn's output. That pattern is fundamentally different from the stateless request/response shape of text or image generation APIs, and it creates routing decisions that most teams have not yet planned for.
What happened
Google DeepMind released Gemini Omni Flash to Google AI Studio and the Gemini API on June 30, 2026. The model is the first in the Gemini Omni family, designed for conversational video generation and editing from combinations of text, image, video, and (soon) audio input.
Key technical facts:
- API model ID:
gemini-omni-flash-preview - API: Gemini Interactions API (session-based, multi-turn; distinct from the standard Responses/Chat endpoint)
- Pricing: approximately $0.10 per second of video output; $1.50 / 1M input tokens, $17.50 / 1M video output tokens; 5,792 tokens per second of 720p video output
- Current preview limit: 10-second video generations per turn
- Input modalities: text, image, video reference; audio references announced but not yet fully shipped
- Safety: SynthID watermarking and C2PA Content Credentials embedded in output
- Availability: Google AI Studio, Gemini API, Gemini Enterprise Agent Platform, Gemini app, Google Flow
The Interactions API differs from stateless generation in a critical way: each turn's context (the prior video state, edits, and scene continuity) is held server-side inside the session. A client sends a new instruction — "dim the lights and add falling snow" — and the model applies it to the existing video state rather than starting from scratch.
Why it matters for AI engineering teams
This is not a text-to-video drop-in. Teams that have integrated wan2.7 via DashScope or Vertex AI Veo for async video jobs will find that Gemini Omni Flash follows a different session lifecycle:
-
Session stickiness: unlike a stateless generation request you can route to any provider replica, an Omni Flash session holds mutable state. Routing must pin the session to a consistent session endpoint for its entire edit loop — or re-upload the full video state on each turn.
-
Billing granularity: video output is billed per second at ~$0.10, meaning a 5-turn edit session generating 10 seconds each time costs ~$0.50 in output tokens before input tokens are counted. Contrast this with async video APIs (e.g. wan2.7 at DashScope) where the job is stateless and the cost is fixed per resolution/duration.
-
Preview-mode budget planning: in public preview, Google may impose lower rate limits than in GA. Teams building on
gemini-omni-flash-previewshould plan for a GA model ID rename when the model leaves preview — the same deprecation pattern that bit teams usinggemini-3.1-flash-image-previewin June. Pin to the preview ID only in non-critical paths, and watch for the GA rename. -
SynthID + C2PA provenance: generated video inherits provenance watermarking automatically. This is an enterprise governance positive (auditable AI-generated content) but also a signal that Google considers Omni Flash output to be declarable as AI-generated even in downstream platforms. Operators in regulated media or ad-tech should verify whether embedded credentials surface in their publishing pipelines.
The router/operator angle
Gemini Omni Flash introduces a new category decision for teams routing media generation workloads:
Stateless async vs. session-stateful video generation. Most current video APIs (wan2.7 DashScope, Veo, Kling, Seedance) accept a single prompt/reference and return a job ID you poll for completion. Omni Flash's Interactions API is stateful by design — edits accumulate on the server across turns. This means your fallback strategy must account for session loss: if a provider error interrupts a multi-turn session, you cannot simply retry on a different endpoint without starting the edit loop over.
Model ID versioning discipline. gemini-omni-flash-preview will eventually be renamed or superseded. A routing config that pins the preview ID without an alias layer will break silently when Google deprecates it. Reference the Gemini image API deprecation pattern from June 2026 as a template: preview IDs can sunset within weeks of a GA release.
Cost model comparison. At ~$0.10/sec ($6.00/min) of video output, Gemini Omni Flash sits above most DashScope async video job pricing for equivalent durations but below premium real-time video synthesis services. Teams routing creative or generative-media workloads can now treat Google as a primary provider for conversational video editing, with DashScope/Wan2.7 as a fallback for stateless batch jobs. The two APIs are not drop-in substitutes — they need to be routed to different handlers based on whether the workload is stateful (edit session) or stateless (bulk generation).
Governance surface. SynthID + C2PA are embedded at generation time, not at delivery. If your routing layer strips or inspects HTTP headers, make sure C2PA content credentials (which may be manifest-embedded) survive your pipeline intact.
What TheRouter users should watch or try
Async video job routing through TheRouter's Async Media endpoint currently routes stateless generation requests. The Gemini Omni Flash Interactions API is session-based — this is a different job lifecycle. Until native Interactions API session routing is available, the safe approach is:
- Use Gemini Omni Flash directly via the Gemini API for multi-turn editing sessions where session continuity is required.
- Use the Async Media endpoint for stateless video generation jobs (wan2.7, Veo, Kling, etc.) where each job is independent.
- Build a routing discriminator in your application layer: if the workload is a stateful multi-turn edit loop, route to Gemini Omni Flash directly; if it is a single-shot generation, route through your standard async media path.
Watch for Gemini Omni Flash GA — that is when the model ID will stabilize, pricing may adjust, and output duration limits are likely to increase beyond 10 seconds per generation.

Gemini 3.5 Live Translate API: Real-Time Speech Translation Enters the Routing Layer
Google's Gemini 3.5 Live Translate is now in public preview via the Gemini API, delivering continuous speech-to-speech translation across 70+ languages. Here is what the new model ID means for voice-agent routing teams.

Google Cloud API Gateway model routing: Spec-first routing and its critical host constraint
Google Cloud API Gateway model routing is now in Public Preview: a serverless OpenAPI-spec layer that routes Gemini, Claude, and OpenAI requests — but only if all your backends live on Vertex AI.

Nano Banana 2 Lite Is Your New Default Gemini Image Endpoint — Here's the Routing Decision Framework
Google's Nano Banana 2 Lite (gemini-3.1-flash-lite-image) landed June 30 at $0.034/1K images and 4-second latency. If you are still routing to gemini-2.5-flash-image, you are on a legacy model. Here is the three-tier routing framework every image pipeline team needs.