Back to Models

Gemini 3.8 Flash

googlegoogle/gemini-3.8-flash

Google's Flash-tier model, released 2026-09-02. 1M-token context, thinking enabled.

Gemini 3.8 Flash is Google's third Flash-tier model in six weeks, released 2 September 2026 alongside Gemini 3.8 Flash Cyber. It builds directly on Gemini 3.7 Flash and is positioned as Google's most capable workhorse Flash model yet β€” Google's own model card describes it as delivering "performance advancements across software engineering and agentic knowledge workflows." The 1M-token context window and 64K output cap are unchanged from the previous Flash generation; thinking is enabled and effort-level controls carry over.

The headline improvement is diligence: Google states the model "works harder" on complex tasks, executing extra reasoning steps and calling tools iteratively. This shows up most clearly on long-horizon coding benchmarks where 3.8 Flash outperforms most larger frontier models at Flash cost. Google warns that the extra effort can mean more output tokens at higher effort levels β€” the model card recommends falling back to Gemini 3.7 Flash for efficiency-first workloads. On TheRouter this model is served over the OpenAI-compatible chat completions API.

Best for
  • β€’ Long-horizon software engineering β€” Google reports 73.7% on DeepSWE v1.1, ahead of GPT-5.6 Sol (72.7%), GPT-5.6 Terra (69.6%), and Claude Sonnet 5 (53.8%), while sitting below Claude Opus 5 (74.0%) by only 0.3 points at a fraction of the cost
  • β€’ Agentic terminal work β€” 89.4% on Terminal-bench 2.1, matching Claude Opus 5 (89.1%) and outperforming GPT-5.6 Sol (88.8%)
  • β€’ Finance and legal agent workflows in production β€” Vals Finance Agent v2 61.4% and Harvey Legal Agent Benchmark 10.0% all-pass-rate, both the highest in Google's comparison table
  • β€’ Long video understanding β€” 87.8% on LVBench (agentic), highest in Google's comparison table
Reach for something else if
  • β€’ Efficiency-first, high-volume pipelines where output token count is the constraint β€” Google explicitly recommends Gemini 3.7 Flash for efficiency-first workloads; 3.8 Flash uses more tokens at higher effort levels
  • β€’ General agent capabilities (non-coding) β€” 19.1% on Terminal-bench 4.0 is far behind Claude Opus 5 (51.8%) and GPT-5.6 Sol (37.3%); the model's strength is squarely in coding and domain-specific agentic benchmarks
  • β€’ Computer use β€” 59.0% on OSWorld-2.0 vs Claude Opus 5's 75.4%; for GUI automation Claude Opus 5 is the clear leader in Google's comparison

How TheRouter serves this differently from the vendor

As the vendor operates it

Google serves Gemini 3.8 Flash through the Gemini API (ai.google.dev) and Vertex AI. The upstream model accepts text, image, audio, and video input and returns text output. Native capabilities include thinking, function calling, structured outputs, code execution, context caching, Batch API, Google Search grounding, Google Maps grounding, and file search. Audio generation, image generation, and the Live API are not supported. Thinking tokens are billed as output tokens at the same rate.

On TheRouter

TheRouter serves Gemini 3.8 Flash as OpenAI-compatible /v1/chat/completions under google/gemini-3.8-flash. Exposed input modalities are text, image, and PDF; audio and video input are not passed through. Function calling and structured outputs work unchanged via the OpenAI tools schema. Thinking tokens appear in usage.completion_tokens. Audio generation, image generation, and the Live API are not available through TheRouter.

Context Length
1.0M
Max Output
66K
Input Priceper 1M tokens
$1.62/ 1M tokens
Output Priceper 1M tokens
$8.10/ 1M tokens

Modalities

textimagepdf→text

Capabilities

Vision

Pricing Breakdown

TypeRate
Input$1.62 / 1M tokens
Output$8.10 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

Release date2026-09-02deepmind.google β†—verified
Knowledge cutoffMarch 2026 (some domains may be limited to January 2025)deepmind.google β†—verified
Input token limit1,048,576deepmind.google β†—verified
Output token limit65,536deepmind.google β†—verified
Native modalities (upstream)Input: text, image, audio, video. Output: text only. TheRouter exposes text, image, and PDF input with text output.deepmind.google β†—verified
Vendor list price (Standard tier)$1.50 / 1M input, $7.50 / 1M output, $0.15 / 1M cached input (Global endpoint, Vertex AI). Introductory 50%-credits-back promotion through 2026-12-31 does not change the billed list rate.cloud.google.com β†—verified
Effort-level controlSupported. Lower effort levels reduce token overhead; Google recommends 3.7 Flash for efficiency-first workloads.blog.google β†—verified

Benchmarks

BenchmarkDistributionScoreSource
DeepSWE v1.1
Long-horizon software engineering. Gemini 3.7 Flash 65.3%; Claude Opus 5 74.0%; GPT-5.6 Sol 72.7%; GPT-5.6 Terra 69.6%; Claude Sonnet 5 53.8%
73.7%%deepmind.google β†—
Terminal-bench 2.1
Agentic terminal coding. Gemini 3.7 Flash 85.8%; Claude Opus 5 89.1%; Claude Sonnet 5 80.4%; GPT-5.6 Sol 88.8%; GPT-5.6 Terra 87.4%
89.4%%deepmind.google β†—
Terminal-bench 4.0
General agent capabilities. Gemini 3.7 Flash 11.2%; Claude Opus 5 51.8%; Claude Sonnet 5 12.4%; GPT-5.6 Sol 37.3%; GPT-5.6 Terra 23.6%
19.1%%deepmind.google β†—
GDPVal-AA v2
Knowledge work. Gemini 3.7 Flash 1482; Claude Opus 5 1824; Claude Sonnet 5 1584; GPT-5.6 Sol 1710; GPT-5.6 Terra 1528
1545Elodeepmind.google β†—
Vals Finance Agent v2
Financial analyst tasks. Highest in Google's table. Gemini 3.7 Flash 59.0%; Claude Opus 5 58.6%; Claude Sonnet 5 53.9%; GPT-5.6 Sol 53.8%; GPT-5.6 Terra 54.4%
61.4%%deepmind.google β†—
Harvey Legal Agent Benchmark
Complex legal workflows, all-pass-rate. Highest in Google's table. Gemini 3.7 Flash 8.8%; Claude Opus 5 6.7%; Claude Sonnet 5 5.0%; GPT-5.6 Sol 2.5%; GPT-5.6 Terra 0.8%
10.0%% (all-pass-rate)deepmind.google β†—
HLE-Verified
Multidisciplinary expert reasoning. Gemini 3.7 Flash 53.6%; Claude Opus 5 54.4%; Claude Sonnet 5 31.0%; GPT-5.6 Sol 54.5%; GPT-5.6 Terra 51.1%
54.9%%deepmind.google β†—
CharXiv (no tools)
Information synthesis from complex charts. Gemini 3.7 Flash 84.5%; Claude Opus 5 83.7%; Claude Sonnet 5 70.1%; GPT-5.6 Sol 85.8%; GPT-5.6 Terra 85.9%
86.2%%deepmind.google β†—
LVBench (agentic)
Long video understanding (agentic). Highest in Google's table. Gemini 3.7 Flash 85.4%; Claude Opus 5 75.4%; Claude Sonnet 5 68.5%; GPT-5.6 Sol 82.1%; GPT-5.6 Terra 78.9%
87.8%%deepmind.google β†—
OSWorld-2.0
Agentic computer use. Gemini 3.7 Flash 50.6%; Claude Opus 5 75.4%; Claude Sonnet 5 42.6%; GPT-5.6 Sol 62.6%; GPT-5.6 Terra 50.2%
59.0%% (partial score, batch tool enabled)deepmind.google β†—
BioMysteryBench (Human Solvable)
Bioinformatics research workflows, human-solvable tier. Gemini 3.7 Flash 87.1%; Claude Opus 5 90.1%; Claude Sonnet 5 87.5%; GPT-5.6 Sol 79.5%; GPT-5.6 Terra 83.8%
88.8%%deepmind.google β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "google/gemini-3.8-flash",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

Standard OpenAI-compatible chat completion. Thinking is enabled by default; its tokens count as output tokens.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemini-3.8-flash","messages":[{"role":"user","content":"Solve this bug in my code"}]}'

More from google

Similar models

Cross-provider sibling models

News & changes

2026-09-02

Google ships Gemini 3.8 Flash and 3.8 Flash Cyber

Google released Gemini 3.8 Flash on September 2, 2026 β€” its third Flash model in six weeks β€” alongside the restricted Gemini 3.8 Flash Cyber for cybersecurity use. The launch blog frames 3.8 Flash as the company's "most intelligent workhorse model," with the key design change being that the model works harder on complex tasks: more reasoning steps and tool calls at higher effort levels. It hits 73.7% on DeepSWE v1.1, within 0.3 points of Claude Opus 5, while remaining at the same $0.75 introductory ($1.50 standard) per-million-input-token price as 3.7 Flash.

re-authored by TheRouterblog.google β†—

Frequently asked

How does Gemini 3.8 Flash compare to Gemini 3.7 Flash?

3.8 Flash is the direct successor and beats 3.7 Flash on every benchmark Google published for the pair. The sharpest gaps are on long-horizon coding: 73.7% vs 65.3% on DeepSWE v1.1, and 89.4% vs 85.8% on Terminal-bench 2.1. The trade-off is token usage β€” 3.8 Flash uses more output tokens on complex tasks at higher effort levels. Price per token is identical.

Why does my usage.completion_tokens seem high?

Thinking is enabled by default and its tokens are billed as output tokens. Additionally, 3.8 Flash deliberately executes extra reasoning steps and tool calls on complex tasks, especially at higher effort levels. Google's model card notes this explicitly. Use a lower effort level or switch to Gemini 3.7 Flash to reduce token overhead.

Is the $0.75/$3.75 introductory price the real billed rate?

No. Google bills at the standard rate of $1.50 per 1M input tokens and $7.50 per 1M output tokens, then returns 50% as account credits through December 31, 2026. TheRouter's pricing reflects the list rate that Google actually charges, not the post-rebate figure.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datedeepmind.google β†—2026-09-19verified
Knowledge cutoffdeepmind.google β†—2026-09-19verified
Input token limitdeepmind.google β†—2026-09-19verified
Output token limitdeepmind.google β†—2026-09-19verified
Native modalities (upstream)deepmind.google β†—2026-09-19verified
Vendor list price (Standard tier)cloud.google.com β†—2026-09-19verified
Effort-level controlblog.google β†—2026-09-19verified
DeepSWE v1.1deepmind.google β†—2026-09-19verified
Terminal-bench 2.1deepmind.google β†—2026-09-19verified
Terminal-bench 4.0deepmind.google β†—2026-09-19verified
GDPVal-AA v2deepmind.google β†—2026-09-19verified
Vals Finance Agent v2deepmind.google β†—2026-09-19verified
Harvey Legal Agent Benchmarkdeepmind.google β†—2026-09-19verified
HLE-Verifieddeepmind.google β†—2026-09-19verified
CharXiv (no tools)deepmind.google β†—2026-09-19verified
LVBench (agentic)deepmind.google β†—2026-09-19verified
OSWorld-2.0deepmind.google β†—2026-09-19verified
BioMysteryBench (Human Solvable)deepmind.google β†—2026-09-19verified
Google ships Gemini 3.8 Flash and 3.8 Flash Cyberblog.google β†—2026-09-19verified
How does Gemini 3.8 Flash compare to Gemini 3.7 Flash?deepmind.google β†—2026-09-19to verify
Why does my usage.completion_tokens seem high?deepmind.google β†—2026-09-19to verify
Is the $0.75/$3.75 introductory price the real billed rate?cloud.google.com β†—2026-09-19to verify
Customer Support