Back to Models

Gemma 4 26B A4B IT

googlegoogle/gemma-4-26b-a4b-it

Google's Gemma 4 26B Mixture-of-Experts instruction-tuned model with 4B active parameters per token β€” open weights under Apache 2.0. Ranks #6 on the open Arena leaderboard. Multimodal text + image input. (Renamed from gemma-4-26b-moe to match official upstream ID gemma-4-26b-a4b-it.)

Gemma 4 26B A4B IT is Google DeepMind's mixture-of-experts member of the Gemma 4 open-model family. The official Gemma 4 model card describes it as a 26B-total-parameter model with roughly 4B active parameters per token, text and image input, text output, thinking modes, native function calling, system-role support, and a context window up to 256K tokens. Google publishes Gemma 4 as open weights under Apache 2.0, with pre-trained and instruction-tuned variants across E2B, E4B, 12B, 26B A4B, and 31B sizes.

For TheRouter, Gemma 4 26B A4B IT is the efficient high-throughput Gemma 4 route when teams want multimodal understanding, agentic/function-calling coverage, and stronger benchmark headroom than the small E-series models without paying for a proprietary frontier model. It should not be treated as a Gemini replacement: the best fit is production chat, image-grounded support, document/PDF parsing, and coding or reasoning tasks where open-weight portability and predictable routing matter more than the absolute frontier ceiling.

Best for
  • β€’ Vision-language extraction, inspection, and image-grounded support flows that need a 26B MoE open-weight model rather than a proprietary frontier model
  • β€’ Multilingual assistants and content workflows where 140+ language coverage matters and model portability is a buying criterion
  • β€’ Cost-sensitive document QA and extraction with long but bounded context, especially when 256K tokens is enough
  • β€’ Teams evaluating open-weight deployment paths before deciding whether to self-host, fine-tune, or keep a managed API route
Reach for something else if
  • β€’ Frontier reasoning, autonomous coding, or hard math where Gemini 2.5 Pro, Gemini 3-class models, Claude, or GPT-5.5-class models have a much higher ceiling
  • β€’ Ultra-cheap routing, classification, or extraction where smaller Gemma, Llama, or Qwen models can meet quality targets at lower latency
  • β€’ Native image, audio, or video generation β€” Gemma 4 26B A4B IT understands images but returns text, so use a dedicated generation model

How TheRouter serves this differently from the vendor

As the vendor operates it

Google publishes Gemma 4 26B A4B IT as open weights for self-hosting, fine-tuning, and deployment across local hardware, Hugging Face, Kaggle, Google AI Studio, Vertex AI, and other runtimes.

On TheRouter

TheRouter serves it as a managed OpenAI-compatible chat-completions route with text/image input and text output; this page does not claim local weight download, fine-tuning, or vendor-native runtime controls on the live TheRouter request path.

Context Length
128K
Max Output
8K
Input Priceper 1M tokens
$0.162/ 1M tokens
Output Priceper 1M tokens
$0.648/ 1M tokens

Modalities

textimage→text

Capabilities

Vision

Pricing Breakdown

TypeRate
Input$0.162 / 1M tokens
Output$0.648 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_pstop

Specifications

Release date2026-07-08Google AI for Developers β†—verified
Architecture / sizeMixture-of-Experts transformer, 26B total parameters with 4B active parameters per tokenGoogle AI for Developers β†—verified
Context windowUp to 256K tokensGoogle AI for Developers β†—verified
Training tokensNot publicly disclosed for the 26B A4B modelGoogle AI for Developers β†—unknown
Input / output modalitiesText and images in; text outGoogle AI for Developers β†—verified
LicenseGemma Terms of Useai.google.dev β†—verified

Benchmarks

BenchmarkDistributionScoreSource
MMLU Pro
Gemma 4 26B A4B, 0-shot, as reported in Google's Gemma 4 model card
82.6%ai.google.dev β†—
LiveCodeBench v6
Gemma 4 26B A4B, 0-shot, code generation benchmark from the official model card
77.1%ai.google.dev β†—
AIME 2026 no tools
Gemma 4 26B A4B, 0-shot, competition-math benchmark from the official model card
88.3%ai.google.dev β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "google/gemma-4-26b-a4b-it",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-26b-a4b-it",
    "messages": [{"role": "user", "content": "Explain adaptive image tiling in two paragraphs."}]
  }'

More from google

Similar models

Cross-provider sibling models

News & changes

2026-07-08

Google documents Gemma 4 with MoE, thinking, and 256K context

Google's launch post presents Gemma 4 as a five-size open-model family with E2B, E4B, 12B, 26B A4B, and 31B variants; the 26B A4B model is a mixture-of-experts option with roughly 4B active parameters per token, image-and-text input, text output, thinking, function calling, and up to 256K context.

re-authored by TheRouterGoogle AI for Developers β†—

Frequently asked

Is Gemma 4 26B A4B IT a good default model for production chat?

Use it when open weights, visual input, and predictable cost matter. For highest reasoning quality, longer context, or managed Google-native tools, choose a Gemini route instead.

Does Gemma 4 26B A4B IT generate images?

No. Gemma 4 26B A4B IT is a vision-language model that can read image input and return text. Use a dedicated image model for generation.

How is TheRouter serving different from Google's open-weight release?

Google's release lets teams download and run weights in multiple environments. TheRouter exposes a managed OpenAI-compatible API route, so integration is simpler but local runtime and fine-tuning controls are outside this route.

Are the API snippets verified?

No. This curator run has no paid API budget, so snippet verification remains unproven until a human operator approves and funds the verification run.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release dateGoogle AI for Developers β†—2026-08-22verified
Architecture / sizeGoogle AI for Developers β†—2026-08-22verified
Context windowGoogle AI for Developers β†—2026-08-22verified
Training tokensGoogle AI for Developers β†—2026-08-22unknown
Input / output modalitiesGoogle AI for Developers β†—2026-08-22verified
Licenseai.google.dev β†—2026-08-22verified
MMLU Proai.google.dev β†—2026-08-22verified
LiveCodeBench v6ai.google.dev β†—2026-08-22verified
AIME 2026 no toolsai.google.dev β†—2026-08-22verified
Google documents Gemma 4 with MoE, thinking, and 256K contextGoogle AI for Developers β†—2026-08-22verified
Help & contact