← All articles

GLM-5.3-FlashX API Integration Guide: 200 Tokens/s Multimodal Speed for Latency-Sensitive Routing

GLM-5.3-FlashX is the speed-optimized variant of Z.ai's GLM-5.3-Flash, delivering up to 200 tokens per second with the same 320B MoE architecture (18B active), 1M context window, and native multimodal input. This guide covers the API, pricing ($0.37/$1.25 per 1M tokens on Z.ai), DashScope availability, and how to route FlashX for latency-sensitive workloads.

· TheRouter

GLM-5.3-FlashX is Z.ai's speed-focused serving variant of GLM-5.3-Flash, launched on DashScope September 21, 2026. It runs on the same 320B-parameter mixture-of-experts architecture with 18B active parameters per token, the same 1M-token context window, and the same native multimodal input (text, images, video, files). The difference is inference speed: Z.ai advertises up to 200 tokens per second for FlashX, compared to roughly 50 tokens/s for the standard Flash variant on Artificial Analysis benchmarks.

The trade-off is price. FlashX costs approximately 2.5x more than Flash per token. For teams routing through an API gateway, the question is whether the speed gain justifies the cost premium for latency-sensitive workloads like interactive coding agents, real-time screenshot analysis, and multi-step agent loops where each round-trip adds user-visible delay.

Note: GLM-5.3-FlashX is not yet listed in TheRouter's model catalog. It is available through Z.ai's direct API (glm-5.3-flashx), DashScope (ZHIPU/GLM-5.3-FlashX), and OpenRouter (z-ai/glm-5.3-flashx). Check the Zhipu provider page for current route availability.

TL;DR: GLM-5.3-FlashX at a Glance

SpecGLM-5.3-FlashXGLM-5.3-Flash
Architecture320B MoE, 18B active320B MoE, 18B active
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Inference speedUp to 200 tok/s (vendor claim)~50 tok/s (Artificial Analysis)
Input modalityText, image, video, fileText, image, video, file
Z.ai input price$0.37 / 1M tokens$0.15 / 1M tokens
Z.ai output price$1.25 / 1M tokens$0.50 / 1M tokens
Z.ai cache read$0.075 / 1M tokens$0.035 / 1M tokens
ReasoningAlways-onAlways-on
DashScope model IDZHIPU/GLM-5.3-FlashXZHIPU/GLM-5.3-Flash

What GLM-5.3-FlashX Is (and Is Not)

FlashX is the same model as GLM-5.3-Flash with optimized serving infrastructure for higher throughput. Z.ai describes it as the "speed-focused" option in the Flash family. Both share the hybrid sparse-plus-linear attention architecture that reduces attention computation by 3x and KV cache by 4.4x compared to the dense GLM-5.3 flagship.

FlashX is not a separate model with different weights or training. The quality of outputs should be functionally identical to Flash for the same prompt and parameters. The 200 tok/s figure is a vendor-advertised maximum, not an independently measured latency result — actual throughput depends on prompt length, concurrent load, and output length.

API Access: Three Routes to FlashX

Z.ai Direct API

The Z.ai platform exposes an OpenAI-compatible endpoint:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_ZAI_API_KEY",
    base_url="https://api.z.ai/api/paas/v4",
)

response = client.chat.completions.create(
    model="glm-5.3-flashx",
    messages=[{"role": "user", "content": "Explain the difference between sparse and linear attention."}],
    stream=True,
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

DashScope (Alibaba Cloud Model Studio)

FlashX appeared on DashScope on September 21, 2026 under the model ID ZHIPU/GLM-5.3-FlashX. DashScope provides an OpenAI-compatible endpoint with multi-region availability:

from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

response = client.chat.completions.create(
    model="ZHIPU/GLM-5.3-FlashX",
    messages=[{"role": "user", "content": "Analyze this screenshot and describe the layout."}],
    stream=True,
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

DashScope regions include China (Beijing), US (Virginia), Germany (Frankfurt), Singapore, and China (Hong Kong). Each model includes 1 million free tokens for initial testing.

OpenRouter

OpenRouter serves FlashX under the model ID z-ai/glm-5.3-flashx:

curl -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "z-ai/glm-5.3-flashx",
  "messages": [{"role": "user", "content": "Hello!"}]
}'

Pricing: The 2.5x Speed Premium

FlashX costs roughly 2.5x more than Flash per token. Whether this premium makes sense depends on your workload pattern.

ProviderModelInput / 1MOutput / 1MCache Read / 1M
Z.aiglm-5.3-flashx$0.37$1.25$0.075
Z.aiglm-5.3-flash$0.15$0.50$0.035
DashScopeZHIPU/GLM-5.3-FlashXVaries by regionVaries by region—
OpenRouterz-ai/glm-5.3-flashx$0.0352 (input)$0.50 (output)—

For a typical 1,000-token input / 300-token output request:

  • Flash: ~$0.00030 per request
  • FlashX: ~$0.00075 per request

Over 100,000 requests per day, that difference adds up to roughly $45/day. The question is whether 4x faster inference saves more than $45/day in user wait time, reduced timeout errors, or tighter agent loop cycles.

Multimodal Input: Images, Video, and Files

Both Flash and FlashX share the same native multimodal capabilities. The model accepts images via image_url content blocks in the OpenAI-compatible format:

response = client.chat.completions.create(
    model="glm-5.3-flashx",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What UI framework does this screenshot use?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/screenshot.png"}},
        ],
    }],
)

Z.ai documents support for video and file input as well, though the exact formats and size limits are best verified against the official API reference.

Capabilities: Reasoning, Tool Use, and Structured Output

FlashX inherits the full GLM-5.3-Flash capability set:

  • Reasoning (always-on): FlashX has reasoning permanently enabled via thinking.type: "enabled". You cannot disable it. The recommended setting is reasoning_effort: max with thinking.clear_thinking: false.
  • Function calling: Multi-step tool sequences are supported. Z.ai recommends enabling tool_stream: true for streaming requests with tool calls.
  • Structured output: JSON response format is supported for seamless system integration.
  • Context caching: Automatic cache reads at reduced rates. On Z.ai, cached input reads cost $0.075/1M tokens (vs $0.37 standard input).
  • Web search: Server-side web search is available as a tool, billed per successful call (~$0.01/call on Z.ai).

The recommended parameter settings from Z.ai documentation: temperature: 1, top_p: 0.95, reasoning_effort: max.

When to Route to FlashX vs Flash vs GLM-5.3

The GLM-5.3 family has three tiers. Here is a decision framework for API routing:

Route to FlashX when:

  • Interactive coding agents where each round-trip adds user-visible latency
  • Real-time screenshot analysis in development workflows
  • Multi-step agent loops where total wall-clock time matters more than per-token cost
  • Streaming responses in user-facing applications where perceived speed drives engagement

Route to Flash when:

  • Batch processing where latency is not the bottleneck
  • Background tasks (document processing, code review) where results are consumed asynchronously
  • High-volume workloads where the 2.5x cost difference is material
  • Tasks that benefit from the same model quality at lower cost

Route to GLM-5.3 (dense flagship) when:

  • Maximum reasoning quality is required (cybersecurity, complex software engineering)
  • Tasks where the dense model's additional capability justifies $1.40/$4.40 per 1M tokens
  • Long-horizon agent workflows where quality per step compounds

TheRouter Configuration: FlashX with Flash Fallback

When TheRouter adds GLM-5.3-FlashX to its model catalog, you could configure a latency-optimized routing tier. TheRouter routes OpenAI-compatible requests through configured providers, so the configuration would look like:

  1. Primary: GLM-5.3-FlashX for latency-sensitive requests
  2. Fallback: GLM-5.3-Flash when FlashX is unavailable or rate-limited
  3. Cost fallback: Qwen3.8-Flash or DeepSeek V4.1 Flash for budget-constrained requests

Since all three models expose OpenAI-compatible chat completions endpoints, switching between them requires only a model ID change. TheRouter's unified API handles provider-level differences in authentication and endpoint URLs.

FlashX vs Competitors: Speed-Tier Context

FlashX competes in the speed-optimized tier of multimodal models. For context:

ModelVendor Speed ClaimInput / 1MOutput / 1MContext
GLM-5.3-FlashX200 tok/s$0.37$1.251M
GLM-5.3-Flash~50 tok/s (AA)$0.15$0.501M
Qwen3.8-Flash—~$0.10~$0.301M
DeepSeek V4.1 Flash—$0.10$0.30128K

FlashX's speed advantage is most meaningful in interactive workflows where the model is the bottleneck. For batch or offline processing, the 2.5x price premium typically does not pay for itself.

FAQ

Is GLM-5.3-FlashX a different model from GLM-5.3-Flash?

No. FlashX uses the same 320B MoE architecture with 18B active parameters and identical weights. The difference is in the serving infrastructure — FlashX is optimized for higher throughput at a higher per-token cost.

Can I disable reasoning in FlashX?

No. Reasoning is always enabled in both Flash and FlashX. The thinking.type parameter only supports enabled. Z.ai recommends reasoning_effort: max and thinking.clear_thinking: false.

Is the 200 tok/s claim verified independently?

Not as of this writing. The 200 tok/s figure comes from Z.ai's documentation. Artificial Analysis measures GLM-5.3-Flash at approximately 50 tok/s and GLM-5.3 (dense) at approximately 66 tok/s, but independent FlashX benchmarks have not been published yet.

Is GLM-5.3-FlashX available on the GLM Coding Plan?

Not currently. Z.ai's documentation states that "GLM-5.3-FlashX is not yet available on the plan." Flash is available on the Coding Plan with 3x the quota compared to GLM-5.3.

Does FlashX support the Responses API?

The DashScope Responses API currently supports only glm-5.2 and glm-5.3 (dense), not the Flash/FlashX variants. FlashX is available through the Chat Completions API.

What regions does DashScope support for FlashX?

DashScope serves GLM models across China (Beijing), US (Virginia), Germany (Frankfurt), Singapore, and China (Hong Kong) via the OpenAI-compatible endpoint.


Sources: Z.ai GLM-5.3-Flash/FlashX documentation (retrieved 2026-10-04), OpenRouter GLM-5.3-FlashX (retrieved 2026-10-04), DashScope GLM integration guide (retrieved 2026-10-04), china-llm.com FlashX pricing (reviewed 2026-09-21), Artificial Analysis GLM-5.3 benchmarks (retrieved 2026-10-04).

Help & contact