ZCode Launches: Z.ai's Official Coding Agent Exposes a New Endpoint Architecture Every Routing Team Must Map
Z.ai launched ZCode on July 2 — a free coding agent built on GLM-5.2. For routing teams: a separate coding endpoint, dual Anthropic/OpenAI protocol paths, third-party BYOK, and a 1.5x quota promotion closing July 31.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

When Z.ai officially launched ZCode on July 2, 2026, the news cycle focused on competitive framing — Cursor, Claude Code, GitHub Copilot, now this. That framing misses what matters operationally: ZCode is the first production coding agent from the GLM family to expose a dedicated coding endpoint separate from the general BigModel API, dual-protocol support (Anthropic and OpenAI), a BYOK channel for third-party models, and time-bounded quota pricing that expires in 25 days.
If your team routes coding-agent traffic, each of those details changes how you configure providers.
What happened
Z.ai (formerly Zhipu AI) shipped ZCode — a free desktop application for macOS, Windows, and Linux — on Wednesday, July 2. The agent is built around GLM-5.2, Zhipu's flagship 1M-context coding model released in June with MIT-licensed open weights. ZCode organizes development around long-horizon "Goals": the agent plans, edits, runs checks, and iterates across a full task rather than completing one-shot prompts.
The product includes remote control via WeChat, Feishu, and Telegram — a differentiator for China-based developer teams — as well as a subscription plan structure ($16.20/month Lite → $144/month Max) that undercuts comparable Claude Code and Cursor tiers.
Why it matters for AI engineering teams
Three architectural decisions in ZCode create direct routing implications:
1. The coding endpoint is not the general endpoint. For teams using API keys rather than OAuth, ZCode documents a dedicated path for Coding Plan subscribers:
OpenAI Base URL: https://open.bigmodel.cn/api/coding/paas/v4
This is distinct from the general BigModel endpoint (https://open.bigmodel.cn/api/paas/v4). Routing GLM-5.2 coding traffic through the wrong base URL will either use general quota incorrectly or fail entirely. This is a config mistake that is easy to make and hard to notice until quota runs dry at the wrong tier.
2. GLM-5.2 supports both Anthropic and OpenAI protocols via BigModel. The ZCode configuration exposes two protocol paths:
- Anthropic protocol:
https://open.bigmodel.cn/api/anthropic - OpenAI protocol:
https://open.bigmodel.cn/api/paas/v4
This matters for teams routing Claude Code or other Anthropic-SDK clients through a gateway. If a team's gateway routes to BigModel via the Anthropic protocol for GLM-5.2, the behavior differences between the two protocols — including thinking-mode activation and message format — can produce inconsistent outputs across providers.
3. BYOK via third-party provider channels. ZCode explicitly supports routing to third-party models through the Anthropic or OpenAI protocols. This means teams can configure a gateway as ZCode's "third-party provider" and route all agent traffic — including non-GLM calls — through a central policy layer. It's a BYOK model, not a closed system.
The router/operator angle
Here's the decision matrix for teams routing ZCode-backed GLM-5.2 traffic:
| Scenario | Recommended configuration |
|---|---|
| Direct coding-heavy workloads, Coding Plan subscriber | Use dedicated coding endpoint /api/coding/paas/v4; do not mix with general endpoint |
| Routing GLM-5.2 via an OpenAI-compatible gateway | Point gateway to /api/paas/v4; treat as any OpenAI-compatible upstream |
| Routing GLM-5.2 via an Anthropic-compatible gateway | Point gateway to /api/anthropic; verify thinking mode behavior for your workload |
| Multi-provider agent routing (GLM + Claude + Codex) | Use ZCode's BYOK "third-party provider" channel; route all traffic through a single policy layer |
The July 31 quota factor. Through the end of July, ZCode applies a 0.67x token-consumption multiplier to Coding Plan subscribers — effectively stretching plan quota by about 50%. Routing teams that are modeling token costs for the next billing cycle need to factor in that this promotional rate disappears August 1. Build cost estimates using the full rate, not the promotional rate.
Geopolitical sourcing note. Z.ai has confirmed that GLM-5.2 was trained entirely on Chinese-manufactured accelerators. For enterprise routing policies that include chip-sourcing or jurisdiction criteria for AI supply chain, GLM-5.2 has a verifiable Chinese-chip provenance that differs from models trained on US or EU hardware. This is relevant for compliance-aware routing policies, not a performance or quality statement.
What TheRouter users should watch or try
Teams already routing via an OpenAI-compatible gateway can add BigModel's OpenAI endpoint as a provider and route GLM-5.2 coding traffic to it with standard model-name routing rules. The key configuration detail: use https://open.bigmodel.cn/api/paas/v4 as base URL for general access, or https://open.bigmodel.cn/api/coding/paas/v4 if you are a Coding Plan subscriber wanting dedicated coding quota.
Three operational checkpoints before the July 31 campaign deadline:
- Audit which endpoint your GLM-5.2 traffic is hitting. General vs. coding endpoints have separate quota pools.
- Verify protocol consistency. If you route some traffic via Anthropic protocol and some via OpenAI protocol, confirm that thinking-mode and message-format behavior is consistent for your workload type.
- Reset cost models for August. The 1.5x effective quota from the 0.67x multiplier ends July 31. Any cost projections for August and beyond should use the standard rate.
ZCode is free to try. The BYOK architecture means a central routing layer is a natural fit — operators can point ZCode to their existing gateway as the "third-party provider" and gain visibility and policy enforcement over all coding-agent model calls in a single place.
Models covered in this article

GLM-5.2 on OpenRouter & BigModel: Zhipu's 1M-Context Coding Agent with 128K Output
Zhipu GLM-5.2 is now accessible through OpenRouter, BigModel API, and DashScope: 1M-token lossless context, 128K max output, reasoning_effort, context caching, and FrontierSWE-class performance for long-horizon coding-agent routing.

A Model-Level 429 From Zhipu Took GLM Flash Models Offline. TheRouter Now Scopes Cooldowns Per Model and Gives Them a Second Route.
On 2026-10-02 Zhipu's API returned a model-level rate-limit response for glm-4.6v-flash and TheRouter parked the whole account. Cooldowns are now per model, and six GLM flash-family models have a second route through Zhipu's international service.

Qwen3.8-Max Is Now DashScope's Top-Tier Model: What the Flagship Upgrade Means for Your Routing Policy
Alibaba's qwen3.8-max lands on DashScope with 2.4T parameters, 1M context, and thinking mode — while qwen3.7-max drops to legacy. Here is what changes for teams routing to Qwen's flagship tier.