A Model-Level 429 From Zhipu Took GLM Flash Models Offline. TheRouter Now Scopes Cooldowns Per Model and Gives Them a Second Route.
On 2026-10-02 Zhipu's API returned a model-level rate-limit response for glm-4.6v-flash and TheRouter parked the whole account. Cooldowns are now per model, and six GLM flash-family models have a second route through Zhipu's international service.
Drafted with AI assistance from the cited sources; reviewed and published by Joe Werner.

On the afternoon of 2026-10-02, Zhipu's API answered requests for zhipu/glm-4.6v-flash with a rate-limit response that applied to that one model, not to the account. TheRouter did not make that distinction. It treated the response as an account-wide throttle and parked the whole Zhipu account for the length of the cooldown, so every Zhipu model on the shared pool stopped receiving traffic at once. For glm-4.6v-flash the outcome was worse still: TheRouter had exactly one route for it, so when that route was parked there was nowhere else to send the request. The incident data in this article is first-party, from TheRouter's own gateway; the vendor reference linked above documents the rate-limit responses themselves.
Two limits that look the same on the wire
Zhipu's error reference lists several conditions that all arrive as HTTP 429: a request-rate limit on the key, a usage-quota limit, and a "temporarily overloaded" response that is tied to the model being asked for. The last of these is the one that fired on 2026-10-02. It means one model is saturated; the same credentials can still call every other model the account has access to.
Before the fix, TheRouter kept a single cooldown per upstream account. Any 429 wrote that cooldown, and every model behind the account waited for it to expire. That is the right behaviour for a key-level limit and the wrong one for a model-level limit, because it turns a throttle on one model into an outage for all of them.
What changed in TheRouter
Two things changed, and both are live on the request path.
First, cooldowns are now scoped. When Zhipu's response says the limit applies to a single model, TheRouter parks only that model and leaves the rest of the account in rotation. The default window is 15 seconds, which is still above the 2-second client retry cadence we saw during the incident; when Zhipu sends a retry-after header, TheRouter honours that instead. A 429 that really is account-wide still parks the whole account, which is the correct response when the vendor is saying the credential itself is exhausted.
Second, the GLM flash family is no longer single-homed. Zhipu runs an international service for the same models, and TheRouter's route for each of the following now has that service as a second option, used only when the first route fails or is parked:
- zhipu/glm-4.6v-flash
- zhipu/glm-4.6v-flashx
- zhipu/glm-4.7-flash
- zhipu/glm-4.7-flashx
- zhipu/glm-4.6v
- zhipu/glm-4.5-air
The original route stays first for all six, so day-to-day latency and price are unchanged. For glm-4.6v-flash and glm-4.7-flash the second route carries no extra cost: Zhipu's international price list shows both at $0, and TheRouter's own price for a model does not depend on which route serves a request.
Not every Zhipu model gained a second route. The GLM-4.1V thinking variants, CogView-4 and embedding-3 are not offered on the international service and stay single-routed; GLM-5 Turbo is offered there but not yet priced, so it was held back. If you depend on those, a model-level 429 still parks them for the cooldown window.
What operators should do
Nothing in the request format changes. If you already send traffic to glm-4.6v-flash or glm-4.7-flash through TheRouter's Zhipu routes, the second route is picked up automatically; you do not change a model ID, a base URL or an API key. What you should expect going forward is that a model-level 429 from Zhipu parks only that model, for 15 seconds or the vendor-supplied retry-after, after which TheRouter retries on the second route for the six models above.
One thing to check in your routing configuration: if you have an explicit provider override pinning requests to a particular route, the override bypasses the fallback entirely. Automatic fallback only applies when TheRouter chooses the route. Pricing for all six models is on the shared rate card and is the same on either route.
The broader lesson is simple: a vendor-level rate limit and a model-level rate limit call for different responses, and treating them identically inflates the blast radius of every model-scoped throttle. The fix is live, and the six GLM flash-family models now have the redundancy that the 2026-10-02 incident showed they lacked.
Models covered in this article

Fable 5 Biology Classifier Fix: The Silent Model-Swap Your API Billing Never Warned You About
Fable 5 biology classifier false-positive fix cuts fallbacks 85%. For API operators, this exposed a silent billing risk: Fable 5 requests were being served by Opus 5 without warning. Here is what to audit before assuming model parity returns.

Anthropic Fallbacks "default" Mode Reaches Opus 5 — and Opus 4.7 Fast Mode Is Now a Hard 400 Error
Anthropic's July 24 API release added fallbacks: "default" to Opus 5 and Opus 4.8 and turned Opus 4.7 speed: "fast" into a hard 400 error. Here is what each change means for your routing policy today.

ZCode Launches: Z.ai's Official Coding Agent Exposes a New Endpoint Architecture Every Routing Team Must Map
Z.ai launched ZCode on July 2 — a free coding agent built on GLM-5.2. For routing teams: a separate coding endpoint, dual Anthropic/OpenAI protocol paths, third-party BYOK, and a 1.5x quota promotion closing July 31.