Claude Code Fallback Model: Avoid Double Failover in API Routing

Claude Code's --fallback-model can now auto-switch after model-not-found errors. Here's how teams using an API router should separate client fallback, gateway fallback, OTEL entrypoint metrics, and plugin governance.

Published via Anthropic / Claude Code GitHub

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Minimal editorial graphic showing routing flow with a primary model and fallback path branching cleanly, matte dark background with restrained blue-gray accents

The operational decision your team should care about in today's Claude Code release is not the quality-of-life CLI improvements — it is the three changes that shift coding-agent reliability and enterprise control one layer down, closer to the API infrastructure where routing decisions are made: automatic --fallback-model switching, pluginSuggestionMarketplaces governance, and disallowed-tools in skill frontmatter. Combined with a new OpenTelemetry app.entrypoint attribute, this release makes Claude Code meaningfully more observable and controllable at the deployment level.

What changed

The release published at 27 May 01:30 UTC on the official anthropics/claude-code GitHub repository includes the following operator-relevant changes:

--fallback-model now auto-activates on model-not-found errors. Previously, if a session's primary model returned a not-found error, every subsequent request failed with the same error for the remainder of the session. As of this release, Claude Code automatically switches to the configured --fallback-model for the rest of the session the moment the primary model is unavailable — no manual intervention, no session restart. The fallback is sticky for the session: once triggered, it stays active.

pluginSuggestionMarketplaces managed setting — a new enterprise policy that lets admins allowlist specific organization plugin marketplaces whose plugins may be suggested to developers via context-aware tips. Without this setting, Claude Code could suggest any available marketplace plugin. With it, only approved marketplaces surface in suggestions, giving IT and security teams direct control over which third-party integrations developers encounter during active sessions.

disallowed-tools in skill and slash-command frontmatter. Skills — Claude Code's on-demand knowledge modules — can now declare a list of tools the model is not permitted to use while the skill is active. This is a tool-level isolation mechanism: a skill scoped to a read-only documentation task can prevent Claude from executing shell commands or writing files for the duration of that skill's activation.

OTEL app.entrypoint metric attribute. Opt-in via OTEL_METRICS_INCLUDE_ENTRYPOINT=true, this attribute attaches the session entrypoint (CLI, SDK, remote, etc.) to OpenTelemetry metrics. Teams exporting OTEL data to Datadog, Grafana, or a similar backend can now slice model-usage and token-spend metrics by how Claude Code was invoked — distinguishing, for example, terminal sessions from CI pipeline runs or remote agent dispatches.

Additional changes in the same release:

  • SessionStart hooks can now set the session title at startup and resume, and can trigger skill directory rescans without a restart.
  • MessageDisplay hook: hooks can now transform or hide assistant message text as it is rendered — enabling output filtering, redaction, or annotation at the session level.
  • /code-review --fix applies review findings directly to the working tree; /simplify now calls it.
  • Auto mode no longer requires opt-in consent.
  • Fixed: cache_creation_input_tokens was reporting as 0 in transcript and result usage; now correctly sourced from the API's nested cache breakdown.

Why it matters for AI engineering teams

--fallback-model changes the session failure contract. Before this release, a model availability problem meant the session was effectively dead. The developer had to notice, stop, restart, and potentially lose accumulated context. The new behavior means Claude Code can silently recover from a transient model outage by switching to the fallback — without breaking the session. This is the same pattern that API routing gateways have provided at the infrastructure level, now built into the client itself.

The implication: teams relying on a routing gateway for automatic failover should be aware that Claude Code now has its own first-hop fallback logic. A request that fails at the gateway level may trigger the gateway's fallback policy; a request that fails before reaching the gateway (e.g., wrong model name, quota hit at the client config level) may now trigger Claude Code's own fallback. Understanding where in the stack fallback fires is important for accurate reliability attribution and for avoiding double-fallback scenarios.

pluginSuggestionMarketplaces fills a real enterprise governance gap. As Claude Code adoption scales inside organizations, the risk of unsanctioned third-party plugins surfacing in developer sessions grows. A developer in an active coding session who receives a suggestion to install an unreviewed marketplace plugin is a shadow-IT risk. The new managed setting gives security and platform teams the same kind of allowlisting control they expect from corporate software distribution — at the plugin suggestion level, not just the installation level.

disallowed-tools enables least-privilege skill design. A skill that has no write access cannot corrupt files; a skill that cannot execute shell commands cannot trigger unintended side effects. This is a significant upgrade for teams building internal skill libraries. It lets authors encode the intended scope of each skill as a security property, not just as documentation guidance.

OTEL app.entrypoint closes a visibility gap in multi-context deployments. Teams running Claude Code in multiple modes — desktop terminal, CI pipeline, cloud remote dispatch, SDK-hosted sessions — previously had no way to segment OTEL metrics by invocation mode. This attribute makes it possible to answer questions like "how much of our Claude Code token spend comes from CI pipeline runs vs. interactive sessions?" — a distinction that directly affects cost allocation, quota planning, and anomaly detection.

The router/operator angle

Fallback-at-client vs. fallback-at-gateway — you now have both. Claude Code's --fallback-model operates at the client level: if the configured model is not found, the session switches locally. An API gateway's fallback operates at the infrastructure level: if a provider returns a specific error code or latency threshold, the gateway routes to an alternate provider or model. These two layers can coexist — but they need to be configured to complement, not conflict with, each other.

Concretely: if your team configures Claude Code with --fallback-model=claude-opus-4-7 and your routing gateway also has a fallback from the same model to a different provider, a single failure event could trigger both. The behavior depends on which failure mode is hit first and at which layer. For clean reliability accounting, define a clear boundary: client-level fallback handles model-not-found errors (wrong model ID, access denied); gateway-level fallback handles provider errors (5xx, rate limit, timeout). Document this boundary explicitly.

Session observability now has an entrypoint dimension. If your team exports OTEL from Claude Code sessions and routes those metrics through a central observability backend, the app.entrypoint attribute enables a new segmentation: CI pipeline spend vs. interactive developer spend vs. remote agent spend. This is the token-cost equivalent of separating your API traffic by application tier. Teams that manage shared API budgets across multiple Claude Code use modes will want to enable OTEL_METRICS_INCLUDE_ENTRYPOINT=true and add entrypoint as a group-by dimension in their spend dashboards.

disallowed-tools as a skill-layer policy primitive. For teams building internal skill catalogs for large Claude Code deployments, disallowed-tools in frontmatter is now a first-class policy mechanism. A read-only documentation skill can declare disallowed-tools: [Bash, Edit, Write] and prevent the model from running commands or writing files regardless of what the session's main permission policy allows. This is composable with existing managed-settings.json policies — the skill-level restriction is additive.

MessageDisplay hook enables output governance. Teams that need to redact sensitive information from Claude Code output — IP addresses, credentials, PII patterns — can now do so at the display layer via the MessageDisplay hook, without modifying the model's responses. This is useful for enterprise deployments where Claude Code is used on sensitive codebases and IT requires that certain patterns not appear in developer-visible output or session transcripts.

What TheRouter users should watch or try

  • Review your fallback model configuration for double-fallback scenarios. If you're using Claude Code with a --fallback-model setting and your routing layer also has fallback policies covering the same models, map out the exact failure modes each layer handles. The goal is complementary coverage, not redundant triggers that obscure which layer actually resolved a failure.

  • Enable OTEL_METRICS_INCLUDE_ENTRYPOINT=true if you run Claude Code in multiple modes. This attribute costs nothing and adds a segmentation dimension that is difficult to reconstruct after the fact. Add it before you scale Claude Code to CI pipelines or remote dispatch, when per-mode cost attribution starts to matter for budget planning.

  • Audit your internal skill library for disallowed-tools opportunities. Any skill scoped to a read-only task — documentation lookup, design review, code explanation — is a candidate for an explicit disallowed-tools declaration. Encoding the intended scope as a security property rather than a trust assumption reduces blast radius if a skill is invoked in an unexpected context.

  • For enterprise deployments, test pluginSuggestionMarketplaces before scaling rollout. The feature gates which marketplaces can surface plugin suggestions to developers. If your organization has an internal plugin marketplace, add it to the allowlist and verify suggestion behavior before rolling Claude Code out to teams who will be actively installing plugins.

TheRouter routes OpenAI-compatible requests and records per-session token usage. For Claude Code workflows specifically, the --fallback-model change makes the client's own fallback logic more visible — teams using TheRouter alongside Claude Code should verify that their gateway fallback policies and Claude Code's --fallback-model setting cover different failure surfaces rather than the same one. OTEL metric export with session ID and entrypoint segmentation remains the most reliable way to get complete session-level spend visibility across the full stack.

Help & contact