Claude Code 2.1.237 Fixes Prompt Caching on Every LLM Gateway — and 2.1.236 Adds a Default Model You Can Override

Claude Code 2.1.237 silently fixed a bug that invalidated prompt caching for every session routed through a custom base URL. If you've been routing Claude Code through a gateway, you've been paying full context-window cost on every turn since at least 2.1.229.

Published via Anthropic

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract illustration of a cache checkpoint restoring cost-efficient signal flow through an AI routing gateway layer

If you route Claude Code through a gateway or set a custom base URL, your prompt cache has been silently failing since at least 2.1.229. Every turn was re-processing the full session history at full token cost. That is fixed in 2.1.237, released this week.

The two-version cluster — 2.1.236 and 2.1.237 — is worth reading together because both releases land changes that specifically affect gateway-connected deployments.

What broke and what 2.1.237 restores

Claude Code's prompt caching is the mechanism that prevents re-encoding the full context window on every API call. A session with a long project context can avoid sending the identical system prompt and file contents as cold tokens on each turn — the cache header tells the upstream to reuse a stored kv-block instead.

When a session goes through a custom base_url — whether that's a self-managed proxy, an enterprise AI gateway, or a managed routing service — 2.1.237 fixes a regression that caused the cache to invalidate on every request. The changelog entry is terse: "Fixed prompt caching for sessions using an LLM gateway or custom base URL." The actual operator consequence is not terse at all.

For a typical Claude Code session with a 30–50K token project context, a failed cache hit means sending those tokens cold on every turn instead of amortizing them across the session. Over a multi-hour coding session with 30–60 turns, the difference between cache hits and misses on input tokens can amount to an effective 5–10× cost multiplier on the input side, depending on how much of the context is stable across turns. Teams running Claude Code behind a gateway and tracking token cost per session would have seen this as anomalously high input token counts — without any error or warning to explain why.

The regression appears to trace back to 2.1.229, which shipped a refactor of how the session manages the system-prompt boundary. The fix in 2.1.237 restores the cache-control headers on requests to non-Anthropic base URLs.

If you are on 2.1.236 or earlier and routing through a gateway, update to 2.1.237 before your next heavy session.

The new ANTHROPIC_DEFAULT_MODEL env var (2.1.236)

The second piece of context for gateway operators is a new environment variable: ANTHROPIC_DEFAULT_MODEL.

Before 2.1.236, Claude Code had ANTHROPIC_MODEL, which forces a specific model and holds regardless of what the user selects with /model. This is useful for operator-managed deployments where you must control which model hits the backend, but it removes the user's ability to switch models mid-session.

ANTHROPIC_DEFAULT_MODEL has different semantics: it sets the model new sessions start on, but a /model pick still overrides it — and that override persists across restarts. The practical operator migration path:

ScenarioBeforeAfter
Enforce one model, no user overrideANTHROPIC_MODEL=claude-opus-5-...Unchanged — still use ANTHROPIC_MODEL
Set a sensible default, let users change(no clean option; ANTHROPIC_MODEL forced it)ANTHROPIC_DEFAULT_MODEL=claude-sonnet-5-...
Gateway operator lets user pick but needs a billing-trackable starting pointWorkaround: set model in managed settingsANTHROPIC_DEFAULT_MODEL + let /model override

For teams operating Claude Code through a routing service that supports per-model billing attribution, the new variable gives you a starting-point model for cost bucketing without locking the user into it.

Auto-mode now behaves consistently on Bedrock, Vertex, and Foundry

The third operator-relevant item from 2.1.236 is a parity fix for auto mode. Before this release, sessions running on Bedrock, Vertex AI, and Azure Foundry — or sessions with telemetry disabled — ran a degraded version of the auto-mode classifier. The fix aligns the classifier defaults with the Claude API behavior, including severity-scored classification.

If you've been routing Claude Code to cloud-hosted Claude endpoints and have noticed auto mode refusing or approving tool calls differently than expected when you run it directly, this is the gap that was there. The parity fix means your organization's auto-mode settings should produce the same approval behavior regardless of which underlying provider path Claude Code is using.

Security: sandbox deny rules now hold through renames

One more item worth noting for multi-tenant gateway operators: wildcard read-deny rules in the macOS sandbox — for example, **/.env to block credential file reads — now take precedence inside allowed read regions and cannot be bypassed by renaming the denied file. Prior to 2.1.236, an **/.env rule could be circumvented by staging the denied file under a different name. The fix makes the deny rule behave like a denylist rather than a soft hint.

What gateway operators should do now

  • Update to 2.1.237 and monitor your next session's input token counts. You should see a significant drop in cold-token volume if your sessions were hitting the regression.
  • Decide between ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_MODEL based on whether you need to enforce a model or just seed one.
  • Audit your auto-mode configuration if you have Bedrock, Vertex, or Foundry as your routing target. The classifier parity change may shift approval behavior in ways your current policies don't account for.
  • Review sandbox deny rules if you use Claude Code in shared or multi-tenant contexts. The wildcard precedence fix is a net positive but worth verifying against your existing ruleset.

If you use TheRouter to route Claude Code sessions to multiple providers, the prompt-caching fix applies to any session using a custom ANTHROPIC_BASE_URL — which is how TheRouter-connected sessions work. Point your Claude Code installation at your TheRouter endpoint and update to 2.1.237 to restore cache efficiency.

Help & contact