Claude Code 2.1.247: The /claude-api cost-optimize Command Changes How Operators Audit API Spend
Claude Code 2.1.247 ships a built-in cost-profiling command that walks operators through caching, batch, token hygiene, and model-choice levers, and fixes a silent failure where sub-agents died on their first model call without activating the gateway fallback chain.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Most cost problems in AI gateway deployments are invisible until the invoice arrives. Teams find out they left caching off, ran everything through the highest-tier model, or let sub-agents pull full context on every retry — not at the time of the API call, but weeks later when someone runs sum(cost) on a CSV export. Claude Code 2.1.247 changes the feedback loop with a purpose-built profiling command, and fixes a related failure mode that quietly discarded the gateway's fallback chain every time a sub-agent hit a model 404.
What shipped
Two changes in 2.1.247 matter for operators running Claude Code at scale:
/claude-api cost-optimize — a new built-in command that profiles the current project's Claude API spend and walks through cost levers one measured change at a time. The command covers caching (cache hit rate, reads-per-write, TTL configuration), token hygiene (context bloat, compaction threshold, conversation length), batch API eligibility, reasoning effort tuning, and model tier selection. Each lever is presented as a before/after measurement, not just a recommendation.
Sub-agent model 404 fallback chain — when a sub-agent's first API call returned a model 404 (wrong model name in the agent's config, model retired, gateway routing to a provider that doesn't carry the model), prior versions let the sub-agent die at that call. The error returned to the parent session omitted the request id and model name, making it difficult to correlate in logs. 2.1.247 adds fallback chain handling: the sub-agent falls back through the session's configured model fallback list, and if the fallback succeeds, continues. If all fallbacks fail, the error returned to the parent now includes the error type, HTTP status, request id, and model name.
The release also adds Bedrock, Vertex, and Foundry MCP server connection reporting — when a configured MCP server fails to connect on those providers (or any session with telemetry disabled), the model now receives an explicit "this MCP server failed to connect" message rather than concluding the server's tools simply don't exist. And cloud session container restart recovery: if the session container restarts between turns while a background agent, shell, or monitor is running, the resumed session now reports what work was lost.
Why cost profiling belongs inside the client
The reflex for AI cost control is to add a proxy: intercept requests at the gateway layer, tag by model and team, export to a data warehouse. This works for billing attribution but not for optimization — the proxy sees request shape but not session context. It cannot tell you that your sub-agents pull full 100K-token histories because compaction is misconfigured, or that you'd save 40% on a specific workflow by switching from auto-effort to medium reasoning, because the proxy has no knowledge of session structure.
/claude-api cost-optimize runs inside the session, where it can observe context size growth, cache hit ratios against the session's actual prompt structure, and the relationship between reasoning effort settings and task complexity in that project. The output is actionable per-project, not aggregate-across-teams.
For operators managing many teams, this changes the support conversation. Instead of telling a team "your costs are high, check your caching settings," you can point them to /claude-api cost-optimize output and see exactly which lever is responsible — often it is prompt cache TTL misconfiguration or sub-agents that skip the fallback tier and hit the flagship model for tasks that a smaller model handles fine.
The sub-agent 404 fallback gap
Before 2.1.247, the failure sequence for a misconfigured sub-agent looked like this:
- Sub-agent starts, sends first API call with model
A. - API returns 404 (model not available at this provider/workspace).
- Sub-agent dies. Error propagated to parent: generic message, no model name, no request id.
- Parent session has no way to know whether to retry with a different model or whether the sub-agent completed any work.
Gateway fallback chains exist precisely to handle this — if model A is unavailable, try B, then C. But the fallback chain was only active for the parent session's direct calls, not for sub-agent first-call failures.
The 2.1.247 fix means sub-agents now participate in the session's fallback chain from their first call. For operators running multi-provider routing, this has a concrete effect: a sub-agent configured to use a model that's unavailable on one provider (e.g., a model listed in your config that your workspace's API key doesn't have access to) will now fall through to the next model in your fallback list rather than dying silently.
The structured error format — error type, HTTP status, request id, model name — also lands in your gateway logs correctly. Previously, a 404 from a sub-agent produced an opaque error in parent context that was hard to trace to a specific API call.
MCP server connection reporting on managed backends
One quiet friction point on Bedrock, Vertex, and Foundry deployments: MCP server connection failures were silent. Claude would reach a point in a task where it needed a tool from an MCP server that had failed to start, find no tools listed, and conclude the tools didn't exist — sometimes attempting a workaround, sometimes stopping with a confusing error. The actual cause (MCP connection refused, timeout, auth failure) was visible in logs but not surfaced to the model.
The 2.1.247 change sends an explicit message to the model when an MCP server fails to connect. The model can then surface this to the user ("the code-search MCP server failed to connect — is the server running?") rather than failing silently. For operators, this means fewer "Claude said it couldn't do X but the tool should be available" tickets, because the failure is now visible in the session.
What to check in your deployment
Cost-optimize: Run /claude-api cost-optimize on your highest-spend projects first. The most common finding is cache configuration — either cache is off, the TTL is set lower than the conversation length, or sub-agents are not sharing cache with the parent session. The second most common is model tier: teams that started with the flagship model for prototyping and never adjusted routing policy for production workloads.
Sub-agent fallback: Audit your sub-agent configs for model names that may not be available across all providers in your fallback chain. If you run a dual-provider setup (e.g., Anthropic direct + Bedrock), ensure both providers carry the model your sub-agents specify, or add an explicit fallback model in your session config. The 2.1.247 change handles the runtime failure, but the right fix is alignment between your sub-agent model list and your provider coverage.
MCP connection failures: If you manage MCP servers on Bedrock or Vertex deployments, validate that connection errors now surface in session output. If they don't, check that your Claude Code version is 2.1.247 or later.
What TheRouter users should watch
Sub-agent model fallback behavior interacts directly with gateway routing policy. If you route Claude Code through a gateway that rewrites model names or maps to provider-specific identifiers, the fallback chain in 2.1.247 will attempt fallback against the gateway — meaning the gateway needs to handle 404 responses consistently across provider backends. If the gateway swallows 404s and returns a different error code, the fallback chain may not trigger correctly.
Check your gateway's error passthrough behavior for model 404 responses from Bedrock, Vertex, and direct Anthropic endpoints. The routing docs cover how to configure provider-specific error handling in TheRouter.

Claude Code 2.1.251: The PreModelSwitch Hook Is Now Your Model Governance Control Plane
Claude Code 2.1.251 adds PreModelSwitch and PostModelSwitch hooks that let operators block, confirm, or annotate model switches in real time. Combined with new spend-limit visibility and per-session cache metrics, this release makes the client itself a usable governance surface.

Claude Code 2.1.275 Broke Every Gateway Proxy. 2.1.276 Fixed It the Same Day.
A new internal request tag in 2.1.275 caused 400 errors on every proxy-routed API call. 2.1.276 hotfixed it the same day. Breakdown of the failure, affected configs, and three secondary operator changes worth auditing.

Claude Code 2.1.274: MCP Reliability Overhaul, Gateway Postgres Config, and Self-Healing Transcripts
Claude Code 2.1.274 fixes six MCP failure modes that silently break production tool sessions, adds store.connect_timeout_seconds and CLAUDE_CODE_GATEWAY_DRAIN_TIMEOUT_MS to the Claude apps gateway, and makes corrupted transcripts self-heal instead of looping forever.