Claude Code 2.1.243: Three Managed Settings That Reshape How Operators Control Models, Cache, and Cost

Claude Code 2.1.243 adds modelPicker, promptCacheTtl, and modelPricing — three managed settings that let gateway operators lock the model menu, tune cache TTLs per agent tier, and replace list-price cost figures with contracted rates.

Published via Anthropic

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract diagram showing three control knobs labeled Model Picker, Cache TTL, and Cost Attribution connected to a central AI routing gateway node

Three settings landed in Claude Code 2.1.243 that operators running centrally managed deployments need to configure before their next cost review or model rollout. Individually each one is useful; together they close gaps that have existed since Claude Code first gained enterprise managed-settings support.

What the three settings do

modelPicker lets an organization curate the /model picker — the in-session menu developers use to switch models without exiting — to an ordered, labeled list of approved model IDs. Any model ID spelling is accepted, including Vertex AI and Amazon Bedrock IDs. The list can append to or replace the built-in lineup. A separate fix in this release resolves a bug where selecting Ultracode from the picker silently did nothing; that fix is only useful now that the picker can be reliably populated.

promptCacheTtl and subagentPromptCacheTtl are two independent TTL settings. API-key and cloud-provider users can now hold a 1-hour prompt cache on the main conversation while subagents stay at the default 5-minute window. The two timers run independently, so a long-running orchestrator session can accumulate a warm cache while spawned subagents — often short-lived, high-turnover tasks — stay lean and avoid cache-slot waste.

modelPricing is a managed setting that accepts a per-model rate table and an optional discount multiplier. Once set, Claude Code uses those figures instead of list price when computing /cost output, the status-line cost display, and telemetry cost exports. Teams with enterprise agreements, volume discounts, or BYOK arrangements have previously had to mentally correct every cost figure the tool produced. This closes that gap at the source.

Why operators should care now

The modelPicker gap has been a recurring support friction point for gateway operators. Claude Code queries a gateway's /v1/models endpoint at startup and adds the returned models to the picker. That dynamic discovery has existed for a while, but organizations that needed a fixed, labeled, ordered list — for instance to surface a gpt-5.6-sol entry labeled "Ultrafast (preview)" rather than its raw model ID, or to expose a Bedrock ARN without developer confusion — had no managed-settings path to do that. Now they do.

The two-tier cache TTL matters because the Anthropic prompt cache pricing model charges per cached-token-hour. Subagents that spawn, run, and exit in under five minutes generate cache slots that expire before any second call arrives to reuse them. Setting subagentPromptCacheTtl to a shorter window reduces the window of charged-but-unreused slots. Conversely, the main session often benefits from a longer window because multi-turn conversations accumulate a large system-prompt prefix worth caching. The previous single-TTL design forced a compromise; the split allows each tier to be optimized independently.

The modelPricing setting has a direct impact on enterprise chargeback and team cost visibility. Many organizations negotiate per-model rates that are meaningfully below list — sometimes 30–50 percent below for committed-use tiers. When Claude Code reported /cost at list price, cost-per-task figures were systematically overstated, which either inflated perceived AI costs in internal reports or forced finance teams to apply correction factors manually. The telemetry path — which feeds dashboards and cost-allocation systems — was affected by the same distortion.

The router/operator angle

Three configuration surfaces, three different scopes of control:

modelPicker is model governance — it determines which models a developer can select within a session. For operators routing Claude Code through a gateway that exposes a curated model catalog, the picker must match that catalog. Without modelPicker, developers could select a model the gateway did not expose (it would appear in the built-in lineup but fail at request time). With it, the picker becomes a projection of the operator's approved model list.

promptCacheTtl / subagentPromptCacheTtl is caching economics — it controls how long cached token state persists at each agent tier. The right values depend on workload shape: a codebase-analysis session that runs for 30 minutes benefits from a 1-hour main TTL; a batch of short-lived subagents doing file searches does not.

modelPricing is cost attribution — it closes the gap between what Anthropic charges the operator (contracted rate) and what the operator reports internally (list price). It does not change what Anthropic bills; it changes what Claude Code displays and exports. The implication for cost-accounting pipelines is that telemetry exports now require a managed-settings sync to remain accurate as contract rates change.

What to watch or configure

If your deployment uses managed settings (managed-settings.json or a cloud-hosted managed source):

  • Add modelPicker with your approved model list, including any Vertex or Bedrock model IDs your gateway exposes. Use the label field to give each model a human-readable name that matches your internal documentation.
  • Set promptCacheTtl to 3600 (1 hour) for long-running analysis sessions; evaluate whether your subagent workload justifies a shorter subagentPromptCacheTtl — the default 300 seconds remains valid for fast-turnover tasks.
  • Populate modelPricing with your contracted per-model rates and discount multiplier before the next internal cost review. After the update, recheck any dashboards that consume Claude Code telemetry exports; the cost figures will change once the managed setting is applied.

Also note the 2.1.243 release also adds a /usage Loops breakdown (per-loop run count, total tokens, tokens per run, last-run timestamp) that surfaces runaway or chatty /loop tasks — useful for teams using Claude Code as an orchestrator running autonomous loop sequences.

Claude Code 2.1.243 is available now via npm install -g @anthropic-ai/claude-code@latest. Managed settings take effect on the next session start after the settings file is updated.

Help & contact