Claude Enterprise Model Entitlements: Per-Role Model Access and Effort Level Caps Land in Beta
Anthropic's new Enterprise model entitlements let admins lock specific Claude models to roles and cap the effort level—directly capping token spend—per role. Here is what AI engineering teams need to configure.

Anthropic quietly shipped a governance feature on July 1, 2026 that changes the model-access calculus for every Enterprise team: model entitlements. Admins can now enforce which Claude models each role can reach, and—separately—how much compute effort each role can apply per request. For AI engineering teams that have been managing model choice and cost through informal norms or brittle managed-settings.json configs, this is the first time the access boundary is enforced at the platform level.
What changed
The new model access controls operate in two layers:
Layer 1 — Organization-level model ceiling. Admins enable or disable specific Claude models for the whole org. If a model is off at org level, no role can turn it back on. Disabling a model that is set as a role's default requires migrating that default first.
Layer 2 — Role-level access and effort cap. Custom roles can be limited to a subset of the org-allowed models. On top of that, admins can cap the maximum effort level each role can select per model—low, medium, high, or max. Members on that role only see effort options at or below the cap in the model menu.
The feature lands in beta in Claude Code (CLI ≥ 2.1.196), Claude Chat, Cowork, and Office Agents. Claude in Chrome, Design, and Security are not yet supported.
Why it matters for AI engineering teams
Effort levels are not cosmetic UI labels. They map directly to token spend: a max-effort response on Sonnet 5 or Opus costs materially more per request than a medium-effort one. Until now, controlling that spend required either trusting users not to drag the slider up or writing custom API-layer policies.
Model entitlements close this gap at the provider layer before the request even reaches your routing infrastructure:
- A junior engineering role can be locked to Claude Sonnet 5 at
mediumeffort—preventing ad-hoc Opus calls during exploratory work. - A data science role can have access to Opus at
highbut notmax—preserving compute budget for batch jobs. - An external-contractor role can be restricted to Flash-equivalent models entirely.
This is particularly relevant for organizations using Claude Code at scale across departments. Claude Code 2.1.196 introduced org default models; model entitlements now let admins enforce which models are reachable rather than just which model is the default.
The router/operator angle
The feature intersects directly with Claude Code's managed-settings.json availableModels field. Per the updated docs, the two systems compose: in Claude Code CLI and IDE, a member sees only models that appear in both availableModels and their model entitlements. The more restrictive of the two wins.
This matters for teams using a gateway to route Claude API traffic:
-
Fallback chain integrity. If your gateway has a fallback from Sonnet 5 → Opus but the requesting user's role doesn't have Opus access, the fallback will fail with a model-unavailable error. Routing teams must audit their fallback chains against role entitlements.
-
Effort-level signaling. The
effortparameter in the Anthropic API (low/medium/high/max) maps to the effort UI. If your routing layer injects a defaulteffort: "high"for all requests, users on roles capped atmediumwill see unexpected failures or silent downgrades depending on how Anthropic resolves the conflict. -
Cost modeling. Role-based effort caps give a more predictable upper bound on per-request spend. Teams that do cost modeling across model × effort combinations now have a platform-enforced ceiling to plan around—rather than relying on post-hoc billing alerts.
Action checklist:
- Audit which roles currently access Opus or high-effort Sonnet calls; decide whether those should be restricted.
- Cross-check your Claude Code
managed-settings.jsonavailableModelslist against the model entitlements you intend to enforce—conflicts create user-visible failures. - Update fallback chain configs if any fallback model is restricted for a given role.
- Set effort-level caps for cost-sensitive roles before the Claude Sonnet 5 introductory pricing window closes on August 31, 2026.
What TheRouter users should watch or try
Teams routing Claude traffic through a gateway should treat role model entitlements as an upstream constraint rather than a routing parameter. Your gateway does not need to replicate this logic—Anthropic enforces it before the request is dispatched—but your routing policy needs to be consistent with it.
If you use TheRouter or any other AI routing layer to fan out requests across multiple Claude roles or user tiers, verify your routing logic against each tier's entitlement profile. Specifically, test that your error-handling path handles model_not_available responses correctly and does not retry against a model that the role cannot reach.
The docs section on routing policies is a useful starting point for teams setting up multi-tier Claude routing.

Anthropic Opens Seoul Office — What the Korean Enterprise Wave Reveals About Regional Routing Architecture
Anthropic's Seoul office and Korean enterprise rollouts surface a pattern every global AI team faces: regulated industries cannot use the direct Claude API for in-region data residency. Here is what the Korea deployments reveal about routing architecture.

Anthropic Inference Hooks Put a Pre-Inference Gate at the Provider Layer: What It Means for Your Routing Architecture
Anthropic's new Inference Hooks let enterprise organizations intercept every governed Claude prompt before the model runs. For teams already filtering at the gateway layer, this creates a dual-gate architecture that changes where enforcement belongs.

Fable 5 Biology Classifier Fix: The Silent Model-Swap Your API Billing Never Warned You About
Fable 5 biology classifier false-positive fix cuts fallbacks 85%. For API operators, this exposed a silent billing risk: Fable 5 requests were being served by Opus 5 without warning. Here is what to audit before assuming model parity returns.