Claude Code Best Practices: Context Management & Engineering — Anthropic Official Guide

Anthropic official Claude Code best practices guide (code.claude.com): context management and context engineering are first-class constraints for routing teams. Context windows fill fast, performance degrades as they fill, coding agents burn tokens 3-10× vs chat.

Published via Anthropic / Claude Code Docs

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract editorial diagram showing a context window budget filling up across parallel coding agent sessions, with routing checkpoints and cost signals

The decision that changes your team's AI cost model is not which model you pick — it is how many tokens a single coding agent session actually burns. Anthropic's official Claude Code best-practices guide, published at code.claude.com/docs/en/best-practices, makes this explicit for the first time with a single governing constraint: context windows fill fast, and LLM performance degrades as they fill. For teams routing Claude Code through an API gateway, this is not a UX recommendation. It is a cost and reliability signal with direct implications for how you model token spend, observe sessions, and route fallback.

What happened

Anthropic published a comprehensive official best-practices guide for Claude Code — their terminal-based agentic coding environment. The guide is authoritative, covers Anthropic's own internal engineering workflows, and is structured around one central resource constraint: context window budget.

Key patterns the guide codifies:

Context window is the primary resource to manage. Every file read, every command output, every message in a Claude Code session accumulates in the context window. A single debugging or exploration session can consume tens of thousands of tokens. As the window fills, the model begins to lose earlier instructions and make more mistakes. The guide explicitly tells engineers to track context usage continuously with a custom status line.

Explore-plan-code phasing reduces wasted context. The recommended four-phase workflow (explore in read-only "plan mode" → generate a plan → implement → commit) exists not just to improve output quality but to reduce the number of file reads and tool calls that accumulate before implementation starts. Each unnecessary read is tokens that never come back.

CLAUDE.md is the operator configuration primitive. The /init command generates a CLAUDE.md file that persists across sessions. The guide is explicit: keep it short, cut anything Claude could infer from the code, and verify by observing whether Claude's behavior actually changes. A bloated CLAUDE.md wastes tokens on every session startup and can cause Claude to ignore the rules that matter.

Parallel sub-agents multiply context cost. The guide endorses running multiple Claude Code agents simultaneously for independent tasks. Each parallel session has its own context window, runs its own model calls, and accumulates its own token spend. Three parallel agents doing large codebase exploration can collectively burn tokens at 3x the rate of a single session — with no cross-session token sharing.

Skills are on-demand context injection. Skills (domain-specific .md files referenced from CLAUDE.md) are loaded only when invoked, not at session start. This is the guide's recommended way to handle large knowledge bases — load only what the current task requires.

Why it matters for AI engineering teams

The Claude Code best-practices guide is effectively a token consumption architecture document disguised as a workflow guide. Every recommendation can be reframed in cost terms:

Verification reduces re-runs. The guide's highest-leverage recommendation is to give Claude a way to verify its own work — tests, expected outputs, linter results. The routing implication: without verification loops, failed sessions require full restarts, each consuming a fresh context window from scratch. A team that runs ten Claude Code sessions to produce one good outcome is paying for ten context budgets per result, not one.

Plan mode is a cheap pre-read. Plan mode prevents Claude from making edits during exploration. From a token perspective, this matters because it avoids the accumulation of tool-call responses (file writes, command confirmations) that expand the context faster than pure reads do. Teams can encourage plan-mode-first as a cost discipline.

Parallel sessions need separate accounting. Most API billing shows aggregate token usage. If your team runs Claude Code with sub-agents, your gateway telemetry needs to attribute token spend per session, not per request. A single user who launches five parallel Claude Code agents generates five simultaneous token streams — each invisible individually unless your observability layer segments by session ID.

CLAUDE.md drift is a hidden cost. If engineers add to CLAUDE.md without pruning, every session startup becomes more expensive. Teams that manage Claude Code at scale should treat CLAUDE.md as a configuration artifact that has a token cost per session — and audit it the way they would a dependency list.

The router/operator angle

For teams using a routing gateway to route Claude Code's API calls:

Context-triggered fallback is a distinct need. Standard routing fallback logic triggers on provider errors, rate limits, or latency thresholds. Claude Code introduces a new trigger to consider: context window exhaustion. When a session's context fills and the model starts degrading, the correct response may not be to retry the same model — it may be to route to a model with a larger context window, or to signal the harness to start a fresh session. This requires the gateway to be context-window-aware, not just error-aware.

Token cost modeling for agentic sessions differs from chat sessions. Chat API calls are typically bounded — a user sends a message, gets a response. Claude Code sessions are unbounded by design: the agent reads files, runs commands, iterates, and can run for minutes or hours. A $0.003/1K-token model that looks cheap in a chat context can generate a $2–$5 bill per coding session if the agent explores a large codebase before implementing. Teams that budget on per-request cost assumptions will undercount actual spend significantly.

Session observability needs to track context fill rate, not just request count. The guide recommends a custom status line that shows context fill percentage continuously. The gateway-level equivalent is recording context window utilization per session as a time series, not just total tokens per request. This data is what enables you to catch runaway sessions before they drain quota — and to alert when a session crosses 70% context fill, where performance begins to degrade.

Parallel agent workflows need per-agent billing attribution. If your team gives each developer a personal API key for Claude Code, you get per-developer spend but not per-session or per-task attribution. If your team routes through a shared gateway key, you get aggregate spend but lose individual attribution. The correct setup is a gateway that segments by both user identity and session ID — so you can answer "which task consumed the most context last week" without manual log parsing.

CLAUDE.md routing rules. The guide's CLAUDE.md pattern has a direct gateway analog: system prompt management. Teams that configure Claude Code with environment-specific CLAUDE.md files (production vs. staging, frontend vs. backend monorepo) are effectively setting different system prompts for different contexts. If you route those sessions through a gateway, the gateway should treat CLAUDE.md-derived configurations as a policy dimension, not just as invisible user-side configuration.

What to watch and try

  • Measure context fill per session, not just tokens per request. If your gateway logs show total tokens but not context-window utilization within sessions, you are missing the primary cost driver for agentic workloads. Add session ID tagging and track context fill as a metric.

  • Set up per-session spend alerts before parallel workflows scale. The guide endorses parallel sub-agents explicitly. Before your team starts running five agents simultaneously, set gateway-level per-user or per-session spend alerts so you catch runaway sessions before they exhaust daily quota.

  • Audit CLAUDE.md token cost quarterly. Run a rough calculation: CLAUDE.md lines × average sessions per day × token cost per token. If this exceeds 5% of total spend, it is worth a pruning pass. The guide recommends keeping CLAUDE.md to only what Claude cannot infer — treat that as a cost discipline, not just a clarity one.

  • Configure a context-overflow routing policy. If your routing layer supports it, define what should happen when a Claude Code session requests a context window that exceeds the model's limit — whether that triggers a model upgrade to a larger context variant, a session reset, or an alert. Context overflow handled gracefully is invisible to the developer; handled badly, it produces silent quality degradation that is hard to debug.

TheRouter routes model API calls and records token-level usage across sessions. For Claude Code workflows, the key telemetry additions are session ID tracking and context-window utilization per session — so that the token cost model reflects the agentic reality that a single developer task can span dozens of API calls within one context budget. Teams using TheRouter can use per-key or per-tag spend tracking to approximate per-session attribution until native session-ID metadata is available.

Help & contact