Claude Haiku 5.5 Ships Effort Controls and Tiered Context Pricing for High-Volume Subagent Work
Anthropic's Claude Haiku 5.5 is the first Haiku with effort controls, priced at $0.10 per million input tokens for prompts up to 100K tokens—built for classification, subagent delegation, and real-time workloads at scale.
Produced with AI assistance from the cited sources and published after automated editorial checks; not individually reviewed by an editor. Editor of record: Joe Werner.

Claude Haiku 5.5, announced by Anthropic on October 7, 2026, is Anthropic's fastest and most efficient small model in the current Claude generation. It is also the first Haiku generation to ship with effort controls—a feature that lets operators trade inference cost against reasoning depth per request.
What changed from Haiku 4.5
Haiku 5.5 is described by Anthropic as a significant step up over Haiku 4.5 across knowledge work, computer use, multidisciplinary reasoning, agentic coding, and visual reasoning, according to the Anthropic announcement. One customer running 8 million document-question calls per week in production reported a score of 0.84 versus 0.76 for Haiku 4.5 on the same query set. Another reported over 30% lower task-completion latency and up to 2.5x faster inference per agent turn compared with the model previously in use.
The model supports a 1 million token context window with a 128K output limit. Prompts up to 100K tokens are priced at $0.10 per million input tokens and $0.50 per million output tokens; prompts that exceed 100K tokens—including cache reads and writes toward that threshold—are billed at the long-context tier.
Effort controls and subagent patterns
Effort controls are the notable addition at the API level. Teams can now set reasoning intensity per call rather than choosing a different model, which makes Haiku 5.5 more useful as a subagent within larger pipelines. Anthropic's announcement notes that Haiku 5.5 works alongside larger Claude models, making it practical to add summarization, classification, and compaction to complex products and agent systems.
For operators running multi-model pipelines through TheRouter, claude-haiku-5-5 is available via the API and can be combined with claude-sonnet-5-5 or claude-opus-5-5 in the same request flow. The model handles text and image input with text output, and supports tools, tool_choice, response_format, and the reasoning parameter.
Pricing tier and context window behavior
The tiered pricing structure is worth noting when sizing pipelines. A single long-document call that exceeds 100K tokens—counting prompt, cached sections, and any tool results—shifts from the standard rate to the long-context rate for the full call. Short-context workloads such as classification, intent detection, and single-turn extraction stay at the lower tier.
Routing operators can use TheRouter's pricing reference to compare token costs across models before committing a workload. The Anthropic provider page lists the full set of models available through the gateway.
Availability
Haiku 5.5 is available now via the Claude Platform API using the model ID claude-haiku-5-5, and through Amazon Web Services, Google Cloud, and Microsoft Azure. Effort controls are configurable per request.
TheRouter's route for anthropic/claude-haiku-5-5 was confirmed live on 2026-10-11, returning a response in 3210 ms consuming 91 total tokens on a test prompt.
Models covered in this article

Claude Sonnet 5.5 Ships Five Breaking API Changes for Operators Running Sonnet 5
Claude Sonnet 5.5 (Sep 28) ships 5 breaking changes: forced tool use returns 400, `thinking:disabled` is rejected, thinking blocks are model-bound, `computer_20251124` is gone on the Claude API, and advisor pairings are now 5.x-only.

A Model-Level 429 From Zhipu Took GLM Flash Models Offline. TheRouter Now Scopes Cooldowns Per Model and Gives Them a Second Route.
On 2026-10-02 Zhipu's API returned a model-level rate-limit response for glm-4.6v-flash and TheRouter parked the whole account. Cooldowns are now per model, and six GLM flash-family models have a second route through Zhipu's international service.

Claude Code 2.1.281: Bedrock Upstreams Get Cross-Account IAM and Guardrail Enforcement
2.1.281 adds assume_role and guardrail to Bedrock upstreams. assume_role swaps long-lived IAM credentials for per-developer STS tokens. guardrail applies a Bedrock guardrail to every request. Both shift the trust boundary in multi-account AWS deployments.