Claude Haiku 5.5
Anthropic's Haiku-tier model released 2026-10-07 β built for high-volume, latency-sensitive classification, routing, extraction and subagent tasks; adaptive thinking with effort control; 1M context, 128K output. Requests whose prompt exceeds 100K tokens (cache reads and writes included) are priced at the long-context tier. Prompt caching (cache_control) is not yet supported for this model on TheRouter β prompt-length-tiered models are excluded from the caching allowlist, so cache_control markers are stripped before the upstream call.
Claude Haiku 5.5 is Anthropic's Haiku-tier model and the successor to Claude Haiku 4.5. Anthropic describes it as "built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks", with adaptive thinking steered by the effort parameter. In TheRouter it is cataloged as anthropic/claude-haiku-5-5, with text and image input, text output, a 1M-token context window and up to 128K output tokens.
Unlike every other current Claude model, this one is priced by prompt length: a request whose prompt is over 100,000 tokens is billed at a higher rate on every dimension, and the prompt length counts all input tokens including cache reads and cache writes. Each request is tiered on its own. TheRouter applies the same rule β the tier is decided on the cache-inclusive prompt total the upstream reports β so a long multi-turn session steps up as soon as its context crosses the threshold. Like the other Claude 5.5 models, it uses the newer tokenizer, so the same text counts roughly 30% more tokens than on Haiku 4.5.
- β’ High-volume classification, routing, extraction and subagent steps, which is the workload Anthropic positions this model for
- β’ Latency-sensitive paths where a 1M context and up to 128K output are still needed (Anthropic lists it as the fastest model in the current lineup)
- β’ Prompts that routinely exceed 100K tokens β every dimension reprices at the higher tier once the cache-inclusive prompt crosses the threshold; compare the per-request cost against Claude Sonnet 5.5 before committing
- β’ Workloads that rely on prompt-cache breakpoints β TheRouter does not forward
cache_controlmarkers for this model yet (its prompt-length tier is not covered by the cache-billing allow-list), so cache reads and writes are not created through this route - β’ Clients that set a manual thinking budget or prefill the assistant turn β Anthropic rejects both with a 400 on this model; use the effort level and end
messageswith a user turn
Modalities
Capabilities
Pricing Breakdown
| Type | Rate |
|---|---|
| Prompts up to 100K tokens | |
| Input | $0.108 / 1M tokens |
| Output | $0.540 / 1M tokens |
| Cached input | $0.0108 / 1M tokens |
| Cache write (5m) | $0.135 / 1M tokens |
| Cache write (1h) | $0.216 / 1M tokens |
| Prompts over 100K tokens | |
| Input | $0.540 / 1M tokens |
| Output | $2.70 / 1M tokens |
| Cached input | $0.054 / 1M tokens |
| Cache write | $0.675 / 1M tokens |
Tiered pricing: when a request's prompt (input, cached tokens included) is over 100K tokens, every token in that request is billed at the long-context rates; exactly 100K is still the base rate.
Supported Parameters
Specifications
| Anthropic model id | claude-haiku-5-5platform.claude.com β | verified |
| Context window | 1,000,000 tokensplatform.claude.com β | verified |
| Maximum output | 128,000 tokensplatform.claude.com β | verified |
| Thinking | Adaptive, on by default; default effort mediumplatform.claude.com β | verified |
| Reliable knowledge cutoff | June 2026platform.claude.com β | verified |
| License | Anthropic Usage Policy (proprietary, API-only)www.anthropic.com β | verified |
| Pricing (input / output) | List price $0.10 / $0.50 per MTok for prompts up to 100,000 tokens and $0.50 / $2.50 per MTok for prompts over 100,000 tokens (prompt length counts cache reads and writes); cache read $0.01 / $0.05 by the same bandsplatform.claude.com β | verified |
Benchmarks
| Benchmark | Distribution | Score | Source |
|---|---|---|---|
Official benchmark table The cited model documentation publishes no benchmark table at curation time. | β | Not publicly disclosed | β |
API Usage Examples
Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.
curl https://api.therouter.ai/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $THE_ROUTER_API_KEY" -d '{
"model": "anthropic/claude-haiku-5-5",
"messages": [
{"role": "user", "content": "Summarize the key points from this input."}
]
}'API guide
Chat completion
Use the OpenAI-compatible chat completions endpoint. Keep prompts under 100K tokens where the base tier is intended.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"anthropic/claude-haiku-5-5","messages":[{"role":"user","content":"Hello"}]}'More from anthropic
Similar models
Cross-provider sibling modelsRelated news
Frequently asked
When does a Claude Haiku 5.5 request move to the higher price tier?
When the prompt is over 100,000 tokens, counting every input token including cache reads and cache writes. The whole request β input, output and cached tokens β is then priced at the higher tier; a request at exactly 100,000 tokens stays on the base tier. TheRouter evaluates the threshold on the cache-inclusive prompt total the upstream reports.
Can I set budget_tokens or prefill the assistant turn on Claude Haiku 5.5?
No. Anthropic returns a 400 for a manual thinking budget and for an assistant-message prefill on this model. Use the effort level to steer thinking depth, and end messages with a user turn.
Fact ledger β every claim on this page traces here
| source | URL | retrieved | |
|---|---|---|---|
| Anthropic model id | platform.claude.com β | 2026-10-10 | verified |
| Context window | platform.claude.com β | 2026-10-10 | verified |
| Maximum output | platform.claude.com β | 2026-10-10 | verified |
| Thinking | platform.claude.com β | 2026-10-10 | verified |
| Reliable knowledge cutoff | platform.claude.com β | 2026-10-10 | verified |
| License | www.anthropic.com β | 2026-10-10 | verified |
| Pricing (input / output) | platform.claude.com β | 2026-10-10 | verified |
| Official benchmark table | platform.claude.com β | 2026-10-10 | unknown |