Back to Models

Claude Haiku 5.5

anthropicanthropic/claude-haiku-5-5

Anthropic's Haiku-tier model released 2026-10-07 β€” built for high-volume, latency-sensitive classification, routing, extraction and subagent tasks; adaptive thinking with effort control; 1M context, 128K output. Requests whose prompt exceeds 100K tokens (cache reads and writes included) are priced at the long-context tier. Prompt caching (cache_control) is not yet supported for this model on TheRouter β€” prompt-length-tiered models are excluded from the caching allowlist, so cache_control markers are stripped before the upstream call.

Claude Haiku 5.5 is Anthropic's Haiku-tier model and the successor to Claude Haiku 4.5. Anthropic describes it as "built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks", with adaptive thinking steered by the effort parameter. In TheRouter it is cataloged as anthropic/claude-haiku-5-5, with text and image input, text output, a 1M-token context window and up to 128K output tokens.

Unlike every other current Claude model, this one is priced by prompt length: a request whose prompt is over 100,000 tokens is billed at a higher rate on every dimension, and the prompt length counts all input tokens including cache reads and cache writes. Each request is tiered on its own. TheRouter applies the same rule β€” the tier is decided on the cache-inclusive prompt total the upstream reports β€” so a long multi-turn session steps up as soon as its context crosses the threshold. Like the other Claude 5.5 models, it uses the newer tokenizer, so the same text counts roughly 30% more tokens than on Haiku 4.5.

Best for
  • β€’ High-volume classification, routing, extraction and subagent steps, which is the workload Anthropic positions this model for
  • β€’ Latency-sensitive paths where a 1M context and up to 128K output are still needed (Anthropic lists it as the fastest model in the current lineup)
Reach for something else if
  • β€’ Prompts that routinely exceed 100K tokens β€” every dimension reprices at the higher tier once the cache-inclusive prompt crosses the threshold; compare the per-request cost against Claude Sonnet 5.5 before committing
  • β€’ Workloads that rely on prompt-cache breakpoints β€” TheRouter does not forward cache_control markers for this model yet (its prompt-length tier is not covered by the cache-billing allow-list), so cache reads and writes are not created through this route
  • β€’ Clients that set a manual thinking budget or prefill the assistant turn β€” Anthropic rejects both with a 400 on this model; use the effort level and end messages with a user turn
Context Length
1M
Max Output
128K
Input Priceper 1M tokens
$0.108/ 1M tokens
Output Priceper 1M tokens
$0.540/ 1M tokens

Modalities

textimage→text

Capabilities

Vision

Pricing Breakdown

TypeRate
Prompts up to 100K tokens
Input$0.108 / 1M tokens
Output$0.540 / 1M tokens
Cached input$0.0108 / 1M tokens
Cache write (5m)$0.135 / 1M tokens
Cache write (1h)$0.216 / 1M tokens
Prompts over 100K tokens
Input$0.540 / 1M tokens
Output$2.70 / 1M tokens
Cached input$0.054 / 1M tokens
Cache write$0.675 / 1M tokens

Tiered pricing: when a request's prompt (input, cached tokens included) is over 100K tokens, every token in that request is billed at the long-context rates; exactly 100K is still the base rate.

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatreasoningstop

Specifications

Anthropic model idclaude-haiku-5-5platform.claude.com β†—verified
Context window1,000,000 tokensplatform.claude.com β†—verified
Maximum output128,000 tokensplatform.claude.com β†—verified
ThinkingAdaptive, on by default; default effort mediumplatform.claude.com β†—verified
Reliable knowledge cutoffJune 2026platform.claude.com β†—verified
LicenseAnthropic Usage Policy (proprietary, API-only)www.anthropic.com β†—verified
Pricing (input / output)List price $0.10 / $0.50 per MTok for prompts up to 100,000 tokens and $0.50 / $2.50 per MTok for prompts over 100,000 tokens (prompt length counts cache reads and writes); cache read $0.01 / $0.05 by the same bandsplatform.claude.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
Official benchmark table
The cited model documentation publishes no benchmark table at curation time.
β€”Not publicly disclosedβ€”

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "anthropic/claude-haiku-5-5",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

API guide

Chat completion

Use the OpenAI-compatible chat completions endpoint. Keep prompts under 100K tokens where the base tier is intended.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-haiku-5-5","messages":[{"role":"user","content":"Hello"}]}'

More from anthropic

Similar models

Cross-provider sibling models

Related news

Frequently asked

When does a Claude Haiku 5.5 request move to the higher price tier?

When the prompt is over 100,000 tokens, counting every input token including cache reads and cache writes. The whole request β€” input, output and cached tokens β€” is then priced at the higher tier; a request at exactly 100,000 tokens stays on the base tier. TheRouter evaluates the threshold on the cache-inclusive prompt total the upstream reports.

re-authored by TheRouter
Can I set budget_tokens or prefill the assistant turn on Claude Haiku 5.5?

No. Anthropic returns a 400 for a manual thinking budget and for an assistant-message prefill on this model. Use the effort level to steer thinking depth, and end messages with a user turn.

re-authored by TheRouter
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Anthropic model idplatform.claude.com β†—2026-10-10verified
Context windowplatform.claude.com β†—2026-10-10verified
Maximum outputplatform.claude.com β†—2026-10-10verified
Thinkingplatform.claude.com β†—2026-10-10verified
Reliable knowledge cutoffplatform.claude.com β†—2026-10-10verified
Licensewww.anthropic.com β†—2026-10-10verified
Pricing (input / output)platform.claude.com β†—2026-10-10verified
Official benchmark tableplatform.claude.com β†—2026-10-10unknown
Help & contact