返回模型列表

Claude Haiku 5.5

anthropicanthropic/claude-haiku-5-5

Anthropic's Haiku-tier model released 2026-10-07 — built for high-volume, latency-sensitive classification, routing, extraction and subagent tasks; adaptive thinking with effort control; 1M context, 128K output. Requests whose prompt exceeds 100K tokens (cache reads and writes included) are priced at the long-context tier. Prompt caching (cache_control) is not yet supported for this model on TheRouter — prompt-length-tiered models are excluded from the caching allowlist, so cache_control markers are stripped before the upstream call.

Claude Haiku 5.5 是 Anthropic 的 Haiku 级模型,是 Claude Haiku 4.5 的后继型号。Anthropic 称其“面向分类、路由、抽取和子代理等高吞吐、对延迟敏感的任务”,并支持由 effort 参数控制的自适应思考。在 TheRouter 中以 anthropic/claude-haiku-5-5 收录,支持文本与图像输入、文本输出、100 万 token 上下文窗口以及最高 128K 输出 token。

与当前其他所有 Claude 模型不同,该模型按提示长度计价:提示超过 100,000 token 的请求在每个计费维度上都按更高档位计费,且提示长度包含全部输入 token(含缓存读取与缓存写入)。每个请求单独分档。TheRouter 采用同一规则——档位由上游上报的含缓存提示总量决定——因此一段长的多轮会话一旦上下文越过阈值就会升档。与其他 Claude 5.5 模型一样,它使用新版分词器,同样的文本比 Haiku 4.5 多出约 30% 的 token。

适合使用
  • • 高吞吐的分类、路由、抽取和子代理步骤——这是 Anthropic 为该模型定位的工作负载
  • • 对延迟敏感、但仍需 100 万上下文与最高 128K 输出的路径(Anthropic 将其列为当前阵容中最快的模型)
不适合使用
  • • 经常超过 100K token 的提示——含缓存的提示一旦越过阈值,每个维度都会按更高档位计价;上线前先与 Claude Sonnet 5.5 的单请求成本比较
  • • 依赖提示缓存断点的工作负载——TheRouter 目前不会为该模型转发 cache_control 标记(其按提示长度分档的定价尚未纳入缓存计费允许列表),因此经由该路由不会产生缓存读写
  • • 设置手动思考预算或预填 assistant 回合的客户端——Anthropic 对该模型的这两种请求均返回 400;请改用 effort 等级,并以 user 回合结束 messages
上下文长度
1M
最大输出
128K
输入价格每百万 tokens
$0.108每百万 Tokens
输出价格每百万 tokens
$0.540每百万 Tokens

模态能力

文本图像→文本

能力

视觉理解

价格明细

类型费率
提示不超过 100K tokens
输入$0.108 每百万 Tokens
输出$0.540 每百万 Tokens
缓存输入$0.0108 每百万 Tokens
缓存写入(5 分钟)$0.135 每百万 Tokens
缓存写入(1 小时)$0.216 每百万 Tokens
提示超过 100K tokens
输入$0.540 每百万 Tokens
输出$2.70 每百万 Tokens
缓存输入$0.054 每百万 Tokens
缓存写入$0.675 每百万 Tokens

分档计价:当一次请求的提示 tokens(输入,含缓存命中)超过 100K 时,该请求的全部 tokens 均按长上下文费率计费;恰好 100K 仍按基础费率。

支持参数

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatreasoningstop

模型规格

Anthropic 模型 idclaude-haiku-5-5platform.claude.com ↗已核实
上下文窗口1,000,000 tokenplatform.claude.com ↗已核实
最大输出128,000 tokenplatform.claude.com ↗已核实
思考自适应,默认开启;默认 effort 为 mediumplatform.claude.com ↗已核实
可靠知识截止2026 年 6 月platform.claude.com ↗已核实
许可Anthropic Usage Policy(闭源,仅 API)www.anthropic.com ↗已核实
定价(输入 / 输出)官方标价:提示不超过 100,000 token 时每百万 token $0.10 / $0.50,超过 100,000 token 时 $0.50 / $2.50(提示长度包含缓存读写);缓存读取按同样两档为 $0.01 / $0.05platform.claude.com ↗已核实

基准成绩

BenchmarkDistributionScoreSource
Official benchmark table
本次整理时,引用的模型文档未公布基准测试表。
—未公开披露—

API 使用示例

所有新集成都应使用下方示例中的全球端点 api.therouter.ai;旧中国加速端点已下线。

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "anthropic/claude-haiku-5-5",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

API 使用指南

Chat 调用

使用 OpenAI 兼容的 chat completions 端点。希望按基础档计费时,请将提示控制在 100K token 以内。

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-haiku-5-5","messages":[{"role":"user","content":"Hello"}]}'

anthropic 其他模型

同类模型

跨供应商的相似能力档位

相关新闻

常见问题

Claude Haiku 5.5 的请求何时会进入更高的价格档?

当提示超过 100,000 token 时——计数包含所有输入 token,含缓存读取与缓存写入。此时整个请求(输入、输出和缓存 token)都按更高档位计价;恰好 100,000 token 的请求仍按基础档计费。TheRouter 以上游上报的含缓存提示总量来判定阈值。

TheRouter 编辑重写
可以在 Claude Haiku 5.5 上设置 budget_tokens 或预填 assistant 回合吗?

不可以。对该模型,Anthropic 对手动思考预算和 assistant 消息预填均返回 400。请用 effort 等级控制思考深度,并以 user 回合结束 messages。

TheRouter 编辑重写
事实档案 — 本页每条断言可在此回溯来源
来源URL采集于
Anthropic 模型 idplatform.claude.com ↗2026-10-10已核实
上下文窗口platform.claude.com ↗2026-10-10已核实
最大输出platform.claude.com ↗2026-10-10已核实
思考platform.claude.com ↗2026-10-10已核实
可靠知识截止platform.claude.com ↗2026-10-10已核实
许可www.anthropic.com ↗2026-10-10已核实
定价(输入 / 输出)platform.claude.com ↗2026-10-10已核实
Official benchmark tableplatform.claude.com ↗2026-10-10未知
帮助与联系