Claude Haiku 5.5
Anthropic's Haiku-tier model released 2026-10-07 — built for high-volume, latency-sensitive classification, routing, extraction and subagent tasks; adaptive thinking with effort control; 1M context, 128K output. Requests whose prompt exceeds 100K tokens (cache reads and writes included) are priced at the long-context tier. Prompt caching (cache_control) is not yet supported for this model on TheRouter — prompt-length-tiered models are excluded from the caching allowlist, so cache_control markers are stripped before the upstream call.
Claude Haiku 5.5 是 Anthropic 的 Haiku 级模型,是 Claude Haiku 4.5 的后继型号。Anthropic 称其“面向分类、路由、抽取和子代理等高吞吐、对延迟敏感的任务”,并支持由 effort 参数控制的自适应思考。在 TheRouter 中以 anthropic/claude-haiku-5-5 收录,支持文本与图像输入、文本输出、100 万 token 上下文窗口以及最高 128K 输出 token。
与当前其他所有 Claude 模型不同,该模型按提示长度计价:提示超过 100,000 token 的请求在每个计费维度上都按更高档位计费,且提示长度包含全部输入 token(含缓存读取与缓存写入)。每个请求单独分档。TheRouter 采用同一规则——档位由上游上报的含缓存提示总量决定——因此一段长的多轮会话一旦上下文越过阈值就会升档。与其他 Claude 5.5 模型一样,它使用新版分词器,同样的文本比 Haiku 4.5 多出约 30% 的 token。
- • 高吞吐的分类、路由、抽取和子代理步骤——这是 Anthropic 为该模型定位的工作负载
- • 对延迟敏感、但仍需 100 万上下文与最高 128K 输出的路径(Anthropic 将其列为当前阵容中最快的模型)
- • 经常超过 100K token 的提示——含缓存的提示一旦越过阈值,每个维度都会按更高档位计价;上线前先与 Claude Sonnet 5.5 的单请求成本比较
- • 依赖提示缓存断点的工作负载——TheRouter 目前不会为该模型转发
cache_control标记(其按提示长度分档的定价尚未纳入缓存计费允许列表),因此经由该路由不会产生缓存读写 - • 设置手动思考预算或预填 assistant 回合的客户端——Anthropic 对该模型的这两种请求均返回 400;请改用 effort 等级,并以 user 回合结束
messages
模态能力
能力
价格明细
| 类型 | 费率 |
|---|---|
| 提示不超过 100K tokens | |
| 输入 | $0.108 每百万 Tokens |
| 输出 | $0.540 每百万 Tokens |
| 缓存输入 | $0.0108 每百万 Tokens |
| 缓存写入(5 分钟) | $0.135 每百万 Tokens |
| 缓存写入(1 小时) | $0.216 每百万 Tokens |
| 提示超过 100K tokens | |
| 输入 | $0.540 每百万 Tokens |
| 输出 | $2.70 每百万 Tokens |
| 缓存输入 | $0.054 每百万 Tokens |
| 缓存写入 | $0.675 每百万 Tokens |
分档计价:当一次请求的提示 tokens(输入,含缓存命中)超过 100K 时,该请求的全部 tokens 均按长上下文费率计费;恰好 100K 仍按基础费率。
支持参数
模型规格
| Anthropic 模型 id | claude-haiku-5-5platform.claude.com ↗ | 已核实 |
| 上下文窗口 | 1,000,000 tokenplatform.claude.com ↗ | 已核实 |
| 最大输出 | 128,000 tokenplatform.claude.com ↗ | 已核实 |
| 思考 | 自适应,默认开启;默认 effort 为 mediumplatform.claude.com ↗ | 已核实 |
| 可靠知识截止 | 2026 年 6 月platform.claude.com ↗ | 已核实 |
| 许可 | Anthropic Usage Policy(闭源,仅 API)www.anthropic.com ↗ | 已核实 |
| 定价(输入 / 输出) | 官方标价:提示不超过 100,000 token 时每百万 token $0.10 / $0.50,超过 100,000 token 时 $0.50 / $2.50(提示长度包含缓存读写);缓存读取按同样两档为 $0.01 / $0.05platform.claude.com ↗ | 已核实 |
基准成绩
| Benchmark | Distribution | Score | Source |
|---|---|---|---|
Official benchmark table 本次整理时,引用的模型文档未公布基准测试表。 | — | 未公开披露 | — |
API 使用示例
所有新集成都应使用下方示例中的全球端点 api.therouter.ai;旧中国加速端点已下线。
curl https://api.therouter.ai/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $THE_ROUTER_API_KEY" -d '{
"model": "anthropic/claude-haiku-5-5",
"messages": [
{"role": "user", "content": "Summarize the key points from this input."}
]
}'API 使用指南
Chat 调用
使用 OpenAI 兼容的 chat completions 端点。希望按基础档计费时,请将提示控制在 100K token 以内。
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"anthropic/claude-haiku-5-5","messages":[{"role":"user","content":"Hello"}]}'anthropic 其他模型
同类模型
跨供应商的相似能力档位相关新闻
常见问题
Claude Haiku 5.5 的请求何时会进入更高的价格档?
当提示超过 100,000 token 时——计数包含所有输入 token,含缓存读取与缓存写入。此时整个请求(输入、输出和缓存 token)都按更高档位计价;恰好 100,000 token 的请求仍按基础档计费。TheRouter 以上游上报的含缓存提示总量来判定阈值。
可以在 Claude Haiku 5.5 上设置 budget_tokens 或预填 assistant 回合吗?
不可以。对该模型,Anthropic 对手动思考预算和 assistant 消息预填均返回 400。请用 effort 等级控制思考深度,并以 user 回合结束 messages。
事实档案 — 本页每条断言可在此回溯来源
| 来源 | URL | 采集于 | |
|---|---|---|---|
| Anthropic 模型 id | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| 上下文窗口 | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| 最大输出 | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| 思考 | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| 可靠知识截止 | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| 许可 | www.anthropic.com ↗ | 2026-10-10 | 已核实 |
| 定价(输入 / 输出) | platform.claude.com ↗ | 2026-10-10 | 已核实 |
| Official benchmark table | platform.claude.com ↗ | 2026-10-10 | 未知 |