Back to Models

Kimi K2.7 Code Highspeed

moonshotmoonshot/kimi-k2.7-code-highspeed

Moonshot AI Kimi K2.7 Code Highspeed β€” accelerated inference variant of K2.7 Code. Same coding-specialist capabilities with higher throughput. 256K context.

Kimi K2.7 Code Highspeed is Moonshot AI's accelerated inference SKU for the K2.7 Code coding-specialist model. It keeps the same text-only coding-agent profile as K2.7 Code on TheRouter β€” 256K context, 65,536 maximum completion tokens, tool calling, and structured-output support β€” but prices and routes it as the latency-oriented sibling.

Use this route when K2.7 Code quality is needed but interactive agent latency is the bottleneck. The tradeoff is explicit: TheRouter's catalog prices Highspeed at 2Γ— the standard K2.7 Code rate, so it should be reserved for human-in-the-loop coding, fast review loops, and agent orchestration where seconds saved matter more than unit cost.

Best for
  • β€’ Latency-sensitive coding agents where K2.7 Code's reasoning and tool-use quality must return faster
  • β€’ Human-in-the-loop code review, patch explanation, and refactor planning where response speed changes the workflow
  • β€’ Agent orchestration systems that can route only the time-critical turns to Highspeed while leaving background turns on the cheaper standard K2.7 Code route
Reach for something else if
  • β€’ Bulk offline refactoring, nightly code analysis, or batch documentation where 2Γ— pricing does not buy user-visible latency gains
  • β€’ Vision or multimodal requests on TheRouter β€” this Highspeed route is cataloged as text input to text output
  • β€’ Simple low-cost tool dispatch where a smaller non-thinking model can satisfy the turn

How TheRouter serves this differently from the vendor

As the vendor operates it

Moonshot presents Highspeed as part of the Kimi K2.7 Code platform family for faster coding-agent inference.

On TheRouter

TheRouter exposes moonshot/kimi-k2.7-code-highspeed as its own text-to-text chat-completions model id with 256K context, 65,536 max completion tokens, and 2Γ— catalog pricing versus moonshot/kimi-k2.7-code.

Context Length
256K
Max Output
66K
Input Priceper 1M tokens
$2.05/ 1M tokens
Output Priceper 1M tokens
$8.64/ 1M tokens

Modalities

text→text

Pricing Breakdown

TypeRate
Input$2.05 / 1M tokens
Output$8.64 / 1M tokens

Moonshot prices -highspeed as a separate SKU at 2x the standard K2.7-code rate (platform.kimi.ai/docs/pricing/chat-k27-code, international USD card, read 2026-07-29).

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

Release date2026-06-12platform.kimi.ai β†—verified
Model familyKimi K2.7 Code accelerated Highspeed variantplatform.kimi.ai β†—verified
Context window256,000 tokensplatform.kimi.ai β†—verified
Maximum completion65,536 tokensplatform.kimi.ai β†—verified
Catalog pricing2Γ— standard K2.7 Code rate on TheRouterplatform.kimi.ai β†—verified
Training data cutoffNot publicly disclosedunknown

Benchmarks

BenchmarkDistributionScoreSource
Kimi Code Bench v2
Same model-quality basis as K2.7 Code; Highspeed is an inference SKU, not a separate benchmark release.
62.0%platform.kimi.ai β†—
Program Bench
Moonshot reports this for K2.7 Code; use it for quality comparison, then use Highspeed only when latency matters.
53.6%platform.kimi.ai β†—
MCP Mark Verified
Moonshot-reported MCP tool invocation score for the K2.7 Code model family.
81.1%platform.kimi.ai β†—
Highspeed throughput target
Throughput target for the accelerated Highspeed route; compare cost-normalized latency, not only benchmark quality.
180 tokens/sectokens/secplatform.kimi.ai β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "moonshot/kimi-k2.7-code-highspeed",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

Use the dedicated Highspeed model id when latency justifies the higher price.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"moonshot/kimi-k2.7-code-highspeed","messages":[{"role":"user","content":"Review this patch for release-blocking bugs."}]}'

More from moonshot

Similar models

Cross-provider sibling models

News & changes

2026-06-17

Kimi K2.7 Code API: 30% Fewer Thinking Tokens and a HighSpeed Variant Change Your Routing Policy

The routing implication is simple: standard K2.7 Code should carry ordinary deep coding turns, while Highspeed should be reserved for latency-sensitive loops where doubled unit price is justified.

re-authored by TheRouterKimi Open Platform β†—

Frequently asked

When should I choose Highspeed instead of standard K2.7 Code?

Choose Highspeed when a human or orchestrator is waiting on the coding-agent result and lower latency changes the experience. For background analysis, scheduled refactors, and batch generation, use the standard K2.7 Code route to avoid paying the doubled rate.

Does Highspeed change the model's context window or output limit?

No. TheRouter catalog data for moonshot/kimi-k2.7-code-highspeed keeps 256K context and 65,536 maximum completion tokens. The difference is routing, throughput, and price rather than a larger memory window.

Is this route multimodal?

No. TheRouter lists moonshot/kimi-k2.7-code-highspeed as text input to text output. Route screenshots, UI images, or video understanding to a cataloged multimodal Moonshot model such as moonshot/kimi-k2.6.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release dateplatform.kimi.ai β†—2026-08-14verified
Model familyplatform.kimi.ai β†—2026-08-14verified
Context windowplatform.kimi.ai β†—2026-08-14verified
Maximum completionplatform.kimi.ai β†—2026-08-14verified
Catalog pricingplatform.kimi.ai β†—2026-08-14verified
Training data cutoffβ€”β€”unknown
Kimi Code Bench v2platform.kimi.ai β†—2026-08-14verified
Program Benchplatform.kimi.ai β†—2026-08-14verified
MCP Mark Verifiedplatform.kimi.ai β†—2026-08-14verified
Highspeed throughput targetplatform.kimi.ai β†—2026-08-14verified
Kimi K2.7 Code API: 30% Fewer Thinking Tokens and a HighSpeed Variant Change Your Routing PolicyKimi Open Platform β†—2026-08-14verified
When should I choose Highspeed instead of standard K2.7 Code?platform.kimi.ai β†—2026-08-14to verify
Does Highspeed change the model's context window or output limit?platform.kimi.ai β†—2026-08-14to verify
Is this route multimodal?platform.kimi.ai β†—2026-08-14to verify
Help & contact