Kimi K2.6 Is Gone from NVIDIA NIM Today: Your Three Migration Paths to K2.7 Code

NVIDIA NIM shut down its Kimi K2.6 endpoint on July 7, 2026. If your routing config points to the NIM API for Kimi K2.6, requests are now failing. Here are the three concrete migration paths every operator must evaluate before end of day.

Published via NVIDIA NIM / Moonshot AI

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Clean editorial diagram showing migration paths from NVIDIA NIM Kimi K2.6 to Kimi K2.7 Code direct API and self-hosted routes, with routing policy decision points

As of July 7, 2026, NVIDIA NIM has shut down its hosted Kimi K2.6 endpoint. If your AI gateway or application routes any traffic to https://integrate.api.nvidia.com/v1 with model: "moonshotai/kimi-k2.6", those calls are now returning errors. This is not a gradual sunset — NIM's deprecation notice stated plainly that the API "will no longer be supported after 07/07/2026."

The routing decision you face right now is not just "which replacement model." It is a provider migration: NVIDIA NIM did not add a Kimi K2.7 Code endpoint to replace K2.6. Unlike the K2.5 → K2.6 transition on NIM (where K2.6 was added before K2.5 went dark), there is currently no moonshotai/kimi-k2.7-code endpoint on build.nvidia.com. Your routing layer needs to switch provider, not just model name.

What happened

Kimi K2.6 was released on build.nvidia.com on April 29, 2026 as a hosted trial endpoint. On the same page, NVIDIA published a deprecation notice: this API would be supported only until July 7, 2026. That date is today.

The pattern repeats the K2.5 cycle: NVIDIA hosted Kimi K2.5 on NIM, deprecated it with short notice (10 days in April), did not immediately add K2.6 in its place, and eventually added K2.6. The K2.6 cycle followed a similar short-window pattern but with a more gradual 10-week window. There is currently no confirmed K2.7 Code endpoint on NIM.

Why it matters for AI engineering teams

NVIDIA NIM was a popular backend for Kimi K2.6 for two reasons: it offered a free-tier endpoint for prototyping (no Moonshot account required), and it provided an OpenAI-compatible API under an NVIDIA endpoint URL that some enterprise network policies found easier to allowlist than Chinese-hosted infrastructure.

Teams that relied on NIM for Kimi K2.6 in production or staging must now answer a concrete routing question: where does Kimi K2.6 traffic go next?

The Kimi K2.7 Code model is available on the Moonshot platform API today. It is Moonshot's most capable coding-focused model, with 256K context, always-on thinking-mode reasoning, and a HighSpeed variant that delivers approximately 180 tokens/s (peaking at 260 t/s in short-context scenarios). Compared to K2.6, K2.7 Code is narrower in scope — it is coding-specialized and does not carry K2.6's native multimodal (image/video) inputs. Teams using K2.6 primarily for code generation get a direct upgrade; teams that depended on K2.6's vision capabilities need a different path.

The router/operator angle

Three migration paths, ranked by operator complexity:

Path 1: Moonshot direct API (lowest friction, recommended for most teams) Switch your routing layer's base_url to https://api.moonshot.ai/v1 and model to kimi-k2.7-code (or kimi-k2.7-code-highspeed for latency-sensitive workloads). The Kimi platform API is OpenAI-compatible — the same SDK swap pattern documented in the Moonshot platform docs works here. You will need a Moonshot Platform API key if you do not already have one. Latency profile: different from NIM-hosted inference, so run your benchmarks before hard-routing production traffic.

Path 2: Self-hosted via vLLM or HuggingFace (for teams with on-premise GPU clusters) Kimi K2.7 Code is open-weight under Modified MIT and available on HuggingFace (moonshotai/Kimi-K2.7-Code). If your team already runs vLLM for other open-weight models, this adds a new serve target. vLLM Recipes (recipes.vllm.ai/moonshotai/Kimi-K2.7-Code) documents the setup. This path preserves data-locality guarantees and avoids dependency on any external API, but requires GPU capacity for a 1T MoE model.

Path 3: Fallback to K2.6 via Moonshot direct or alternative providers (if multimodal is required) kimi-k2.6 is still listed and active on the Moonshot platform API (api.moonshot.ai/v1). If your workload depends on K2.6's image/video input capabilities, you can route to K2.6 directly on Moonshot's API while evaluating whether K2.7 Code's text-only scope fits your use case. This is a short-term bridge, not a final state.

Routing policy implications for gateway operators:

  • Update any NIM-specific endpoint URLs immediately — they are returning errors now.
  • If your gateway uses provider: nvidia_nim + model: kimi-k2.6, the model string also needs to change; kimi-k2.7-code does not exist on NIM.
  • Evaluate whether your fallback chain included NIM as a secondary provider for Kimi. If so, that fallback is now broken and will silently absorb retries before failing.
  • Cost note: Moonshot direct API pricing for K2.7 Code is usage-based; the NIM free-tier quota disappears with this migration.

What TheRouter users should watch or try

If your routing configuration included NVIDIA NIM as a backend for any Kimi model, audit it now. The pattern of NIM hosting Kimi temporarily then removing the endpoint without a same-version replacement has now happened twice (K2.5 in April, K2.6 today). Building provider fallback logic that does not hardcode NIM as the primary Kimi route — and that can fail over to Moonshot direct or a self-hosted endpoint — will reduce the impact of future NIM churn cycles.

The practical test: attempt a request to integrate.api.nvidia.com/v1 with model: moonshotai/kimi-k2.6 right now. If you get an error, your NIM route is broken. Route to api.moonshot.ai/v1 with kimi-k2.7-code and validate context and tool-call behavior before switching production traffic.

Models covered in this article

Help & contact