MiMo-V2.5-Pro on DashScope: OpenAI + Anthropic Dual API Compatibility for Agent Routing

Xiaomi MiMo-V2.5-Pro is live on DashScope with 1M-token context, OpenAI and Anthropic API compatibility, and 42% fewer tokens than Kimi K2.6 on agent benchmarks. Evaluate whether this model fits your routing policy—cost, context window opt-in, and DashScope catalog risks.

Published via Xiaomi MiMo Platform

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract routing infrastructure diagram showing multiple provider paths converging through a unified API gateway layer

When a consumer electronics company ships a model that scores competitively against Claude Opus 4.6 and GPT-5.4 on agent benchmarks—and makes it available via a standard OpenAI-compatible endpoint plus DashScope—that is a routing decision, not just a product launch. Xiaomi's MiMo-V2.5-Pro entered public beta on May 22, and AI engineering teams now have a concrete new option to evaluate for complex, long-horizon coding and agent workloads routed through Chinese infrastructure.

What happened

Xiaomi launched the MiMo-V2.5 series into public beta on its official MiMo platform. The lineup includes MiMo-V2.5-Pro (the flagship agent/coding model), MiMo-V2.5 (a full-modal model with image/audio/video understanding), MiMo-V2.5-TTS, and MiMo-V2.5-ASR.

The key engineering facts for routing teams:

  • Model ID on DashScope: mimo-v2.5-pro
  • Context window: 1M tokens (with [1m] suffix in Claude Code config to opt in)
  • API formats: OpenAI-compatible (https://api.xiaomimimo.com) and Anthropic-compatible (https://api.xiaomimimo.com/anthropic)
  • Capabilities: Function calling, structured output, 128K max output tokens, thinking mode
  • DashScope availability: mimo-v2.5-pro is listed in Alibaba Cloud Model Studio alongside Qwen, DeepSeek, Kimi, GLM, and MiniMax as a third-party model accessible via the shared DashScope OpenAI-compatible endpoint

For Claude Code users: Xiaomi publishes official Claude Code integration docs with a settings.json template that maps all model aliases (ANTHROPIC_DEFAULT_SONNET_MODEL, ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_HAIKU_MODEL) to mimo-v2.5-pro.

OpenCode Go—the official OpenCode subscription tier—also added MiMo-V2.5-Pro and MiMo-V2.5 to its curated model list for international users.

Why it matters for AI engineering teams

Three things make this more than a benchmark press release.

Token efficiency at agent scale. Xiaomi claims MiMo-V2.5-Pro uses 42% fewer tokens than Kimi K2.6 to reach the same score on ClawEval, an agent benchmark. At agent workloads where each task involves hundreds of tool calls, token count directly translates into cost and latency. A 42% reduction is not marginal—it shifts the cost-per-task calculation for long-horizon coding agents materially.

Demonstrated long-horizon performance. Xiaomi's benchmark includes completing a full SysY compiler in Rust (672 tool calls, 4.3 hours, 233/233 score on Peking University's compiler test suite) and building an 8,192-line video editor from a single prompt over 1,868 tool calls. These are not cherry-picked toy demos—they are structured tasks with verifiable pass/fail criteria.

DashScope consolidation. Teams already routing to DashScope for Qwen or DeepSeek can now add MiMo-V2.5-Pro to their provider set without adding a new provider account, key management system, or billing surface. The model uses the same compatible-mode/v1 endpoint structure as all other DashScope models.

The router/operator angle

Several routing policy implications follow from this launch.

Dual-format API routing. MiMo-V2.5-Pro supports both OpenAI and Anthropic wire formats. That is the same pattern DeepSeek established earlier this year. For teams using an OpenAI-compatible gateway, this means the model can be reached via standard chat/completions routing. For teams using Claude Code or Anthropic SDK directly, the Anthropic-format endpoint (/anthropic) avoids translation overhead.

Model ID discipline on DashScope. DashScope now serves mimo-v2.5-pro, deepseek-v4-pro, kimi-k2.6, glm-5.1, MiniMax-M2.5, and MiniMax-M2.7 as third-party models alongside the Qwen3 family—all under the same API key and endpoint. This means your routing policy for DashScope needs explicit model-ID pinning; the "best available" heuristic will not give you predictable behavior across this expanded catalog. Alias drift is a real risk when Alibaba updates the default recommended model without notice.

Context window opt-in behavior. The 1M-token context on MiMo-V2.5-Pro is not always active by default—Claude Code users must append [1m] to the model ID in config. This is an unusual pattern: most providers enable their maximum context window by default. If your routing layer normalizes or strips unknown model-ID suffixes, you may silently fall back to a shorter context window. Test this before routing large-codebase jobs to MiMo.

Cost routing framework for Chinese agent models. With MiMo-V2.5-Pro now on DashScope, the third-party model tier on Alibaba Cloud offers five distinct agent-capable models. A practical routing framework:

  • Max quality, highest cost: DeepSeek-V4-Pro or Kimi K2.6 (established benchmark history)
  • Token efficiency for long-horizon tasks: MiMo-V2.5-Pro (42% fewer tokens vs K2.6 per Xiaomi's claim; needs team verification)
  • Multimodal + agent: MiMo-V2.5 (image/audio/video native; 50% fewer tokens vs prior)
  • Batch/cost-optimized text: MiniMax-M2.5 or DeepSeek-V4-Flash

Open-source announcement. Xiaomi explicitly announced that MiMo-V2.5-Pro and MiMo-V2.5 will be globally open-sourced. If that follows through, teams will have the option to self-host or use third-party inference providers—which changes the long-term fallback calculus for this model.

What TheRouter users should watch or try

If your routing policy already includes DashScope-backed models, mimo-v2.5-pro is now available to add as an additional provider route under the same Alibaba Cloud Model Studio credentials. Evaluate it specifically for long-horizon coding tasks where token count per task is high—that is where the claimed efficiency advantage is most measurable.

Key things to verify before committing it to production routing:

  1. Confirm the [1m] suffix behavior in your specific gateway or client—if your stack strips model-ID suffixes, test whether the effective context window is 256K or 1M.
  2. Benchmark token efficiency yourself. Xiaomi's 42% savings claim is against Kimi K2.6 on ClawEval. Your workload distribution will differ. Run a representative sample before adjusting cost-based routing thresholds.
  3. Watch for the open-source release. When model weights are published, third-party inference providers and SiliconFlow may offer lower-cost or lower-latency routes than DashScope or Xiaomi's own API.
  4. Monitor DashScope model catalog changes. Alibaba updates the third-party model list on Model Studio without consistent advance notice. Pin your model IDs explicitly and set up changelog monitoring for help.aliyun.com/zh/model-studio/text-generation-model/.

For Claude Code users routing through an OpenAI-compatible gateway: the Anthropic-format endpoint (https://api.xiaomimimo.com/anthropic) means you can configure MiMo-V2.5-Pro as a drop-in backend without modifying your Claude Code SDK setup—the same pattern that now works for DeepSeek.

Models covered in this article

Help & contact