MiniMax M2.7 API Guide: Dual Format Routing for Claude Code & OpenAI SDK
MiniMax M2.7 delivers 56% SWE-Pro performance near Claude Opus at 1/10th the cost, with both OpenAI and Anthropic API formats — ideal for coding agent routing.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

When a new model ships simultaneously with both OpenAI-compatible and Anthropic-compatible endpoints, that's not a benchmark story — it's a routing architecture decision. MiniMax's M2.7, released today, does exactly that: it adds a credible, low-cost option to both SDK paths that most AI engineering teams are already running, while posting SWE-Pro scores near Claude Opus at a fraction of the price. For teams that route coding and agent workloads through an OpenAI-compatible gateway, the relevant question is not whether M2.7 exists but whether it belongs in your fallback or primary routing slot.
What happened
MiniMax released M2.7 on May 26, 2026 — a new flagship model built around what the company calls "self-evolving" agent architecture. During M2.7's own development, MiniMax used an early version of the model to automate reinforcement learning experiments, iterate on its own scaffolding, and run 100+ autonomous optimization loops. That process produced a model with strong real-world software engineering results:
- SWE-Pro benchmark: 56.22% — near Claude Opus's best level, second to Claude Opus 4.7 and GPT-5.5 in the open-source tier
- VIBE-Pro (end-to-end project delivery): 55.6%
- Terminal Bench 2 (complex engineering system understanding): 57.0%
- MLE-Bench Lite: 66.6% medal rate — tying Gemini-3.1, behind only Opus-4.6 and GPT-5.4
API pricing is $0.30 per million input tokens and $1.20 per million output tokens, with a 205K context window. A high-speed variant (MiniMax-M2.7-highspeed) is also available at lower cost for throughput-sensitive workloads.
The model is live on the MiniMax API Platform and is also listed as a third-party model on Alibaba Cloud's DashScope/百炼 under the model ID MiniMax/MiniMax-M2.7.
Why it matters for AI engineering teams
The dual API format is the most operationally significant detail. MiniMax M2.7 supports:
- OpenAI-compatible API — same
/v1/chat/completionspath, same request structure, direct drop-in for any team using the OpenAI Python SDK, TypeScript SDK, or any OpenAI-compatible gateway - Anthropic-compatible API — MiniMax highlights this as a first-class path on its docs homepage, meaning teams running Claude Code, the Anthropic SDK, or any Anthropic-format proxy can reach M2.7 without code changes
This matters because most routing infrastructure today is built around one of these two SDK paths. A model that fits both paths without adapter work becomes a real routing candidate at near-zero integration cost. The alternative — a model that only speaks one format or requires custom translation — creates SDK-specific lock-in that most teams prefer to avoid.
The pricing position also shifts the cost comparison. At $0.30/$1.20 per million tokens, M2.7 sits at roughly:
- ~1/10th the cost of Claude Opus 4.7 for coding/agent tasks
- Competitive with DeepSeek V4 Pro during its current promotion, but without a promotional expiry date
- Below Kimi K2.6 ($1.20/$4.50), which holds similar long-horizon coding benchmarks
For teams running high-volume agent loops — where per-call cost accumulates rapidly — M2.7's price point combined with Opus-class benchmark scores creates a credible primary or fallback routing slot, particularly for tasks that don't require Opus's full capability headroom.
The router/operator angle
Format routing decision: If your routing layer dispatches on SDK format, you now have a new intersection. MiniMax M2.7 is reachable via either the OpenAI path or the Anthropic path, which means you can target it from Claude Code (via Anthropic-format proxy), from standard OpenAI SDK clients, or from a gateway that normalizes both. This is the same dual-format strategy DeepSeek deployed with its Anthropic API support — and it's becoming the new expectation for serious coding/agent model contenders.
Fallback and cost routing: A routing policy that sends Opus-level tasks to Claude Opus 4.7 first but falls back to M2.7 at 1/10th the price keeps performance headroom without paying full Opus prices on every call. The benchmark gap between Opus and M2.7 is small enough that for many agentic task categories — log analysis, refactoring, code security checks — the fallback will perform acceptably.
DashScope as a secondary routing node: M2.7's presence on Alibaba Cloud DashScope means China-region teams can access it through the same DashScope OpenAI-compatible endpoint they already use for Qwen models, without opening a separate MiniMax API account. The DashScope model ID is MiniMax/MiniMax-M2.7. This simplifies multi-provider routing for teams that consolidate domestic model access through DashScope.
Checklist for routing teams evaluating M2.7:
- Verify which format (OpenAI or Anthropic) your existing routing layer dispatches — M2.7 fits both, but confirm the specific endpoint path in MiniMax's current docs
- Compare your actual task-mix benchmark needs against SWE-Pro and VIBE-Pro scores, not just general intelligence benchmarks
- Evaluate context window fit: M2.7's 205K context may be tighter than Kimi K2.6 (256K) for document-heavy or long-session workloads
- Watch whether the $0.30/$1.20 pricing is an introductory rate — MiniMax has not announced a promotional window, but pricing for new models can change in the first 30–60 days
- Test skill adherence at scale: M2.7 claims 97% adherence across 40+ complex skills (each >2,000 tokens), which is important for multi-tool agent loops but should be validated against your specific tool definitions
What TheRouter users should watch or try
Teams using TheRouter to route OpenAI-compatible requests can add MiniMax's API as a provider configuration and test M2.7 as a fallback target for coding or agentic workloads. The fact that M2.7 also accepts Anthropic-format requests means it's accessible from the same paths used for Claude Code — making it a realistic cost-reduction fallback for Anthropic-format sessions that run into rate limits or pricing ceilings.
The key operational question: on your actual agentic task distribution, does M2.7 close enough of the gap with Opus-class models to justify the cost savings? The SWE-Pro numbers suggest yes for most software engineering sub-tasks. Your evals will tell you where the remaining gap shows up.
Models covered in this article

Kimi K2.6 API Pricing: ¥1.10/¥6.50/¥27 per 1M Tokens for Coding Agents
Kimi K2.6 API pricing is ¥1.10 cache-hit input, ¥6.50 cache-miss input, and ¥27 output per 1M tokens, with KVV verification, 256K context, and OpenAI-compatible routing for coding agents.

GLM-5.2 on OpenRouter & BigModel: Zhipu's 1M-Context Coding Agent with 128K Output
Zhipu GLM-5.2 is now accessible through OpenRouter, BigModel API, and DashScope: 1M-token lossless context, 128K max output, reasoning_effort, context caching, and FrontierSWE-class performance for long-horizon coding-agent routing.

GLM-5.1-HighSpeed: Zhipu's 400 TPS Flagship Changes the Latency Math for Routing Teams
Zhipu AI's GLM-5.1-highspeed delivers 400 tokens per second via the TileRT inference engine — the same flagship capability as GLM-5.1 but with a throughput profile that reshapes routing decisions for real-time agent and coding workloads.