Step 5 Preview API Integration Guide: StepFun's 600B Agentic Flagship via DashScope and TheRouter
Step 5 Preview packs 600B MoE parameters into 27B active per token, ships a 1M-token context window with video input, and prices output at $2.70 per million tokens — reasoning included. This guide covers the API setup, DashScope access, pricing context, benchmark positioning, and how to route Step 5 Preview through TheRouter.
StepFun released Step 5 Preview on September 20, 2026, and the pricing alone makes it worth paying attention to. At $2.70 per million output tokens — reasoning included — it undercuts every model in its benchmark class by a wide margin. Kimi K3 charges $15.00 for the same unit of output. GPT-6 Astra and Claude Opus 5 are multiples higher still. Step 5 Preview does not match those frontier models on every benchmark row, but it lands close enough that the cost gap turns heads.
The model runs a sparse mixture of experts with 600B total parameters and only 27B active per token. It accepts text, images, and video, provides a 1M-token context window with up to 64K output tokens, and ships with tool calling, JSON Schema output, and three reasoning-effort levels. The API is OpenAI-compatible, and the model is available both on StepFun's own platform and through DashScope.
This is the first StepFun coverage on TheRouter's blog. The guide walks through the full integration path.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Quick Start: Step 5 Preview in 3 Minutes
Step 1 — Get a StepFun API key. Register at platform.stepfun.com and generate an API key from the developer console.
Step 2 — Install the OpenAI SDK:
pip install --upgrade openai
Step 3 — Make your first Step 5 Preview request:
from openai import OpenAI
client = OpenAI(
api_key="your-stepfun-api-key",
base_url="https://api.stepfun.com/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[
{"role": "user", "content": "Explain the MoE architecture in Step 5 Preview."}
],
)
print(response.choices[0].message.content)
The endpoint follows the OpenAI Chat Completions format. If your codebase already targets OpenAI, the change is base_url and api_key.
What Step 5 Preview Is
Step 5 Preview is StepFun's flagship foundation model, built for agentic software engineering and professional knowledge work. The architecture trades total parameter count for inference efficiency.
| Spec | Value |
|---|---|
| Architecture | Sparse MoE, 600B total, 27B active per token |
| Context window | 1M tokens input, 64K tokens max output |
| Input modalities | Text, images (up to 60 per request), video (MP4, QuickTime, Matroska) |
| Reasoning | Three effort levels (low, medium, high), set per request |
| Features | Tool calling, streaming, JSON Mode, JSON Schema, prompt caching |
| Model ID | step-5-preview |
| Open weights | Scheduled October 15, 2026 |
| Output speed | 73.9 tokens/second (Artificial Analysis measurement, StepFun endpoint) |
The 27B active parameter count is the smallest of any model that StepFun benchmarks itself against. For comparison, Kimi K3 activates 104B of 2.8T total parameters, and DeepSeek V4.1 Flash activates 8B–16B of 552B. Step 5 Preview sits between them in active size and well below K3 in output price.
Pricing
Output tokens include reasoning tokens, so the $2.70 figure covers the full chain-of-thought.
| Token type | Price per 1M tokens |
|---|---|
| Input (cache miss) | $1.00 |
| Input (cache hit) | $0.05 |
| Output (incl. reasoning) | $2.70 |
For context, here is how that $2.70 output price stacks up against the models StepFun positions Step 5 Preview alongside.
| Model | Output per 1M tokens |
|---|---|
| DeepSeek V4.1 Flash | $0.60 (off-peak) |
| Step 5 Preview | $2.70 |
| Kimi K3 | $15.00 |
| GPT-6 Astra | higher |
| Claude Opus 5 | higher |
The pricing claim is straightforward: comparable intelligence class at a fraction of the output cost. Whether "comparable" holds depends on your workload, and the benchmarks in the next section are the evidence StepFun provides.
Benchmark Positioning
All figures below are from StepFun's own benchmark table, comparing Step 5 Preview at the "High" reasoning effort against competitors at their maximum reasoning settings. Benchmark data is vendor-reported unless otherwise noted.
| Benchmark | Step 5 Preview | GLM 5.3 | Kimi K3 | GPT-6 Astra | Claude Opus 5 |
|---|---|---|---|---|---|
| GPQA Diamond | 93.5% | 91.7% | 93.5% | 96.1% | 93.2% |
| DeepSWE v1.1 | 67.7% | 66.9% | 67.5% | 74.1% | 74.0% |
| Terminal-Bench v4 | 33.3% | 41.9% | 12.6% | 57.9% | 52.3% |
| FrontierFinance | 66.4% | 64.1% | 62.6% | 55.0% | 69.7% |
| SWE-Marathon v1.1 | 72.7% | 67.4% | 84.4% | 77.3% | 85.6% |
| BrowseComp | 88.7% | — | 91.2% | 91.5% | 90.2% |
| ALE-CLI | 29.5% | 28.6% | 27.6% | 33.3% | 28.6% |
The pattern: Step 5 Preview matches or edges the Chinese open-weight flagships (GLM 5.3, Kimi K3) on most rows while trailing GPT-6 Astra and Claude Opus 5 by 6–7 points on the hardest coding benchmarks. FrontierFinance is the one row where Step 5 Preview leads most competitors, consistent with StepFun's positioning around financial knowledge work.
Three caveats apply. First, StepFun's comparison runs High effort against competitors' Max. Second, six benchmarks in the full table are StepFun's own (including StepCodeBench). Third, independent evaluations from Artificial Analysis are still emerging.
Accessing Step 5 Preview Through DashScope
Step 5 Preview appeared on Alibaba Cloud Model Studio (DashScope) on September 21, 2026. The DashScope endpoint is OpenAI-compatible, so the SDK integration looks nearly identical.
from openai import OpenAI
client = OpenAI(
api_key="your-dashscope-api-key",
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
response = client.chat.completions.create(
model="STEPFUN/step-5-preview",
messages=[
{"role": "user", "content": "Summarize Q3 revenue trends."}
],
)
print(response.choices[0].message.content)
DashScope pricing for third-party models may differ from StepFun's direct pricing. Check the DashScope pricing page for current rates.
Using DashScope gives you access to StepFun's model alongside Qwen, DeepSeek, GLM, and other models through a single API key and billing account, which simplifies multi-model operations.
Multimodal Input: Images and Video
Step 5 Preview accepts up to 60 images per request and supports video input in MP4, QuickTime, and Matroska formats. The vision capability follows the standard OpenAI image_url content format.
response = client.chat.completions.create(
model="step-5-preview",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe the architecture diagram."},
{
"type": "image_url",
"image_url": {"url": "https://example.com/diagram.png"},
},
],
}
],
)
Video input follows a similar pattern, with the video file passed as a URL reference in the message content.
Tool Calling and Agentic Workflows
Step 5 Preview supports function calling through the standard OpenAI tools interface. Combined with the 1M-token context window and 64K output limit, it can handle long-horizon agentic tasks where the model needs to read extensive codebases, call tools iteratively, and produce substantial outputs.
tools = [
{
"type": "function",
"function": {
"name": "search_codebase",
"description": "Search the repository for files matching a pattern",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"},
"file_pattern": {"type": "string", "description": "Glob pattern"}
},
"required": ["query"]
}
}
}
]
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Find all API route handlers."}],
tools=tools,
)
StepFun highlights three long-horizon demonstrations: a 24-hour H100 MLA kernel optimization run reaching 508 TFLOPS, automated post-training lifting a Qwen3-30B-A3B base model by 6.7 points on AIME24, and 3,000+ turns of Pokemon Red gameplay.
Reasoning Effort
Step 5 Preview exposes three reasoning-effort levels: low, medium, and high. Lower effort trades accuracy for speed and cost; higher effort enables deeper chain-of-thought.
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Prove the Cauchy-Schwarz inequality."}],
extra_body={"reasoning_effort": "high"},
)
The reasoning tokens count toward the output token total, so at $2.70 per million output tokens, the thinking cost is bundled rather than billed separately.
Routing Step 5 Preview Through TheRouter
TheRouter routes OpenAI-compatible requests through configured providers. To add Step 5 Preview as a routing option, point your SDK at TheRouter and configure the provider chain.
from openai import OpenAI
client = OpenAI(
api_key="your-therouter-api-key",
base_url="https://api.therouter.ai/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Optimize this SQL query."}],
)
A practical routing pattern for agentic coding workloads: route to Step 5 Preview for long-context tasks where $2.70/M output tokens beats the cost of frontier models, and fall back to DeepSeek V4.1 Flash at $0.60/M for simpler completions where flash-tier speed matters more than reasoning depth.
For financial analysis tasks, Step 5 Preview's FrontierFinance benchmark lead (66.4% vs GPT-6 Astra's 55.0%) suggests it may be worth routing financial-domain queries to this model specifically.
When to Choose Step 5 Preview
Step 5 Preview makes sense when:
- You need flagship-class reasoning at a fraction of frontier pricing
- Financial knowledge work is a core use case (FrontierFinance benchmark strength)
- Long-context agentic tasks require 1M-token input windows
- Multimodal input (images, video) is part of the workflow
- You want a single model that bundles reasoning cost into output tokens
Consider alternatives when:
- You need the absolute highest coding benchmark scores (GPT-6 Astra, Claude Opus 5 lead by 6–7 points on DeepSWE)
- Flash-tier latency and sub-$1 output pricing is the priority (DeepSeek V4.1 Flash, Qwen3.8-Flash)
- You need an established ecosystem with extensive third-party tooling
What to Watch
Step 5 Preview is labeled a preview for a reason. Open weights are scheduled for October 15, 2026, and no license has been named yet. There is no published technical report, and independent evaluations from Artificial Analysis are still in progress.
The model is live on StepFun's API and DashScope today. OpenRouter availability has not been announced. When the weights land, self-hosted deployment economics will shift the cost comparison further in Step 5 Preview's favor.
FAQ
What is the model ID for Step 5 Preview?
step-5-preview on StepFun's API. On DashScope, use STEPFUN/step-5-preview.
Is Step 5 Preview OpenAI-compatible?
Yes. It supports the Chat Completions endpoint format. Set base_url to https://api.stepfun.com/v1 with your StepFun API key.
How does Step 5 Preview pricing compare to Kimi K3? Step 5 Preview charges $2.70 per million output tokens (reasoning included). Kimi K3 charges $15.00 per million output tokens. That is 18% of K3's price.
Does Step 5 Preview support tool calling? Yes. Function calling follows the OpenAI tools format with standard tool_choice options.
When will Step 5 Preview open weights be released? StepFun announced October 15, 2026 as the open-weight release date.
Can I access Step 5 Preview through DashScope? Yes. Step 5 Preview is listed on Alibaba Cloud Model Studio (DashScope) as a third-party model.
Sources: StepFun announcement (retrieved 2026-10-04), CellCog analysis (retrieved 2026-10-04), MarkTechPost (retrieved 2026-10-04), Artificial Analysis (retrieved 2026-10-04), DashScope newly released models (retrieved 2026-10-04), Hugging Face (retrieved 2026-10-04).