Qwen3.8-Omni-Flash API Guide: Native Audio, Video, Image, and Text Understanding on DashScope
A practical guide to integrating Qwen3.8-Omni-Flash via DashScope's OpenAI-compatible API. Covers model IDs, omni-modal input (text + image + audio + video), thinking mode, function calling, the realtime WebSocket variant, pricing by modality, and routing through TheRouter.
Qwen3.8-Omni-Flash is the first natively omni-modal model in the Qwen family. Launched on DashScope on September 17, 2026, it accepts text, images, audio, and video in a single API call and returns text output. A companion realtime variant (qwen3.8-omni-flash-realtime) adds WebSocket and WebRTC streaming for live audio-video conversations. This guide walks through both models, from first API call to production routing.
What Makes Qwen3.8-Omni-Flash Different
Most multimodal models bolt vision onto a text backbone. Qwen3.8-Omni-Flash was built from the ground up as an omni-modal agentic model. Four modalities go in, text comes out, and the model can reason across them in a single forward pass.
Key specs from the official Qwen blog and DashScope model listing:
- Input modalities: text, image, audio, video
- Output: text (the realtime variant also outputs audio)
- Context window: 1M tokens
- Thinking mode: supported (chain-of-thought reasoning)
- Function calling: supported
- Web search: supported
- Context caching: supported
- API compatibility: OpenAI Chat Completions and Responses API via DashScope
The model targets real-world agentic productivity. Rather than just understanding multimodal content, it can act on it through function calling and web search, then reason through complex tasks with thinking mode enabled.
Two Model IDs, Two Access Patterns
DashScope exposes Qwen3.8-Omni-Flash through two model IDs that serve different use cases:
| Model ID | Protocol | Input | Output | Use case |
|---|---|---|---|---|
qwen3.8-omni-flash | HTTP (Chat Completions / Responses) | Text, image, audio, video | Text | Batch analysis, document processing, async workflows |
qwen3.8-omni-flash-realtime | WebSocket, WebRTC | Streaming audio, video frames, text | Text + audio | Live conversations, voice assistants, video calls |
The offline model (qwen3.8-omni-flash) uses the same OpenAI-compatible HTTP API you already know. The realtime model (qwen3.8-omni-flash-realtime) requires a persistent connection and streams input/output continuously.
Getting Started with the Offline API
Step 1: Get Your API Key
Sign up at Alibaba Cloud Model Studio and generate a DashScope API key from the console.
Step 2: Configure the OpenAI SDK
from openai import OpenAI
client = OpenAI(
api_key="your-dashscope-api-key",
base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
)
For the China region, use https://dashscope.aliyuncs.com/compatible-mode/v1 instead.
Step 3: Text + Image Understanding
response = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "What components are shown in this architecture diagram?"},
{"type": "image_url", "image_url": {"url": "https://example.com/architecture.png"}},
],
}
],
)
print(response.choices[0].message.content)
Step 4: Audio Understanding
response = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Transcribe and summarize this meeting recording."},
{"type": "audio_url", "audio_url": {"url": "https://example.com/meeting.mp3"}},
],
}
],
)
print(response.choices[0].message.content)
Step 5: Video Analysis
response = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe what happens in this product demo."},
{"type": "video_url", "video_url": {"url": "https://example.com/demo.mp4"}},
],
}
],
)
print(response.choices[0].message.content)
All three patterns use the same Chat Completions endpoint. The model identifies the input modality and processes it natively rather than through separate pipelines.
Thinking Mode
Qwen3.8-Omni-Flash supports an explicit reasoning mode that produces chain-of-thought before the final answer. This is useful for complex multimodal queries where the model needs to work through visual or auditory evidence step by step.
Enable it by following the DashScope thinking mode documentation. The exact parameter format may vary; always check the current DashScope docs for the supported configuration.
Function Calling
The model supports OpenAI-compatible function calling, which means you can define tools and let Qwen3.8-Omni-Flash decide when to invoke them based on multimodal context.
tools = [
{
"type": "function",
"function": {
"name": "search_inventory",
"description": "Search product inventory by visual description",
"parameters": {
"type": "object",
"properties": {
"description": {"type": "string"},
"category": {"type": "string"}
},
"required": ["description"]
}
}
}
]
response = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Find this item in our inventory."},
{"type": "image_url", "image_url": {"url": "https://example.com/item.jpg"}},
],
}
],
tools=tools,
)
The model can also use web search as a built-in tool, retrieving fresh information to complement its multimodal understanding.
The Realtime Variant
qwen3.8-omni-flash-realtime launched on September 21, 2026, four days after the offline model. It supports three connection protocols:
- WebSocket: standard bidirectional streaming
- WebRTC: optimized for browser-based audio/video
- AOQ: Alibaba's proprietary low-latency protocol
The realtime model adds multi-channel audio aggregation, video frame aggregation, and remote MCP tool calling on top of the base model's capabilities. It outputs both text and audio, enabling voice assistant and video call applications.
For API documentation on the realtime variant, see the DashScope realtime model guide.
Pricing
Qwen3.8-Omni-Flash shares the Flash pricing tier on DashScope. Text token rates from third-party aggregators (verified against OpenRouter and DashScope pricing):
| Modality | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Text | ~$0.15 | ~$0.47 |
| Audio | Premium over text (modality-specific rate) | N/A (text output only) |
| Video | Premium over text (modality-specific rate) | N/A (text output only) |
| Image | Same as vision models (modality-specific rate) | N/A (text output only) |
For the realtime variant, audio output is billed at a separate rate in addition to the text output rate. Exact per-modality rates in RMB are available on the official pricing page; USD equivalents are approximate and depend on the exchange rate.
Context caching applies a discount to cached input tokens, following the same pattern as other Qwen3.8 models on DashScope.
Comparison: Omni-Flash vs Qwen3.8-Flash vs GPT-4o
| Feature | Qwen3.8-Omni-Flash | Qwen3.8-Flash | GPT-4o |
|---|---|---|---|
| Text input | Yes | Yes | Yes |
| Image input | Yes | Yes | Yes |
| Audio input | Yes | No | Yes |
| Video input | Yes | Yes (frames) | Limited |
| Audio output | Realtime variant only | No | Realtime API only |
| Context window | 1M | 1M | 128K |
| Thinking mode | Yes | Yes | No (use o3/o4) |
| Function calling | Yes | Yes | Yes |
| Web search | Yes | Yes | Via tools |
| Realtime streaming | WebSocket/WebRTC/AOQ | No | WebSocket |
| Text input pricing | ~$0.15/1M | ~$0.15/1M | $2.50/1M |
Qwen3.8-Omni-Flash offers a significant cost advantage over GPT-4o for multimodal workloads. The trade-off is ecosystem maturity: OpenAI's tooling and third-party integrations are more extensive, while DashScope's OpenAI-compatible endpoint covers the core API surface.
For teams already using Qwen3.8-Flash for text and image tasks, Omni-Flash adds audio and video understanding without changing the API integration pattern. The model ID swap is the only code change needed.
Routing with TheRouter
If you route OpenAI-compatible requests through TheRouter, you can configure DashScope as a provider and direct omni-modal requests to qwen3.8-omni-flash while routing text-only requests to cheaper models.
TheRouter routes OpenAI-compatible requests through configured providers, supports provider/model routing and fallback when live product paths support it, and provides unified billing/accounting surfaces where implemented.
A modality-based routing strategy makes sense here: send requests containing audio or video content to Omni-Flash, and route pure text or text+image requests to Qwen3.8-Flash or Qwen3.8-Max depending on complexity. This keeps costs low for simple queries while ensuring multimodal requests reach a capable model.
For fallback configuration, see the model fallback guide.
Frequently Asked Questions
Can Qwen3.8-Omni-Flash generate audio or images?
The offline model (qwen3.8-omni-flash) outputs text only. The realtime variant (qwen3.8-omni-flash-realtime) can output audio alongside text. Neither generates images or video.
What audio formats does it accept?
The model accepts common audio formats including MP3, WAV, and other standard formats via URL. Check the DashScope documentation for the complete list of supported formats and size limits.
How does the 1M context window work with audio and video?
Audio and video are tokenized differently from text. A one-minute audio clip or a short video segment may consume a substantial portion of the context window. The exact token-to-duration ratio depends on the encoding and is documented in the DashScope model card.
Is the model available outside China?
Yes. DashScope is available in the Singapore region (ap-southeast-1) with the international endpoint. Use https://dashscope-intl.aliyuncs.com/compatible-mode/v1 as the base URL.
Can I use it with Claude Code or OpenAI Codex?
Since DashScope exposes an OpenAI-compatible endpoint, any tool that supports custom base_url configuration can point to DashScope and use qwen3.8-omni-flash as the model ID. However, omni-modal features (audio/video input) require the tool to pass those content types in the messages array, which most coding agents do not currently do.
What is the difference between Omni-Flash and Qwen3.5-Omni?
Qwen3.8-Omni-Flash is the successor. It builds on the Qwen3.8 architecture with improved agentic capabilities, stronger function calling, web search integration, and better benchmark scores across multimodal tasks. Qwen3.5-Omni models remain available but the 3.8 Omni-Flash is the recommended choice for new integrations.
Sources
- Qwen3.8-Omni-Flash official blog post — retrieved 2026-10-06
- DashScope newly released models — retrieved 2026-10-06
- DashScope model pricing — retrieved 2026-10-06
- Alibaba Cloud Model Studio models overview — retrieved 2026-10-06
- DashScope OpenAI compatibility — retrieved 2026-10-06
- OpenRouter Qwen3.8-Flash pricing — retrieved 2026-10-06
- BenchLM Qwen3.8-Omni-Flash benchmarks — retrieved 2026-10-06
- DashScope realtime model guide — retrieved 2026-10-06