Back to Models

text-embedding-v4 is the latest generation of Qwen's general-purpose multilingual text embedding. It is the recommended choice for new RAG indices, semantic search, clustering, and classification workloads β€” higher retrieval quality than v3 at the same headline price.

Same OpenAI-compatible `/v1/embeddings` surface as v3, so migration is a model-ID swap plus a one-off reindex. TheRouter routes between bailian-cn and bailian-sg per request based on cost.

Latest generation
Default Qwen embedding model for new projects β€” improved retrieval quality over v3.
Multilingual
Chinese, English, and other major languages share a consistent vector space.
8K input window
Embed long passages without aggressive chunking β€” covers most RAG document slices.
Dual-region routing
Selector picks bailian-cn or bailian-sg per request based on cost.
When to use
New RAG indices, semantic search, clustering, classification, deduplication β€” any project starting today that needs multilingual embeddings.
When not to use
Existing indices built on text-embedding-v3 unless you can afford the reindex cost β€” vectors are not comparable across embedding generations.
Pricing: $0.12 per MTok of input. TheRouter routes to the cheaper of bailian-cn / bailian-sg per request.
Context Length
8K
Max Output
--
Input Priceper 1M tokens
$0.0756/ 1M tokens

Modalities

text→embedding

Pricing Breakdown

TypeRate
Input$0.0756 / 1M tokens

Supported Parameters

inputdimensionsencoding_format

Specifications

Release date2025-06-05qwenlm.github.io β†—verified
Parameter count4Bhuggingface.co β†—verified
Transformer layers36huggingface.co β†—verified
Maximum embedding dimensions2560 (MRL: 32–2560)huggingface.co β†—verified
Context length (TheRouter)8,192 tokenstherouter.ai β†—verified
Upstream context length (native)32,768 tokenshuggingface.co β†—verified
Supported languages100+ (natural languages + programming languages)github.com β†—verified
MRL (Matryoshka) supportYes β€” output dimensions configurable from 32 to 2560huggingface.co β†—verified
Instruction-aware promptingYes β€” prepend task description to query for 1–5% retrieval gaingithub.com β†—verified
LicenseApache 2.0github.com β†—verified
Training cutoffNot publicly disclosedunknown

Benchmarks

BenchmarkDistributionScoreSource
MTEB Multilingual (Qwen3-Embedding-8B, for context)
8B variant; ranked #1 on MTEB multilingual leaderboard as of June 5, 2025. The 4B variant (text-embedding-v4) scores lower but remains competitive in its size class.
70.58scoreqwenlm.github.io β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "qwen/text-embedding-v4",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

API guide

Embeddings request

Generate embeddings via the OpenAI-compatible /v1/embeddings endpoint. Supports configurable output dimensions (32–2560, Matryoshka MRL) and instruction-aware query prompting for 1–5% retrieval gain. TheRouter auto-routes between Alibaba Bailian bailian-cn and bailian-sg.

cURL
curl https://api.therouter.ai/v1/embeddings \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/text-embedding-v4",
    "input": "The quick brown fox jumps over the lazy dog",
    "encoding_format": "float",
    "dimensions": 2560
  }'

# MRL: reduce dimensions for lower storage cost (range 32-2560)
# "dimensions": 512

# Instruction-aware query (prepend task prefix to query, NOT to documents):
# "input": "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: What is the capital of France?"

More from qwen

Similar models

Cross-provider sibling models

News & changes

2025-06-05

Qwen3 Embedding series open-sourced under Apache 2.0

Alibaba Qwen released the Qwen3-Embedding-0.6B/4B/8B and Qwen3-Reranker-0.6B/4B/8B series on Hugging Face and ModelScope under Apache 2.0. The 8B embedding model claimed the #1 spot on the MTEB multilingual leaderboard with a score of 70.58. The release includes a technical report (arXiv 2506.05176) and open-source training code on GitHub.

re-authored by TheRouterqwenlm.github.io β†—

Frequently asked

Is text-embedding-v4 the same as Qwen3-Embedding-4B?

Yes. qwen/text-embedding-v4 is the API product name Alibaba Bailian uses for Qwen3-Embedding-4B. The TheRouter model ID qwen/text-embedding-v4 routes to this model. The 8B variant is not separately available through this ID.

re-authored by TheRouterbailian.console.aliyun.com β†—
Can I migrate from text-embedding-v3 without reindexing?

No. v4 and v3 vectors live in different embedding spaces and are not comparable. A full reindex of your vector database is required. If reindexing is not immediately feasible, keep calling qwen/text-embedding-v3 for existing indices until you are ready to rebuild.

re-authored by TheRoutertherouter.ai β†—
How do I use MRL to reduce vector dimensions?

Pass the dimensions parameter in your embeddings request (e.g. dimensions: 512). The valid range is 32–2560. Smaller dimensions reduce storage and ANN query cost with a moderate quality trade-off; benchmark on your specific dataset before choosing a reduced dimension for production.

re-authored by TheRouterhuggingface.co β†—
Should I add an instruction prefix to my documents as well as my queries?

No. The recommended pattern is to prefix only your query with the task instruction (e.g. Instruct: Given a web search query, retrieve relevant passages...\nQuery: <your query>). Passage/document inputs should be sent without any instruction prefix. Prepending instructions to documents typically hurts retrieval quality.

re-authored by TheRoutergithub.com β†—
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release dateqwenlm.github.io β†—2026-06-09verified
Parameter counthuggingface.co β†—2026-06-09verified
Transformer layershuggingface.co β†—2026-06-09verified
Maximum embedding dimensionshuggingface.co β†—2026-06-09verified
Context length (TheRouter)therouter.ai β†—2026-06-09verified
Upstream context length (native)huggingface.co β†—2026-06-09verified
Supported languagesgithub.com β†—2026-06-09verified
MRL (Matryoshka) supporthuggingface.co β†—2026-06-09verified
Instruction-aware promptinggithub.com β†—2026-06-09verified
Licensegithub.com β†—2026-06-09verified
Training cutoffβ€”β€”unknown
MTEB Multilingual (Qwen3-Embedding-8B, for context)qwenlm.github.io β†—2026-06-09verified
Qwen3 Embedding series open-sourced under Apache 2.0qwenlm.github.io β†—2026-06-09verified
Is text-embedding-v4 the same as Qwen3-Embedding-4B?bailian.console.aliyun.com β†—2026-06-09to verify
Can I migrate from text-embedding-v3 without reindexing?therouter.ai β†—2026-06-09to verify
How do I use MRL to reduce vector dimensions?huggingface.co β†—2026-06-09to verify
Should I add an instruction prefix to my documents as well as my queries?github.com β†—2026-06-09to verify
Help & contact