Tencent Hy3 Lands on TokenHub: What the 295B MoE Model's Three Reasoning Modes Mean for Your API Routing Policy

Tencent officially launched Hy3 on July 6, setting live API access on TokenHub and rolling out to multi-provider AI gateways globally. The Tencent Hy3 API routing operator decision now involves three reasoning modes and a price point below $0.15 per million input tokens.

Published via Tencent

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Editorial diagram showing three routing lanes for Tencent Hy3 - fast, slow, and hybrid reasoning modes - with cost and latency tradeoffs on a clean muted-color policy decision map

On July 6, 2026, Tencent officially launched Hy3 - the production-grade successor to Hy3 preview - with API access now live on Tencent Cloud TokenHub and a confirmed rollout to a broad set of global third-party AI gateway and coding-agent platforms. For routing teams, the Tencent Hy3 API routing operator decision is no longer theoretical: there is a real endpoint, real pricing, and a three-mode reasoning architecture that requires deliberate lane assignment in your routing policy.

What happened

Hy3 is a Mixture-of-Experts (MoE) model with 295 billion total parameters and 21 billion active parameters - placing it in the same efficiency class as other sparse-activation models where you pay compute for the activated fraction, not the full parameter count. It supports a 256K context window (262,144 tokens) and exposes three reasoning modes: a fast path for low-latency interactions, a slow-thinking path for complex reasoning, and a hybrid mode that blends both.

The official launch brings several improvements over the April Hy3 preview: enhanced reinforcement learning, better data quality and diversity, and improved stability at scale. Since the preview launch, average daily token consumption grew twenty-fold and the number of WorkBuddy users selecting Hy3 preview grew sixfold - indicating real production load, not just benchmark traffic.

API access is now live on Tencent Cloud TokenHub. The model is also open-source under Apache 2.0 on Hugging Face and ModelScope as tencent/Hy3. Third-party AI gateway platforms - including Hermes, Kilo, Cline, OpenCode, and Cherry Studio - are integrating Hy3 progressively. Self-hosted deployment is supported via vLLM and SGLang using the OpenAI-compatible API.

Why it matters for AI engineering teams

Hy3 addresses a gap that matters for routing teams: a hybrid-thinking model that can vary its reasoning depth per request rather than being locked to a single compute tier. Most production routing decisions involve a tradeoff between fast, cheap inference for high-frequency interactions and slow, expensive inference for complex tasks. Hy3 externalizes that tradeoff as a first-class API parameter rather than forcing you to maintain two separate model routes.

The pricing makes the operator math interesting. Tencent cut the input price for the official Hy3 relative to the preview tier. The preview was priced at approximately 1.2 yuan per million tokens for requests under 16K tokens; Caixing reports the official launch cuts input cost to 1 yuan per million tokens (approximately $0.14/MTok at current rates). That puts Hy3 below the input price floor of most Western frontier models and competitive with Chinese domestic alternatives like DeepSeek-V4-Flash on TokenHub (also 1 yuan/MTok input). Output pricing on TokenHub is 4 yuan per million tokens for short-context requests.

From an agent-workload perspective, Tencent validated Hy3 preview on real agentic use: it reliably powered agent workflows of up to 495 steps across document processing, data analysis, knowledge retrieval, and MCP tool orchestration in Tencent's own products. The Time to First Token dropped 54% versus the prior Hy2 generation, and end-to-end response time dropped 47%. Success rate exceeded 99.99% in production workloads. These are meaningful reliability signals, not just benchmark claims.

The Tencent Hy3 API routing operator angle

Three reasoning modes, three routing lanes

Hy3's hybrid fast-and-slow-thinking architecture means you can route at the mode level, not just the model level. Concretely:

  • Fast mode: low-latency lane for high-frequency chat, autocomplete, or retrieval-augmented generation queries where time-to-first-token matters more than reasoning depth. Pair this with latency-budget routing rules.
  • Slow (extended reasoning) mode: high-accuracy lane for complex multi-step tasks, code generation, financial modeling, or long-horizon planning. Treat this like a premium compute tier - route selectively and account for longer TTFR in your timeout settings.
  • Hybrid mode: the default operating point, where the model dynamically chooses reasoning depth. Use this when you cannot classify incoming requests cleanly or when your workload is mixed.

If your gateway supports per-request parameter overrides, you can route the same tencent/hy3 model ID across all three lanes and control compute cost through the reasoning mode parameter rather than maintaining three separate provider entries.

Provider endpoint options

Teams have three access paths for Hy3 today:

  1. Tencent Cloud TokenHub (https://api.hunyuan.cloud.tencent.com/v1) - the canonical vendor endpoint. Supports OpenAI-compatible Chat Completions with model: "hunyuan-hy3" (verify the exact model string in your TokenHub console, as naming may differ from the open-weight checkpoint name). Requires a Tencent Cloud account and TokenHub API key. Best for teams already in the Tencent Cloud ecosystem.

  2. Third-party inference platforms (Novita AI, and others as rollout progresses) - use model: "tencent/hy3" against OpenAI-compatible endpoints. Novita AI already lists tencent/hy3 as available. These platforms handle capacity, allowing you to skip TokenHub account setup, at the cost of provider dependency and potentially different SLA terms.

  3. Self-hosted via vLLM / SGLang - download from tencent/Hy3 on Hugging Face and serve under any OpenAI-compatible runtime. This gives you data-locality guarantees and removes external provider dependency, at the cost of GPU capacity for a 295B parameter model (21B active, but full weight hosting is still resource-intensive).

Routing policy checklist for teams evaluating Hy3

  • Audit your current cost-quality frontier: at $0.14/MTok input, Hy3 fits below most frontier-model input prices while claiming performance competitive with models 2-5× larger. If your routing policy currently sends all complex tasks to a Western frontier model at $3-15/MTok, Hy3 is a candidate for a secondary cost-reduction lane.
  • Verify model ID before routing production traffic: the open-weight checkpoint is tencent/Hy3 on Hugging Face, but TokenHub and third-party platforms may use different strings. Mismatch between what your routing config sends and what the endpoint accepts causes silent model errors, not 4xx failures.
  • Set reasoning-mode-aware timeout defaults: slow-thinking requests have materially longer TTFR than fast-mode requests. A blanket 30-second gateway timeout may terminate valid slow-mode responses. Segregate timeout budgets by lane.
  • Evaluate MCP tool-call compatibility: Hy3 preview was validated for MCP orchestration in production workflows inside Tencent's own products. For teams routing Hy3 into agent pipelines that use MCP servers, run tool-call compatibility tests before putting it on your primary agent model lane.
  • Watch multi-platform pricing divergence: TokenHub direct and third-party platforms will develop their own pricing over time. If you route Hy3 through a third-party AI gateway that adds a per-call markup, the $0.14/MTok floor may not apply.

256K context: a practical routing signal

The 256K context window (262,144 tokens) is actionable for document-heavy workloads. Teams routing long-document summarization, retrieval augmentation over large corpora, or multi-turn coding sessions with large codebases can configure Hy3 as a fallback or co-primary for requests that exceed 128K tokens. Most frontier models cap usable context at 128K or charge significant premiums for extended context; Hy3's 256K is available at the base token price.

What TheRouter users should watch or try

The Hy3 launch introduces a new tier in the Chinese domestic model roster: a MoE model with hybrid reasoning, validated at real production scale, with pricing that competes on both sides of the China-global frontier. If your routing policy currently includes DeepSeek-V4-Flash for cost-efficiency tasks and a Western frontier model for quality tasks, Hy3 is a candidate to evaluate in the quality-at-cost slot - particularly for coding, long-document, and agent-orchestration workloads where the 256K context and multi-step validation matter.

Watch the rollout timeline on the AI gateway platforms you currently route through. When Hy3 appears on your preferred multi-provider backend, running an A/B routing test at 5-10% of eligible traffic is the fastest way to validate whether Hy3's latency and quality profile fits your production SLA. Browse the TheRouter model catalog to compare Hy3 against other available providers and cost tiers before committing to a routing lane.

For teams not yet routing through any Chinese domestic provider: Hy3's Apache 2.0 license and open-weight availability on Hugging Face remove the enterprise IP concern that has historically blocked adoption of Chinese models in some compliance environments. The self-hosted path is now a credible option for organizations with the GPU capacity.

Keep the model-ID naming discipline tight: lock to tencent/Hy3 (open-weight / third-party platforms) or to whatever string TokenHub’s console assigns, and do not mix the two in the same routing config entry. If you’re setting up a new provider lane, the TheRouter quickstart guide walks through adding a new provider endpoint and validating routing behavior end-to-end.

Help & contact