Tencent Opens Hy3 to Global Users, Claims Top Spot on OpenRouter Within a Week

Tencent announced global availability of its Hy3 large language model on August 5, expanding access beyond China through three channels: the WorkBuddy AI workspace, the Miora creative studio, and the Tencent Cloud TokenHub model-as-a-service platform. The rollout follows Hy3's initial release on July 6 and comes with a free access period on WorkBuddy through August 31. Hy3 uses a hybrid fast-and-slow-thinking Mixture-of-Experts architecture with 295 billion total parameters and 21 billion activ

2 min
Tencent Opens Hy3 to Global Users, Claims Top Spot on OpenRouter Within a Week

Tencent announced global availability of its Hy3 large language model on August 5, expanding access beyond China through three channels: the WorkBuddy AI workspace, the Miora creative studio, and the Tencent Cloud TokenHub model-as-a-service platform. The rollout follows Hy3's initial release on July 6 and comes with a free access period on WorkBuddy through August 31.

Hy3 uses a hybrid fast-and-slow-thinking Mixture-of-Experts architecture with 295 billion total parameters and 21 billion active parameters, supporting a context length of up to 256,000 tokens. Tencent claims the model performs comparably to flagship models with two to five times as many active parameters on reasoning, instruction following, code generation, and agent tasks.

Usage metrics suggest strong early demand. Tencent reports Hy3 generated more than 68 times the API call volume of its predecessor and reached the top position on OpenRouter's global LLM usage leaderboard within one week of launch. The model is available under the Apache 2.0 license and has been distributed through Hugging Face, ModelScope, and third-party developer platforms including Cline, Kilo, and OpenCode.

Pricing on OpenRouter starts at $0.1288 per million input tokens and $0.5336 per million output tokens, positioning Hy3 as a cost-competitive option against comparable models from OpenAI, Anthropic, and Google.

Tencent is pairing the model with its existing product ecosystem. WorkBuddy, which Tencent describes as China's most widely used AI agent workspace, achieved a task success rate above 90 percent in internal evaluations when running on Hy3, while reducing average task completion time by 34 percent compared to the previous model generation. Miora, Tencent's AI-native creative studio, connects Hy3's reasoning capabilities into workflows spanning graphics, video, 3D, and UI design.

On the enterprise side, Tencent Cloud TokenHub serves as a multi-model gateway with intelligent routing, letting organizations switch between Hy3 and third-party models through a single API. Regional partners including South Korea's Cafe24 and Japan's Metelix are integrating Hy3 into their respective AI platform services.

The global expansion of Hy3 adds another major Chinese model to an increasingly crowded international market, following recent launches from Alibaba's Qwen, ByteDance's Seed, and Moonshot AI's Kimi series.

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min