ByteDance Trains 10 Trillion-Parameter AI Model to Rival Anthropic's Mythos

ByteDance is pretraining a large model with up to 10 trillion parameters, a scale the Financial Times reports could put it in the same class as Anthropic's most advanced systems. The model, still in early pretraining, would be more than three times the size of Moonshot AI's Kimi K3, currently the largest Chinese model at 2.8 trillion parameters. Three people familiar with the project told the FT the model is in pretraining, a phase that typically lasts three to six months before full training a

2 min
ByteDance Trains 10 Trillion-Parameter AI Model to Rival Anthropic's Mythos

ByteDance is pretraining a large model with up to 10 trillion parameters, a scale the Financial Times reports could put it in the same class as Anthropic's most advanced systems. The model, still in early pretraining, would be more than three times the size of Moonshot AI's Kimi K3, currently the largest Chinese model at 2.8 trillion parameters.

Model scale comparison showing ByteDance's reported 10T-parameter target against current Chinese and US frontier models

Three people familiar with the project told the FT the model is in pretraining, a phase that typically lasts three to six months before full training and evaluation. Parameter counts are a rough measure of capacity rather than a guarantee of capability, since output quality also depends on training data and methods.

Industry estimates cited by the FT place Anthropic's Mythos 5 at roughly 8 trillion parameters and its Fable 5 at around 5 trillion. Neither Anthropic nor OpenAI discloses parameter counts for its flagship systems, which makes direct comparisons with private US models difficult. ByteDance's reported target sits close to the estimated size of Mythos.

The project sits with ByteDance's Seed research team, the roughly 2,000-person lab behind the company's frontier model work. Founder Zhang Yiming has told the team to aim for world-leading model capabilities over the long term, according to one of the sources. The source also said ByteDance has avoided distillation, training on outputs from other companies' models, for over a year.

xAI is pursuing comparable scale. Elon Musk has said the company is training Grok variants with six and ten trillion parameters on its Colossus 2 cluster.

China's largest AI model is being developed at ByteDance, and the company is publicly positioning around scale. It launched SeedRealtime, a full-duplex audio-visual LLM, in early August, and its Doubao assistant app has one of the largest user bases among Chinese AI products. A 10 trillion-parameter pretraining run, if completed, would mark a step up in the scale of Chinese frontier models.

Sources

Financial Times via The Decoder: China's Largest AI Model Is Being Developed at Bytedance

Reuters via Yahoo Finance: ByteDance targets mega AI model that could match Mythos scale

The News International: ByteDance trains 10 trillion-parameter AI model to rival Anthropic's Mythos

Written by

More to read

  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min
  • AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization

    AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization Inference costs have become the second-largest line item in enterprise AI budgets, trailing only talent spend according to RapidData's State of Enterprise AI 2026. This shift represents a fundamental inversion from the 2021-2023 era when training dominated AI expenditure. The compounding nature of serving costs—accumulating every hour as long as users hit the API—means that even modest producti

    1 min