Nvidia's Nemotron 4 aims for a trillion parameters, but China is already there

Nvidia is building Nemotron 4, a new family of open-weight AI models meant to challenge the strongest freely available models, according to reporting by The Information cited by The Decoder. The largest model in the family is planned to reach at least one trillion parameters, twice the size of Nvidia's current Nemotron 3 Ultra. To train it, Nvidia has tripled its cloud spending on in-house model development to 28 billion dollars through 2031. The earliest the models could ship is this fall. C

1 min
Nvidia's Nemotron 4 aims for a trillion parameters, but China is already there

Nvidia is building Nemotron 4, a new family of open-weight AI models meant to challenge the strongest freely available models, according to reporting by The Information cited by The Decoder. The largest model in the family is planned to reach at least one trillion parameters, twice the size of Nvidia's current Nemotron 3 Ultra.

To train it, Nvidia has tripled its cloud spending on in-house model development to 28 billion dollars through 2031. The earliest the models could ship is this fall.

China is already there

Illustration of the AI model scale gap between US and Chinese labs

A trillion parameters would be a milestone for a US open model, but not for the field. China's labs already operate at that scale and beyond. Moonshot AI's Kimi K3 uses 2.8 trillion parameters, and DeepSeek V4 Pro uses 1.6 trillion. On the Artificial Analysis Intelligence Index, Nvidia's Nemotron 3 Ultra scores 38 points and trails Moonshot's Kimi K3 at roughly 60.

A tension with its own customers

The move also puts Nvidia in tension with its own customers. The chipmaker recently announced a 10 gigawatt infrastructure partnership with OpenAI, yet Nemotron 4 would compete directly with OpenAI's own models in the open-weight arena. Nvidia is also among the signatories of a petition opposing new regulation of open models, a position that aligns with selling more GPUs as companies choose to self-host.

Sources

The Decoder: https://the-decoder.com/nvidias-nemotron-4-aims-for-one-trillion-parameters-a-scale-chinese-labs-already-surpassed/

The Information (via The Decoder): https://www.theinformation.com/articles/nvidia-open-source-ai

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min