MiniMax Ships H3, an Open-Weights Video Model with Native Stereo Audio

MiniMax's H3 generates 2K video with native stereo audio in a single forward pass, and the open weights are already on Hugging Face.

2 min
MiniMax Ships H3, an Open-Weights Video Model with Native Stereo Audio

MiniMax released H3 on July 31, a multimodal generation model that produces 2K video with stereo audio in a single forward pass. The company published the weights to Hugging Face on August 3 under the MiniMax H3 Community License, making it the most capable open-weights video model currently available.

H3 is not a text-to-video model with audio bolted on afterward. It is a single transformer that accepts text, images, video, and audio as unified input context and generates 4-to-15-second clips with native stereo sound. A user can drop a product photo, a motion reference clip, and a voice recording into one prompt, describe the relationship between them, and H3 resolves the cross-modal work itself. The model supports up to 9 reference images, 3 video clips, and 3 audio tracks per generation.

The architectural change that makes this practical is H3-VAE, a rewritten tokenizer with a compression ratio that MiniMax claims yields roughly a 4x gain in effective sequence length. This is what makes native 2K output economically viable. At 2K, MiniMax says per-second pricing is less than a third of competing models; at 768p, less than half the price of comparable 720p offerings.

ComfyUI shipped day-zero support. The model runs locally on an RTX 3060, which puts open-weights video generation with native audio on consumer hardware for the first time. RunPod published a guide the same day on GPU sizing for self-hosted inference.

On the Artificial Analysis video leaderboard, H3 ranks first in video editing but trails competitors in text-to-video and image-to-video. The open weights may shift those rankings as the community fine-tunes.

MiniMax is billing this as a commercial tool for advertising, branding, e-commerce, and game cinematics. The model is the third generation in the Hailuo line, following Hailuo 01 and 02, and the first with open weights. The company has also launched MiniMax M3, a new LLM, and MiniMax Speech 2.8 in recent weeks, signaling an acceleration across modalities.

Sources

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min