Alibaba Launches Wan 3.0 AI Video Model with Native 30-Second Generation and Document Inputs

Alibaba Tongyi Lab has launched a public beta of Wan 3.0, the latest iteration of its video generation model family. Available on Alibaba Cloud Model Studio and Qwen Cloud under the model identifier wan3.0-video, the model produces up to 30 seconds of continuous video in a single pass at resolutions up to 1080p. Unlike predecessor models such as Wan 2.7, which capped single-pass output at 15 seconds, Wan 3.0 consolidates video synthesis into a unified architecture and expands supported input mo

2 min
Alibaba Launches Wan 3.0 AI Video Model with Native 30-Second Generation and Document Inputs

Alibaba Tongyi Lab has launched a public beta of Wan 3.0, the latest iteration of its video generation model family. Available on Alibaba Cloud Model Studio and Qwen Cloud under the model identifier wan3.0-video, the model produces up to 30 seconds of continuous video in a single pass at resolutions up to 1080p.

Unlike predecessor models such as Wan 2.7, which capped single-pass output at 15 seconds, Wan 3.0 consolidates video synthesis into a unified architecture and expands supported input modalities beyond prompt text and still images.

Multimodal Video Synthesis Architecture

Omni-Reference Multimodal Inputs

The primary architectural addition in Wan 3.0 is support for structured document ingestion alongside traditional media references. The model accepts:

  • Office documents, including PDF files, PowerPoint slide decks, and spreadsheets.
  • Web pages, URL references, and raw HTML structures.
  • Audio tracks and spoken-word voice clips for synchronized lip movement.
  • Reference video clips, character images, and style templates.

By ingesting slide decks or structured reports directly, Wan 3.0 extracts sequential semantic information to generate explanatory or promotional video sequences without requiring manual prompt decomposition.

Visual Continuity and Resolution Tiers

To mitigate common temporal artifacts such as character drift, flickering, and background deformation across extended generation horizons, Wan 3.0 implements enhanced cross-frame attention mechanisms. Alibaba reports improved stability across facial micro-expressions, user interface elements, and text rendering in synthesized scenes.

Generation outputs are available in three resolution profiles:

  • 480p standard definition for rapid prototyping.
  • 720p high definition for general web playback.
  • 1080p full high definition for final production delivery.

Alibaba has not released open weights for Wan 3.0, restricting access to API endpoints and cloud hosting services. Previous open-weight models in the series, such as Wan 2.1, remain available on open model repositories.

Enterprise and Robotics Applications

Beyond creative content production and marketing workflows, Alibaba is positioning Wan 3.0 as a simulation engine. The model is capable of generating synthetic video data to train autonomous driving perception models and humanoid robotics vision systems, simulating diverse lighting conditions, dynamic obstacles, and complex physical interactions.

Access is currently open in public beta through Alibaba Cloud Model Studio and Qwen Cloud, with commercial API pricing structured on a per-second video generation metric.

Sources

Written by

More to read

  • Generalized Advantage Estimation: How Exponential Weighting Balances Bias and Variance in Policy Optimization

    Policy gradient algorithms form the theoretical backbone of modern policy optimization, ranging from continuous robotic control to reinforcement learning from human feedback (RLHF) in frontier large language models. A persistent challenge in policy optimization is variance: estimating the gradient of expected cumulative reward over stochastic trajectories generates high-variance Monte Carlo signals that require massive sample sizes and risk destabilizing gradient updates. Generalized Advantage

    1 min
  • IBM Unveils 2nm Dual-Architecture Mainframe Processor with Native Arm and On-Chip AI Acceleration

    IBM unveiled the industry's first dual-architecture mainframe processor at the annual Hot Chips conference, detailing custom silicon capable of natively executing both IBM Z (s390x) and Arm (Arm64) instruction set architectures on the exact same physical cores. Fabricated on a leading-edge 2-nanometer process node, the upcoming processor is engineered to bridge traditional enterprise transaction processing with the modern, Arm-dominated software ecosystem, particularly containerized AI framewor

    1 min
  • ByteDance Consolidates TRAE and Coze into Doubao, Readies 'Doubao Work' Enterprise Brand

    ByteDance has initiated an internal organizational restructuring to consolidate its enterprise AI developer tools and agent platforms under the Doubao ecosystem. The company is merging the teams and technologies behind its AI programming suite TRAE and agent-building platform Coze into Doubao, preparing to launch a unified enterprise AI suite branded "Doubao Work." The restructuring concentrates disparate AI tools into a single corporate pillar to compete against domestic rivals, notably Tencen

    1 min