ByteDance Launches SeedRealtime, a Full-Duplex Audio-Visual LLM

ByteDance's Seed research team released SeedRealtime on August 5, a native audio-visual full-duplex large language model that fuses sound, vision, and text within a single unified architecture. Unlike cascaded systems that chain separate modules for speech recognition, vision, and text-to-speech, SeedRealtime runs perception, understanding, and response generation in parallel over continuous multimodal streams. The model has already been deployed in ByteDance's Douyin and Doubao consumer apps,

1 min
ByteDance Launches SeedRealtime, a Full-Duplex Audio-Visual LLM

ByteDance's Seed research team released SeedRealtime on August 5, a native audio-visual full-duplex large language model that fuses sound, vision, and text within a single unified architecture. Unlike cascaded systems that chain separate modules for speech recognition, vision, and text-to-speech, SeedRealtime runs perception, understanding, and response generation in parallel over continuous multimodal streams.

The model has already been deployed in ByteDance's Douyin and Doubao consumer apps, which the company claims is the first large-scale rollout of full-duplex audio-visual technology.

Three capabilities distinguish SeedRealtime from earlier approaches. First, joint audio-visual understanding lets the model resolve ambiguities by cross-referencing what it hears with what it sees -- for example, disambiguating homophones by analyzing visual context, or interpreting temporal references like "this one" by tracking gestures and scene changes. Second, proactive interaction means the model monitors the environment continuously and can speak unprompted when it detects a relevant change, such as spotting an object the user asked to be reminded about. Third, the model manages conversational timing natively, distinguishing background chatter from the primary speaker and deciding when to interject without relying on an external voice activity detector.

ByteDance published end-to-end human evaluation results showing SeedRealtime cuts conversational pacing failures by roughly half compared to cascaded pipelines. The company reports fewer instances of the model being cut off mid-sentence, responding sluggishly after pauses, or being falsely triggered by background noise.

The launch follows OpenAI's GPT-Live announcement in late July and places ByteDance directly in competition with other full-duplex voice efforts from OpenAI, Microsoft, and Google. SeedRealtime's immediate consumer deployment through Douyin gives it a distribution advantage that lab-stage competitors currently lack.

ByteDance also operates Volcano Engine, its cloud AI platform, which suggests an API path for enterprise use, though the company has not yet announced developer access details for SeedRealtime.

The release marks a shift in the voice AI landscape away from turn-based interaction and toward models that watch, listen, and speak simultaneously -- a technical threshold that multiple labs are now crossing within the same quarter.

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min