ByteDance Launches SeedRealtime, a Full-Duplex Audio-Visual LLM

ByteDance's Seed research team released SeedRealtime on August 5, a native audio-visual full-duplex large language model that fuses sound, vision, and text within a single unified architecture. Unlike cascaded systems that chain separate modules for speech recognition, vision, and text-to-speech, SeedRealtime runs perception, understanding, and response generation in parallel over continuous multimodal streams. The model has already been deployed in ByteDance's Douyin and Doubao consumer apps,

1 min
ByteDance Launches SeedRealtime, a Full-Duplex Audio-Visual LLM

ByteDance's Seed research team released SeedRealtime on August 5, a native audio-visual full-duplex large language model that fuses sound, vision, and text within a single unified architecture. Unlike cascaded systems that chain separate modules for speech recognition, vision, and text-to-speech, SeedRealtime runs perception, understanding, and response generation in parallel over continuous multimodal streams.

The model has already been deployed in ByteDance's Douyin and Doubao consumer apps, which the company claims is the first large-scale rollout of full-duplex audio-visual technology.

Three capabilities distinguish SeedRealtime from earlier approaches. First, joint audio-visual understanding lets the model resolve ambiguities by cross-referencing what it hears with what it sees -- for example, disambiguating homophones by analyzing visual context, or interpreting temporal references like "this one" by tracking gestures and scene changes. Second, proactive interaction means the model monitors the environment continuously and can speak unprompted when it detects a relevant change, such as spotting an object the user asked to be reminded about. Third, the model manages conversational timing natively, distinguishing background chatter from the primary speaker and deciding when to interject without relying on an external voice activity detector.

ByteDance published end-to-end human evaluation results showing SeedRealtime cuts conversational pacing failures by roughly half compared to cascaded pipelines. The company reports fewer instances of the model being cut off mid-sentence, responding sluggishly after pauses, or being falsely triggered by background noise.

The launch follows OpenAI's GPT-Live announcement in late July and places ByteDance directly in competition with other full-duplex voice efforts from OpenAI, Microsoft, and Google. SeedRealtime's immediate consumer deployment through Douyin gives it a distribution advantage that lab-stage competitors currently lack.

ByteDance also operates Volcano Engine, its cloud AI platform, which suggests an API path for enterprise use, though the company has not yet announced developer access details for SeedRealtime.

The release marks a shift in the voice AI landscape away from turn-based interaction and toward models that watch, listen, and speak simultaneously -- a technical threshold that multiple labs are now crossing within the same quarter.

Written by

More to read

  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min
  • AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization

    AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization Inference costs have become the second-largest line item in enterprise AI budgets, trailing only talent spend according to RapidData's State of Enterprise AI 2026. This shift represents a fundamental inversion from the 2021-2023 era when training dominated AI expenditure. The compounding nature of serving costs—accumulating every hour as long as users hit the API—means that even modest producti

    1 min