Speech2 articles

Speech

Articles

  • Google Releases Gemini 3.5 Transcribe with Disfluency Filtering and Task Delegation

    Google has launched Gemini 3.5 Transcribe, a dedicated speech-to-text model designed for real-time streaming, automated disfluency cleanup, and agentic task delegation. The release introduces two API interfaces alongside integration across Google developer tooling and consumer operating system surfaces. Dual API Architecture for Live and Batch Audio Gemini 3.5 Transcribe is split into two operational endpoints tailored for distinct latency profiles: * Real-time streaming (gemini-3.5-transcr

    1 min
  • NVIDIA Releases Magpie Multilingual TTS: 364M Open-Weight Model for Sub-200ms Voice Agents

    NVIDIA has released Magpie Multilingual TTS, a 364-million parameter open-weights text-to-speech model engineered for low-latency conversational AI agents. Released under the NVIDIA Open Model License, the model is available as open checkpoints on the Hugging Face Hub and as an optimized microservice container within NVIDIA NIM. The release expands language support to 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean,

    1 min