Speech-to-Text2 articles

Speech-to-Text

Articles

  • IBM Releases Granite Speech 5.0 with 12,600x Real-Time CTC Conformer Architecture

    IBM has released Granite Speech 5.0, a pair of compact 470-million-parameter automatic speech recognition (ASR) models capable of transcribing over 3.5 hours of audio in one second on modern datacenter silicon. In benchmark evaluations, IBM demonstrated aggregate throughput exceeding 12,600x real-time (12,600 RTFx) on a single NVIDIA H200 GPU. The release includes two variants: Granite Speech 5.0 TurboCTC under the permissive Apache 2.0 license, and an extended research checkpoint licensed unde

    1 min
  • Speech-to-Text Serving in Production: Comparing Faster-Whisper, Moonshine, SenseVoice, and NeMo Canary Architecture, Streaming Latency, and GPU Economics

    In conversational voice AI and real-time agentic workflows, the speech-to-text (STT) layer sets the hard lower bound on system responsiveness. Human conversational cadence expects turn-taking latencies between 200ms and 500ms. When an AI pipeline must accommodate downstream large language model (LLM) time-to-first-token generation (100ms to 250ms) and text-to-speech (TTS) audio synthesis (100ms to 200ms), the automatic speech recognition (ASR) stage cannot exceed 100ms to 150ms of processing ove

    1 min