Groq2 articles

Groq

Articles

  • NVIDIA Enters Full Production on Groq 3 LPX, Hitting 3,400 Tokens per Second in Benchmarks

    NVIDIA has moved its Groq 3 LPX dedicated inference accelerator into full commercial production. Announced at Hot Chips 2026, the rack-scale accelerator system is designed as a purpose-built extension for NVIDIA's Vera Rubin NVL72 data center platform, targeting the compounding decode latency bottlenecks created by multi-step autonomous AI agents. European neocloud provider Nebius Group N.V. has committed as the first cloud infrastructure customer to deploy the accelerators, integrating them in

    1 min
  • NVIDIA Groq 3 LPX Enters Full Production Delivering 3,400 Tokens per Second for Agentic Inference

    At Hot Chips 2026, NVIDIA announced that the Groq 3 LPX dedicated inference accelerator has entered full commercial production. Built as a workload-optimized extension to the Vera Rubin data center architecture, the accelerator addresses the decode bottlenecks inherent in multi-step agentic AI workloads by offloading token generation from primary GPU clusters onto specialized Language Processing Units (LPUs). In independent benchmarking conducted by Artificial Analysis, the Groq 3 LPX delivered

    1 min