GPU1 article

GPU

Articles

  • French startup Kog says smarter software can pull 30x faster inference out of stock GPUs

    The race for faster AI inference has pushed labs toward custom silicon, but a French startup argues the cheapest speedup is already sitting in servers companies bought. Kog says it can reach roughly thirty times faster LLM inference on standard data-center GPUs using software alone. Kog demonstrated three thousand tokens per second per request on AMD MI300X and NVIDIA H200 GPUs in May, running a small two-billion-parameter model it has since open-sourced as Laneformer 2B. Chief executive Gael D

    1 min