Hardware8 articles

Hardware

Articles

  • Alibaba Demonstrates Native Qwen 3.8 27B Inference on XuanTie C950 RISC-V CPU at 30 Tokens per Second

    Alibaba's semiconductor division, T-Head, announced day-zero native inference support for its latest open-weight model, Qwen 3.8 27B, running directly on the XuanTie C950 RISC-V server processor. Operating without discrete graphics processing units, the 64-core RISC-V chip delivered sustained decode throughput of 30 tokens per second alongside a time-to-first-token latency of 1.9 seconds. The benchmark demonstrates how architectural extensions on general-purpose open instruction sets can handle

    1 min
  • Local LLM Inference on Apple Silicon: Architecture, Unified Memory, and Serving Benchmarks for MLX, llama.cpp, and Ollama

    Local large language model (LLM) serving on consumer hardware has historically faced a hard trade-off between memory capacity and execution bandwidth. Discrete consumer GPUs offer high memory bandwidth (up to 1,008 GB/s on an Nvidia RTX 4090) but are capped at 24 GB of VRAM, requiring model sharding or quantization to fit models beyond 14 billion parameters. Apple Silicon platforms bypass this capacity ceiling through a Unified Memory Architecture (UMA), where the CPU, GPU, and Apple Neural Eng

    1 min
  • Velaura AI Raises 10M Series A at B Valuation for Low-Power AI Silicon

    Velaura AI Raises $110M Series A at $1B Valuation for Low-Power AI Silicon Velaura AI has closed a $110 million Series A funding round at a valuation exceeding $1 billion. The financing was led by Seligman Ventures, with participation from Capricorn Investment Group alongside existing backers including Samsung Catalyst Fund, StepStone Group, Maverick Silicon, Celesta Capital, and Mayfield. The capital will fund the commercialization and deployment of Velaura's silicon IP and physical design te

    1 min
  • Etched in Talks to Raise 00M Led by Jane Street at 1B Valuation

    AI inference chip startup Etched is in negotiations to raise $700 million in a new financing round led by existing investor Jane Street, according to reporting from The Wall Street Journal. The proposed financing would value the San Jose-based semiconductor company at approximately $21 billion. The planned round follows a $300 million Series C led by Sequoia Capital that valued the startup at $10.3 billion, effectively doubling its valuation within weeks amid intensifying enterprise demand for

    1 min
  • LG Partners with Nvidia on 10,000-Square-Meter Robot Data Factory Targeting 100,000 Training Hours

    LG Electronics hosted senior Nvidia leadership at its Yangjae R&D Campus in Seoul on August 18, 2026, advancing a joint initiative to build physical AI training pipelines and target 100,000 hours of embodied robotics data by the end of the year. The site review took place five days after LG Group and Nvidia signed a strategic memorandum of understanding at Nvidia headquarters in Santa Clara on August 13. The accelerated timeline reflects LG's effort to convert its industrial manufacturing infra

    1 min
  • FlashAttention: How IO-Aware Tiling and Online Softmax Solved Transformer Memory Bottlenecks

    Standard multi-head attention is the fundamental computational primitive of modern autoregressive language models. While mathematically straightforward, the operation introduces a severe operational bottleneck as context windows scale. Naive implementations of scaled dot-product attention exhibit quadratic memory complexity $O(N^2)$ and quadratic memory access costs, bounding sequence lengths and leaving modern GPU tensor cores severely underutilized. FlashAttention, introduced by Tri Dao, Dani

    1 min
  • AMD Acquires Taalas to Build Chips Hard-Wired for Specific AI Models

    AMD announced on August 6 that it has reached a definitive agreement to acquire Taalas, a Toronto-based startup that builds processors customized around individual AI models. The deal adds specialized inference silicon to AMD's accelerator portfolio, targeting a class of workloads where a model is fixed and the goal is maximum throughput at minimum cost. Taalas, founded in 2023, produces what it calls Hardcore Models: chips whose circuitry is laid out for one specific model's weights. The custo

    1 min
  • Anthropic assembles in-house chip design team for Claude

    Anthropic assembles in-house chip design team for Claude Anthropic confirmed on August 5 that it is building an internal team to design custom silicon for its Claude AI models, joining a growing list of AI labs that have concluded off-the-shelf chips are no longer sufficient for the scale they need. The company is recruiting engineers with chip design experience for what it calls a "custom silicon team," according to a job listing. The group will co-design hardware and AI models in tandem, aim

    1 min