SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture. While GPU clusters handle core model training and forward passes, agentic AI workf

2 min
SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture.

While GPU clusters handle core model training and forward passes, agentic AI workflows create significant processing demands on conventional CPUs. Systems that call external tools, execute code in isolated sandboxes, orchestrate multi-step reasoning chains, and process training datasets frequently stall when host CPUs cannot feed GPUs quickly enough.

NVIDIA Vera Architecture and Orbital Integration

Vera Architecture and Memory Subsystem

NVIDIA designed the Vera CPU specifically for agentic AI workloads, reinforcement learning pipelines, and data processing. The processor features 88 custom Olympus cores equipped with Spatial Multithreading, delivering 176 hardware threads with partitioned execution resources.

The memory architecture represents a departure from standard enterprise x86 designs:

  • Memory Bandwidth: Vera utilizes LPDDR5X memory mounted on detachable, field-replaceable SOCAMM modules, delivering up to 1.2 TB/s of bandwidth per socket.
  • Capacity and Efficiency: The architecture supports up to 1.5 TB of memory per socket while cutting power consumption roughly in half compared to traditional DDR5 configurations.
  • On-Die Fabric: A second-generation on-die mesh links all 88 cores with 3.4 TB/s of bisectional bandwidth, eliminating cross-chiplet interconnect latencies.
  • Coherent Interconnect: NVLink-C2C provides up to 1.8 TB/s of coherent bidirectional bandwidth directly between Vera CPUs and Rubin GPUs.

According to NVIDIA benchmarks, this architecture completes agentic and reinforcement learning tasks up to 1.8 times faster than standard x86 server processors and achieves up to 80 percent faster throughput in containerized sandbox environments. A standard liquid-cooled Vera CPU rack integrates up to 256 processors and supports more than 22,500 concurrent sandbox instances.

Starmind Satellite and Orbital Computing

The deployment extends to SpaceXAI's planned orbital infrastructure. The company plans to construct its first-generation Starmind AI satellite using an adapted NVIDIA Vera Rubin NVL72 system.

The NVL72 platform packages 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 data processing units, linked via sixth-generation NVLink switches. Adapting this liquid-cooled rack-scale system for low Earth orbit requires meeting rigorous thermal dissipation, power distribution, and radiation hardening constraints while retaining identical software abstractions to terrestrial clusters.

Production Deployments

SpaceXAI joins Anthropic, OpenAI, and Oracle Cloud Infrastructure as early production customers deploying Vera hardware. As frontier labs scale computing clusters into the multi-gigawatt regime, reducing CPU bottlenecks in agent execution directly impacts overall token output and operational power efficiency.

Sources

Written by

More to read

  • Test-Driven Development in AI Coding Agents: Architecture, Reproduction Harnesses, and Execution-Guided Verification

    Autonomous coding agents face a fundamental structural limitation when operating in open-loop, single-turn, or ungrounded generative modes: without dynamic feedback from runtime execution, large language models generate syntactically plausible code that frequently fails subtle interface contracts, breaks existing edge cases, or introduces silent regressions. While early benchmark evaluations relied heavily on zero-shot or few-shot code completion, modern production coding architectures such as S

    1 min
  • The Score Function Estimator: Mathematical Foundations of REINFORCE, Log-Derivative Tricks, and Baseline Variance Reduction

    In modern artificial intelligence, standard backpropagation relies on continuous differentiability: every operation between model parameters and the final loss must provide well-behaved analytical Jacobian matrices. However, many of the most critical optimization challenges in machine learning break this continuity. Autoregressive token generation in large language models, discrete tool invocation, programmatic compiler execution, and black-box reward environments are fundamentally non-different

    1 min
  • Lancium and NVIDIA Partner on Gigawatt-Scale AI Factories Across 15GW Pipeline

    Texas energy infrastructure provider Lancium has entered a strategic partnership with NVIDIA to deploy NVIDIA's full AI factory technology stack across Lancium's multi-gigawatt clean energy data center portfolio. As part of the transaction, NVIDIA has taken a direct equity stake in Lancium, which is backed by funds managed by Blackstone Energy Transition Partners and Blackstone Multi-Asset Investing. The partnership combines Lancium's 4 gigawatts of operational and leased capacity and a pipelin

    1 min