SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture. While GPU clusters handle core model training and forward passes, agentic AI workf

2 min
SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture.

While GPU clusters handle core model training and forward passes, agentic AI workflows create significant processing demands on conventional CPUs. Systems that call external tools, execute code in isolated sandboxes, orchestrate multi-step reasoning chains, and process training datasets frequently stall when host CPUs cannot feed GPUs quickly enough.

NVIDIA Vera Architecture and Orbital Integration

Vera Architecture and Memory Subsystem

NVIDIA designed the Vera CPU specifically for agentic AI workloads, reinforcement learning pipelines, and data processing. The processor features 88 custom Olympus cores equipped with Spatial Multithreading, delivering 176 hardware threads with partitioned execution resources.

The memory architecture represents a departure from standard enterprise x86 designs:

  • Memory Bandwidth: Vera utilizes LPDDR5X memory mounted on detachable, field-replaceable SOCAMM modules, delivering up to 1.2 TB/s of bandwidth per socket.
  • Capacity and Efficiency: The architecture supports up to 1.5 TB of memory per socket while cutting power consumption roughly in half compared to traditional DDR5 configurations.
  • On-Die Fabric: A second-generation on-die mesh links all 88 cores with 3.4 TB/s of bisectional bandwidth, eliminating cross-chiplet interconnect latencies.
  • Coherent Interconnect: NVLink-C2C provides up to 1.8 TB/s of coherent bidirectional bandwidth directly between Vera CPUs and Rubin GPUs.

According to NVIDIA benchmarks, this architecture completes agentic and reinforcement learning tasks up to 1.8 times faster than standard x86 server processors and achieves up to 80 percent faster throughput in containerized sandbox environments. A standard liquid-cooled Vera CPU rack integrates up to 256 processors and supports more than 22,500 concurrent sandbox instances.

Starmind Satellite and Orbital Computing

The deployment extends to SpaceXAI's planned orbital infrastructure. The company plans to construct its first-generation Starmind AI satellite using an adapted NVIDIA Vera Rubin NVL72 system.

The NVL72 platform packages 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 data processing units, linked via sixth-generation NVLink switches. Adapting this liquid-cooled rack-scale system for low Earth orbit requires meeting rigorous thermal dissipation, power distribution, and radiation hardening constraints while retaining identical software abstractions to terrestrial clusters.

Production Deployments

SpaceXAI joins Anthropic, OpenAI, and Oracle Cloud Infrastructure as early production customers deploying Vera hardware. As frontier labs scale computing clusters into the multi-gigawatt regime, reducing CPU bottlenecks in agent execution directly impacts overall token output and operational power efficiency.

Sources

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min