AWS and NVIDIA Expand AI Partnership to Deploy 2 Million Additional Blackwell Ultra and Rubin GPUs

Amazon Web Services (AWS) and NVIDIA have announced a major expansion of their cloud infrastructure partnership, committing to deploy two million additional high-end NVIDIA GPUs across AWS global data centers in 2027 and 2028. The deployment expands on AWS's previous commitment from GTC 2026 to add one million GPUs starting in 2026, bringing total forward allocations across the multi-year cycle to three million units. The upcoming capacity will comprise NVIDIA Blackwell Ultra, Rubin, and Rubin

2 min
AWS and NVIDIA Expand AI Partnership to Deploy 2 Million Additional Blackwell Ultra and Rubin GPUs

Amazon Web Services (AWS) and NVIDIA have announced a major expansion of their cloud infrastructure partnership, committing to deploy two million additional high-end NVIDIA GPUs across AWS global data centers in 2027 and 2028. The deployment expands on AWS's previous commitment from GTC 2026 to add one million GPUs starting in 2026, bringing total forward allocations across the multi-year cycle to three million units.

The upcoming capacity will comprise NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra architectures. The hardware is designated to support frontier training, large-scale agentic AI systems, enterprise automation, and physical AI workloads.

Hardware Architecture and Next-Generation Deployments

The expanded infrastructure program introduces several architectural shifts across compute nodes, memory topologies, and networking fabrics:

  • Vera CPU Integration: AWS will introduce instances powered by NVIDIA's standalone ARM-based Vera CPUs, designed to work alongside next-generation Rubin accelerators.
  • Rubin and Rubin Ultra Silicon: The Vera Rubin platform succeeds Blackwell, delivering up to 50 petaflops of FP8 inference per accelerator when paired with Vera processors and supporting up to 288 GB of high-bandwidth memory (HBM4). Rubin systems will deploy in rack-scale NVL144 configurations.
  • NVLink Fusion with NVHBM: The deployment will incorporate NVIDIA NVLink Fusion interconnects featuring custom NVIDIA high-bandwidth memory (NVHBM) architectures.
  • Nitro System and EFA Binding: Accelerators will interface with AWS Nitro System virtualization engines and Elastic Fabric Adapter (EFA) networking to provide line-rate throughput and isolation across multi-tenant partitions.
  • Workstation Expansion: AWS is also expanding Blackwell capacity for Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs.
NVIDIA and AWS Next-Generation AI Infrastructure Architecture

Dedicated AI Factories and Federal Capacity

As part of the expanded agreement, AWS and NVIDIA will construct specialized "AI factories" tailored for enterprise and public-sector compute demands.

The initiative includes provisioning a dedicated pool of 100,000 GPUs deployed across sovereign, secure AWS infrastructure specifically designated for United States government agencies and national security research workloads. These isolated clusters will support classified and restricted data processing pipelines while operating under federal compliance frameworks.

Cloud Allocation Visibility

Securing guaranteed delivery schedules for two million advanced accelerators across 2027 and 2028 provides AWS with multi-year hardware visibility in an environment where advanced packaging and high-bandwidth memory constraints continue to dictate hyperscaler capacity limits. By locking in Blackwell Ultra and Vera Rubin silicon allocations early, AWS aims to ensure continuous instance availability as client model parameters scale into tens of trillions.

Sources

Written by

More to read

  • Multi-Agent Orchestration Frameworks in Production: Comparing LangGraph, AutoGen, CrewAI, and LlamaIndex Workflows

    Production AI agent architectures have evolved past single-prompt loops and linear chains into complex multi-agent systems. When systems require multiple specialized models, tools, and validation gates to collaborate, selecting an orchestration framework determines the application's runtime latency, fault tolerance, state persistence, and debugging overhead. Four major frameworks dominate modern production multi-agent design: LangGraph, Microsoft AutoGen, CrewAI, and LlamaIndex Workflows. Each

    1 min
  • Group Relative Policy Optimization (GRPO): Mathematical Foundations, Group Baseline Advantage, Critic-Free Policy Gradients, and Reasoning Scaling

    Reinforcement learning from human feedback (RLHF) and reinforcement learning with verifiable rewards (RLVR) have become central to post-training large language models. For years, the default policy optimization algorithm in LLM alignment was Proximal Policy Optimization (PPO). While PPO offers stable policy updates through clipped surrogate objectives and Generalized Advantage Estimation (GAE), it introduces severe computational and architectural overhead when scaled to hundred-billion-parameter

    1 min
  • SandboxAQ Launches Switch to Coordinate Multi-Framework AI Agents in Slack, Teams, and Discord

    SandboxAQ has launched Switch, a framework-agnostic coordination layer designed to connect AI agents into existing enterprise chat environments, including Slack, Microsoft Teams, and Discord. The software is publicly available at no cost for self-hosted deployment on internal infrastructure. Switch addresses the operational fragmentation caused by disparate agent development frameworks. Rather than isolating autonomous assistants within bespoke web interfaces or terminal windows, the platform e

    1 min