The AWS NVIDIA 2 million GPU Vera Rubin infrastructure expansion announced on August 26, 2026 is less a capacity story than a structural one. Amazon Web Services and NVIDIA used the announcement to outline a full-stack buildout that goes well beyond renting more accelerators: it folds NVIDIA’s new Vera CPU, custom NVHBM memory, and Annapurna Labs silicon into a single fabric alongside Blackwell Ultra and Rubin GPUs, and it commits 100,000 of those GPUs to a federal AI factory line item that now stands as the largest non-Stargate AI procurement disclosure of 2026.
The headline figure is two million additional NVIDIA GPUs scheduled to come online across AWS’s global infrastructure in 2027 and 2028. That sits on top of the more than one million GPUs AWS already committed to deploy starting in 2026 at GTC 2026. Combined, AWS is now publicly on the hook for roughly three million NVIDIA accelerators across a roughly three-year window. The silicon mix for the new tranche spans NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra. AWS framed the move as a response to demand that has already outpaced the GTC forecast, with workloads spanning agentic AI, scientific discovery, enterprise automation, and physical AI moving from pilot to production.
The more consequential change is heterogeneous. For the first time, a hyperscaler is integrating a custom NVIDIA CPU alongside its own Trainium and Annapurna ARM silicon, all behind one Nitro System and Elastic Fabric Adapter (EFA) fabric. NVIDIA Vera, purpose-built for the next generation of AI, is positioned as an additional compute option for agentic workloads that need high-performance CPU cycles next to accelerated infrastructure. AWS CEO Matt Garman framed the partnership in terms of customer choice: “Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together.”
The memory and interconnect layer is where the architecture gets specific. At re:Invent 2025, AWS announced support for NVIDIA NVLink Fusion on next-generation Trainium. The new announcement extends NVLink Fusion to work with NVIDIA’s custom high-bandwidth memory, branded NVHBM, and opens the door for Annapurna Labs — Amazon’s in-house silicon unit — to tap that memory and scale-up fabric. The practical effect is that Trainium chips and NVIDIA GPUs can now share a common rack-scale architecture, which removes one of the long-standing friction points in mixing in-house and merchant silicon inside a single training or inference deployment.
On the federal side, the disclosure is unusually concrete. AWS and NVIDIA said they plan to build AI factories for the U.S. government, with 100,000 GPUs running on secure AWS infrastructure for federal and national-security workloads classified at Impact Level 6 (IL6) and above. For context, IL6 covers controlled unclassified information and a wide band of defense and intelligence workloads that have historically resisted commercial-cloud migration. Pairing 100,000 GPUs with that classification tier is a signal that the procurement bar has shifted from pilots to industrial-scale deployment.
For commercial customers, the most immediate product is the new Amazon EC2 G7 instance, the first major cloud offering accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. AWS is also expanding existing NVIDIA Blackwell capacity on the platform. The performance claims are specific: G7 delivers 4.6x AI inference performance and 2.1x graphics performance versus the previous-generation G6. That puts a workstation-class GPU behind an EC2 SKU for the first time at meaningful scale, which matters for digital twins, simulation, robotics visualization, and any workload that has been bottlenecked on single-node GPU memory.
The data plumbing got its own set of numbers. GPU-accelerated Amazon EMR using the NVIDIA cuDF library is now up to 3.7x faster with 30% better price-performance, while GPU-accelerated vector indexing on Amazon OpenSearch, powered by NVIDIA cuVS CUDA-X libraries, runs up to 9x faster at one-quarter the cost. For Bedrock customers, NVIDIA Nemotron open models remain available as fully managed serverless options and are also deployable through Amazon SageMaker, which preserves model flexibility as enterprises mix proprietary and open weights. AWS NVIDIA 2 million GPU Vera Rubin infrastructure expansion.
Robotics closes the loop. Amazon Robotics is adopting NVIDIA’s full-stack physical AI platform, including Jetson for edge inference, Omniverse for simulation, and Isaac for open robotics development. That ties the same Vera-Rubin-NVHBM stack that trains frontier models back to the fleet optimization and route planning problems inside Amazon’s own warehouses, and it gives NVIDIA a marquee physical-AI reference customer on AWS. NVIDIA founder and CEO Jensen Huang tied the threads together: “We are expanding our partnership across the full stack — GPUs, CPUs, networking, open models and software — to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver.” The AWS NVIDIA 2 million GPU Vera Rubin infrastructure expansion is, in effect, the contract that turns that sentence into a delivery schedule.
Source: NVIDIA Newsroom

