IN Brief:
- AWS plans to deploy two million additional NVIDIA GPUs during 2027 and 2028.
- The programme adds Vera CPUs, NVLink Fusion with custom NVIDIA HBM, networking, and secure government infrastructure.
- Amazon Robotics will use NVIDIA's physical AI stack for simulation, training, functional safety, and robot validation.
Amazon Web Services and NVIDIA plan to deploy two million additional NVIDIA GPUs across AWS infrastructure during 2027 and 2028, extending their collaboration into CPUs, high-bandwidth memory, scale-up interconnects, networking, data processing, and robotics.
The planned deployment will include Blackwell Ultra, Rubin, and Rubin Ultra accelerators across AWS global infrastructure and AI factories. AWS had already announced plans to add more than one million NVIDIA GPUs beginning in 2026, but the companies say subsequent demand has led to a further two-million-unit expansion for the following two years.
The size of the GPU fleet is only one part of the engineering programme. AWS and NVIDIA are also working to introduce Vera CPU-based infrastructure, while Amazon’s Annapurna Labs will extend its use of NVLink Fusion to NVIDIA’s custom high-bandwidth memory technology. The objective is to combine Trainium and NVIDIA GPUs within common rack-scale architectures rather than treating custom and merchant silicon as separate compute estates.
That integration reflects the limits of scaling accelerator numbers alone. Large AI clusters depend on memory bandwidth, scale-up links inside the rack, scale-out networking between racks, host CPUs, storage, software, power, and cooling. Adding GPUs faster than those surrounding systems can expand merely relocates the bottleneck from compute throughput to data movement or infrastructure capacity.
The collaboration includes plans for 100,000 GPUs on secure AWS infrastructure for US federal and national-security workloads. NVIDIA GPU and AWS Trainium EC2 instances will continue to use the AWS Nitro System and Elastic Fabric Adapter, providing a common security and networking layer as the underlying processors and interconnects change.
Data processing is also being pushed onto accelerators. AWS and NVIDIA say GPU-accelerated processing on Amazon EMR using NVIDIA cuDF can deliver up to 3.7 times faster processing and 30% better price performance than the CPU configurations used for their comparison. Amazon OpenSearch is separately gaining GPU-accelerated vector indexing, with the companies claiming index construction up to nine times faster and at a quarter of the cost.
Those figures are configuration and workload dependent, but the direction is significant because GPUs are being applied before a model reaches inference. Data preparation, feature engineering, retrieval, vector indexing, and analytics can consume substantial resources around an AI workload, so accelerating those stages changes the proportion of the infrastructure that may eventually require GPU or other specialised processing.
The semiconductor demand extends beyond GPUs. NVLink Fusion and custom HBM require memory, packaging, interconnect, substrate, and power-delivery capacity, while high-density clusters need switches, optical modules, network interface hardware, voltage regulation, and cooling controls. A two-million-GPU deployment programme therefore creates requirements across a much wider electronics supply chain than the headline accelerator count suggests.
NVIDIA has separately linked AI compute expansion to financing platforms intended to support very large infrastructure projects. The AWS commitment provides a more specific technology path, with defined accelerator generations and a two-year deployment period alongside work on CPUs, memory, networking, and software.
The agreement reaches edge hardware through Amazon Robotics. The operation will use NVIDIA Jetson, Omniverse libraries, and the Isaac robotics platform across simulation, synthetic data generation, training, route optimisation, functional safety, and real-to-simulation validation, with much of the computational work running on GPU-accelerated EC2 infrastructure.
That creates a feedback path between central compute and physical systems. A robot may execute inference locally against sensors and actuators, but simulation, synthetic data, model training, testing, and fleet-level validation can consume far larger compute resources away from the machine. Edge AI consequently drives both embedded processor demand and central accelerator demand rather than shifting workload entirely from one environment to the other.
The deployment spans several NVIDIA generations, so AWS will not be installing two million identical devices into a static architecture. Memory, interconnect, server design, networking, and power characteristics will change as Blackwell Ultra gives way to Rubin and Rubin Ultra, forcing infrastructure qualification to continue through the programme.
The arithmetic remains formidable. Two million accelerators arriving over two years require matching packaging, memory, server manufacturing, networking, power, and data centre capacity on a similar timescale. The commercial announcement is large; the harder electronics problem is ensuring every layer around those GPUs scales quickly enough that expensive compute does not spend its time waiting for memory, network bandwidth, or electricity.


