Amazon Web Services (AWS) and NVIDIA have officially escalated their 16-year alliance, announcing a blockbuster expansion aimed at crushing the surging global demand for AI compute. In a landmark update, the duo confirmed plans to inject an additional 2 million NVIDIA GPUs into AWS’s global data center infrastructure between 2027 and 2028. This massive scale-up, coupled with deep co-engineering across CPUs, networking, and open models, solidifies AWS as the premier cloud destination for running NVIDIA’s AI stack, catering to everyone from frontier AI labs to top-tier government agencies.
Surging Demand Drives Unprecedented GPU Capacity
AI workloads are scaling at breakneck speed, transitioning from experimental pilots to full-scale production across agentic AI, scientific discovery, and enterprise automation. To keep pace, AWS is expanding beyond its previously announced 1 million GPUs. The new commitment includes deploying NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs throughout 2027–2028.
This expansion will support AI factories and mission-critical workloads requiring massive parallel processing power. Notably, AWS is also the first major cloud provider to offer NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs via Amazon EC2 G7 instances, which deliver a staggering 4.6x improvement in AI inference performance and 2.1x better graphics performance over prior generations. Additionally, the companies are collaborating on NVIDIA Spectrum networking to further optimize large-scale AI training clusters.
Full-Stack Innovation: Vera CPUs and NVLink Fusion
Beyond sheer GPU count, the partnership is redefining heterogeneous computing. AWS and NVIDIA are actively integrating NVIDIA Vera CPU-based infrastructure into AWS, offering customers high-performance CPU options for agentic AI tasks that require robust parallel processing alongside accelerated hardware.
Furthermore, the collaboration on NVIDIA NVLink Fusion technology is expanding to include custom NVIDIA high-bandwidth memory (NVHBM). By partnering with memory suppliers, Amazon’s Annapurna Labs can now leverage this custom memory tech, enabling Trainium chips to access faster, more power-efficient memory. This integration ensures seamless co-existence of Trainium and NVIDIA GPUs within a unified rack-scale architecture, optimizing performance for complex AI workloads.
Powering Federal AI with Top-Tier Security
National security demands cutting-edge, secure AI. Recognizing this, AWS and NVIDIA are constructing dedicated AI factories for the U.S. Government, planning to deliver 100,000 GPUs on AWS’s highly secure infrastructure. This initiative supports workloads classified at Impact Level 6 (IL6) and above, positioning the partnership as a cornerstone of federal AI advancement.
To guarantee security and reliability across all deployments, every NVIDIA GPU-based and Trainium-based EC2 instance will continue to leverage the AWS Nitro System and the Elastic Fabric Adapter (EFA). Together, these technologies ensure production-grade network performance, isolation, and the low-latency scaling required for enterprise and government AI workloads.
Accelerating Data Pipelines and Open Model Access
Data remains the lifeblood of AI. The expanded collaboration brings GPU-accelerated data processing to Amazon EMR using the NVIDIA cuDF library, delivering up to 3.7x faster processing and 30% better price-performance than CPU-only setups. Concurrently, GPU-accelerated vector indexing on Amazon OpenSearch Service—powered by NVIDIA cuVS—achieves up to 9x faster index construction at a quarter of the cost, dramatically improving RAG pipelines and semantic search.
Additionally, NVIDIA’s Nemotron open models are now available on Amazon Bedrock and SageMaker, giving customers flexible, managed access to NVIDIA’s latest open-source AI models with the robust security and operational tooling of AWS.
The partnership extends into the physical world through Amazon Robotics, which is adopting NVIDIA’s full-stack physical AI platform—including Jetson, Omniverse, and Isaac. This collaboration focuses on massive simulation, synthetic data generation, robot training, and real-to-sim validation, accelerating the development of next-generation warehouse automation and intelligent robots.
Executive Insights on the AI Gold Rush
Highlighting the strategic importance of the collaboration and the focus on seamless interoperability, AWS CEO Matt Garman stated:
“Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together,” said Matt Garman, CEO of AWS.
“That’s why we’ve invested deeply with NVIDIA to make AWS the best place to run NVIDIA AI technologies, optimizing performance across our infrastructure from networking and security to deployment. This expanded collaboration gives frontier labs, enterprises, and governments even more ways to build and deploy AI on AWS.”
NVIDIA CEO Jensen Huang echoed this sentiment, emphasizing the unprecedented historical scale of the partnership and the urgency of demand:
“NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast,” said Jensen Huang, founder and CEO of NVIDIA.
“For 16 years, we have scaled NVIDIA computing in the cloud together.
Now we are expanding our partnership across the full stack—GPUs, CPUs, networking, open models and software—to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver.
This expansion reflects customers’ demand for NVIDIA’s platform on AWS.”
With a clear roadmap for 2 million new GPUs, enhanced custom silicon integrations via NVLink Fusion, specialized federal AI factories, and a relentless focus on data acceleration, the AWS-NVIDIA juggernaut is set to redefine the boundaries of artificial intelligence.
As demand runs ahead of every forecast, this expanded collaboration ensures that customers—from global enterprises to national security agencies—have the resilient, high-performance infrastructure required to build and deploy the next generation of intelligent applications at scale.
