SAN FRANCISCO — Amazon Web Services met with engineers in May, directing them to reduce CPU waste as internal demand for EC2 instances now causes multi-day waits — up from hours.
One engineer said the delays are unlike anything seen in years at Amazon.
The strain traces directly to AI agents. One coding agent at Amazon ran up $1.8 million in token costs last month, exceeding its development budget by 860 percent.
Traditional data center infrastructure ran GPU-to-CPU ratios of eight-to-one or four-to-one, with CPUs primarily feeding data to GPUs for inference. That ratio is moving toward parity.
Agentic workloads drive the shift. They rely on frequent tool calls that execute on the CPU and require complex orchestration of GPU inference, pushing CPUs into a central role.
Chipmakers are moving to capture the demand. AMD recently unveiled its Zen 6 "Venice" CPU, its first data center-first architecture launch in decades. Nvidia is now promoting its Vera CPU as agentic AI infrastructure becomes a larger share of cloud spending.
AWS deploys CPUs from AMD and Intel alongside its proprietary Graviton5 chip. Graviton5 uses an Arm-based architecture, as does Nvidia's Vera CPU.
The capacity shortages are concentrated in spot instances. A consultant said contracted capacity has not seen similar shortages, meaning enterprises with long-term agreements face fewer disruptions.
Intel, AMD and other chipmakers report that companies are accepting whatever CPU supply is available to meet compute demand.
