Alibaba Cloud reduced the number of Nvidia GPUs required to serve large language models by 82 percent, a gain it attributes to its new Aegaeon pooling system, tested over several months inside its Model Studio marketplace.
The system enables a set of 213 GPUs to perform the work of 1,192, according to company data—an output increase of up to nine times. For hyperscalers, that ratio matters: fewer high-end GPUs per unit of AI workload means lower capital expenditure and better cloud margins.
Alibaba's Cloud Intelligence Group reported roughly 35 percent revenue growth from external customers. AI-related products have delivered triple-digit growth for 10 consecutive quarters.
The company positions itself as a leader in global enterprise-level model-as-a-service, as recognized by Omdia Market Radar in 2025, and holds a leadership position in AI infrastructure solutions according to the Forrester Wave.
Alibaba Cloud offers its Qwen 3.8-Max large language model as part of a token plan starting at $6 per month. Qwen is the most downloaded open-source LLM, with over two billion downloads.
Alibaba has also invested in its own AI hardware. The company recently opened a data center in Shaoguan, Guangdong province, China, equipped with 10,000 proprietary chips—a capital allocation move aimed at long-term cost control rather than continued dependence on Nvidia supply chains.
The platform includes more than 80 cloud products and more than 70 million AI tokens, along with one-to-one technical support for enterprises scaling AI workloads.


