Alibaba's plan to use fewer Western AI chips is a margin story. It is also the first read on what AI compute costs when Chinese and global clouds run on different silicon.
Alibaba Cloud plans to use fewer Western AI chips in its data centers, CFO Toby Xu told analysts on the company's August 20 earnings call. Alibaba frames the shift as a margin move. The number a regional-AI reader should keep is the unit cost of Chinese AI compute, because the silicon underneath Alibaba Cloud's customers is no longer the same silicon the rest of the world runs on.
The June quarter 2026 results, released August 20, are the cleanest print yet on what regional AI compute costs. Group revenue was RMB268.95B (~$39.6B), up 9% year over year. The new AI Cloud and Compute Services segment, which combines Cloud Intelligence Group and chip arm T-Head, posted RMB48.44B (~$7.1B) in revenue, up 45% year over year; adjusted EBITA of RMB5.63B (~$830M); adjusted EBITA margin around 12%, up from about 7% a year ago. The new AI Labs and Applications segment posted RMB3.34B (~$492M) in revenue, up 16% year over year, with an adjusted EBITA loss of RMB13.86B (~$2.0B), a 330% widening year over year and about 2.46 times the AI Cloud unit's adjusted EBITA profit. Capex for the quarter was RMB67.68B (~$10B), up 75% year over year, driven by uneven customer purchase timing, CPU compute capacity additions, and higher chip-component prices. CNBC's recap and ts2.tech's segment breakdown put those numbers in the same frame: a Chinese cloud unit scaling AI revenue faster than peers, while its in-house AI lab burns capital at a rate that now exceeds the cloud unit's profit.
The "fewer Western chips" line is not a procurement tweak. On the call, Xu spelled out the unit economics: Alibaba runs servers for five years; AI servers generate enough revenue to cover their costs in three years; years four and five produce free cash flow. CEO Eddie Wu added that 2018- and 2020-vintage servers carrying Nvidia V100 and A100 accelerators "are still being used by customers at near full capacity."
That detail should give a non-beat reader pause. Eight-year-old V100s, on Nvidia's older Volta architecture, are still doing real AI work for Alibaba Cloud's customers. The reason is software: the CUDA SDK, the Nvidia programming stack most AI models and frameworks target, locks workloads to Nvidia hardware, and the older cards still work for inference even when newer domestic accelerators are available. V100 silicon is itself export-controlled under the US regime, but it shipped before the strictest controls landed; the practical point is that Alibaba Cloud's customers have not moved off it.
The newer chip on Alibaba's roster is T-Head's Zhenwu M890 GPU, now formally disclosed and described by management as in scaled mass production. The M890 is positioned against Nvidia's H100, the Hopper-generation part that US export controls bar from sale to Chinese customers, rather than against Nvidia's newer Blackwell generation. CryptoBriefing reports that T-Head has cumulatively shipped 470,000 AI chips, with more than 60% going to external customers across internet, finance, and autonomous-vehicle sectors. T-Head is preparing an IPO.
The competitive picture: T-Head is targeting the workload class Nvidia can no longer sell into China, while the CUDA-locked V100 and A100 fleet keeps running older inference jobs. Both reduce Alibaba Cloud's exposure to the dollar-denominated Nvidia bill. Group gross margin, around 34.5% in the quarter, is the line Alibaba is selling.
In March 2026, Alibaba Cloud raised prices on AI compute and storage by as much as 34%. The August print lands the unit economics that make that price hike stick. AI Cloud adjusted EBITA margin moved from roughly 7% to 12% in a year. The three-year AI and cloud investment plan is RMB380B (~$56B), and about half, roughly RMB190B (~$28B), had been deployed by the end of June 2026.
The regional-AI read is straightforward: a Chinese cloud customer buying tokens from Alibaba Cloud is now buying compute whose silicon mix is partially T-Head, partially older Nvidia, and increasingly not the newest Nvidia at any price. A global cloud customer buying the same model from AWS, Azure, or Google Cloud is buying compute off the latest Nvidia or AMD accelerators, because the export controls that bar H100 sales to China do not apply elsewhere.
That is the multi-polar AI compute stack. The price gap between the two stacks is the number a regional-AI reader should track across Alibaba's next four quarterly prints, and the same number will reappear when Tencent, Baidu, ByteDance, or the Indian hyperscalers report. Alibaba's wire-line margin story is the surface. The unit-cost divergence underneath it is the story that survives the next regional cloud print.