At Huawei Connect 2026, Huawei conceded the per chip AI race against Nvidia and reorganized its Ascend AI chip family around SuperClusters of up to 512,000 AI processors.
Huawei is not trying to catch Nvidia on the chip. It is trying to make the chip the wrong unit of competition.
At Huawei Connect 2026, deputy chairman David Wang pulled the next two chips in the Ascend 960 family forward by three to nine months and openly reorganized the company's AI roadmap around cluster-scale architecture. The Ascend 960DT will ship in Q1 2027, three quarters earlier than previously planned. The Ascend 960R will ship in Q3 2027, three months ahead of schedule. Both are neural processing units (NPUs), Huawei's AI accelerators in the same category as Nvidia's GPUs.
Wang outlined SuperPoDs of up to 4,096 optically linked NPUs, and SuperClusters of up to 512,000 NPUs built by stitching SuperPoDs together. SuperPoDs are rack-scale blocks of tightly wired accelerators; SuperClusters are systems of those blocks. The idea: instead of one chip fighting one chip, Huawei is selling a system where thousands of weaker chips are connected so densely that the cluster behaves like one large computer, with the optical links between them doing the work that single-chip performance would otherwise have to do.
"We acknowledge that Huawei is still lagging behind Nvidia in single-chip processing power," he said, per Network World. The strategy is to win on the layer above the chip: the optical fabric, the interconnect, the rack-scale engineering that decides whether thousands of NPUs can train a frontier model without bottlenecking on each other. Per-chip FLOPS, the standard measure of AI accelerator performance, stop being the headline number; per-cluster throughput, the amount of useful training a wired-together system can sustain, becomes the headline metric.
That bet has two parts. The first is technical. Optical links between NPUs, the cables that let one chip talk to its neighbors with minimal latency, become the bottleneck once a system grows past a few thousand accelerators. Huawei is treating that fabric as the product, not a side feature, and Wang used the keynote to initiate what TechTimes described as a global AI interconnect standard: an open specification that any vendor could build to, a move that signals Huawei wants to set the rules for cluster-scale AI rather than win customers one at a time.
The second part is geopolitical. Outside China, Huawei faces US export controls that limit access to leading-edge manufacturing. A product framed as "catching up to Nvidia per chip" invites tighter attention from Washington, while one framed as "a different theory of how AI infrastructure scales" leaves more room to sell to overseas customers who want an alternative without directly competing on the metric Washington polices. Network World suggested that Huawei may have deliberately understated per-chip capability to keep that margin; the framing reads as analyst speculation, not confirmed corporate intent, and Huawei has not confirmed it.
The pull-in is real and dated. The strategy is named in the keynote, not inferred from product launches. The numbers, 4,096 NPUs per SuperPoD and 512,000 per SuperCluster, are upper-bound ceilings, not deployed scale. What the bet does not yet have is independent benchmarks showing that cluster-level training throughput per dollar actually beats Nvidia systems, or confirmation that customers outside China are buying SuperClusters at the announced scale. The architecture is a public bet. The receipts are still pending.
Watch item: a production SuperPoD deployed near 4,096 NPUs with disclosed throughput and pricing would be the first hard evidence that the cluster bet is engineering rather than marketing.