At WAIC 2026 (the World Artificial Intelligence Conference in Shanghai), twelve Chinese AI chip executives agreed on the same shift: the competition is no longer about peak specs, but how cheaply a system can produce tokens, the basic units AI
For two years, every AI chip launch has been a spec fight. The contest was peak FLOPs (floating-point operations per second, a raw measure of compute throughput), transistor count, gigabytes of memory bandwidth, and the number you could put in the biggest font on the slide. That number is no longer the one that matters. The new scoreboard is tokens produced per dollar per watt, measured at the system level. A token is the basic unit of text an AI model reads or writes, roughly three-quarters of a word in English, and the cost of producing one has become the single figure that buyers compare.
At WAIC 2026 in Shanghai (the World Artificial Intelligence Conference, China's largest annual AI industry gathering), twelve executives who actually sell AI compute to Chinese data centers said this in nearly the same words. The roundtable, hosted by Leiphone (雷峰网), produced an unusual consensus: inference, not training, is the dominant AI workload now, and Agent applications are the main reason. Deloitte's 2026 TMT Predictions forecast that inference will be roughly two-thirds of all AI compute this year, the macro anchor for what the execs were describing.
The unit of competition has moved from the chip to the supernode, a tightly coupled cluster of hundreds or thousands of chips acting as one compute unit, with shared memory and high-speed interconnect. NVIDIA introduced the template at GTC 2024 with the GB200 NVL72, 72 GPUs wired into a single rack-scale system. Huawei followed a year later at WAIC 2025 with a 384-card Ascend supernode and this year showed a 1024-card Ascend 950 supernode. The trajectory is one clean doubling cycle per year, and the supernode concept has crossed from a single-vendor story into an industry path. Moore Threads (摩尔线程) showed its MTT C256 supernode design, and XingIn (奕行智能) launched what it called the first RISC-V AI compute supernode, an open-ISA play that attacks the "one card, one software stack" fragmentation problem. Ten other domestic vendors showed their own designs at the show.
The exec voices make the metric shift concrete. Moore Threads VP Ma Jian put the new competitive frame bluntly: buyers now ask how many useful tokens a system produces, not what the peak single-chip FLOPs are. Several execs framed the change as a memory-wall problem, where the bottleneck is no longer raw compute but the speed at which data can move between chips and memory, which is why tightly coupled supernodes beat loosely connected clusters at the same nominal FLOPs. The Agent workload is what makes this load-bearing: an Agent calls the model many times in a loop, so the cost of each token compounds across a session, and small per-token savings turn into the difference between a product that ships and one that does not.
The shift is partly honest and partly strategic. When single-chip peaks are dominated by one company, the industry is happy to redefine the metric. System-level tokens-per-dollar is a number where Chinese vendors, working on mature process nodes and optimizing for the inference workload, can compete on different terms than the frontier-training leader. Dongfang Suanxin (东方算芯) explicitly bet on a mature-node strategy, using software-defined 3D near-memory computing to run a different race than the one measured in transistor count. The framing is not a denial of the leader's lead. It is an attempt to change what the leader is being measured against.
The edge side is still waiting. The same WAIC roundtable, weighted toward cloud and infrastructure vendors, heard the edge and embedded silicon side described as still in a "waiting for a hit product" state. The supernode story is a cloud story; the consumer and embedded silicon that would put an LLM on a phone, a car, or a home appliance has not yet produced a product that sells at the scale investors want. The honest read for now is "cloud races ahead on supernodes, edge waits for its iPhone moment."
The test case is whether the supernode consensus holds when the first commercial deployments report real cost-per-token numbers. Huawei's 1024-card Ascend 950 is the load-bearing example. If it produces tokens at the price the execs are claiming, the scoreboard change is permanent. If it does not, the next round of WAIC will find a different metric to argue about.