Alibaba says its Zhenwu M890 super node, 64 in house chips linked, now runs the 2.4 trillion parameter Qwen3.8 model on its Bailian cloud; the 'first in China' and '1.5x' claims are its own measurements.
Alibaba has wired its own chip, its own flagship model, and its own cloud into a single live inference stack. The company says the Zhenwu M890 super-node is now running Qwen3.8, a roughly 2.4-trillion-parameter mixture-of-experts (MoE) model, on Bailian, its model-as-a-service platform. Each instance links 64 M890 chips through Alibaba's ICN Switch 1.0 interconnect at 800 gigabytes per second, pooling about 9 terabytes of high-bandwidth memory.
Alibaba describes the M890 as a "training-inference unified" chip that natively supports FP32 down to FP4 precision, the lower formats letting inference run faster on the same silicon. With chip, model, and cloud all in-house, the company can co-design operator and hardware rather than depend on a third-party accelerator roadmap. The 9-terabyte memory pool and 800-gigabyte-per-second links are sized to hold a 2.4-trillion-parameter model in a single instance, with the chips sharding the MoE's experts so only a fraction activates per query.
Two load-bearing claims are Alibaba's own. The company says the M890 super-node is the first domestic super-node to run a 2-trillion-parameter model, and that the stack delivers up to 1.5x faster inference in agentic workloads through operator-plus-hardware co-design. Neither figure has been independently benchmarked. The text is a release Alibaba provided to 量子位, which reposted it with authorization; it is not a third-party evaluation.
Per a techstrong.ai writeup of SemiAnalysis analyst Myron Xie, Alibaba-designed chips are gaining traction with Chinese enterprise buyers, but the M890's advertised memory capacity and bandwidth still trail Western benchmarks and key compute performance metrics have not been publicly disclosed. Other Chinese outlets that re-reported the announcement added no new facts.
Two things will tell whether the closed loop holds. First, whether customers can actually hit the M890 stack on Bailian with a defined workload and a public price, rather than through a single internal demo. Second, whether the "first in China" claim survives contact with competing Chinese accelerator stacks from Huawei Ascend, Cambricon, or the newer domestic players. Alibaba may have set the bar; the question is who else clears it.