Zhongcheng Hualong is selling cluster systems, not chips, as China's AI infrastructure race pivots from training to running models at scale.
A Beijing chipmaker is selling AI infrastructure one rung up the stack: not a single chip, but a 10,000-GPU system as one product.
Zhongcheng Hualong used its 2026 GPU launch event in Beijing on August 21 to unveil the HL200 inference chip and a super-node cluster that scales from 8 to 10,240 cards. Each cabinet holds 64 GPUs, a single node stacks 1,024 cards vertically, and the system expands horizontally to more than 10,000 cards (QbitAI).
The unit shift matters because the AI industry's bottleneck is moving. Training a frontier model is largely done; running one for millions of users, the inference workload that powers chatbots, is where latency, throughput, and cost-per-token get decided. Zhongcheng Hualong is positioning for that serving market with an inference-first chip: native FP4 and FP8 low-precision math (faster number formats for running models), 5.12 TFLOPS per watt (trillions of math operations per second per watt), and an OpenAI-compatible API. Mydrivers carried the same release (Mydrivers).
Chairman Dr. Wang Jiacheng framed the posture: "We are not followers of parameters, we are definers of inference efficiency" (QbitAI). The company says HL200 was tested against DeepSeek V4 Flash, DeepSeek V4 Pro, and Zhipu's GLM 5.2, with a stated edge on Prefill (time to first token) and decode latency over unnamed "international peers of the same tier." It announced partnerships with China Energy Engineering's CPE, Inspur, Unicloud-Zhixuan, and Guanghuan Cloud, and pointed to existing China Mobile and China Telecom deployments in 20+ provinces.
The load-bearing numbers, 5.12 TFLOPS/W, the 10,240-card ceiling, and the partnership scale, are company-stated, not independently benchmarked. Whether the cluster ships at this scale is the open question.