The 4 bit quantized model fits on an RTX 3090, putting a 29B parameter office assistant within reach of any team willing to run it on prem.
A 29-billion-parameter open-source model from China Telecom can run an office workload on a single consumer graphics card, after 4-bit quantization compresses a footprint that would otherwise exceed most desktop GPUs.
The model, Xing4.0-29B-A4B, is a mixture-of-experts design with roughly 4 billion parameters active per inference, QbitAI reports. At full precision, the model occupies about 60GB of GPU memory; 4-bit quantization cuts that to roughly 15GB, which is what lets it fit on a single RTX 3090 or 4090. A 256K-token context window, extendable to 512K, is cited for whole-codebase and long-document work.
The release targets on-prem enterprise use: long-document review, multi-file code work, and tool-using agent tasks that chain weather, calendar, and travel services. A 200-page bid-document review is cited as taking about 20 minutes across technical, commercial, and price dimensions, all on local hardware. The company also reports a customer-service deployment with a 90%-plus success rate on complex business handling, though that figure is not independently verified.
Distribution runs through GitHub, Hugging Face, Gitee, ModelScope, and Modelers, with API access available. Training used Huawei Ascend 910C hardware and the MindSpore / MindFormers stack, per the same QbitAI report, which frames the release as a "full domestic" stack. At publication, the model ranked fourth on Hugging Face's trending list.
What remains untested: independent benchmarks against equivalent open models, and whether the 15GB footprint holds for workloads beyond the company's own demos.