Xiaomi's new Xring O3 chip puts 44MB of on chip memory, more parallel math units than most laptop processors, and wide data processing lanes on a die that fits in your pocket.
Xiaomi's Xring O3 puts roughly 44MB of on-chip cache, 21 execution ports, and 128-bit SIMD lanes on a die that fits in your pocket. The story is what those specs add up to: the architectural playbook that built laptop CPUs is now in mass-market phone silicon, and the next year of on-device AI, multitasking, and mobile games will be written in width, cache, and matrix extensions rather than clock speed.
44MB is more on-chip cache than most current Intel and AMD laptop CPUs carry, and it is the single spec that explains why a phone SoC can start behaving like a small laptop. Larger caches reduce the number of times the cores have to leave the chip and talk to slower DRAM, which is the bottleneck that has kept mobile CPUs from doing serious desktop-class work. Xiaomi's design adds 21 execution ports, the datapaths a CPU can use in a single cycle to issue work, and routes six of them through 128-bit-wide SIMD units, the kind of parallel-arithmetic hardware that used to be a desktop-only feature. The chip also ships with ARM's SME2 matrix extensions and SVE2 data-parallel SIMD, both aimed at on-device AI and the kind of bulk number-crunching phones have not had to do before.
The benchmark numbers floating around put the Xring O3 at 3,945 single-core and 15,221 multi-core in Geekbench 6, reported by leaker @universeice and read by computer-science professor Daniel Lemire as roughly matching Apple's cores per thread and pulling meaningfully ahead on multithreaded work. Xiaomi separately claims an AnTuTu score above 5 million, framed as the fastest smartphone SoC on the market.
Geekbench is one workload family, and a Hacker News thread on the same scores flags that under real phone thermals the C1-Ultra cores settle closer to 3,300 single-core, a real-world drop that matters more for daily use than a lab run. The Xring O3 is also genuinely hard to find in shipping phones today, which means there is no independent reviewer benchmark yet. Apple can erase the single-threaded parity in its next A-series cycle. The category lead is the more durable story than any one quarter's scoreboard.
The "Xiaomi's own CPU" framing understates how the Xring O3 was built. The big cores are C1-Ultra, the same core IP MediaTek licenses to partners and uses in its own Dimensity 9500, with TSMC's 3nm process and ten cores total wrapped around it. Xiaomi's contribution is the SoC integration, the cache and uncore design, and the willingness to push cache and port counts beyond what MediaTek ships in its own retail chip. That is real engineering work, but it is a different story from a fully homegrown mobile core.
The architectural direction is the part that survives the next benchmark cycle. Wider out-of-order cores, more cache, more parallel arithmetic per cycle, and dedicated matrix extensions are the same moves that took laptop CPUs from adequate to dominant over the last fifteen years. Phone chips are running the same playbook now, which changes what phones can do: serious on-device language models without round-tripping to the cloud, heavier multitasking without the UI stutter, and mobile games that can afford to do real physics and AI per frame instead of pre-baked animations. The Xring O3 is the cleanest data point yet that this shift has crossed from flagship experiments into mass-market silicon, and the next twelve months of phone launches will be the first cycle where every chip announcement reads through the same lens: more cache, more lanes, more matrix units, and a clock speed that matters less than it used to.