At China's flagship AI conference in Shanghai, Qualcomm sketched a four chip stack for an "agent phone" that runs an AI assistant 24/7, and skipped the math on battery and consent.
On Sunday at the World AI Conference in Shanghai, Xu Hao walked reporters through the four-chip stack that is supposed to make that possible, and skipped the part of the math that falls on the buyer.
Qualcomm calls the result the "agent phone": a handset that runs an AI assistant continuously in the background, not just when you tap it, and that can plan, sense, and act without round-tripping to a data center. The architecture laid out at WAIC's on-device AI forum, captured by Lei Feng Wang, Sina Tech, and The Paper, is specific. The "agent" part depends on four pieces of silicon working in parallel.
The first is a new CPU that handles task planning and control: the part of the agent that decides which app to open, which sensor to read, when to interrupt you. The second is a low-power sensor hub that runs 24 hours a day, fusing motion, location, and audio so the agent can react before you ask. Two examples given: muting a phone when the user walks into a theater, adjusting cabin temperature and seat position when a known driver gets in. The third is a neural processing unit (NPU) tuned for 2-3 billion parameter language models, small enough to fit in a phone, large enough to hold a multi-turn conversation. Xu Hao named ModelBest's MiniCPM family as the kind of model the NPU is sized for; Qualcomm and ModelBest already collaborate on end-side models for phones and vehicles. The fourth piece is the modem: 5G, and eventually 6G, for the moments when the agent hands a hard question to the cloud and waits for an answer.
The current flagship, the fifth-generation Snapdragon 8 Elite, already ships most of this stack. The next platform Qualcomm teased for later in 2026 carries no name, no CPU part number, no NPU tops in trillions of operations per second, and no ship date. The WAIC talk treats the unnamed platform as the one that delivers the always-on agent experience. The current chip, on the evidence presented, is not asked to.
A 2-3 billion parameter model in your pocket has to run often. Xu framed token demand on a roughly 10x curve per agent generation: a single-turn chatbot burns on the order of 10,000 tokens, a multi-turn conversation 100,000, an agent scheduling other agents 1,000,000. "10x per stage" is Qualcomm's illustration, not an audited benchmark, but the direction is the point: the NPU is the part of the phone asked to do more work each year, and it is the part that draws the most power per useful token. Thermal envelopes on phones are not kind to that curve. A 2-3B model running multi-turn agent tasks without throttling has not, on the evidence presented Sunday, been independently demonstrated.
The sensor hub is always on, by design. A low-power chip that watches your motion, location, and audio all day is, by definition, an always-listening device. The examples given—auto-mute in a theater and in-cabin personalization—are friendly ones. The privacy posture for a 24/7 sensor pipeline on a consumer phone is not solved by friendly examples. Apple's Private Compute Cloud pattern, where raw sensor input is processed on-device and only anonymous intent tokens leave the phone, is one credible answer. Qualcomm did not point to a comparable architectural commitment on Sunday, and the ModelBest collaboration on end-side models does not, on the public record, address it.
Xu Hao told the audience that in one to two years, on-device models will match the capability of cloud models running on roughly ten times the compute, and that the trend will pull "more large models" onto end-side hardware in smaller form factors. That is a stated roadmap, not a settled outcome. ModelBest has made versions of the same claim from the model side; Sunday was the first time a major handset chipmaker repeated it from the silicon side at a public conference. The curve has pointed that way for three years. The WAIC talk assumes the always-on, multi-turn, sensor-aware agent is the same product as the 2-3B model the NPU is sized for today. Independent reporting has not yet tested that assumption.
Four numbers are missing from the WAIC talk. They are the four a buyer should ask any phone vendor for before the next upgrade. First, what the unnamed 2026 platform ships with, in concrete tops-per-second and battery figures, not "stronger." Second, whether the sensor hub runs end-to-end on-device with intent-only egress, or whether raw sensor data leaves the phone under any agent workflow. Third, what the 5G/6G hand-off actually hands off: the question, the context, or the raw sensor window. Fourth, whether an independent benchmark can run a 2-3B model through a multi-turn agent task on a shipping phone without thermal throttling in the first fifteen minutes.