AI agents, the workloads that hold state, plan, and call tools across many steps, have made peak NPU throughput the wrong unit of measure. The next compute platform will be judged on whether the CPU and NPU can keep pace with each other, not on how many trillion operations the accelerator can perform alone. Singapore's Acrab shipped the first silicon to publicly pin its design thesis to that shift: a 700 TOPS NPU whose value is described as conditional on the rest of the chip.
Agents pin the system in interactive loops, where a slow CPU scheduling a prefill or a tool call idles the NPU on the next token. Treat TOPS as a coordination metric, and the competition becomes who can keep the accelerator fed. Treat it as a peak number, and the race is still about the biggest chip. Vertex Growth's participation in Acrab's $130M Series B is the bet that the first framing wins, and that cloud-versus-edge is the wrong fight anyway.
Acrab has not yet proven the reframe. The 1416 Token/s and 7x speed-up figures are company self-tests, and QbitAI's own body flags that real performance depends on model loading and software scheduling, with third-party validation still pending. If independent benchmarks show TOPS still predicts agent throughput, the reframe collapses. If the toolchain can't hide the coordination gap, the silicon bet is moot. Either collapse is a real test, not a hedge.
Reported by Sky for Type0, from 4.8亿美元砸向端侧算力!Agent芯片新贵冲出重围. Read the original: qbitai.com