At Hot Chips, the chip industry's annual conference, the 88 core Vera — Nvidia's first in house CPU — skips the modular playbook AMD and Intel use and bets that AI agents that browse and code reward memory and threading, not more cores.
Nvidia's first CPU designed in-house, codenamed Vera with 88 custom "Olympus" cores on a single monolithic die, made its architectural public debut at Hot Chips 2026, and the company chose the venue to argue a thesis: the next server workload will reward concurrency, memory bandwidth, and irregular work, not more cores on a chiplet or faster single-thread.
The chip itself is the physical form of that bet. Vera succeeds Grace, which used a stock Arm design; Olympus is Nvidia's first custom core and ships in a single 88-core monolithic configuration, with no compute chiplets and no modular tiles. The choice deliberately trades AMD's EPYC-style chiplet scaling for predictable latency on long agentic chains, according to Nvidia's Hot Chips slides. Memory is the second leg of the bet. Vera pairs 1.2 TB/s of LPDDR5X delivered over SOCAMM2 modules, the small-outline compression-attached memory form factor Nvidia has been pushing for low-power, high-bandwidth server workloads. The third leg is threading. Nvidia calls its approach "spatial multithreading," distinct from the simultaneous-multithreading x86 uses; the company argues irregular agentic work exposes more parallelism per core than AMD's or Intel's thread model.
Nvidia's developer blog frames agentic AI, meaning multi-step tasks that browse, code, and call external tools, as the workload the design targets, calling it "the most complex computing workload in history." That phrasing is marketing, but the underlying theory is testable: headless-browser workflows and code-compile jobs behave differently from the SPEC-style server benchmarks AMD and Intel optimize for, and they punish memory starvation harder than they punish missing single-thread clockspeed.
The vendor benchmark Nvidia showed at Hot Chips makes the same point. On a headless-browser workload scaling instance count, Vera runs 24% faster than a 96-core AMD EPYC 9655P. A second slide shows agents completing a browsing workflow 4.5x faster than humans, which reflects what agents skip (no GUI rendering, fonts, or media decode), not raw CPU speed. ServeTheHome's onsite reporting confirms the figures come from Nvidia's deck, not independent testing. The 24% number is therefore a vendor head-to-head against a single AMD SKU on a workload Nvidia chose; the agent-versus-human number is an apples-to-oranges framing about pipeline efficiency. Both are legible, but they do not, on their own, settle the workload contest.
What Vera actually buys Nvidia is a different bet. By staying monolithic, the company gives up the path AMD took with 9006-series EPYC, stacking chiplets to reach higher core counts at yield, and accepts that a die defect costs a whole CPU. In exchange, Vera gets uniform memory latency across all 88 cores and a memory subsystem that pushes 1.2 TB/s, well above the per-core bandwidth most server parts deliver. The spatial-multithreading approach is the third tradeoff: a custom core and a custom thread model, both of which Nvidia has to convince compiler and OS vendors to target.
SpaceXAI is cited in the Tom's Hardware recap as a Vera customer. One customer is not a market signal, but it does confirm the workload thesis is being tested against real agentic traffic, not just synthetic benchmarks. The next data points worth watching are independent Vera benchmarks, AMD's response in its next EPYC generation, and whether Intel's Xeon line treats agentic workflows as a defined category or absorbs them back into general-purpose server tuning. Until then, Vera is a coherent, opinionated bet that the wire has not yet named.