Built from the IP Nvidia licensed in its largest ever chip deal, the Groq 3 LPX, Nvidia's new long context inference chip, posted its first published benchmark running one model at one input length — a best case ceiling, not a verdict on the bet.
Nvidia's first third-party benchmark on Groq 3 LPX landed this week on the company's developer blog, republishing Artificial Analysis's run of Google's Gemma 4 31B at a 100,000-token context (Nvidia developer blog). The LPX is the inference accelerator Nvidia built from the IP it paid roughly $20B in cash to license from Groq in December 2025 (CNBC). The deal brought Groq founder Jonathan Ross and other senior leaders into Nvidia. The company itself, and GroqCloud, stay independent.
The 3,431 number is a best-case ceiling. The LPX is built for small-batch, long-context interactive inference, the workload where a single user's prompt and a single model's reply need to feel instant, and the Nvidia blog names a target use case, multiagent systems running 2T+ parameter models, that is not yet standard fleet practice. On the same Gemma 4 31B, an 8x H100 NVLink-4 rack peaks at 3,050 tokens per second at 8K context, with 85 ms time to first token. That is a different test, but close enough that "LPU beats GPU" is not yet a defensible read from one benchmark.
The next round, more models, more context lengths, and a deployment read from inside an NVL72 rack, is the one that moves 3,431 from ceiling to verdict.