The 'Ultrafast' tier is powered by Cerebras's single wafer AI processor and is initially limited to a select group of customers.
OpenAI has added a new "Ultrafast" tier to its API, powered by Cerebras, that runs GPT-5.6 Sol at up to 750 output tokens per second. The tier is initially limited to a select group of customers, with broader access promised as capacity scales.
Cerebras builds wafer-scale processors that pack more compute and memory onto a single piece of silicon than a conventional GPU, which is why the company is positioned for high-throughput inference rather than training. OpenAI is previewing the tier on its own index page and markets GPT-5.6 Sol as its best model for legal briefs, financial models, and engineering reports.
The speed numbers come from Cerebras and from Artificial Analysis, a third-party benchmark outfit. Cerebras says the AA output-speed comparison shows Ultrafast running roughly 11× faster than a leading rival frontier model and 5× faster than the same rival on its Fast mode. On Humanity's Last Exam, a 2,500-question benchmark of PhD-level problems, Cerebras reports GPT-5.6 Sol on Ultrafast answered all 2,500 in 11 hours 11 minutes, versus 78 hours 27 minutes for the rival, a 7× gap it calls "comparable accuracy."
It is still a vendor benchmark on a closed tier. A Hacker News thread flags that the AA suite may not have been re-run and the comparison windows do not overlap, so the 11×, 5×, and 7× figures should be read as Cerebras's claim until the methodology is confirmed.