Jalapeño, OpenAI's first in house chip co designed with Broadcom, only handles the half of AI work that runs trained models, not the half that trains them.
OpenAI's first in-house AI chip, codenamed Jalapeño, doesn't train models. It runs them.
Inference, the step that turns a trained model into a response, is the workload Jalapeño is built for. Training new models still happens on Nvidia hardware.
According to OpenAI's first-results post, Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency than rival chips across three models tested on the SemiAnalysis InferenceX benchmark. For highly interactive AI agent workloads, the company said the advantage climbed to as much as 4.1 times. The chip was co-developed with Broadcom.
Those are vendor-reported numbers on a benchmark curated by SemiAnalysis, with no independent third-party replication yet visible, per SemiAnalysis's own writeup of the announcement. OpenAI plans small-volume deployment by the end of 2026 and meaningful scale in 2027.
The move puts OpenAI alongside Google's TPU line and Amazon's custom silicon, the other two major cloud players that already run their own accelerators. OpenAI has used Nvidia GPUs since its founding; Jalapeño is a first step into in-house silicon, and a narrow one. Training, the more compute-hungry half of the AI workload, stays on Nvidia.