OpenAI has published detailed benchmark results for Jalapeño, its first custom-built inference chip, claiming it can serve more AI work per unit of power while also responding faster, a combination the company says existing hardware typically cannot achieve without sacrificing one for the other.
Testing on InferenceX, a public benchmark run by SemiAnalysis, showed Jalapeño delivering between 1.5 and 1.9 times more AI work per watt at peak throughput, and between 1.7 and 3.6 times lower end-to-end latency, than comparison systems across three models: GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. For highly interactive workloads, the advantage rose to between 2.1 and 4.1 times. OpenAI said the chip is rated at 700 watts but ran at or below 550 watts on the workloads tested.
The company said Jalapeño's gains stem from "designing the chip, memory, network, software, and rack-scale system together around real language-model workloads," minimising the data movement between components that typically slows other systems down. AI tools were also used in the chip's own design, helping the team move from initial design to tapeout in nine months, and were later used to help optimise chip performance for models not part of the chip's original production plan.
OpenAI plans to begin deploying Jalapeño within its own compute infrastructure by the end of the year, describing it as the first of a multigenerational roadmap, with a second generation already in development and a third taking shape. The company said it will continue to rely on Nvidia and other external accelerator partners alongside its own chip.
