OpenAI has published detailed benchmark results for Jalapeño, its first custom-built inference chip, claiming it can serve more AI work per unit of power while also responding faster, a combination the company says existing hardware typically cannot achieve without sacrificing one for the other.

Testing on InferenceX, a public benchmark run by SemiAnalysis, showed Jalapeño delivering between 1.5 and 1.9 times more AI work per watt at peak throughput, and between 1.7 and 3.6 times lower end-to-end latency, than comparison systems across three models: GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. For highly interactive workloads, the advantage rose to between 2.1 and 4.1 times. OpenAI said the chip is rated at 700 watts but ran at or below 550 watts on the workloads tested.

The company said Jalapeño's gains stem from "designing the chip, memory, network, software, and rack-scale system together around real language-model workloads," minimising the data movement between components that typically slows other systems down. AI tools were also used in the chip's own design, helping the team move from initial design to tapeout in nine months, and were later used to help optimise chip performance for models not part of the chip's original production plan.

OpenAI plans to begin deploying Jalapeño within its own compute infrastructure by the end of the year, describing it as the first of a multigenerational roadmap, with a second generation already in development and a third taking shape. The company said it will continue to rely on Nvidia and other external accelerator partners alongside its own chip.


Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.
Share this post
The link has been copied!