OpenAI disclosed the first benchmark results for Jalapeño, its custom inference chip, at the Hot Chips conference on Tuesday, and the numbers put the part ahead of Nvidia's Blackwell on SemiAnalysis' InferenceX benchmark. Jalapeño registered more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors, according to results OpenAI presented alongside its hardware head Richard Ho. The chip was first announced in October of last year and is being co-developed with Broadcom.
Deployment is still more than a year out. Ho said Jalapeño will reach production at the end of 2026 in very small volumes, with more significant deployment coming in 2027. That timeline matters: the comparison point today is Blackwell, but by the time Jalapeño ships in scale, Nvidia will have moved on to newer silicon.
“The bottom line is that the results show a very, very significant performance advance over state of the art.”— Richard Ho, OpenAI head of hardware
The pitch is efficiency at scale. OpenAI's argument is that Jalapeño delivers both higher tokens-per-user and higher throughput-per-kilowatt than the Blackwell systems currently serving frontier workloads. Those two metrics tend to trade off against each other on general-purpose accelerators, and OpenAI is claiming gains on both simultaneously — the practical translation being lower latency for end users and lower cost per served query for the operator.
Key facts
- 01OpenAI's Jalapeño chip beat Nvidia Blackwell on SemiAnalysis' InferenceX benchmark in both tokens-per-user and throughput-per-kilowatt.
- 02Jalapeño will ship in very small volumes at the end of 2026, with more significant deployment in 2027.
- 03The chip was co-developed with Broadcom and first announced in October 2025, with OpenAI's own models assisting in the design process.
- 04Jalapeño is architected to minimize delays in the prefill and communication phases of inference, which OpenAI identifies as bottlenecks.
- 05OpenAI plans Jalapeño as a multigenerational platform, co-developing models, chips, and memory together.
Jalapeño was designed alongside the models it will run, and OpenAI used its own models in the development process. Ho described the platform as multigenerational, meaning products, models, chips, and memory will be planned together rather than the model team adapting to whatever silicon Nvidia ships next. That vertical alignment is the core rationale for building custom silicon at all, and it echoes the approach Google has taken with TPUs and Amazon with Trainium.
The specific engineering choices target inference bottlenecks. OpenAI says Jalapeño is architected to minimize delays in the prefill phase and the communication phase, two stages that frequently cap throughput on general-purpose GPUs. The KV cache — the running state a model holds while generating a response — can be explicitly placed and kept local, avoiding round-trips that eat latency budget.
Broadcom is the manufacturing and design partner, continuing a pattern of hyperscalers pairing with Broadcom for custom accelerators. The arrangement gives OpenAI a co-designed silicon path without needing to build a full chip-design organization from scratch, and it gives Broadcom another anchor customer for its AI ASIC business, which has become one of the fastest-growing segments in the semiconductor industry.
The Nvidia comparison is the most-scrutinized part of any custom-silicon announcement, and OpenAI is being careful. Ho acknowledged that Blackwell is today's reference point but that Nvidia's roadmap will have advanced by the time Jalapeño ramps. Nvidia has already begun shipping Vera Rubin-generation parts to hyperscalers, and the pricing on those systems has climbed more than 15% amid a memory-cost surge — a dynamic that gives OpenAI a stronger economic case for owning its inference stack even if the peak per-chip performance gap narrows.
OpenAI's motivations here are as much about supply as speed. Model deployment is bottlenecked by chip availability, and every hyperscaler that has committed to Nvidia at scale has run into allocation limits. A first-party inference chip gives OpenAI a hedge — a second source for the workload that most directly affects ChatGPT's response times and unit economics. It also lets OpenAI negotiate harder on the Nvidia parts it will still buy in volume through the transition.
The counterweight is execution risk. Getting a new accelerator into production is difficult, and the gap between benchmark numbers presented at Hot Chips and stable, high-utilization deployment at datacenter scale has historically been where custom silicon projects stumble. Ho's phrasing — very small volumes at the end of 2026 — suggests OpenAI itself is treating 2027 as the year that matters. The InferenceX result is a proof point, not a shipping product.
For the AI compute market, Jalapeño is another signal that the largest model operators intend to control their own silicon destiny. Google runs TPUs, Amazon runs Trainium and Inferentia, Meta is deploying MTIA, and now OpenAI has a benchmarked inference part with a firm deployment window. Nvidia will keep the training market and much of the inference market for years, but the share of frontier inference cycles running on non-Nvidia hardware is set to climb meaningfully starting in 2027 — and every point of share shift there reprices the AI infrastructure stack.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




