OpenAI on Tuesday unveiled the first public benchmark results for its in-house Jalapeño custom inference chip at the Hot Chips conference, marking the most concrete disclosure yet of how the company’s silicon program compares against the Nvidia accelerators that dominate AI infrastructure today — The release marks the first public benchmarks for the OpenAI Jalapeño custom inference chip, a chip OpenAI has been developing with Broadcom since 2024. The presentation, delivered by members of OpenAI’s hardware team and Broadcom co-engineers, put hard numbers behind a project that has been rumored since 2023 and formally acknowledged last year, and set the stage for a measured ramp into production before the end of 2026.
The headline figures concern efficiency rather than peak throughput. Across three large open-weights models tested at full precision, GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T, Jalapeño delivered between 1.5x and 1.9x more AI work per watt than Nvidia’s Blackwell generation at peak throughput. End-to-end latency, measured from prompt submission to final generated token, came in 1.7x to 3.6x lower than the best commercially available systems OpenAI used as a reference. For interactive workloads, where time-to-first-token and sustained streaming speed dominate the user experience, the chip ran between 2.1x and 4.1x faster across the same model suite.
OpenAI also released per-user throughput numbers that translate the efficiency gains into something engineers can reason about. On GPT-OSS, Jalapeño sustained roughly 1,400 tokens per second per user, a figure that puts very long context or tool-using sessions in reach on a single accelerator. On Deepseek R1, the chip held above 700 tokens per second on a single concurrent request, meaning that one Jalapeño device can serve one heavyweight reasoning job at speeds that previously required multi-GPU sharding. These were not cherry-picked marketing slides, OpenAI engineers argued during the Q&A, but configurations taken from internal production traces.
What makes the numbers more interesting is who helped design the silicon. Jalapeño was co-developed with Broadcom, which contributed its SerDes IP, packaging expertise, and a portion of the physical design flow. The project started in mid-2024 and the final design went to fabrication in November 2025, a 16-month program overall and a roughly nine-month chip-design-to-tape-out sprint once the architectural specification was frozen. OpenAI said it used its own frontier models to assist with parts of the place-and-route, verification, and power-grid optimization work, a notable admission for a company that has otherwise been a consumer of chips rather than a producer of them.
The industry read of those disclosures arrived quickly. Independent semiconductor analyst firm SemiAnalysis published a note on Tuesday evening declaring that Jalapeño “smokes every other chip” on the efficiency metric that matters most for inference, and arguing more provocatively that Nvidia’s CUDA software moat is “potentially dead” now that a frontier model lab has shipped a competitive accelerator and is willing to port its own workloads onto it. The firm cautioned that the comparison is not symmetrical: Jalapeño was tuned for OpenAI’s specific serving stack, while Nvidia’s stack supports a far broader universe of models and frameworks. Still, the framing matters because SemiAnalysis has been a leading voice on AI infrastructure economics, and its assessment will move how institutional investors think about Nvidia’s pricing power heading into 2027.
One point of friction in the benchmark conversation is memory bandwidth. OpenAI compared Jalapeño favorably against Nvidia’s announced Vera Rubin platform, which uses HBM4 and is expected to ship commercially next year. Jalapeño does not match Vera Rubin on raw HBM bandwidth per accelerator, and OpenAI was candid about that during the session. Where Jalapeño still wins, the company said, is on output tokens per megawatt, the metric that ultimately governs data-center operating cost. On total cost of ownership per token, OpenAI’s own internal modeling puts the two platforms roughly even, with Jalapeño’s efficiency gains offsetting the smaller memory envelope when serving the model’s own workloads.
What the company is not claiming is that Jalapeño will replace Nvidia inside its fleet. CFO Sarah Friar, speaking to investors the same day, framed the chip as one input into a diversified compute strategy that still includes Nvidia, AMD, AWS Trainium, Cerebras, and CoreWeave. The plan is for small Jalapeño volumes to land in OpenAI data centers by the end of 2026, with a broader ramp through 2027 as yields stabilize and the software stack matures. Friar described the rollout as additive rather than substitutionary, a hedge that reflects how much revenue OpenAI still expects to send Nvidia’s way for training workloads and for inference on models outside the company’s first-party catalog.
The competitive backdrop is harder to ignore. AMD’s MI455X is shipping in limited quantities, AWS has refreshed Trainium 3 with claims of 2x generational gains, and a handful of startups including Groq, Tenstorrent, and Etched are pursuing specialized inference silicon of their own. What distinguishes Jalapeño is that it comes from a frontier model operator with first-party workloads large enough to anchor a custom chip, and with the engineering depth to optimize both silicon and serving stack together. If the Hot Chips numbers hold up in third-party hands later this year, the more interesting question for the rest of the industry will be how quickly other model labs follow the same playbook, and how Nvidia adjusts its pricing and product cadence in response. For now, the OpenAI Jalapeño custom inference chip has moved from rumor to measurable silicon, and the rest of 2026 will determine whether that measurement translates into shipped volume.
Source: the-decoder.com

