OpenAI’s Jalapeño Chip Crushes Throughput & Efficiency Tests at Hot Chips

At the Hot Chips conference today, OpenAI unveiled benchmark data for its in-house “Jalapeño” chip, showcasing performance gains in inference throughput and power efficiency over today’s top AI accelerators. Built with Broadcom and aided by OpenAI’s own models in its design, the chip promises to shift the economics of large-scale model serving. Deployment begins late 2026—with broader rollout through 2027.

What sets Jalapeño apart

In tests run on SemiAnalysis’s InferenceX benchmark, Jalapeño delivered both higher tokens per user and greater throughput per kilowatt compared to leading inference processors. This includes beating an Nvidia Blackwell system on those metrics. OpenAI’s hardware chief, Richard Ho, emphasized that Jalapeño is optimized to deliver AI work with much lower latency and better power efficiency—so it can support large workloads while also serving individual users fast. pitting its metrics against Blackwell sets a high bar implicitly.

The chip’s architecture is tightly integrated into OpenAI’s full stack—chip, memory, networking, models—all co-designed. That lets Jalapeño reduce key inefficiencies that commonly slow inference tasks, especially during the “prefill” and communication phases where model state (including KV caches used for context) is moved across components. By placing the model state locally and matching compute, memory, and networking more precisely to each phase, OpenAI claims Jalapeño cuts latency and saves energy.

Timeline & limitations

OpenAI first revealed Jalapeño in October, with Broadcom contributing to its physical implementation. The chip is intended to be a generational platform, meaning subsequent versions will build on this design philosophy. However, full deployment won’t happen overnight. OpenAI expects “very small volumes” of Jalapeño infrastructure by late 2026, with wider scale coming in 2027. That means competition—especially from Nvidia and other vendors—may catch up or introduce new offerings during that timeframe.

The comparisons to Blackwell are impressive, but OpenAI acknowledges they’re against current Blackwell systems, not whatever comes next. Upcoming improvements in competing inference architectures could narrow the gap. OpenAI also note that workload specifics—model size, context length, concurrency—can affect whether Jalapeño’s advantages fully manifest.

What this means: Jalapeño isn’t just another chip—it’s part of a broader move toward inference-specialized silicon optimized end-to-end. Firms are recognizing that training speed and peak FLOPs are only part of the puzzle; inference latency, interconnect delays, memory movement, and power efficiency matter even more in deployment. Jalapeño shows that co-designing hardware, software, and models can yield meaningful gains.

What to watch: Whether Jalapeño’s performance holds up in real-world settings with diverse models and workloads; how its energy efficiency translates in hyperscale settings; and how competitors respond with alternative architectures or optimizations. Also relevant are costs—not just chip manufacturing but system cooling, power delivery, and data center integration. The inference arms race just got hotter.