OpenAI's head of hardware, Richard Ho, presented the first public benchmark results for Jalapeño, the company's custom inference chip built with Broadcom and manufactured at TSMC, at the Hot Chips conference on Aug. 25, 2026.

Tested on SemiAnalysis' InferenceX benchmark, Jalapeño outperformed Nvidia Blackwell on two core metrics: tokens per user and throughput per kilowatt. "The results show a significant performance advance over state of the art," Ho said. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly."

That dual claim—high throughput and low latency—is the core tension in inference chip design. Most architectures optimize for one at the expense of the other. OpenAI's approach treats memory placement as a first-order design constraint: the KV cache, the memory structure tracking what a model has generated in a response, stays local rather than moving across interconnects. Keeping it on-chip cuts latency at the hardware level.

Jalapeño carries a 700-watt thermal design power rating, though Ho said the chip ran at or below 550 watts on tested workloads. OpenAI plans to deploy the chip in racks of 128 units, with a full pod comprising 2,048 ASICs—a configuration indicating the scale at which the company intends to run its own inference infrastructure.

The chip was first announced in October 2025 and publicly unveiled June 24, 2026, at which point OpenAI claimed roughly 50 percent lower cost per inference token compared to Nvidia GPU clusters. The Hot Chips presentation adds independent benchmark data to that claim.

Deployment at scale is not imminent. Ho said Jalapeño would ship in "very small volumes" by the end of 2026, with meaningful deployment pushed to 2027. That timeline carries material weight: by the time Jalapeño reaches volume production, Nvidia's roadmap will have advanced. Blackwell is the current generation; Nvidia's next-generation architecture is expected before Jalapeño hits full deployment.

OpenAI is treating Jalapeño as the first generation of a multigenerational platform. The company's stated plan is to develop AI products, models, chips, and memory together—a full-stack approach enabling co-optimization across layers in ways customers buying off-the-shelf Nvidia hardware cannot achieve. That vertical integration is the core strategic argument for the capital and engineering investment the chip program requires.

The economics are straightforward. Inference is now the dominant cost driver for large-scale AI deployments. Training a model is a one-time expense; serving it runs continuously. A chip delivering 50 percent lower cost per token at the inference layer directly reduces the variable cost of every API call OpenAI sells. At the volumes OpenAI operates, that arithmetic matters more than benchmark rankings.

Nvidia's stock rose 2.2 percent on Aug. 25 to $213.05, suggesting the market does not read Jalapeño as an existential threat to Nvidia's data center business. OpenAI's chip is internal deployment infrastructure, not merchant silicon sold to third parties, which limits direct competitive impact on Nvidia's revenue. The indirect pressure is real, though: every chip OpenAI runs internally is a rack it does not buy from Nvidia.