OpenAI unveiled Jalapeño, its first custom silicon, on June 24, 2026—and the first-generation chip is not a learning exercise. Built with Broadcom and manufactured on TSMC's 3nm process, the custom inference accelerator matches the performance of Nvidia's Blackwell chips and Google's Tensor Processing Units, according to Broadcom CEO Hock Tan.
The economics behind the decision are direct. Early testing shows Jalapeño delivers roughly 50 percent lower inference cost compared to GPU-based alternatives. For a company running ChatGPT at the scale OpenAI does, that margin difference compounds fast across millions of daily queries.
The tape-out—the final step in chip design before mass production—took nine months. That timeline is aggressive for any chip program and unusually fast for a company building its first silicon from scratch. OpenAI plans to deploy Jalapeño through Azure before the end of 2026.
The chip is purpose-built for inference, not training. Training large language models requires the brute-force parallelism that Nvidia's GPUs do exceptionally well, and OpenAI is not walking away from that dependency. Jalapeño targets the inference workload—running the model to generate responses—which is where ongoing operating costs live at production scale.
Broadcom's role as the manufacturing and design partner matters for context. Broadcom has become the go-to custom chip partner for hyperscalers: it built Google's TPU line and Apple's transition away from Intel processors demonstrated what custom silicon tuned to specific workloads can do for performance and unit economics. OpenAI is following the same logic.
When Apple moved to its own M-series chips, it did not simply cut its Intel bill—it gained architectural control over how software and silicon interact. OpenAI's stated goal with Jalapeño is the same: hardware tuned to the specific demands of running large language models, not general-purpose GPU architecture adapted to that use case.
OpenAI joins Google, Apple and SpaceX in building custom silicon as a hedge against single-supplier risk. None of these companies describe the move as a clean break from Nvidia. The framing from OpenAI is deliberate diversification—more control over the inference stack, better performance per dollar on a specific workload, and reduced exposure to one vendor's supply chain and pricing.
Nvidia's stock rose 1.5 percent to $211.56 on Aug. 25, which reflects the market's read that Jalapeño does not displace Nvidia's training dominance. The training market—where Nvidia's H100 and Blackwell architecture have no credible short-term competitor—remains intact. What changes is OpenAI's leverage in the inference market and, over time, its cost structure.
The competitive pressure on Nvidia is specific. Custom inference chips from Amazon (Inferentia), Google (TPU), and now OpenAI all attack the same seam: inference at scale is a volume business where cost per query determines margin. Nvidia's GPUs are not optimized for that economics problem the way a purpose-built ASIC is. An ASIC—application-specific integrated circuit—is a chip designed for one task rather than general computation, which allows designers to strip out unused circuits and reduce power and cost.
For Broadcom, Jalapeño is another data point in a growing custom silicon franchise. The company designs and manufactures custom ASICs for several of the largest technology companies in the world. Each new hyperscaler or AI lab that decides to build its own chip rather than buy off the shelf is a potential Broadcom engagement.
The nine-month tape-out timeline and Blackwell-competitive performance on the first try will draw attention across the industry. First-generation custom chips typically trade performance for lower cost and improved thermal efficiency—matching an established GPU architecture on raw performance while also cutting cost by half is not the expected result from a first attempt. Hock Tan's confirmation of performance parity with both Blackwell and Google's TPU gives the claim a named, primary source rather than OpenAI's own benchmarks alone.
OpenAI has not disclosed production volumes or the number of Jalapeño units running in its labs. The Azure deployment target of end-2026 is the next concrete milestone. Whether the chip scales to carry a meaningful share of ChatGPT's inference load—or remains a smaller complement to GPU capacity—will determine how much the 50 percent cost figure actually moves OpenAI's operating economics.
