Quick Answer: OpenAI published its first independent benchmark results for Jalapeño, its custom AI inference chip built with Broadcom, on August 25, 2026, at the Hot Chips conference. In tests run on SemiAnalysis’s public InferenceX benchmark across three large open-weight models, the 700-watt chip delivered 1.5 to 1.9 times more throughput per kilowatt and up to 3.6 times lower latency than Nvidia’s current-generation Blackwell (GB200 and GB300) systems, with even bigger gains on highly interactive workloads. The results landed one day before Nvidia’s quarterly earnings and mark the clearest sign yet that OpenAI is serious about reducing its reliance on Nvidia hardware, though the chip won’t ship in meaningful volume until 2027 and hasn’t yet been tested against Nvidia’s upcoming Vera Rubin generation.

What’s Actually Happening
On Tuesday, August 25, 2026, at the Hot Chips conference in Silicon Valley, OpenAI presented the first measured performance numbers for Jalapeño, the custom inference chip it has been developing with Broadcom. The company’s head of hardware, Richard Ho, described the results as “a very, very significant performance advance over state of the art” on a press call. The timing wasn’t subtle: the announcement landed roughly 24 hours before Nvidia reported its fiscal second-quarter earnings, a scheduling choice that several outlets noted without OpenAI directly confirming any intent behind it.
The tests were run on InferenceX, a public benchmarking framework built by SemiAnalysis that measures the full process of serving an AI request, not just raw chip throughput. OpenAI compared Jalapeño against Nvidia’s current Blackwell-generation GB200 and GB300 systems, since those are what’s actually shipping today. Across three open-weight models, GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 at one trillion parameters, Jalapeño posted consistent efficiency and latency advantages.
What Is Jalapeño?
Jalapeño is OpenAI’s first custom AI accelerator, an inference-only chip designed to serve already-trained models rather than train new ones. It was first announced in October 2025 as part of a broader partnership with Broadcom, formally unveiled in June 2026, and is described by both companies as an “Intelligence Processor” rather than a traditional GPU. The chip was designed from scratch by OpenAI’s hardware team, now roughly 40 engineers under Richard Ho, drawing on the company’s own understanding of how large language models actually get served in production. Broadcom handled the silicon implementation and networking, manufactured on TSMC’s 3nm (N3) process, while systems integrator Celestica built the board and rack-level hardware around it.
Each Jalapeño package uses 216 GiB of HBM4 memory across six stacks, compared to 288 GB on Nvidia’s GB300, a deliberately smaller memory footprint that matters a great deal given how constrained global HBM4 supply currently is. The chip runs at 700 watts, roughly half the 1,400-watt power draw of the GB300 system it was benchmarked against.
A Quick Timeline: How We Got Here
- October 2025: OpenAI and Broadcom announce a multi-year deal to co-develop and deploy 10 gigawatts of custom AI accelerators through 2029.
- June 24, 2026: OpenAI and Broadcom formally unveil the Jalapeño chip, describing it as an “Intelligence Processor” and the first step in a multi-generation compute platform.
- August 25, 2026: OpenAI presents the first independently measured Jalapeño benchmark results at the Hot Chips conference, run on SemiAnalysis’s public InferenceX suite.
- August 26, 2026: Nvidia reports fiscal second-quarter earnings of $96.2 billion, roughly 24 hours after the Jalapeño results were published.
The Benchmark Numbers, Explained
OpenAI normalized all results by each chip’s published power rating, arguing that performance per watt is a more meaningful comparison than raw throughput, especially at a scale where data-center power availability, not chip supply, is often the binding constraint. Here’s how the headline numbers broke down model by model:
| Model Tested | Peak Throughput/kW Advantage | Latency Advantage |
|---|---|---|
| GPT-OSS 120B (vs GB200) | ~1.9x (85,448 vs 44,960) | 1.03s vs 1.80s |
| DeepSeek R1 670B (vs GB300) | ~1.7x | 1.65s vs 5.99s |
| Kimi K2.5 1T (vs GB300) | ~1.5x | 3.4x lower |
On the most interactive, low-latency workloads, the kind of chatty, back-and-forth traffic ChatGPT actually generates, OpenAI reported gains climbing as high as 2.1 to 4.1 times faster than the Nvidia systems tested. On DeepSeek R1 with a single user connected, Jalapeño reportedly cleared 700 tokens per second. Notably, these figures came from an untuned, single-token configuration, without the Multi Token Prediction (MTP) optimization that the Nvidia comparison chips were using in their best-performing configs, according to SemiAnalysis’s independent write-up of the results.

How Jalapeño Compares to Nvidia’s Blackwell
It’s worth being precise about what was actually compared here. OpenAI benchmarked Jalapeño against Nvidia’s current-generation Blackwell systems, the GB200 and GB300, because those are the chips actually shipping and available today. Critics have pointed out that this makes the comparison somewhat uneven: Jalapeño uses newer HBM4 memory, while the specific Blackwell configurations tested were running older-generation memory. A more apples-to-apples test would pit Jalapeño against Nvidia’s upcoming Vera Rubin platform, which hasn’t shipped yet and against which Jalapeño has not been benchmarked.
Yole Group technology analyst Adrien Sanchez told CNBC that a hyperscaler-designed chip being able to “match or beat Nvidia’s Blackwell-class GPUs on inference efficiency” represents a real shift, even while noting that Nvidia still controls the “vast majority” of AI compute and benefits from deep ecosystem lock-in through its CUDA software platform. Sanchez characterized Jalapeño specifically as a “threat to Nvidia’s inference margins,” which he called the fastest-growing segment of the AI compute market right now, rather than a threat to Nvidia’s overall dominance.
The Broader Broadcom Deal: 10 Gigawatts Through 2029
Jalapeño is just the first visible product of a much larger agreement. OpenAI and Broadcom’s October 2025 partnership committed the two companies to co-developing and deploying 10 gigawatts of custom AI accelerators across OpenAI’s data centers and partner facilities through 2029, spanning both current 3nm chips and a future 2nm generation. For context, 10 gigawatts is roughly comparable to a meaningful slice of entire national power grids, underscoring just how much electricity modern AI inference is expected to consume at scale.
Some reporting has suggested Microsoft could take roughly 40% of the first-phase Jalapeño chips, which would mean the accelerator isn’t just meant to power OpenAI’s own services but to ship in volume to a major cloud provider as well. Broadcom, meanwhile, has projected its annual AI semiconductor revenue will reach $56 billion in fiscal 2026, with some Wall Street analysts forecasting that figure could climb as high as $116 billion in fiscal 2027 as deals like this one scale up.
Why the Timing Matters
OpenAI released the Jalapeño benchmark results roughly 24 hours before Nvidia reported fiscal second-quarter earnings after market close. Several outlets flagged the proximity without concluding it was deliberate, but the effect was clear either way: for a brief window, the AI hardware conversation shifted from Nvidia’s earnings to a rival’s efficiency claims. As it turned out, Nvidia’s results were strong regardless, with the company reporting quarterly revenue of $96.2 billion, more than double the prior year, which somewhat blunted whatever narrative pressure the Jalapeño announcement was intended to create.
There’s also a financing angle. OpenAI has roughly a decade of expensive data-center construction ahead of it, and unlike Nvidia or Broadcom, it has no chips to actually sell. A strong, independently measured efficiency number speaks directly to the investors and partners underwriting that infrastructure buildout, which is part of why some analysts read the benchmark release as much as a financing pitch as a technical announcement.
The Catches: What Jalapeño Doesn’t Solve Yet
For all the strong headline numbers, there are real limits worth keeping in view. First, volume: Jalapeño is targeting only low-volume production by the end of 2026, with meaningful deployment at scale not expected until 2027. Second, the comparison gap: these results measure Jalapeño against Nvidia’s current Blackwell hardware, not the newer Vera Rubin platform Nvidia has coming, so the efficiency lead could narrow or disappear once that comparison becomes possible. Third, OpenAI itself has been clear that Jalapeño isn’t a Nvidia replacement. Richard Ho described it as one part of a broader compute strategy that still relies on what he called “very, very good partners” at both Nvidia and Cerebras, and the company expects to keep buying large volumes of Nvidia accelerators alongside its own chips.
There’s also a supply-chain wrinkle that cuts both ways. All three major HBM4 suppliers, Samsung, SK Hynix, and Micron, are currently fully allocated, with SK Hynix’s CEO publicly warning that 2027 could be the worst year yet for the HBM shortage. Jalapeño’s smaller 216 GiB memory footprint per package could make it somewhat easier to produce in a supply-constrained environment than memory-hungrier alternatives, but OpenAI’s 10-gigawatt deployment commitment also makes it a major new competitor for that same scarce HBM4 allocation, the identical bottleneck that’s been driving up prices across the broader GPU market this year.
What This Means for AI Costs
Inference, the ongoing cost of actually running trained models in production rather than training them in the first place, is where AI bills accumulate over time, and it’s the specific area Jalapeño targets. If OpenAI’s efficiency advantage holds up once the chip moves from lab benchmarks into real production traffic (a transition that has historically eroded the advantages of other custom hyperscaler chips), the company could serve meaningfully more tokens from the same power budget. That doesn’t automatically mean cheaper prices for developers or ChatGPT subscribers, though. OpenAI’s own pricing history is mixed: per-token costs for older model tiers have fallen substantially since 2023, but pricing on its newest, most capable models has moved down more slowly. The efficiency gain gives OpenAI the option to lower prices; it doesn’t guarantee that outcome.
How This Fits Into the Broader AI Chip Race
Jalapeño arrives amid a broader scramble among AI labs and hyperscalers to reduce dependence on any single chip supplier, a trend that’s been reshaping the wider hardware market throughout 2026. It follows a similar playbook to Google’s TPUs and Amazon’s Trainium chips, both of which exist to give their respective companies leverage against Nvidia’s pricing and supply terms. The stakes are only getting higher given how central Nvidia has become to essentially every corner of the AI stack: our coverage of Nvidia’s reported $12.9 billion Hugging Face acquisition covers a separate but related move by Nvidia to extend its reach beyond chips into the software layer where AI models actually get shared and distributed.
The memory shortage underpinning both stories is also the same one driving consumer hardware prices higher this year. Our breakdown of the 2026 GPU price hike traces how the same HBM and DRAM shortage squeezing Jalapeño’s supply chain has pushed RTX 50-series graphics card prices well above their launch MSRPs, a reminder that the AI chip race and the consumer electronics market are pulling on the same limited pool of memory manufacturing capacity.
What Analysts Are Saying
Broadcom CEO Hock Tan has described Jalapeño as comparable in capability to both Nvidia’s Blackwell chips and Alphabet’s tensor processing units, positioning it as a legitimate third option in a market that has been dominated by essentially one supplier. Independent analysts have been more measured. SemiAnalysis, which built and ran the InferenceX benchmark OpenAI used, called the results a credible first step but cautioned that “the next hurdles, AgentX results, independent third-party benchmarking, and deployment at scale, will determine whether the lab lead becomes a fleet advantage.” AgentX is SemiAnalysis’s newer benchmark suite specifically designed to measure long-context, multi-turn agentic workloads, the kind of sustained back-and-forth reasoning tasks that increasingly define how AI gets used in real products, as opposed to the simpler prompt-response pattern InferenceX measures.
Where the Name Comes From
OpenAI hasn’t published an official explanation for the “Jalapeño” codename, and the company has generally used food-themed internal project names for its hardware efforts, a pattern that predates this specific chip. What is confirmed is that Jalapeño represents the first shipped result of a project that was reportedly known internally as “Titan” during its earlier development phases, before the Jalapeño branding was applied at the June 2026 public unveiling. Engineering samples of the chip have already been shown running real machine learning workloads in OpenAI’s labs at production target frequency and power, including one of OpenAI’s own models, GPT-5.3-Codex-Spark, according to the companies’ joint announcement.
The Katsu Rack: What Surrounds the Chip
Jalapeño doesn’t operate alone. The systems built around it include a CPU host rack that SemiAnalysis’s reporting identifies by the internal name “Katsu,” which houses 16 CPU trays, each equipped with two Turin-class AMD EPYC processors, 1.5 terabytes of standard DRAM, local NVMe and M.2 solid-state storage, and 400-gigabit frontend networking. The rack-level design, along with the Ethernet-based networking OpenAI chose for the broader system, reflects a deliberate effort to build an inference platform without paying Nvidia’s margins on either the core accelerator chip or the networking gear that connects racks of chips together, both of which represent significant costs in a large-scale AI deployment.
What Happens Next
The immediate next milestones to watch are straightforward: whether OpenAI publishes AgentX results on more complex, agentic workloads; whether independent third parties are able to verify these numbers outside OpenAI’s own testing environment; and how Jalapeño’s efficiency claims hold up once the chip moves from engineering samples running in a lab to actual production deployment at scale, expected to begin in small volumes by the end of 2026 and ramp meaningfully through 2027. Nvidia’s response also bears watching, particularly how its upcoming Vera Rubin platform performs once it ships and becomes available for the more direct comparison that today’s benchmarks couldn’t yet provide.
Nvidia’s Response: A Record Quarter Anyway
If the Jalapeño announcement was meant to overshadow Nvidia’s earnings, the numbers Nvidia posted the following day made that difficult. The company reported quarterly revenue of $96.2 billion, more than double what it posted a year earlier, and its shares climbed following the results despite the fresh competitive pressure from OpenAI’s benchmarks. That reaction underscores a point several analysts have made about Jalapeño’s actual competitive threat: even a genuinely more efficient inference chip from a single major customer doesn’t meaningfully dent demand for Nvidia hardware in the near term, given how much of the broader AI infrastructure buildout still runs on Nvidia’s ecosystem and how deeply CUDA is embedded across the industry’s software stack. Nvidia has separately been approached for comment specifically on the Jalapeño benchmarks and, as of this writing, has not issued a detailed public response to the specific performance claims.
Why Custom Silicon Keeps Attracting Hyperscalers
OpenAI is far from the first AI company to build its own chips. Google has developed custom Tensor Processing Units (TPUs) for years specifically to reduce its dependence on Nvidia GPUs for both training and inference, and Amazon has pursued a similar strategy with its Trainium and Inferentia chip lines for AWS customers. What makes OpenAI’s move notable is less the strategy itself, which is well established among hyperscalers with the balance sheets to fund custom silicon development, and more the fact that a company whose entire business model until now has centered on renting compute rather than manufacturing it is now positioning itself as a chip designer in its own right. Whether that shift proves as durable and cost-effective as OpenAI hopes will likely take another full product generation, and a genuine head-to-head test against Nvidia’s next-generation hardware, to determine with any confidence.
Frequently Asked Questions
What is OpenAI’s Jalapeño chip?
Jalapeño is OpenAI’s first custom AI accelerator, an inference-only chip built with Broadcom and Celestica, designed to serve trained AI models more efficiently than general-purpose GPUs.
Does Jalapeño really beat Nvidia’s Blackwell chips?
In OpenAI’s published benchmarks on the InferenceX test suite, Jalapeño delivered 1.5 to 1.9 times more throughput per watt and up to 3.6 times lower latency than Nvidia’s current GB200 and GB300 Blackwell systems. It has not yet been tested against Nvidia’s upcoming Vera Rubin platform.
When will Jalapeño actually ship?
OpenAI plans low-volume production by the end of 2026, with meaningful deployment at scale expected during 2027.
Is OpenAI dropping Nvidia because of Jalapeño?
No. OpenAI’s head of hardware said Jalapeño is one part of a broader compute strategy that still relies heavily on Nvidia and Cerebras, and the company expects to keep purchasing large volumes of Nvidia accelerators alongside its own chips.
Who makes the Jalapeño chip?
OpenAI designed the chip itself, with Broadcom handling silicon implementation, networking, and manufacturing (via TSMC’s 3nm process), and Celestica building the board and rack-level systems.
How much memory does Jalapeño have?
Each Jalapeño package uses 216 GiB of HBM4 memory across six stacks, compared to 288 GB on Nvidia’s GB300, a smaller footprint that may help with production given the ongoing global HBM4 shortage.
Will Jalapeño make ChatGPT cheaper?
Not necessarily. The efficiency gains give OpenAI the option to lower prices, but the company’s pricing history is mixed, with older model tiers getting cheaper over time while frontier model pricing has moved down more slowly.
How big is OpenAI’s chip deal with Broadcom?
OpenAI and Broadcom’s October 2025 agreement covers co-developing and deploying 10 gigawatts of custom AI accelerators across OpenAI’s infrastructure through 2029.



