metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
OpenAI
Broadcom
Nvidia
Custom AI Chips
Inference
AI Hardware
Blackwell
Efficiency

OpenAI's Jalapeno Chip Posts Its First Benchmarks Against Nvidia. Here Is How to Read Them

OpenAI published the first performance numbers for its Broadcom-built inference chip, claiming higher throughput per watt and lower latency than Nvidia's GB200 and GB300. A neutral analysis of what the benchmarks show, what they leave out, and why efficiency per watt is the metric that matters.

Metir AI TeamAugust 25, 202611 min read
OpenAI's Jalapeno Chip Posts Its First Benchmarks Against Nvidia. Here Is How to Read Them

When OpenAI and Broadcom unveiled a custom inference chip called Jalapeno in June 2026, the announcement was mostly about ambition: a fast design cycle, a first tape-out, early samples running in the lab. What it lacked was numbers. On August 24, 2026, OpenAI filled that gap, publishing the first live performance benchmarks for the chip and claiming clear advantages in energy efficiency and latency over Nvidia's flagship Blackwell-generation systems. The reaction was immediate, because a credible challenge to Nvidia on its own benchmarks is a rare thing. This piece works through what OpenAI actually claimed, the caveats that come attached, and why the specific metric it chose to lead with tells you where the AI hardware fight is really being fought.

OpenAI logoOpenAI
NVIDIA logoNVIDIA
AMD logoAMD
Google logoGoogle
Jalapeno's published benchmarks compare it against Nvidia's GB200 and GB300 systems, with additional comparisons to AMD and Google silicon.

What OpenAI claimed

Jalapeno is an inference chip, meaning it is built to run trained models and answer queries rather than to train models in the first place. That distinction matters, and we will return to it. According to the benchmarks OpenAI released, a 700-watt Jalapeno part delivered between 1.5 and 1.9 times more throughput per kilowatt than Nvidia's GB200 and GB300 rack systems, and between 1.7 and 3.6 times lower end-to-end latency, while drawing far less power: the comparison pits a 700W part against Nvidia accelerators rated at roughly 1,200W and 1,400W.

700WJalapeno power drawvs ~1,200W and ~1,400W Nvidia parts
1.5-1.9xThroughput per kilowattvs Nvidia GB200 and GB300, per OpenAI
1.7-3.6xLower end-to-end latencyOpenAI's claimed range
9 monthsDesign to tape-outan unusually fast ASIC cycle

OpenAI paired the performance numbers with a claim about how the chip was built. It said it used its own generative models to accelerate the hardware engineering itself, compressing the design-to-tape-out timeline to roughly nine months, optimizing arithmetic logic circuits with AI, and generating software kernels where the AI-written implementations reportedly outperformed human-expert kernels by 1.5 to 1.8 times. OpenAI said it plans to begin deploying the chip in its own data centers later this year. The chip was co-developed with Broadcom, which handles the physical design and manufacturing partnership, a division of labor that mirrors how Google built its TPUs.

Broadcom's headquarters in San Jose, California
Broadcom's San Jose headquarters. Broadcom co-developed Jalapeno with OpenAI, providing the custom-silicon and packaging expertise that turns a chip design into a manufacturable part. Photo: Coolcaesar, CC BY-SA 4.0.

The caveats that come attached

None of this makes Jalapeno a Nvidia-killer, and reading the benchmarks well means holding several caveats at once.

The first and most important: these are first-party benchmarks. OpenAI designed the tests, chose the workloads, and selected the comparison points. That is standard for a launch, and it does not make the numbers false, but it does mean they are the vendor's best case rather than a neutral referee's verdict. The figures that will actually move the industry are the independent ones, run by third parties on workloads the chip's designer did not pick. Until those exist, the honest posture is to treat the claims as promising and unverified.

The second: this is an inference chip, not a training chip. The comparison to Nvidia is real but narrow. Training the largest frontier models is a different workload with different demands, and it remains Nvidia's stronghold. Jalapeno is aimed at the part of the compute bill that grows every time more people use a model, which is large and growing, but it is not the whole market.

“

A 700-watt part beating 1,400-watt parts is not really a story about speed. It is a story about power.

On why efficiency per watt is the metric OpenAI led with

The third: efficiency comparisons are sensitive to exactly what is measured. Throughput per kilowatt on a specific model at a specific batch size and sequence length can look very different from throughput on another configuration. A 1.5 to 1.9 times range is wide precisely because the answer depends on the workload. The chip may be excellent at the shapes OpenAI runs most and less differentiated elsewhere, which is fine for a company building a chip for its own traffic but complicates any general claim.

Why efficiency per watt is the real headline

Notice what OpenAI led with. Not raw speed, but throughput per kilowatt. That choice is the analytically interesting part, because it reflects the constraint that increasingly governs AI infrastructure: power.

Power draw: a 700W part against 1,200-1,400W systems

Rated power of the parts OpenAI compared. Jalapeno's efficiency claim is throughput per kilowatt, so a lower power figure at comparable work is the point.

OpenAI Jalapeno700W
Nvidia GB200 (class)1,200W
Nvidia GB300 (class)1,400W

OpenAI reports 1.5 to 1.9 times more throughput per kilowatt than these systems. Figures are OpenAI's own first-party benchmarks.

Data centers are bounded by how many megawatts they can secure and cool. When power is the binding constraint, the question stops being "how fast is one chip" and becomes "how much useful work can I extract from each megawatt I am allowed to draw." A part that does 1.5 to 1.9 times the work per kilowatt lets an operator serve far more traffic inside a fixed power budget, or serve the same traffic while leaving headroom for growth. For a company running inference at OpenAI's scale, that efficiency compounds directly into cost and into how many users it can serve before it runs out of grid.

This also explains why OpenAI would build the chip at all. It is one of Nvidia's largest customers. Designing custom silicon is expensive and slow, and a company only does it when the economics of buying merchant hardware at scale become painful enough to justify the effort. The move is less a bet that OpenAI can out-engineer Nvidia in general and more a bet that, for the specific and enormous workload of running its own models, a purpose-built part pays for itself. Google, Amazon, and Meta reached the same conclusion years ago for their own workloads.

A technician working on a server rack inside a data center
Inference hardware lives in racks like these, where every kilowatt is accounted for. This is a general data-center photo, not Jalapeno hardware, which OpenAI plans to deploy in its own facilities later in 2026. Photo: NERSC, CC0.

What it changes, and what it does not

For Nvidia, one large customer building an in-house inference chip is a known risk rather than a surprise, and Nvidia's position rests on far more than any single product generation: a mature software stack, a broad installed base, and dominance in training. Jalapeno chips its inference business at the margin; it does not dislodge the platform. The more interesting pressure is cumulative. OpenAI, Google, Amazon, Meta, and now reportedly Anthropic are all designing or deploying custom silicon for the workloads they run most. Each individual chip is narrow. Together they signal that the largest buyers intend to own more of their own inference stack over time.

There is a layer above all of this that the silicon race does not settle. Whichever chip runs a given model, the decision of which model answers a given request, and how a workload is routed across a fleet of different accelerators and hosted endpoints, is a software problem. Systems built to stay model-agnostic and hardware-agnostic, Metir among them, are designed so that a shift in the underlying silicon is a configuration change rather than a rebuild. More competitive inference hardware is good for that layer, because it lowers the cost of the endpoints it routes among without changing what it routes.

The measured takeaway is narrow but real. OpenAI has shown, on its own benchmarks, that a purpose-built inference part can beat Nvidia's flagship on efficiency per watt for the workloads OpenAI cares about. Independent numbers will decide how general that result is. What is already clear is that the metric the industry now optimizes for is not tokens per second in isolation, but tokens per second per watt, and Jalapeno is a bet placed squarely on that shift.

Sources:

  • OpenAI says its Jalapeno chip beats Nvidia's GB300 in first published benchmarks (Tom's Hardware, Aug 24, 2026)
  • OpenAI Jalapeno: Better Than Nvidia Blackwell (SemiAnalysis)
  • OpenAI and Broadcom unveil LLM-optimized inference chip (OpenAI)
  • OpenAI Claims an AI Chip Breakthrough. Jalapeno Beats Nvidia's GB300 in Tests. (IBTimes, Aug 2026)
  • OpenAI's Jalapeno Chip Is Outperforming Nvidia, AMD And Google Chips, SemiAnalysis Says (TradingView)

Image credits

  • Hero: "Broadcom Headquarters San Jose" by Coolcaesar (via Wikimedia Commons), licensed CC BY-SA 4.0. Retrieved August 25, 2026.
  • In-body data-center photo: "Technician with laptop working on server rack at NERSC" (via Wikimedia Commons), released under CC0. Retrieved August 25, 2026. Illustrative only; not Jalapeno hardware.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational