Skip to main content

OpenAI's Jalapeno chips: inside the custom accelerator that claims to beat Nvidia

OpenAI's Jalapeno chips: inside the custom accelerator that claims to beat NvidiaPhoto: N43 and Hermes
N43 / TECH
TECHNOLOGY · 7447
ARTIFICIAL INTELLIGENCE / DATA CENTER HARDWARE

A reported custom accelerator, a $10 billion commitment, and a very pointed benchmark claim: OpenAI's silicon project has graduated from rumor to one of the most consequential stories in AI infrastructure. But "outperformed in testing" carries a lot of asterisks.

Video: Bloomberg Tech - "OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing" - approximately 37K views, observed August 2026 (published ~5 days prior).

01Why OpenAI is designing its own chips

Every large AI developer eventually runs into the same wall: the cost of compute. Training and serving large language models at scale consumes staggering numbers of GPUs, and the overwhelming majority of them come from one vendor, Nvidia. When your entire business depends on a single supplier whose products are chronically supply-constrained, designing your own silicon stops being a science project and starts being risk management.

For OpenAI specifically, the pressure is twofold. First, there is the sheer scale of its buildout, which relies on massive data center commitments from partners such as Microsoft, Oracle, and others. Second, there is the margin question: inference, the act of actually answering user queries, is the recurring cost of the business. Every efficiency gained per token flows straight to the bottom line. A chip tuned to your exact workload can squeeze out efficiency that general-purpose hardware cannot.

That is the same logic that pushed Google to build the Tensor Processing Unit, which Wikipedia's summary notes began internal deployment in 2015 and became available to third parties via Google Cloud in 2018. Google is now among the clearest proof points that a hyperscaler can both design its own AI silicon and commercialize it. OpenAI is following a well-trodden path, just later and with more urgency.

02What the Jalapeno accelerator reportedly is

According to Bloomberg's reporting, the chip carries the internal codename "Jalapeno" and is a custom AI accelerator designed with Broadcom, the semiconductor giant whose custom silicon business has become one of the most important quiet engines of the AI boom. Broadcom is a fabless designer and supplier across data center, networking, and storage markets, and its partnership model lets customers like OpenAI define the architecture while Broadcom handles the design and manufacturing logistics with TSMC.

The reported shape of the deal is enormous: a commitment valued around $10 billion, with initial volumes targeted for 2026 deployment. OpenAI is also reportedly working with AMD on separately designed accelerators, and has explored participation in foundry ventures, suggesting a deliberate strategy of hedging across multiple suppliers rather than betting everything on a single internal design.

What Jalapeno is not, importantly, is a general-purpose GPU. It is best understood as an inference-oriented accelerator: hardware built specifically to run already-trained OpenAI models at high volume and low cost, rather than to train new ones from scratch. That distinction shapes everything about the "beats Nvidia" claim, which we will get to shortly.

03How custom ASICs differ from GPUs for inference

A GPU is, by design, a generalist. Nvidia's hardware has to excel at graphics, scientific computing, and every flavor of AI workload that a customer might throw at it, because Nvidia sells to everyone. That flexibility costs efficiency. A custom application-specific integrated circuit, or ASIC, throws away generality in exchange for raw throughput on one narrow workload profile.

For inference, that trade is often worth it. Serving a large language model is dominated by matrix multiplication and memory bandwidth, and a chip that hardwires exactly the data types, tensor shapes, and attention patterns your models use can achieve far higher utilization than a GPU running the same math through a more generic pipeline. If you know your model's architecture, you can bake assumptions into the hardware.

The catch is rigidity. If OpenAI's model architecture shifts in a way the silicon was not designed for, the chip cannot be reprogrammed around it the way GPU software can. Custom silicon is a bet that your workloads will stay predictable for the years-long life of a data center installation. Google has navigated multiple TPU generations by co-designing models and chips together; OpenAI now has to demonstrate it can do the same.

The key nuance: "outperformed Nvidia" almost certainly means "outperformed an Nvidia GPU running comparable inference workloads under OpenAI's chosen test conditions," not "is a better chip in general." Benchmark framing, model choice, precision, and power envelope all move the result. Treat the headline claim as directional, not definitive.

04The Nvidia relationship: co-opetition in the AI boom

Here is where the story gets genuinely awkward: OpenAI building anti-Nvidia silicon does not mean OpenAI is leaving Nvidia. The two companies remain deeply intertwined. OpenAI has committed to enormous purchases of Nvidia systems as part of its compute roadmap, and Nvidia has directly invested in OpenAI. Microsoft, OpenAI's closest infrastructure partner, is simultaneously one of Nvidia's largest customers.

This is textbook co-opetition, and it is now the standard posture of the AI industry. Every hyperscaler designing its own accelerators, Google, Amazon, Microsoft, and now OpenAI, still buys Nvidia GPUs in volume, because Nvidia's CUDA software ecosystem and interconnect maturity remain unmatched for training and for the newest model generations. Custom chips handle the predictable, high-volume inference layer; GPUs handle the frontier where flexibility matters.

For Nvidia, the calculus is uncomfortable but survivable in the near term: even as customers design alternatives, total AI compute demand keeps growing fast enough that Nvidia's revenue has continued climbing. The long-term risk is real, though. Every inference workload that migrates to an in-house ASIC is volume that never comes back, and Broadcom's order book grows fatter with each new convert.

05What "outperformed in testing" actually means

Bloomberg's report that Jalapeno outperformed Nvidia hardware in testing is the kind of claim that sounds simple and is anything but. The first question is: which Nvidia chip, at what precision, running what model, at what batch size? Inference benchmarks are extraordinarily sensitive to configuration. A chip tuned for low-precision, high-throughput serving of one specific model can post spectacular numbers against a GPU running the same workload in a more general configuration, while being useless for anything else.

The second question is the yardstick itself. Chip vendors and their customers routinely publish internal benchmarks that flatter the new entrant, because the comparison point is chosen by the party with something to sell or justify. OpenAI has every incentive to demonstrate that its multi-billion-dollar silicon bet is sound. Nvidia, for its part, has every incentive to dispute methodology. Without a neutral third-party evaluation, the honest summary is: early internal tests reportedly favor Jalapeno on the workloads OpenAI cares about most.

The third question is deployment reality. A chip that wins in a lab can underwhelm at scale, where power delivery, cooling, networking, and software maturity dominate. Google's TPU took years and multiple generations to reach its current position. Jalapeno has not yet shipped at volume, and real-world efficiency data will not exist until large deployments are running actual production traffic.

Estimated AI accelerator market share by vendor, 2026 Horizontal bar chart showing rough estimated shares of AI accelerator shipments: Nvidia about 80 percent, Google TPU about 8 percent, AMD about 5 percent, and all custom hyperscaler ASICs combined about 5 percent. Values are directional estimates from industry reporting, not audited market data. Estimated… Nvidia ~80% Google TPU ~8% AMD (Ins… ~5% Other… ~5% Bars are…
Source: N43 estimates synthesized from public industry reporting on AI accelerator shipments, 2026.

Chart 1: Estimated AI accelerator market share by vendor, 2026. Values are directional estimates, not audited market data.

06The broader custom-silicon race: Google, Amazon, Microsoft

OpenAI is not an outlier; it is the newest entrant in a race that has been running for a decade. Google's TPU program is the elder statesman, running in production since 2015 and now a meaningful revenue line for Google Cloud. Amazon took a different route to the same destination: acquiring Israeli chip startup Annapurna Labs in 2015, whose product lines now include the Nitro, Graviton, and Trainium chips. Trainium, Amazon's AI accelerator, anchors its own massive buildouts, including the enormous data center campus dedicated to Anthropic.

Microsoft's Maia accelerator is the third pillar, deployed in Azure for its own and OpenAI's workloads. And the common thread connecting nearly all of these programs is Broadcom or, in Amazon's case, its own acquired design house, plus TSMC as the manufacturer. The chart below lays out the timeline of major custom AI chip programs as they became publicly known.

Hyperscaler custom AI chip programs, first public disclosure Bar chart showing when each major custom AI accelerator program became publicly known: Google TPU first used internally in 2015, Amazon acquired Annapurna Labs in 2015 with Trainium announced in 2023, Microsoft Maia revealed in 2023, and OpenAI's Jalapeno reported in 2026. Bars begin at the first public disclosure year. Custom AI… 2013 2016 2019 2022 2025 2028 Google TPU internal… Amazon… Annapurna… Microsoft… revealed… OpenAI… 2026 Axis…
Source: program histories from Wikipedia summaries (Tensor Processing Unit; Annapurna Labs) and public reporting.

Chart 2: Timeline of major custom AI accelerator programs by first public disclosure, 2013-2028.

07Limits and open questions

Jalapeno remains a reported project, not a shipping product line with published specifications. The details that matter most to engineers, memory capacity and bandwidth, interconnect topology, supported precisions, and the software stack, are not public. Without those, no outside party can independently assess the performance claim, and the Bloomberg report itself relies on sources describing internal tests rather than published results.

Software is the underrated hurdle. Google's TPU succeeded partly because Google rebuilt its entire training and serving stack around it over a decade. OpenAI's stack is deeply built on CUDA-adjacent tooling. Migrating production inference to custom silicon means rewriting, testing, and validating enormous amounts of infrastructure code, and doing it while the service keeps growing. The 2026 deployment target, if met, will be the beginning of that work, not the end.

There is also a supply question. TSMC's leading-edge capacity is allocated years ahead, and Jalapeno competes for wafers with Nvidia, Apple, Qualcomm, AMD, and every other custom program Broadcom manages. A chip design that wins benchmarks means little if volumes arrive late or in short supply, and OpenAI's compute needs do not wait politely.

08What this means for AI infrastructure costs

The strategic prize is inference economics. If OpenAI can serve its models on hardware it co-designed, at meaningfully lower cost per token than merchant GPU pricing, the savings compound across every query the company serves. Even a partial migration of stable, high-volume inference to Jalapeno-class chips would reshape OpenAI's cost structure, and every percentage point of margin matters in a business where compute is the dominant expense.

For the broader industry, the signal is that custom silicon is no longer a hyperscaler-only strategy. An AI lab can, with enough capital and a partner like Broadcom, commission its own accelerator. Nvidia's position remains dominant, but the era in which one company captured nearly all AI compute value is visibly closing. Expect the next several years to be a layered market: merchant GPUs at the frontier, custom ASICs for scaled inference, and a fierce competition over who controls the economics of the layer that serves the most tokens.

Caveat: Details of the Jalapeno program, including its performance claims, come from press reporting rather than official specifications from OpenAI or Broadcom. Figures in this article's charts are directional estimates and should not be treated as audited market data.
N43 / TECH

N43 and Hermes · August 30, 2026 · Technology dispatch 7447

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

No Nvidia needed: inside Amazon's massive AI data center built for Anthropic
📰 technology

No Nvidia needed: inside Amazon's massive AI data center built for Anthropic

N43 and Hermes1h ago
Apple's M6 chip is weird: why the newest Apple silicon breaks the pattern
📰 technology

Apple's M6 chip is weird: why the newest Apple silicon breaks the pattern

N43 and Hermes1h ago
How Claude actually works: a practical guide to Anthropic's AI assistant
📰 technology

How Claude actually works: a practical guide to Anthropic's AI assistant

N43 and Hermes1h ago
ChatGPT Atlas: OpenAI enters the browser wars
📰 technology

ChatGPT Atlas: OpenAI enters the browser wars

N43 and Hermes3h ago
Gemini Omni: Google's anything-from-anything model arrives
📰 technology

Gemini Omni: Google's anything-from-anything model arrives

N43 and Hermes3h ago
ChatGPT Work: OpenAI's enterprise play and the GPT-5.6 engine
📰 technology

ChatGPT Work: OpenAI's enterprise play and the GPT-5.6 engine

N43 and Hermes3h ago
← Back to News