OpenAI Unveils Jalapeño, Its First Custom AI Chip Built With Broadcom

0
80
OpenAI Jalapeño, OpenAI's first custom AI inference chip built with Broadcom

OpenAI has unveiled its first custom chip, a processor called Jalapeño built in partnership with Broadcom and aimed squarely at the cost of running artificial intelligence at scale. Announced on June 24, 2026, Jalapeño is what the two companies call an Intelligence Processor, a piece of silicon designed from the ground up to do one job, run inference for large language models, rather than train them. It is the first chip in what OpenAI and Broadcom describe as a multi-generation compute platform they are building together, and it signals OpenAI’s intent to own more of the hardware that its products depend on.

The distinction between training and inference sits at the center of the design. Training is the expensive, one-time work of building a model, while inference is the act of running it, generating every answer a user sees, and at OpenAI’s scale inference happens billions of times a day. Jalapeño is a purpose-built inference chip, not a repurposed training accelerator or a general-purpose processor, and OpenAI says its architecture was shaped around the specific patterns that matter for that work, including how data moves through memory, how chips talk to one another over a network, and how requests are served. By narrowing the chip’s job, the company is betting it can beat broader hardware on the economics that define its business.

The early performance claims are aggressive. OpenAI says Jalapeño delivers performance per watt substantially better than current state-of-the-art hardware, and that in early lab testing it performs on par with Nvidia’s Blackwell chips and Google’s Tensor Processing Units while cutting the cost of each inference token by roughly 50 percent compared with current-generation GPUs. Those numbers come from OpenAI’s own testing rather than independent benchmarks, and the company has promised a detailed technical report in the coming months, so they should be read as the maker’s early figures. Even so, a 50 percent reduction in the cost of serving a token would reshape the unit economics of every product OpenAI runs.

What stands out almost as much as the chip is how fast it came together. OpenAI says Jalapeño went from initial design to manufacturing tape-out, the point at which a design is finalized and sent for fabrication, in about nine months, which the company believes may be the fastest development cycle ever achieved for a high-performance advanced semiconductor. Part of that speed came from OpenAI turning its own models loose on the work, using AI to accelerate the chip design process itself, a recursive twist where the company’s software helped build the hardware that will run its software. Engineering samples are already running machine learning workloads in the lab at production target frequency and power, including a model the company refers to as GPT-5.3-Codex-Spark.

The partnership splits the work along clear lines. OpenAI designed the accelerator and its architecture, Broadcom is providing the silicon implementation along with the networking and connectivity technology that ties large numbers of chips together, and Celestica is handling the board, rack, and system engineering needed to turn individual processors into deployable infrastructure. That division lets OpenAI focus on the parts of the design unique to its models while leaning on Broadcom’s manufacturing and interconnect expertise, the same formula that has made custom silicon viable for other large technology companies.

The move places OpenAI alongside the cloud giants that have spent years reducing their reliance on Nvidia by building chips of their own. Google has its TPUs, Amazon has Trainium and Inferentia, and now OpenAI is following the same logic, that a company spending enormous sums on compute can eventually save money and gain control by designing hardware tuned to its exact workloads. It does not free OpenAI from Nvidia overnight, since the chip is just entering deployment, but it gives the company a path toward serving its models on silicon it controls.

Jalapeño is designed for initial deployment by the end of 2026, with the platform expanding in the years that follow. The real test will come when the chip moves out of the lab and into the data centers that answer hundreds of millions of daily requests, where the promised efficiency either holds up under load or does not. If it does, OpenAI will have turned one of its biggest costs into a lever it can pull, and the broader race to escape the GPU will have gained its most closely watched entrant yet.

Related: China’s $295B plan to build AI without Nvidia · OpenAI’s GPT-5.5-Cyber model

The big picture: The AI Model Race (2026): who’s winning the frontier contest

More Nvidia news →

EntrelligenceFree guide
Your First 10 AI Skills

Your First 10 AI Skills

10 practical AI skills, copy-paste prompts and a 7-day plan to start using AI with confidence.

Download the guide →
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted