On June 24, 2026, OpenAI put its name on a piece of silicon for the first time. The OpenAI Jalapeño chip is a custom processor built specifically to run large language models after they have already been trained, and it arrives roughly eight months after the company signed a sprawling hardware pact with Broadcom. TechCrunch reported that the part was designed around the narrow demands of OpenAI's own inference systems rather than adapted from a general-purpose accelerator, which is a meaningful shift for a company that until now has leaned almost entirely on Nvidia hardware bought through partners.
- Jalapeño is OpenAI's first in-house chip, aimed at inference rather than model training.
- It was designed with Broadcom, fabricated by TSMC, and integrated into systems by Celestica.
- Large-scale deployment is slated for late 2026 at gigawatt scale, with Microsoft reportedly committed to a large share of early supply.
- OpenAI claims a sizable performance-per-watt advantage, but those figures are self-reported and unverified.
What OpenAI actually announced
The reveal centered on a single product OpenAI is calling an "Intelligence Processor." According to The Decoder, Jalapeño was engineered from the start for language-model inference instead of being a modified version of an existing design. That framing matters because inference and training place very different loads on a chip. Training chews through enormous batches of data over weeks or months and rewards raw throughput. Inference, the work of generating a response once a user sends a prompt, is latency-sensitive, runs constantly, and now represents the bulk of the day-to-day cost of operating a service like ChatGPT.
OpenAI President Greg Brockman framed the project as a hunt for places where standard hardware leaves value on the table. "We have a deep understanding of the workload," he told TechCrunch, adding that the team went looking for "specific workloads that are underserved" by what is currently available. The company has spent years optimizing models and products while renting compute; Jalapeño is the first time it has tried to push those lessons down into the metal itself.
The unveiling did not come out of nowhere. OpenAI and Broadcom announced a strategic collaboration in October 2025 to deploy 10 gigawatts of OpenAI-designed accelerators, with OpenAI handling the design of the chips and systems and Broadcom developing and deploying them. The press release described Broadcom racks of accelerator and network systems starting in the second half of 2026 and completing by the end of 2029. Jalapeño is the first concrete product to emerge from that agreement.
How the OpenAI Jalapeño chip is built
The supply chain behind the chip reads like a who's who of the contract-silicon world, with each company taking a defined slice of the work. OpenAI owns the architecture and design. Broadcom contributes the silicon manufacturing expertise and the networking technology that ties many chips together, including its Tomahawk networking line, as The Decoder noted. Fabrication itself runs through Taiwan Semiconductor Manufacturing Company, the same foundry that produces Nvidia's flagship accelerators and most of the industry's leading-edge logic. Canada-based Celestica handles the board and rack integration, turning bare chips into the dense server cabinets a data center can actually plug in.
That division of labor is the standard recipe for a custom AI ASIC in 2026. A model company brings deep knowledge of its own workloads, a partner like Broadcom supplies the physical-design muscle and intellectual property, and a foundry prints the result. What stands out here is the speed. The Decoder reported that the design phase took about nine months, which is quick for a high-performance processor, and that OpenAI used its own models to accelerate parts of the design process. Engineering samples are already running real machine-learning workloads, including the company's GPT-5.3-Codex-Spark model, so the part is past the paper stage.
A nine-month design cycle is unusually compressed for a leading-edge accelerator, where eighteen months to two years is more typical. OpenAI attributes part of that pace to using its own models as design aids, though it has not published details on exactly which steps were automated.
The inference-only design choice
The decision to build an inference chip rather than a training chip is the most telling part of the announcement. Training the next frontier model is still expected to run on Nvidia hardware, and TechCrunch reported that more performance-intensive tasks like pre-training will continue to use Nvidia GPUs. Inference, by contrast, is where the economics have become hardest to manage. Every prompt a user sends costs money to answer, and as products like ChatGPT and Codex along with the API business scale into hundreds of millions of users, the cumulative bill for serving responses has grown into one of the largest line items in the company's compute budget.
By tailoring Jalapeño to that single job, OpenAI can strip out the parts of a general-purpose GPU it does not need and spend the saved silicon area and power on what it does. The Decoder quoted the company's claim that the architecture "cuts data movement and pushes utilization closer to its theoretical max." Data movement is a useful thing to attack, because shuttling weights and activations between memory and compute units burns a large fraction of the energy in modern accelerators. A chip that keeps more data local and keeps its arithmetic units busy can, in principle, do more useful work per joule.
TechCrunch added a specific use case: the chip is optimized for low operating cost when running real-time coding models. That tracks with OpenAI's heavy investment in Codex and agentic coding tools, which generate long streams of tokens and run for extended sessions. If a workload is both high-volume and predictable, it is exactly the kind of target a fixed-function accelerator can serve well, since the hardware can be tuned to the shape of the traffic.
There is a deeper reason inference deserves its own silicon. A training run is a one-time capital event for a given model, but inference is a recurring operating cost that scales with every new user and every new feature. As models grow more capable and agentic systems make many calls to answer a single request, the number of tokens generated per user has climbed sharply. A coding agent that plans, writes, tests, and revises can emit tens of thousands of tokens to complete one task, far more than a single chat reply. Multiply that by a large user base and the inference bill starts to dominate the budget. Hardware that lowers the cost per token, even modestly, compounds into very large savings at OpenAI's volume, which is precisely the lever Jalapeño is meant to pull.
Performance claims and what stays unverified
OpenAI says Jalapeño delivers performance per watt that is "substantially better" than the best hardware available today, with early testing showing what TechCrunch described as significantly better performance-per-watt than current alternatives. Those are striking numbers if they hold up at scale, because performance per watt is the metric that ultimately decides how much it costs to run a model and how much compute a fixed power budget can buy.
The caveat is that the figures are self-reported. The Decoder was direct about this, noting that the claims remain unverified pending a technical report, and that the testing conditions, the competitive benchmarks used, and the specific tasks measured are not yet clear. "Substantially better" is a comparison without a published baseline. It is common for first-party silicon announcements to lead with favorable internal numbers, and it is equally common for the picture to look more nuanced once independent benchmarks arrive. Until OpenAI publishes methodology or third parties get hands-on access, the responsible read is that the chip looks promising on the company's own tests and that the real comparison against Nvidia's latest inference parts is still pending.
"Performance per watt that is substantially better than the best hardware available today" remains, for now, a first-party claim awaiting a technical report and independent benchmarks.
The Microsoft share and the deployment timeline
One of the more revealing operational details came from The Decoder, which reported that Microsoft is expected to purchase 40 percent of the initial production run, and that this commitment was a requirement Broadcom imposed before it would greenlight manufacturing. That arrangement says a lot about how the risk of a first-generation custom chip gets shared. Broadcom wanted a guaranteed buyer for a meaningful chunk of early volume, and Microsoft, which holds a deep commercial relationship with OpenAI and runs much of its infrastructure, was positioned to provide it.
On timing, large-scale deployment is set to begin in late 2026 at gigawatt scale, according to The Decoder. That lines up with the broader Broadcom partnership timeline, which targeted the second half of 2026 for the first racks and the end of 2029 for completion of the full 10-gigawatt buildout. Gigawatt scale is the unit the industry now uses to talk about AI data centers, and reaching it on a first-generation in-house chip would be a significant logistical feat, dependent on TSMC capacity, Celestica's integration throughput, and the availability of power and networking.
The Microsoft commitment also sheds light on how the costs and rewards of the project are distributed. OpenAI gets a tailored chip and the strategic independence that comes with it. Broadcom books a large multi-year hardware program and deepens its position as the go-to partner for hyperscale custom silicon. Microsoft locks in supply of a part optimized for the exact models it serves through its own cloud and copilot products. Each side is hedging a different risk, and the fact that Broadcom asked for a purchase guarantee before committing to manufacturing shows that even a chip with OpenAI's name on it has to clear the same financial gates as any other large semiconductor program. First silicon is expensive to bring up, mask sets cost millions, and a foundry slot reserved for a part that does not sell is a costly mistake. The 40 percent commitment removes much of that downside before the first wafer ships.
Why custom silicon, and the Nvidia question
The strategic logic behind Jalapeño is vertical integration. OpenAI argues that designing its own chips and systems lets it embed what it has learned from building frontier models directly into the hardware, running models faster, more reliably, and at lower cost. That is the same argument cloud providers have made for years, and it rests on a simple observation: when you control both the model and the metal, you can co-design them so that neither wastes effort accommodating the other.
The unspoken second motive is reducing dependence on Nvidia. TechCrunch placed Jalapeño squarely in that context, framing it as designed to reduce reliance on Nvidia GPUs. Nvidia's accelerators remain the default for serious AI work, and demand has consistently outstripped supply, giving the chipmaker enormous pricing power. A buyer the size of OpenAI has every incentive to build an alternative for at least part of its fleet, both to lower cost and to gain leverage. It is worth being precise about the scope here, though. Jalapeño is an inference part. It does not replace Nvidia for training, and OpenAI has signaled it will keep buying GPUs for the most demanding workloads. The likely outcome is a mixed fleet, with custom silicon handling steady high-volume inference and Nvidia hardware reserved for the frontier.
Where Jalapeño sits in the custom-ASIC race
OpenAI is late to a party that has been going for years. Google has run its Tensor Processing Units in production since the mid-2010s and now uses them for both internal workloads and Google Cloud customers. Amazon has built out its Trainium and Inferentia lines for training and inference respectively. Meta has shipped its MTIA accelerators for recommendation and, increasingly, generative workloads. A Tom's Hardware survey of the custom-ASIC landscape earlier in 2026 mapped how Broadcom in particular has become the connective tissue behind many of these efforts, supplying the design services and networking IP that turn a hyperscaler's ambitions into shipping hardware.
What separates OpenAI from those incumbents is that it is not a cloud provider with a decade of chip experience. It is a model company that, until this year, defined itself by software. Building competitive silicon on a nine-month design cycle, even with a partner as capable as Broadcom, is an aggressive move, and the proof will be in sustained production rather than a launch-day announcement. The companies that succeeded with custom AI chips did so over multiple generations, learning from each one. Jalapeño is generation one.
The competitive picture also explains why Broadcom keeps appearing in these stories. Rather than selling finished accelerators the way Nvidia does, Broadcom sells the ability to build them, packaging design services, high-speed networking, and the deep relationships with foundries that a software company lacks. As more large AI buyers conclude that off-the-shelf parts leave money on the table, that business has become one of the most valuable in the industry. OpenAI joining the list of Broadcom's custom-silicon customers is as much a statement about Broadcom's position as it is about OpenAI's ambitions. For Nvidia, the trend is worth watching but not yet alarming: its accelerators still dominate training, its software ecosystem remains the default, and even buyers building their own inference chips keep ordering GPUs in volume. Custom silicon chips away at the edges of that dominance rather than confronting it head-on, at least for now.
What to watch as Jalapeño scales
The announcement establishes intent and a credible supply chain, but the hard part is still ahead. The numbers that matter most are the ones OpenAI has not yet published: real performance per watt against Nvidia's current inference accelerators under comparable conditions, yield and cost at TSMC, and how quickly Celestica can integrate systems at gigawatt scale. The Microsoft purchase commitment suggests Broadcom needed a financial backstop to commit to volume, which is a reminder that first-generation silicon carries real risk even for a company with OpenAI's resources.
It is also worth remembering how young this effort is. OpenAI has shipped exactly one design, and the gap between a working engineering sample and a part deployed across gigawatts of data center capacity is wide. Yield rates at TSMC, the reliability of the systems Celestica assembles, and the maturity of the software stack that compiles OpenAI's models down to the new hardware all have to come together before the chip earns its keep. None of those are guaranteed on a first attempt.
If Jalapeño performs close to its claims and ships on schedule, it would give OpenAI a measure of control over its single largest variable cost and a bargaining chip in its relationship with Nvidia. If it slips or underdelivers, the company still has its Nvidia supply to fall back on, which is part of why an inference-only first chip is a sensible place to start. Either way, the broader signal is clear: the largest AI labs no longer see buying chips off the shelf as sufficient, and the line between who builds models and who builds the hardware to run them is getting harder to draw. The next year of benchmarks, technical reports, and shipped racks will show whether OpenAI's first attempt at silicon lives up to the bet it just made.