OpenAI’s Jalapeño chip aims to speed up inference workloads while reducing compute costs.

OpenAI has unveiled its first custom AI accelerator, called Jalapeño, marking the company’s move into chip design as it looks to reduce the cost and improve the efficiency of running large language models (LLMs).
Developed in partnership with Broadcom and Celestica, the chip is designed specifically for AI inference — the process of generating responses from trained AI models. OpenAI said early testing shows Jalapeño delivers significantly better performance per watt than current state-of-the-art AI accelerators, though detailed benchmarks will be released later.
The announcement expands OpenAI’s efforts to control more of the infrastructure behind its products. In addition to building models and applications such as ChatGPT and Codex, the company is now designing the hardware that powers them. Engineering samples of Jalapeño are already running machine learning workloads in the lab, including GPT-5.3-Codex-Spark, at production target frequency and power levels, according to the company.
Custom chip, faster AI
Unlike general-purpose AI accelerators adapted for multiple workloads, Jalapeño was built specifically for LLM inference. OpenAI said the architecture was designed around the compute, memory, networking, and serving requirements of modern AI models.
The company claims the chip reduces data movement while balancing compute, memory, and networking resources to improve hardware utilization. Broadcom contributed silicon implementation and networking technologies, including its Tomahawk networking platform.
“Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems,” said Greg Brockman, President and Co-Founder of OpenAI.
Richard Ho, who leads OpenAI’s hardware program, said the accelerator was optimized around the workloads most important for frontier AI systems.
“Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits,” Ho said. The chip is also intended to support future LLMs across the broader AI industry, not just OpenAI’s own models.
Nine-month design sprint
According to the companies, Jalapeño was developed from initial design to manufacturing tape-out in just nine months. OpenAI described the effort as potentially the fastest ASIC development cycle achieved for a high-performance advanced semiconductor.
The development process involved extensive software-hardware co-design between OpenAI and Broadcom engineers. OpenAI also said its own AI models were used to accelerate portions of the chip design and optimization workflow.
“Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI,” said Hock Tan, President and CEO of Broadcom.
The companies plan to deploy the accelerator at gigawatt-scale data centers beginning in 2026. Jalapeño is the first product in what OpenAI describes as a multi-generation compute platform that will combine OpenAI-designed accelerators with Broadcom networking and connectivity technologies and Celestica’s system integration expertise.
OpenAI said improvements in inference efficiency could translate into faster ChatGPT responses, lower AI operating costs, and more reliable access to advanced AI services as demand continues to grow.
Recommended Articles
Get the latest in engineering, tech, space & science - delivered daily to your inbox.
With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.

















