惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园_首页
雷峰网
雷峰网
IT之家
IT之家
I
InfoQ
酷 壳 – CoolShell
酷 壳 – CoolShell
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
B
Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 【当耐特】
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
Hugging Face - Blog
Hugging Face - Blog
A
About on SuperTechFans
月光博客
月光博客
P
Proofpoint News Feed
博客园 - 三生石上(FineUI控件)
J
Java Code Geeks
G
Google Developers Blog
小众软件
小众软件
宝玉的分享
宝玉的分享
Jina AI
Jina AI
V
Visual Studio Blog

Latest from Tom's Hardware in News

Analytics group signals possible delays at 40% of AI data center construction sites — companies deny schedule holdups, but satellite imagery indicates otherwise Intel hires tenured Samsung exec to lead Foundry Services — signals company focus on winning business from potential Foundry suitors Elon Musk pushing forward with Terafab at Microsoft's April patch puts Windows domain controllers into reboot loops — third known issue from KB5082063… AMD Ryzen 9 9950X3D2 appears on Amazon with $1,000 pre-order price — AMD confirms recommended pricing is still set… AMD's market cap hits all-time high, Intel hits 25-year high on Agentic AI's insatiable demand for CPUs Meta raising Quest headset prices due to AI-driven RAM shortage — Quest 3 to cost $600, Quest 3S $350 from April… Elegoo announces the Jupiter 2 resin 3D printer for $949, early bird price of $849 — new model offers massive print volume but is still physically smaller than previous models Pragmata PC performance tested: 18 GPUs take us to the Moon Chinese fabs import record volumes of US chipmaking equipment via Singapore and Malaysia — homegrown tool makers booked record 2025 revenues as price competition squeezes margins US appeals court restarts $3 billion patent infringement lawsuit against Intel — VLSI case from 2017 returns after… Two US citizens get combined 16 years in prison for running North Korean laptop farms — fake remote IT work scheme netted DPRK $5 million in around three years Intel launches Wildcat Lake as Core Series 3 for value laptops and edge systems — six consumer SKUs built on 18A promise Legendary Qualcomm, Apple, and Nuvia alumni form new CPU startup — Nuvacore promises to Bambu updates its 3D printers to print unique hues or gradients using two or three filaments — company acknowledges OrcaSlicer-FullSpectrum fork as the basis for the color prediction part of the new feature Broadcom to supply Meta with custom silicon through 2029 — Broadom CEO Hock Tan departs Meta Anonymous perps behind 86 million files scraped from Spotify hit with $322 million court judgement — Anna Non-functioning counterfeit Samsung 990 Pro SSDs are circulating in Europe — Despite convincing packaging, blue… Oklahoma farmer arrested and jailed for trespassing during AI data center town hall — removed by officers after going a few seconds over allotted speaking time, trying to hand paperwork to counselors Virginia voter support for new data centers collapses from 69% in 2023 to 35% in new poll — Multi-gigawatt, 37-building Digital Gateway project abandoned Struggling shoemaker and apparel brand Albird pivots to AI data centers, stock jumps 580% in a single day — sells core business and leveraging $50 million in financing to become a GPU-as-a-Service and AI cloud solutions provider IPv6 usage reaches historic 50% across Google services, matching IPv4 — increased usage eases pressure on the IPv4 address market as 'new' protocol designed in 1998 finally hits its stride Engineer open-sources DIY radar system that's 95% cheaper than $250,000 commercial offerings, has 20 kilometer range — Moroccan engineer designs Aeris-10 radar, shares it on GitHub Elon Musk demonstrates first sample of Tesla AI5 processor, accidentally thanks TSC rather than TSMC  — claims 40X performance boost over the predecessor Valve might be adding a 30-day price tracker to Steam — feature is already available in some EU countries to spoof… Memory cards and flash drives prices rocket 124%, some products peak at 261% jump — increases from 2025 driven by AI chip shortage across a range of formats and capacities Netgear secures conditional approval from the FCC following router ban — company can continue importing foreign-made routers through October 2027 China tests deep-sea electro-hydrostatic actuator that can cut undersea cables at a depth of 3,500 meters — state hails successful trial and hints at deployment readiness Iran reportedly bought an in-orbit Chinese satellite to target US military sites in the Middle East — purchase agreement included ongoing ground control services based in China Our lifestyle tech colleagues at Tom's Guide have overhauled their site for smarter shopping — more video and access to experts make it 'the biggest relaunch in our history'
Broadcom and OpenAI unveil custom-built Jalapeño inferenc...
https://www.tomshardware.com/author/anton-shilov · 2026-06-25 · via Latest from Tom's Hardware in News
OpenAI Jalapeño
(Image credit: OpenAI)

OpenAI and Broadcom have introduced Jalapeño, a custom-built inference processor designed specifically for modern large language models and future agentic AI workloads, which is designed to deliver performance per watt they claim is higher than today's leading-edge hardware. OpenAI considers its hardware project a strategic one and envisions Jalapeño to be the first generation of its inference hardware.

Not another AI accelerator

OpenAI stresses that Jalapeño is a purpose-built inference ASIC and not a repurposed training accelerator or a general-purpose AI processor. OpenAI says the architecture of Jalapeño was designed based on its understanding of LLM behavior and is meant to address practical bottlenecks that matter for inference at scale, including costly data movement, balance between compute and memory resources, networking efficiency, and overall behavior. OpenAI also states that the design of the processor is meant to wed high throughput with low latency (which is why it uses a huge compute chiplet and HBM memory and not cheaper types of DRAM like many other inference accelerators), which will be particularly handy for reasoning and agentic workloads.

In addition, OpenAI and Broadcom claim the processor is built to deliver higher effective utilization than conventional AI accelerators and deliver performance that is close to the theoretical maximum, which means very high efficiency both in terms of costs and in terms of power. Meanwhile, the companies did not disclose performance targets for their Jalapeño ASIC, so these claims should be taken with a grain of salt.

Engineering samples are already operating in the lab at target clock speed and power (though Broadcom and OpenAI do not disclose details about this, either), and OpenAI says it is running machine learning workloads, such as GPT-5.3-Codex-Spark.

The two companies also claim that early internal testing indicates that Jalapeño's performance-per-watt is substantially better than 'current state-of-the-art hardware,' although no hard numbers, benchmarks, memory configuration, or other details are disclosed, so again, we will have to take the claims with a grain of salt. In addition, one must bear in mind that while Jalapeño can purportedly beat existing AMD's Instinct MI350-series and Nvidia's Blackwell-based accelerators, it remains to be seen how competitive it will be against AMD's Instinct MI400-series and Nvidia's Rubin-based offerings.

"Jalapeño was designed from the ground up for LLM inference using detailed insights from our close collaboration with OpenAI researchers," said Richard Ho, who leads OpenAI's hardware program. "We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits."

A massive chip with six HBM modules

While Broadcom and OpenAI did not disclose specifications of Jalapeño, they did show its wafer and packaging, so we can do a brief analysis. The package appears to contain one large compute chiplet surrounded by six HBM modules and another chiplet that likely packs input/output interfaces and is surrounded by two structural dummy dies.

Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.

OpenAI Jalapeño

(Image credit: OpenAI)

The wafer image does look like a Broadcom-style systolic-array-heavy accelerator, in the sense that it shows a very regular, repeated, columnar floorplan with what looks like replicated compute regions and fixed infrastructure macros. Yet, keep in mind that we are speculating, and the image is not clean enough to say that this is definitely Broadcom's standard TPU-like systolic array template with some perks from OpenAI,

From the image alone, it is impossible to tell whether Jalapeño uses a true 2D systolic array, a set of 1D/2D matrix engines, a collection of vector or tensor tiles, or some other inference datapath. All we can say is that the die has a highly repetitive floorplan consistent with several kinds of tiled AI accelerator architectures.

OpenAI Jalapeño

(Image credit: OpenAI)

What we can tell from the image is the approximate die size of Jalapeño's compute chiplet based on the size of HBM3/4 packages (10.975 mm × 10.975 mm) that surround it. From what we can tell, the chiplet measures 25.46 mm (width) × 33 mm (height), which means that its die size is around 840 mm2, which is very close to the reticle size of EUV lithography systems (858 mm2). Given that the quality of the shot is poor, the die size we estimate cannot be 100% accurate, but we suspect it is close enough.

The die size of Jalapeño's compute chiplet implies that it packs quite a lot of compute oomph, though, of course, we cannot make performance estimates based on this metric. Yet, it is safe to say that Jalapeño's compute die is considerably bigger than compute dies of other inference accelerators on the market and more resembles processors for AI training. Speaking of processors for AI training, we increasingly see multi-chiplet designs for these workloads as companies like AMD and Nvidia want to pack as much performance as possible. Meanwhile, the fact that OpenAI and Broadcom chose to go with a large compute chiplet possibly indicates that they wanted to reduce latencies by as much as possible.

Designed in nine months

The companies say the chip reached tape-out in just nine months and is slated for deployment beginning in late 2026, which represents an extremely fast turnaround time in ASIC design. It is unclear whether Broadcom and OpenAI extensively used artificial intelligence to define and then develop Jalapeño, though the companies admitted that they used OpenAI's models to speed up parts of the chip's design and optimization work. Typically, it takes 1.5 – 2 years to design an ASIC from scratch, so AI can shrink the development cycle. Another means to accelerate the design cycle is Broadcom's extensive reuse of its logic across different custom designs to deliver new chips faster than other companies.

It is noteworthy that, according to the announcement, Jalapeño is designed to support not only OpenAI's own workloads but also present and future LLMs across the industry, which potentially lets OpenAI sell its hardware to third parties, assuming that it can get enough supply from Broadcom and TSMC. Meanwhile, the chief executive of Broadcom indicates that Jalapeño will be deployed at gigawatt-scale data centers with Microsoft and other partners starting this year, though it is unclear whether the processor will be used exclusively for OpenAI workloads or will be available for other tenants as well.

"Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI," said Hock Tan, President and CEO, Broadcom. "This is just the beginning of a multi-generation roadmap. By co-developing our industry-leading silicon directly with OpenAI, we are enabling the deployment of gigawatt-scale data centers with Microsoft and other partners beginning in 2026."

Google Preferred Source

Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

Anton Shilov is a contributing writer at Tom’s Hardware. Over the past couple of decades, he has covered everything from CPUs and GPUs to supercomputers and from modern process technologies and latest fab tools to high-tech industry trends.