惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LangChain Blog
B
Blog RSS Feed
阮一峰的网络日志
阮一峰的网络日志
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
H
Help Net Security
MyScale Blog
MyScale Blog
WordPress大学
WordPress大学
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
小众软件
小众软件
大猫的无限游戏
大猫的无限游戏
Martin Fowler
Martin Fowler
Vercel News
Vercel News
S
SegmentFault 最新的问题
M
MIT News - Artificial intelligence
Microsoft Security Blog
Microsoft Security Blog
G
Google Developers Blog
Last Week in AI
Last Week in AI
Hugging Face - Blog
Hugging Face - Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
Google DeepMind News
Google DeepMind News
Engineering at Meta
Engineering at Meta
云风的 BLOG
云风的 BLOG

FourWeekMBA

Musk vs Altman: The $90B Fight That Will Define AI’s Future Why DeepMind’s $1.1B Bet Signals the End of Human-Trained AI The AI Orchestrator's Leverage Points AI & The Harness Theory Why AI Companies Are Selling Fiction as Partnership Strategy Google’s $40B Anthropic Bet Reveals AI Infrastructure Wars Anthropic’s Agent Economy Signals End of Human-Mediated Commerce Claude OS: The AI Strategy Skill That Turns Claude Into Your Analyst Agent Harness OS: Build AI-Augmented Strategic Operations 🔥 AI & The Harness Theory 🔥 The Harnessing Players Map of AI 🔥 The Business Engineer’s Claude Code OS 🔥 Skills as the Architecture of the Personal OS Google's $40B Anthropic Bet Exposes Big Tech's AI Desperation Google's $40B Anthropic Bet Signals Platform Wars 2.0 20 Mental Models For AI Business Google's TPU Gambit: Why Hardware Will Crown the AI King LinkedIn Business Model: How LinkedIn Makes Money (2026) Netflix Organizational Structure: The Culture of Freedom (2026) Amazon Pricing Strategy: How Amazon Uses Price to Win Amazon Supply Chain: The Logistics Empire (2026) Apple Supply Chain: How Apple Built the World’s Best Supply Chain Tesla Supply Chain: Vertical Integration Strategy (2026) Anthropic Business Model: How Anthropic Makes Money (2026) OpenAI Business Model: How OpenAI Makes Money (2026) Meta (Facebook) Organizational Structure 2026 Google's Agentic TPUs Signal the Death of Traditional SaaS Google's $40B Anthropic Bet Signals The End of AI Independence The OpenAI–Anthropic Convergent Bets Google’s $40B Anthropic Bet Signals the End of Open AI Innovation
OpenAI Unveils Jalapeño — Its First AI Chip, Built With B...
FourWeekMBA · 2026-06-24 · via FourWeekMBA

OpenAI just unveiled its first custom AI chip — Jalapeño — built with Broadcom in nine months. Purpose-built for LLM inference, it delivers 50% cost savings vs GPUs. OpenAI is no longer just a model company. It’s building the full stack — from products to models to silicon.

Jalapeño — First Numbers

50%

Cost savings vs current GPUs

9 mo

Design to tape-out — fastest ASIC cycle ever

2026

Initial deployment by end of year

LLM

Purpose-built for inference, not training

OpenAI designed Jalapeño from the ground up — a custom inference accelerator architected specifically for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. Broadcom handled manufacturing. The chip went from initial design to tape-out in nine months — what OpenAI calls the fastest ASIC development cycle ever in high-performance semiconductors.

Early testing shows 50% cost savings compared to current GPUs and substantially better performance per watt. First samples are being tested now. Initial deployment is planned for end of 2026, with this being the first chip in a multi-generation compute platform.

The key insight: OpenAI burns $3.7 billion per quarter. Inference is the largest cost. A chip that cuts inference cost by 50% doesn’t just save money — it changes the entire economics of serving 400M+ ChatGPT users. And it’s the answer to the $100B ad business: cheaper inference = lower floor for free-tier users = more ad impressions.

The Full Stack Play

OpenAI explicitly framed this as building the “full stack” — products (ChatGPT, Codex) → models (GPT series) → infrastructure (Jalapeño). This is the same vertical integration playbook running across the industry this week:

SpaceX: Models (xAI) + Compute (Colossus) + Dev Tools (Cursor) + Robotics (Tesla) + Connectivity (Starlink)

Anthropic: Models (Claude) + Memory supply (Micron deal) + Government trust (Glasswing) + Series H capital

Google: Models (Gemini) + Chips (TPUs) + Cloud (GCP) + Distribution (Search/Android) + Content (A24)

OpenAI (now): Models (GPT) + Chips (Jalapeño) + Dev Tools (Codex) + Distribution (ChatGPT 400M users) + Revenue (Ads + Subs)

The Structural Read

NVIDIA’S PRICING POWER JUST GOT CHALLENGED

OpenAI is Nvidia’s largest customer. A custom chip that cuts inference costs 50% is a direct shot at Nvidia’s margin structure. OpenAI won’t stop buying Nvidia GPUs for training — but every dollar of inference that moves to Jalapeño is a dollar Nvidia doesn’t get. This is why Nvidia absorbed Groq’s IP for $20B — to prevent exactly this kind of competition from emerging.

THE IPO NARRATIVE JUST GOT ANOTHER CHAPTER

This week OpenAI revealed: $100B ad target, GPT-5.5-Cyber + Daybreak, and now custom silicon. Each announcement builds the IPO story: not a model company, but a full-stack AI platform with its own chips, its own ad engine, and its own security program. That’s the pitch to public markets.

THE INFERENCE ECONOMICS CHANGE EVERYTHING

50% cheaper inference means: more free-tier users (more ad impressions), cheaper API (more developers), faster agents (more agentic products), and a viable path to profitability. The $3.7B quarterly burn was always an inference cost problem. Jalapeño is the structural answer — not more revenue, but less cost per query.

The Bottom Line

OpenAI just announced it built a chip in nine months that cuts inference costs in half. Nine months. From a company that didn’t have a hardware team two years ago. Jalapeño isn’t a moonshot — it’s a business necessity. When you’re burning $3.7 billion a quarter serving 400 million users, you either cut the cost of serving them or you run out of money. OpenAI chose to build the solution rather than rent it from Nvidia. That’s the most consequential strategic decision the company has made since launching ChatGPT — and it happened in nine months.

Sources: OpenAI, CNBC, Bloomberg — June 24, 2026