惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
Cloudbric
Cloudbric
云风的 BLOG
云风的 BLOG
A
About on SuperTechFans
The GitHub Blog
The GitHub Blog
IT之家
IT之家
F
Full Disclosure
B
Blog RSS Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
B
Blog
H
Help Net Security
The Cloudflare Blog
Recorded Future
Recorded Future
P
Proofpoint News Feed
P
Proofpoint News Feed
C
Cisco Blogs
T
Tailwind CSS Blog
P
Palo Alto Networks Blog
D
Docker
爱范儿
爱范儿
Know Your Adversary
Know Your Adversary
博客园 - 聂微东
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Y
Y Combinator Blog
雷峰网
雷峰网
AWS News Blog
AWS News Blog
D
DataBreaches.Net
博客园 - 司徒正美
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
博客园 - Franky
C
Cybersecurity and Infrastructure Security Agency CISA
Blog — PlanetScale
Blog — PlanetScale
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Latest news
Latest news
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
MongoDB | Blog
MongoDB | Blog
C
CERT Recently Published Vulnerability Notes
阮一峰的网络日志
阮一峰的网络日志
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
C
CXSECURITY Database RSS Feed - CXSecurity.com
酷 壳 – CoolShell
酷 壳 – CoolShell
C
Cyber Attacks, Cyber Crime and Cyber Security
腾讯CDC
小众软件
小众软件
G
Google Developers Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
Scott Helme
Scott Helme
O
OpenAI News

Pierce Freeman

A browser for agents | Pierce Freeman The grey market of podcast appearances The way I travel | Pierce Freeman Fixing slow AWS uploads | Pierce Freeman Local tools should still use vaults We solved scratch content first Starting a podcast in 2025 Being late but still being early Automating our home video imports Adding my parents to tailscale A deep dive on agent sandboxes Language servers for AI | Pierce Freeman My simple home podcast studio We need centralized infrastructure | Pierce Freeman Coercing agents to follow conventions using AST validation My unified theory of social selling My personal backup strategy | Pierce Freeman July updates to the homelab How the KV Cache works httpx is the right way to do web requests in Python Reputation is becoming everything | Pierce Freeman Building a (kind of) invisible mac app Updated knowledge in language models Making an ascii animation | Pierce Freeman How speculative decoding works | Pierce Freeman Under the hood of Claude Code Doing things because they're easy, not hard Speeding up sideeffects with JIT in mountaineer Firehot for hot reloading in Python Misadventures in Python hot reloading How text diffusion works | Pierce Freeman The tenacity of modern LLMs The ergonomics of rails | Pierce Freeman How language servers work | Pierce Freeman Just add eggs | Pierce Freeman Unfortunately SEO still matters | Pierce Freeman The futility of human-only web requirements Setting up Input Leap | Pierce Freeman Checking in on Waymo | Pierce Freeman The react revolution | Pierce Freeman Speeding up many small transfers to a unifi nas Quick notes on swift libraries San Francisco | Pierce Freeman Debugging a mountaineer rendering segfault Local network config on macOS Building our home network | Pierce Freeman Introducing Envelope.dev | Pierce Freeman Legacy code and AI copilots Typehinting from day-zero | Pierce Freeman Generating database migrations with acyclic graphs Lofoten | Pierce Freeman Mountaineer v0.1: Webapps in Python and React Constraining LLM Outputs | Pierce Freeman Passthrough above all | Pierce Freeman Accuracy in kudos | Pierce Freeman How quick we are to adapt The curious case of LM repetition Costa Rica | Pierce Freeman Debugging chrome extensions with system-level logging Speeding up runpod | Pierce Freeman Inline footnotes with html templates Parsing Common Crawl in a day for $60 An era of rich CLI All or nothing with remote work The Next 10 Years | Pierce Freeman Adding wheels to flash-attention | Pierce Freeman LLMs as interdisciplinary agents | Pierce Freeman New Zealand | Pierce Freeman Representations in autoregressive models | Pierce Freeman Let's talk about Siri | Pierce Freeman Minimum viable public infrastructure | Pierce Freeman Reasoning vs. Memorization in LLMs Automatically migrate enums in alembic Greater sequence lengths will set us free On learning to ski | Pierce Freeman Dolomites | Pierce Freeman Using grpc with node and typescript Opportunity years | Pierce Freeman Buzzword peaks and valleys | Pierce Freeman Buenos Aires | Pierce Freeman Network routing interaction on MacOS Independent work: November recap | Pierce Freeman Debugging slow pytorch training performance The provenance of copy and paste Debugging tips for neural network training Patagonia | Pierce Freeman Santiago | Pierce Freeman My 2022 digital travel kit AWS vs GCP - GPU Availability V2 Independent work: October recap | Pierce Freeman Planning Patagonia | Pierce Freeman Relationship modeling | Pierce Freeman The power of status updates A new chapter | Pierce Freeman Give my library a coffee shop AWS vs GCP - GPU Availability V1 Switzerland | Pierce Freeman Headfull browsers beat headless | Pierce Freeman Webcrawling tradeoffs | Pierce Freeman Copenhagen | Pierce Freeman
AI engineering is a different animal
2025-03-15 · via Pierce Freeman

John Gruber has written a lot this week on how Apple missed the Siri deadline:

When Apple showed a feature, you could bank on that feature being real. When they said something was set to ship in the coming year, it would ship in the coming year. In the worst case, maybe that “year” would have to be stretched to 13 or 14 months. You can stretch the truth and maintain credibility, but you can’t maintain credibility with bullshit. And the “more personalized Siri” features, it turns out, were bullshit.

I have no inside baseball on this particular setback. But I can tell you, Apple wouldn't be the first company to get bitten by slotting AI research into engineering planning windows.

For 30 years, ever since CPUs were pretty well understood animals and memory was more than a few kilobytes, software engineering has been a deterministic profession. You design the interface, you write the spec, you figure out constraints, then you write the code. Rinse and repeat. Great engineers can see most of the major roadblocks in the way because they can visualize how to get from Point A to Point B. You make progress day by day. You pull some all-nighters. And before you know it you have enough features to make a product.

Engineering can have measurable progress week over week and month over month. Deadlines are still missed but you can typically see them before they happen. Half of the whole scrum/agile movement was built with the goal of project tracking.1 If someone is constantly behind their target launch dates on each individual feature, you're probably going to miss the final milestone. If someone is on-track throughout, they can usually wrap it up.

AI is not like that. Deep transformer architectures perhaps the least of all. I'm reminded of this xkcd comic that rings true even now. Back when I was leading an ML research team, I always had to tell other departments: until we've done it, we haven't done it. You can make 0% progress for months only to one day hit the right combination of data, architecture, and training configurations that solve the problem. You might even get to 95% of the solution only for that last 5% to be impossible.

But even that term - impossibility. It feels so imprecise, doesn't it? Is it really impossible? So let's say we mean that it's beyond the capabilities of any probabilistic model right now. There's some missing sauce that just isn't there. But the problem is you don't know that. You can't know that. It's impossible to prove that models can't do something. It's only possible to prove that they can. The only thing we can do is keep trying, however fruitless that may seem.

You see this when you look at the release cadence of the major AI shops. Most deploy updates incrementally: updating the system prompt or adding a new tool call comes near daily for OpenAI, Anthropic, and Perplexity. Their researchers modify, A/B test, measure, and go back to the drawing board. These small changes are occasionally bundled and make a splashy media release, but the cadence is unannounced and unexpected. They release a thing when they're ready to release. That's it.

WWDC is a yearly deadline. It falls on roughly the same dates every year, and Apple wants to make a splash. It's in their DNA going back since macOS became a platform: you give developers a sneak peek of things and they help take it to the next level. I get that impulse. A deadline is motivating. And indeed, a WWDC keynote slot is perhaps the ticket to get most features to the finish line.

So: what should they have done? They could have waited. Wait until it was really polished, wait until they had gotten to that 100% (or at least to a good 90). Wait until the next WWDC or wait until a new iPhone hardware drop to do a joint presentation. I think that's what Gruber is advocating in his piece. Pre-announcing and then an indefinite delay can crush years of credibility. It's hard to earn those back.

But there's something deeper to this story. Something that speaks more about the core DNA of what makes a company culture. If Apple wants to win in the AI game, it has to rid itself of the mindset of deadline-based deliveries. It also may have to rid itself of the mindset of immediate perfection. Researchers need to be given the flexibility to ship things periodically, mess up, and try again. LLMs were embarrassing before they got good.2 We don't hear nearly as much about hallucinations as we did three years ago.

It's hard to have any tolerance for imperfection when you're one of the most valuable companies in the world. But Apple also doesn't have to train frontier foundational models. They can easily partner with research labs that do, or just provide the framework SDKs to have tighter operating system integration. If they're going to play, they need to play to win. I think that goes deeper than just a missed deadline.

  1. Agile tries to move some of the pain of big features into the planning stage. You break down big features into chunks both so different people can work on them, but also so PMs can track % completion to the overall goal. I'd debate whether it succeeds. But the hours spent backlog grooming or doing t-shirt sizing sure try to scope properly. ↩

  2. They still are embarrassing sometimes. But importantly, the big labs mostly shrug and admit that's a cost of doing state-of-the-art research. The utility outweighs the embarrassment. And never did their brands meaningfully suffer for it. ↩