惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Tor Project blog
博客园 - 聂微东
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 【当耐特】
G
Google Developers Blog
J
Java Code Geeks
The Cloudflare Blog
Attack and Defense Labs
Attack and Defense Labs
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
Cisco Talos Blog
Cisco Talos Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
I
Intezer
Jina AI
Jina AI
T
Tenable Blog
P
Palo Alto Networks Blog
Project Zero
Project Zero
D
DataBreaches.Net
Hugging Face - Blog
Hugging Face - Blog
The Hacker News
The Hacker News
F
Full Disclosure
Cloudbric
Cloudbric
量子位
H
Heimdal Security Blog
K
Kaspersky official blog
有赞技术团队
有赞技术团队
罗磊的独立博客
V
Vulnerabilities – Threatpost
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
阮一峰的网络日志
阮一峰的网络日志
Vercel News
Vercel News
Recent Announcements
Recent Announcements
WordPress大学
WordPress大学
GbyAI
GbyAI
S
SegmentFault 最新的问题
M
MIT News - Artificial intelligence
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
I
InfoQ
Recorded Future
Recorded Future
Security Archives - TechRepublic
Security Archives - TechRepublic
AI
AI
Webroot Blog
Webroot Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
爱范儿
爱范儿
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
The Exploit Database - CXSecurity.com
Apple Machine Learning Research
Apple Machine Learning Research
C
Cybersecurity and Infrastructure Security Agency CISA
H
Hacker News: Front Page
Latest news
Latest news

Towards AI

Building AI Agents in Rust — part 4 | Towards AI Building AI Agents in Rust — part 5 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
How to Build a Self-Improving Company with AI
Editorial Team · 2026-06-03 · via Towards AI

Author(s): Naresh Idiga

Originally published on Towards AI.

Most founders are using AI wrong. They’re adding it on top of existing company structures. That’s the old model.

How to Build a Self-Improving Company with AI
This image was created using an AI image creation program.

This article was written with the assistance of an AI writing program, based on the author’s notes and analysis of Tom Blomfield’s YC batch talk.

Tom Blomfield co-founded Monzo — took it from zero to 5M+ customers and $900M raised. Now he’s a General Partner at Y Combinator. In a recent batch talk, he made one core argument, built partly on ideas he borrowed from Jack Dorsey’s tweets and a talk by fellow YC partner Diana:

Most founders are using AI wrong. They’re adding it on top of existing company structures. That’s the old model.

The new model: your company is a set of recursive, self-improving AI loops.

Here’s the framework — and what it means for every developer, builder, and startup founder right now.

TL;DR

  • Most companies bolt AI onto old hierarchies. Wrong model.
  • The right model: recursive, self-improving AI loops that get better while you sleep.
  • Key shifts: record everything → make your company legible to AI → burn tokens not headcount → cut middle management → keep humans at the edge.
  • The constraint is shifting from headcount to token spend. The companies that win will maximize what AI can do, not how many people they employ.

Table of Contents

  1. The Roman Legion Problem
  2. The Self-Improving AI Loop — All 5 Layers
  3. The Holy Shit Moment at YC
  4. Real Examples: Product, Support, and Ops Loops
  5. Four Things to Do Right Now
  6. Where Humans Still Matter
  7. What This Means for Developers and Builders

1. The Roman Legion Problem

Blomfield opened with a sharp analogy — one he credits to Jack Dorsey’s recent tweets.

The Roman legions were built to project power from Rome to the far edges of the empire through nested hierarchies: named individuals with fixed spans of control, passing orders down and sending information back up.

Most companies today are still organized like Roman legions. Humans are the conduit for information flowing up and down. That’s not a productivity problem. It’s a structural assumption baked into how we think about organizations.

And AI doesn’t just improve this structure — it breaks the assumption underneath it.

The old mental model: add AI co-pilots to your existing workflows. Make engineers 20% more productive. Ship more code. Blomfield’s critique: that’s like putting a more powerful engine onto a horse-drawn cart. You’re taking the old way of working and adding a more powerful engine onto it.

The new mental model: redesign what a company is.

This image was created using an AI image creation program.

2. The Self-Improving AI Loop — All 5 Layers

Instead of thinking about AI as a tool bolted onto your org, think about your company as a set of recursive, self-improving loops. Each loop has 5 layers. If every single step runs with minimal human intervention, your system gets better and better while you’re sleeping.

Here are all five layers:

Layer 1 — Sensor Layer This is where real-world data comes in. Emails from customers, support tickets, code changes, people cancelling their subscription, product telemetry. Think of it as the sensory layer that picks up signals from the outside world.

Layer 2 — Decision Layer (Policy Layer) This is where rules live. What can the system act on autonomously? What does it have to log? What must it ask a human permission for? This is the policy and decision layer — guardrails for what the AI can and cannot do on its own.

Layer 3 — Tool Layer This is where code executes. Deterministic APIs — query a database, look at a calendar, send an email, update a record. These are the tools the AI can call. Blomfield describes this as the “skills and code” layer — the actual mechanisms of execution.

Layer 4 — Quality Gate Before any action is committed, it passes through a quality gate. This might be eval checks, safety filters, or human review for high-risk actions. It’s the checkpoint between “the AI decided to do something” and “the thing actually happens.”

Layer 5 — Learning Mechanism This is what makes the whole thing self-improving. The system interacts with the real world, picks up where it didn’t work, and loops back to the top again. Failures become training signal. The loop tightens over time.

This image was created using an AI image creation program.

The key insight: if you can run every single step of that loop without human intervention — or with minimal human supervision — your system gets better and better overnight. Automatically.

3. The Holy Shit Moment at YC

Blomfield described exactly when this clicked for him at YC.

They built a simple internal agent — a tool their partners could ask questions like “when did I last have a meeting with this founder?” It worked well. Partners got answers faster. Useful, but nothing groundbreaking. Think of it as a smarter search bar for their internal database.

Then they put a monitoring agent on top of it.

This agent watched every query every YC employee made. It tracked when queries succeeded and when they failed. When they failed, it asked: Why? What would have made this work? Do we need different deterministic tools? A new database view? A different index?

Then — overnight — it wrote the fix, opened a pull request to the YC codebase, had another agent review it, and merged and deployed it. By the time a human came in the next morning to ask the same query, it worked.

No human involved.

This image was created using an AI image creation program.

“For me, that was like the holy shit moment. That’s not just AI making you 20 or 30% more valuable. It is the AI going through this loop to figure out how to self-improve.”

That’s the shift. Not a productivity gain. A system that detects its own failures and fixes them.

4. Real Examples: Product, Support, and Ops Loops

Blomfield didn’t stop at the YC internal example. He described how this same loop structure applies to every business function.

The self-optimizing product loop: Imagine an agent that goes through your product analytics to figure out which part of your sales funnel has the highest friction. It researches best practices, puts in place an A/B test, runs it for a week, picks the best-performing version, and deploys it. Then does that again and again and again. A self-optimizing product loop — running continuously, without a product manager needing to be involved in each cycle.

The automated customer support loop: Customer suggestions come in constantly. An AI “CPO/CTO” agent makes judgment calls: this suggestion doesn’t fit the roadmap, discard it. This one does — write the code, deploy it, ship it to the customer. No human involved in the decision, the build, or the delivery. The loop closes overnight.

These aren’t speculative. Blomfield was describing loops that exist or are being built right now inside YC companies.

This image was created using an AI image creation program.

5. Four Things to Do Right Now

1. Make your company legible to AI

If an interaction isn’t recorded, it doesn’t exist for your AI system. That means every email needs to be in a searchable database. Every Slack message. Every meeting recorded. Every decision logged.

YC took 2,000 hours of recorded office hours from the last 3 months and used AI to regenerate their entire user manual — a living document that now updates itself with every new piece of advice given. That’s what “legible to AI” means in practice.

Practical step: Audit your company. What information lives only in someone’s head, in a chat that expires, or in a meeting that was never recorded? That’s where you’re leaking intelligence.

2. Burn tokens, not headcount

The constraint is shifting. You will hit limits on token usage before you hit limits on hiring. The directional measure of a high-performing team member right now: how much are they actually using AI? Who is token-maxing?

Practical step: Before hiring for a function, ask whether that function could run as a well-designed AI loop with a human supervisor.

3. Eliminate middle management

Middle management exists to solve a coordination problem. AI solves that problem better and faster.

What remains: individual contributors (ICs) who build and operate, and a directly responsible individual (DRI) for every meaningful outcome. Blomfield: “I just don’t think you need middle management for this coordination problem. I think AI should be doing it.”

Practical step: Map your org chart. For every manager layer, ask: is this person primarily coordinating information flow? That’s a candidate for AI replacement.

4. Treat software as ephemeral, context as valuable

Models improve every few months. The dashboard you built six months ago? Regenerate it. The internal tool? Rewrite it with the latest model and better instructions.

What you should never throw away: the business context. The domain knowledge. The reasoning behind decisions. Blomfield: “The valuable part is the comprehension inside people’s heads of how the function works. The software on top of it is ephemeral.”

Practical step: Store everything as text — markdown, structured logs, plain-language runbooks. Your business knowledge should be model-agnostic and always ready to be fed into the next generation of AI.

This image was created using an AI image creation program.

6. Where Humans Still Matter

Blomfield is clear: humans don’t disappear. They move to the edge — where your intelligence system makes contact with reality.

That includes:

  • Novel situations the models haven’t encountered before
  • High-stakes, high-emotion moments — a founder thinking about breaking up with their co-founder, a difficult client negotiation
  • Ethical judgment calls that require genuine moral reasoning
  • Sales conversations — Blomfield says these remain human for the next 20 years
This image was created using an AI image creation program.

The company brain — all your data, emails, DMs, skills, and institutional knowledge — sits in the middle. Humans wrap around the outside, interfacing with the real world.

The Roman legion had humans everywhere. The AI-native company has AI everywhere in the middle, and humans precisely placed at the edges where it matters most.

7. What This Means for Developers and Builders

If you’re a developer, startup founder, or indie builder, this talk isn’t just interesting — it’s a forcing function.

The old valuable skill: writing code fast.

The new valuable skill: designing systems that improve themselves.

This image was created using an AI image creation program.

Old approach New approach Write code fast Design self-improving systems Add logging later Build with observability from day one Fix bugs when reported Failures feed back into the loop automatically Scale by hiring Scale by designing loops Ship and forget Ship and monitor the quality gate

The developers who win in the next 5 years won’t be the ones who prompt the most cleverly. They’ll be the ones who build companies and systems that learn.

Start with one loop. Pick one part of your product, ops, or content pipeline where you can close the feedback cycle — detect failure, analyze it, fix it, verify. Make that one loop self-improving.

Then do it again next week.

Blomfield ended with a question worth sitting with:

“If you were building your company today, would you start it in this shape?”

For most of you reading this — you’re small enough to build it right from the start. That’s the advantage.

Based on Tom Blomfield’s batch talk at Y Combinator. Watch the original (13 minutes): youtube.com/watch?v=t-G67yKAHBQ

This article was written with the assistance of an AI writing program.

Published via Towards AI