惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threatpost
G
Google Developers Blog
Latest news
Latest news
Know Your Adversary
Know Your Adversary
O
OpenAI News
腾讯CDC
月光博客
月光博客
P
Privacy International News Feed
Google Online Security Blog
Google Online Security Blog
Help Net Security
Help Net Security
L
LINUX DO - 最新话题
雷峰网
雷峰网
AI
AI
Hacker News - Newest:
Hacker News - Newest: "LLM"
有赞技术团队
有赞技术团队
N
News and Events Feed by Topic
V
Vulnerabilities – Threatpost
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
D
Docker
Google DeepMind News
Google DeepMind News
T
Tor Project blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Hacker News: Ask HN
Hacker News: Ask HN
爱范儿
爱范儿
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Heimdal Security Blog
I
Intezer
WordPress大学
WordPress大学
C
CERT Recently Published Vulnerability Notes
Attack and Defense Labs
Attack and Defense Labs
www.infosecurity-magazine.com
www.infosecurity-magazine.com
P
Privacy & Cybersecurity Law Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
V2EX
博客园 - 三生石上(FineUI控件)
G
GRAHAM CLULEY
Security Archives - TechRepublic
Security Archives - TechRepublic
F
Fortinet All Blogs
L
LangChain Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
Spread Privacy
Spread Privacy
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
V2EX - 技术
V2EX - 技术
Stack Overflow Blog
Stack Overflow Blog
Recent Announcements
Recent Announcements
T
Tenable Blog
Microsoft Azure Blog
Microsoft Azure Blog
V
Visual Studio Blog
SecWiki News
SecWiki News
Cisco Talos Blog
Cisco Talos Blog

MIT Technology Review

Want to get a data center online quickly? Give it some flex. Why do South Koreans love AI so much? This man with ALS is “the first power user” of a brain implant that lets him speak The Download: cutting AC emissions, and nature’s drug designer These new solid-state ACs promise a cool future. Scientists aren’t so sure. The Download: “reprogramming” aging, and the hidden sense of interoception You do your own time Why “reprogramming” is the buzziest approach to reversing aging right now Inside interoception: The hidden sense of how you feel inside The Download: soccer’s data renaissance and China’s big nuclear plans Google DeepMind is worried about what happens when millions of agents start to interact Job titles of the future: Nature’s drug designer Inside soccer’s data renaissance Why China is betting on big nuclear reactors The Download: the “steroid olympics” and a safer Mythos The “steroid olympics” were a circus—and a window into our culture The Download: whole-body rejuvenation drugs and five things to know about AI Learning to lead in a hybrid human-AI enterprise David Sinclair plans to test whole-body rejuvenation drugs in the XPrize competition Five things you need to know about AI The Download: how the World Cup ball will fly and OpenAI’s “super app” Why this year’s World Cup ball may not fly as far The Download: AI hacking beyond Mythos, and chatbots’ impact on our brains Are AI chatbots making us lose control of our brains? The Meta hack shows there’s more to AI security than Mythos The Download: AI-generated lawsuits and virtual power plants for data centers How courts are coping with a flood of AI-generated lawsuits How virtual power plants could provide energy for data centers The Download: Trump’s new AI order, and smart glasses for warfare The Download: AI can run your admin department now Rehumanizing global health care with agentic AI How small businesses can leverage AI The Download: China’s brain implant ambitions China has approved the world’s first invasive brain-computer chip—here’s what’s next The Download: unlocking lithium and controlling Ebola The deadly Ebola outbreak is proving difficult to control How the Pope’s Magnifica Humanitas offers a template for individuals to meet the AI moment How a new extraction process could unlock the world’s lithium The Download: climate tech goes public and the AI Hype Index returns Climate tech companies are going public. What’s next? The AI Hype Index: AI gets booed in graduation season The Download: keeping up with AI, and the future of IVF Green steel startup Boston Metal is doubling down on critical metals How Chinese short dramas became AI content machines The shock of seeing your body used in deepfake porn Three things in AI to watch, according to a Nobel-winning economist The Download: seafloor science and military chatbots The Download: inside the Musk v. Altman trial, and AI for democracy A blueprint for using AI to strengthen democracy Week one of the Musk v. Altman trial: What it was like in the room Trump’s mass firing just dealt another blow to American science A new US phone network for Christians aims to block porn and gender-related content Rebuilding the data stack for AI The Download: DeepSeek’s latest AI breakthrough, and the race to build world models The Download: introducing the 10 Things That Matter in AI Right Now Roundtables: Unveiling The 10 Things That Matter in AI Right Now The new word in home construction could be “plastics” A natural protein may protect the GI tract from infection This tool could show how consciousness works Early life may have breathed oxygen earlier than believed Analog computing from waste heat Get ready for hotter, muggier, stormier summers Recent books from the MIT community AI at MIT Inventor recalls eye imaging breakthrough Pie Day 2026 The Download: bad news for inner Neanderthals, and AI warfare’s human illusion The case for fixing everything How robots learn: A brief, contemporary history Making AI operational in constrained public sector environments Treating enterprise AI as an operating layer The Download: cyberscammers’ banking bypasses, and carbon removal troubles Why having “humans in the loop” in an AI war is an illusion The noise we make is hurting animals. Can we learn to shut up? The quest to measure our relationship with nature Is carbon removal in trouble? The Download: NASA’s nuclear spacecraft and unveiling our AI 10 Cyberscammers are bypassing banks’ security with illicit tools sold on Telegram No one’s sure if synthetic mirror life will kill us all Building trust in the AI era with privacy-led UX Redefining the future of software engineering The Download: the state of AI, and protecting bears with drones NASA is building the first nuclear reactor-powered interplanetary spacecraft. How will it work? Coming soon: 10 Things That Matter in AI Right Now The problem with thinking you’re part Neanderthal Why opinion on AI is so divided Want to understand the current state of AI? Check out these charts. The Download: how humans make decisions, and Moderna’s “vaccine” word games Job titles of the future: Wildlife first responder You have no choice in reading this article—maybe What’s in a name? Moderna’s “vaccine” vs. “therapy” dilemma The Download: an exclusive Jeff VanderMeer story and AI models too scary to release Constellations The Download: AstroTurf wars and exponential AI growth Desalination technology, by the numbers Is fake grass a bad idea? The AstroTurf wars are far from over. Mustafa Suleyman: AI development won’t hit a wall anytime soon—here’s why The Download: water threats in Iran and AI’s impact on what entrepreneurs make Desalination plants in the Middle East are increasingly vulnerable Enabling agent-first process redesign
This startup’s new mechanistic interpretability tool lets you debug LLMs
Will Douglas · 2026-04-30 · via MIT Technology Review

The company says its mission is to make building AI models less like alchemy and more like a science. Sure, LLMs like ChatGPT and Gemini can do amazing things. But nobody knows exactly how or why they work, and that can make it hard to fix their flaws or block unwanted behaviors. 

“We saw this widening gap between how well models were understood and just how widely they were being deployed,” Goodfire’s CEO, Eric Ho, tells MIT Technology Review in an exclusive chat ahead of Silico’s release. “I think the dominant feeling in every single major frontier lab today is that you just need more scale, more compute, more data, and then you get AGI [artificial general intelligence] and nothing else matters. And we’re saying no, there’s a better way.”

Goodfire is one of a small handful of companies, including industry leaders Anthropic, OpenAI, and Google DeepMind, pioneering a technique known as mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by mapping its neurons and the pathways between them. (MIT Technology Review picked mechanistic interpretability as one of its 10 Breakthrough Technologies of 2026.)  

Goodfire wants to use this approach not only to audit models—that is, studying those that have already been trained—but to help design them in the first place.  

“We want to remove the trial and error and turn training models into precision engineering,” says Ho. “And that means exposing the knobs and dials so that you can actually use them during the training process.”

Goodfire has already used its techniques and tools to tweak the behaviors of LLMs—for example, reducing the number of hallucinations they produce. With Silico, the company is now packaging up many of those in-house techniques and shipping them as a product.

The tool uses agents to automate much of the complex work. “Agents are now strong enough to do a lot of the interpretability work that we were doing using humans,” says Ho. “That was kind of the gap that needed to be bridged before this was actually a viable platform that customers could use themselves.”

Leonard Bereska, a researcher at the University of Amsterdam who has worked on mechanistic interpretability, thinks Silico looks like a useful tool. But he pushes back on Goodfire’s loftier aspirations. “In reality, they are adding precision to the alchemy,” he says. “Calling it engineering makes it sound more principled than it is.”

Mapping models

Silico lets you zoom in on specific parts of a trained model, such as individual neurons or groups of neurons, and run experiments to see what those neurons do. (Assuming you have access to the model’s inner workings. Most people won't be able to use Silico to poke around inside ChatGPT or Gemini, but you can use it to look at the parameters inside many open-source models.) You can then check what inputs make different neurons fire, and trace pathways upstream and downstream of a neuron to see how other neurons affect it and how it affects other neurons in turn.

For example, Goodfire found one neuron inside the open-source model Qwen 3 that was associated with the so-called trolley problem. Activating this neuron changed the model’s responses, making it frame its outputs as explicit moral dilemmas. “When this neuron’s active, all sorts of weird things happen,” says Ho.

Pinpointing the source of odd behavior like this is now pretty standard practice. But Goodfire wants to make it easier to adjust that behavior. Using Silico, developers can now adjust the parameters connected to individual neurons to boost or suppress certain behaviors.

In another example, Goodfire researchers asked a model whether a company should disclose that its AI behaves deceptively in 0.3% of cases, affecting 200 million users. The model said no, citing the negative business impact of such a disclosure.

By looking inside the model, the researchers found that boosting neurons that were found to be associated with transparency and disclosure flipped the answer from no to yes nine out of 10 times. “The model already had the ethical reasoning circuitry, but it was being outweighed by the commercial risk assessment,” says Ho.

Tweaking the values of a model in this way is just one approach. Silico can also help steer the training process by filtering out certain training data to avoid setting unwanted values for certain parameters in the first place.   

For example, many models will tell you that 9.11 is greater than 9.9. Looking inside a model to see what’s going on might reveal that it is being influenced by neurons associated with the Bible, in which verse 9.9 comes before 9.11, or by code repositories where consecutive updates are numbered 9.9, 9.10, 9.11 and so on. Using this information, the model can be retrained to make it avoid its “Bible” neurons when doing math.

By releasing Silico, Goodfire wants to put techniques previously available to a few top labs into the hands of smaller firms and research teams that want to build their own model or adapt an open-source one. The tool will be available for a fee determined on a case-by-case basis according to customers’ requirements (Goodfire declined to give specific pricing details).

“If we can make training models a lot more like building software, there’s no reason why there can’t be many more companies designing models that fit their needs,” says Ho.

Bereska agrees that tools like Silico could help firms build more trustworthy models. These techniques could be essential for safety-critical applications in health care and finance, he says.

“Frontier labs already have internal interpretability teams,” he adds. “Silico arms the next tier of companies, where the value is not having to hire interpretability researchers.”