惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

罗磊的独立博客
小众软件
小众软件
The Cloudflare Blog
博客园 - 【当耐特】
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
WordPress大学
WordPress大学
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
V
Visual Studio Blog
量子位
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
美团技术团队
S
SegmentFault 最新的问题
宝玉的分享
宝玉的分享
博客园 - 叶小钗
月光博客
月光博客
Apple Machine Learning Research
Apple Machine Learning Research
T
Tailwind CSS Blog
博客园 - 聂微东
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
J
Java Code Geeks
Y
Y Combinator Blog
D
Docker
Microsoft Azure Blog
Microsoft Azure Blog

Goodfire Research

Models know when they’re reward hacking — and we can catch them at scale - Goodfire Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation Steering Along Manifolds to Control Neural Networks Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Meandering on Manifolds: The Neural Geometry of Stories Over Time The Neural Geometry Series Can SAEs Capture Neural Geometry? A Geometric Calculator Inside a Neural Network Meandering on Manifolds: The Neural Geometry of Stories Over Time Interpreting Language Model Parameters Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts The Shape of Stories Inside Neural Networks Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Can SAEs Capture Neural Geometry? Steering Along Manifolds to Control Neural Networks A Geometric Calculator Inside a Neural Network The Neural Geometry Series The World Inside Neural Networks The World Inside Neural Networks Verbalized Eval Awareness Inflates Measured Safety Verbalized Eval Awareness Inflates Measured Safety Paper Summary: Interpreting Language Model Parameters Paper Summary: Interpreting Language Model Parameters
The Circuits Research Landscape: Results and Perspectives
Jack Lindsey, · 2025-11-29 · via Goodfire Research

Research

This page should redirect you to the post on Neuronpedia.

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train

June 11, 2026

Logits as a new monitor for evaluation awareness

June 4, 2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

June 1, 2026