惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
Blog — PlanetScale
Blog — PlanetScale
腾讯CDC
aimingoo的专栏
aimingoo的专栏
Microsoft Azure Blog
Microsoft Azure Blog
A
About on SuperTechFans
J
Java Code Geeks
G
Google Developers Blog
N
Netflix TechBlog - Medium
Vercel News
Vercel News
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
I
InfoQ
Stack Overflow Blog
Stack Overflow Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
T
The Blog of Author Tim Ferriss
罗磊的独立博客
GbyAI
GbyAI
小众软件
小众软件
大猫的无限游戏
大猫的无限游戏
WordPress大学
WordPress大学
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More

Goodfire Research

Models know when they’re reward hacking — and we can catch them at scale - Goodfire Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation Steering Along Manifolds to Control Neural Networks Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Meandering on Manifolds: The Neural Geometry of Stories Over Time The Neural Geometry Series Can SAEs Capture Neural Geometry? A Geometric Calculator Inside a Neural Network Meandering on Manifolds: The Neural Geometry of Stories Over Time Interpreting Language Model Parameters Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts The Shape of Stories Inside Neural Networks Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Can SAEs Capture Neural Geometry? Steering Along Manifolds to Control Neural Networks A Geometric Calculator Inside a Neural Network The Neural Geometry Series The World Inside Neural Networks The World Inside Neural Networks Verbalized Eval Awareness Inflates Measured Safety Verbalized Eval Awareness Inflates Measured Safety Paper Summary: Interpreting Language Model Parameters Paper Summary: Interpreting Language Model Parameters
Understanding Sparse Autoencoder Scaling in the Presence ...
Eric J. Michaud, Liv Gorton, Thomas McGrath, · 2025-12-05 · via Goodfire Research

Research

This page should redirect you to the arXiv paper.

Research

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train

June 11, 2026

Logits as a new monitor for evaluation awareness

June 4, 2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

June 1, 2026