惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
小众软件
小众软件
博客园 - 叶小钗
宝玉的分享
宝玉的分享
博客园_首页
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
SegmentFault 最新的问题
B
Blog RSS Feed
Engineering at Meta
Engineering at Meta
N
Netflix TechBlog - Medium
Google DeepMind News
Google DeepMind News
U
Unit 42
F
Fortinet All Blogs
IT之家
IT之家
Y
Y Combinator Blog
Martin Fowler
Martin Fowler
T
The Blog of Author Tim Ferriss
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
The GitHub Blog
The GitHub Blog
Stack Overflow Blog
Stack Overflow Blog
Blog — PlanetScale
Blog — PlanetScale
酷 壳 – CoolShell
酷 壳 – CoolShell

Goodfire Research

Models know when they’re reward hacking — and we can catch them at scale - Goodfire Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation Steering Along Manifolds to Control Neural Networks Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Meandering on Manifolds: The Neural Geometry of Stories Over Time The Neural Geometry Series Can SAEs Capture Neural Geometry? A Geometric Calculator Inside a Neural Network Meandering on Manifolds: The Neural Geometry of Stories Over Time Interpreting Language Model Parameters Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts Logits as a new monitor for evaluation awareness Predicting Rare LLM Failures with 30× Fewer Rollouts The Shape of Stories Inside Neural Networks Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Can SAEs Capture Neural Geometry? Steering Along Manifolds to Control Neural Networks A Geometric Calculator Inside a Neural Network The Neural Geometry Series The World Inside Neural Networks The World Inside Neural Networks Verbalized Eval Awareness Inflates Measured Safety Verbalized Eval Awareness Inflates Measured Safety Paper Summary: Interpreting Language Model Parameters Paper Summary: Interpreting Language Model Parameters
Mixing Mechanisms: How Language Models Retrieve Bound Ent...
Fundamental ResearchLink post · 2025-12-05 · via Goodfire Research

Research

This page should redirect you to the arXiv paper.

Research

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train

June 11, 2026

Logits as a new monitor for evaluation awareness

June 4, 2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

June 1, 2026