惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
V2EX
C
Check Point Blog
博客园_首页
B
Blog
D
Docker
U
Unit 42
量子位
I
InfoQ
有赞技术团队
有赞技术团队
Martin Fowler
Martin Fowler
GbyAI
GbyAI
L
LangChain Blog
云风的 BLOG
云风的 BLOG
博客园 - Franky
美团技术团队
T
The Blog of Author Tim Ferriss
阮一峰的网络日志
阮一峰的网络日志
月光博客
月光博客
Vercel News
Vercel News
Recent Announcements
Recent Announcements
雷峰网
雷峰网
大猫的无限游戏
大猫的无限游戏
小众软件
小众软件
Google DeepMind News
Google DeepMind News

AI Alignment Forum

Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI — AI Alignment Forum CoT controllability evals seem very under-elicited — AI Alignment Forum An operationalization of opaque serial depth — AI Alignment Forum Proposal for tracking the effects of architecture on monitorability — AI Alignment Forum Astra can do a concerning amount with no chain of thought — AI Alignment Forum How good are slop-vestigators? — AI Alignment Forum A Conceptual Framework for Reasoning about Exploration Hacking — AI Alignment Forum Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting — AI Alignment Forum Training on probes: Research ideas — AI Alignment Forum Training on probes: What's going on The Alignment Journal: Organization, Personnel, and Scope — AI Alignment Forum Training a Misaligned Reward Seeker — AI Alignment Forum Value generalisation Theory of Change: putting it into practice — AI Alignment Forum Value generalisation theory of change: the theory behind the approach — AI Alignment Forum Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident — AI Alignment Forum Debate Training Reduces Reward Hacking in RLAIF — AI Alignment Forum Does DiffusionGemma do latent reasoning? — AI Alignment Forum AI swarms are starting to pose indirect takeover risk — AI Alignment Forum An anytime algorithm for mixing the computable measures — AI Alignment Forum Misaligned AIs could use killer robots to take over — AI Alignment Forum Four LLM loss functions → four flavors of LLM misalignment — AI Alignment Forum Why do models task game? — AI Alignment Forum User awareness in frontier models — AI Alignment Forum R-lens: Making J-lens More Faithful on Early Layers — AI Alignment Forum Returning to ARC — AI Alignment Forum Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — AI Alignment Forum Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — AI Alignment Forum AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026) — AI Alignment Forum The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026) — AI Alignment Forum OpenAI has already ended an internal pause — AI Alignment Forum
Shallow Beliefs: Midtraining does not inoculate against E...
Jozdien · 2026-09-16 · via AI Alignment Forum

x