惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Jina AI
Jina AI
博客园 - 司徒正美
大猫的无限游戏
大猫的无限游戏
博客园 - 三生石上(FineUI控件)
J
Java Code Geeks
博客园 - 聂微东
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
美团技术团队
腾讯CDC
博客园 - Franky
MyScale Blog
MyScale Blog
人人都是产品经理
人人都是产品经理
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
月光博客
月光博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
aimingoo的专栏
aimingoo的专栏
博客园_首页
V
V2EX
Martin Fowler
Martin Fowler
T
The Blog of Author Tim Ferriss

AI Alignment Forum

Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking — AI Alignment Forum Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI — AI Alignment Forum CoT controllability evals seem very under-elicited — AI Alignment Forum An operationalization of opaque serial depth — AI Alignment Forum Proposal for tracking the effects of architecture on monitorability — AI Alignment Forum Astra can do a concerning amount with no chain of thought — AI Alignment Forum How good are slop-vestigators? — AI Alignment Forum A Conceptual Framework for Reasoning about Exploration Hacking — AI Alignment Forum Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting — AI Alignment Forum Training on probes: Research ideas — AI Alignment Forum Training on probes: What's going on The Alignment Journal: Organization, Personnel, and Scope — AI Alignment Forum Training a Misaligned Reward Seeker — AI Alignment Forum Value generalisation Theory of Change: putting it into practice — AI Alignment Forum Value generalisation theory of change: the theory behind the approach — AI Alignment Forum Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident — AI Alignment Forum Debate Training Reduces Reward Hacking in RLAIF — AI Alignment Forum Does DiffusionGemma do latent reasoning? — AI Alignment Forum AI swarms are starting to pose indirect takeover risk — AI Alignment Forum An anytime algorithm for mixing the computable measures — AI Alignment Forum Misaligned AIs could use killer robots to take over — AI Alignment Forum Four LLM loss functions → four flavors of LLM misalignment — AI Alignment Forum Why do models task game? — AI Alignment Forum User awareness in frontier models — AI Alignment Forum R-lens: Making J-lens More Faithful on Early Layers — AI Alignment Forum Returning to ARC — AI Alignment Forum Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — AI Alignment Forum Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — AI Alignment Forum AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026) — AI Alignment Forum The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026) — AI Alignment Forum
Risk reports need to address deployment-time spread of mi...
Alex Mallen · 2026-05-16 · via AI Alignment Forum
Risk reports commonly use pre-deployment alignment assessments to measure misalignment risk from an internall…