惯性聚合
高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文
在惯性聚合中打开
即将跳转到惯性聚合
3
在聚合应用中查看完整内容和互动
立即跳转
取消
推荐订阅源
博
博客园 - 叶小钗
J
Java Code Geeks
奇客Solidot–传递最新科技情报
阮一峰的网络日志
爱范儿
量
量子位
N
Netflix TechBlog - Medium
博
博客园 - 聂微东
博
博客园 - Franky
aimingoo的专栏
The Cloudflare Blog
T
The Blog of Author Tim Ferriss
MyScale Blog
Google DeepMind News
小众软件
博
博客园 - 三生石上(FineUI控件)
C
Check Point Blog
钛媒体:引领未来商业与生活新知
B
Blog
Engineering at Meta
Microsoft Azure Blog
博
博客园_首页
H
Hackread – Cybersecurity News, Data Breaches, AI and More
腾
腾讯CDC
AI Alignment Forum
Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking — AI Alignment Forum
Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI — AI Alignment Forum
CoT controllability evals seem very under-elicited — AI Alignment Forum
An operationalization of opaque serial depth — AI Alignment Forum
Proposal for tracking the effects of architecture on monitorability — AI Alignment Forum
Astra can do a concerning amount with no chain of thought — AI Alignment Forum
How good are slop-vestigators? — AI Alignment Forum
A Conceptual Framework for Reasoning about Exploration Hacking — AI Alignment Forum
Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting — AI Alignment Forum
Training on probes: Research ideas — AI Alignment Forum
Training on probes: What's going on
The Alignment Journal: Organization, Personnel, and Scope — AI Alignment Forum
Training a Misaligned Reward Seeker — AI Alignment Forum
Value generalisation Theory of Change: putting it into practice — AI Alignment Forum
Value generalisation theory of change: the theory behind the approach — AI Alignment Forum
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident — AI Alignment Forum
Debate Training Reduces Reward Hacking in RLAIF — AI Alignment Forum
Does DiffusionGemma do latent reasoning? — AI Alignment Forum
AI swarms are starting to pose indirect takeover risk — AI Alignment Forum
An anytime algorithm for mixing the computable measures — AI Alignment Forum
Misaligned AIs could use killer robots to take over — AI Alignment Forum
Four LLM loss functions → four flavors of LLM misalignment — AI Alignment Forum
Why do models task game? — AI Alignment Forum
User awareness in frontier models — AI Alignment Forum
R-lens: Making J-lens More Faithful on Early Layers — AI Alignment Forum
Returning to ARC — AI Alignment Forum
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — AI Alignment Forum
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — AI Alignment Forum
AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026) — AI Alignment Forum
The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026) — AI Alignment Forum
Synthetic document finetuning for instilling positive tra...
CallumMcDoug
·
2026-06-16
·
via
AI Alignment Forum
x
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。
原文来自
— 版权归原作者所有。