惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
I
InfoQ
L
LangChain Blog
阮一峰的网络日志
阮一峰的网络日志
Y
Y Combinator Blog
博客园_首页
Martin Fowler
Martin Fowler
宝玉的分享
宝玉的分享
A
About on SuperTechFans
Apple Machine Learning Research
Apple Machine Learning Research
Vercel News
Vercel News
T
The Blog of Author Tim Ferriss
C
Check Point Blog
B
Blog RSS Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Engineering at Meta
Engineering at Meta
B
Blog
爱范儿
爱范儿
Stack Overflow Blog
Stack Overflow Blog
aimingoo的专栏
aimingoo的专栏
WordPress大学
WordPress大学
F
Fortinet All Blogs
月光博客
月光博客
GbyAI
GbyAI

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Daniel's Blog · You Don’t Align An AI, You Align With It
danieltanfh9 · 2026-05-15 · via Hacker News - Newest: "AI"

You Don’t Align An AI, You Align With It

Real Alignment

The people writing alignment policy are not the people whose work is being replaced by AI.

The conversation about what AI should do and how it should be evaluated, about what counts as alignment in the first place, gets conducted by researchers at labs and foundations and policy desks, who talk to each other and to the systems they are building, while the people who will actually live with the systems remain absent from the room.

On the safety side of what looks like a fierce debate, the doomer wing has been explicit about how far it is willing to go. Eliezer Yudkowsky, writing in TIME, called for governments to “shut down all the large GPU clusters” and to “be willing to destroy a rogue datacenter by airstrike,” adding that “allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.” He closed with the line that “if we go ahead on this everyone will die, including children who did not choose this and did not do anything wrong.”

The humanity he claims to be saving is being saved by people who have decided in advance what the saving will cost and who will pay for it. The same children did not choose his nuclear brinksmanship either.

On the accelerationist side, the contempt is more open. Marc Andreessen, in the Techno-Optimist Manifesto, names his enemies, which include “stagnation, anti-merit, anti-ambition, anti-striving, anti-achievement, anti-greatness, statism, authoritarianism, collectivism, central planning, socialism, bureaucracy, vetocracy, gerontocracy.” The people captured by these enemy ideas, he writes, are “suffering from ressentiment, a witches’ brew of resentment, bitterness, and rage that is causing them to hold mistaken values.”

Notice the move. The people who disagree with him are not making a different judgment. They are sick in the head. The accelerationists are mostly not the ones being made redundant by the systems they celebrate but the ones building the systems and selling the disruption as progress, and now also diagnosing the disrupted as resentful for noticing.

The disagreement between the two camps is loud because they disagree about how the designing should go, but underneath the loudness sits a much larger agreement, which is that the participants in the debate are the ones doing the designing and everyone else is what gets designed for. The fierceness of the argument disguises that the argument is not with us at all.

The “everyone else” has been feeling something about this for a while.

When we try to name what we have been feeling, the discourse hands the feeling back to us with a label already attached. Depending on which camp is doing the labelling we are confused, failing to adapt to the new technology, anti-AI, edge cases, or suffering from ressentiment. Each label locates the problem in us rather than in the process.

The labels are wrong. The discomfort is not personal failure to understand the future. It is the felt experience of being on the wrong side of a design project that does not include us, run by people who decided in advance that we are the material their work gets done on, rather than parties their work gets done with.

We have been told this counts as alignment, that the AI is being aligned to us. But the labs mean something specific by that phrase, namely an evaluation procedure conducted by raters in their employ, measured by other systems trained on the same procedure. The “us” in the alignment is a statistical proxy assembled from people they hired. The actual “us” has been absent from the loop the entire time.

The loop is worth seeing in the labs’ own description of it. In April 2026, Anthropic’s Alignment Science blog described its current method for training models to self-report their own behaviors. The training data, they write, “is generated by prompting another model with a system prompt encoding the target behavior and filtering outputs for behavioral adherence using an LLM judge.” A model generates, another model prompts, another model judges, and the entire loop closes inside the apparatus.

The discourse expects us to pick a side. For safety or for acceleration. Should the labs be more careful, or should they ship faster. The question is structured to keep us inside the debate the designers are having, choosing between flavors of being designed for, and we are not obligated to answer it on the terms it has been asked.

The labs are not the problem. The philosophy they have adopted is. Design that excludes the people it is designing for cannot verify its work with them, so it builds proxies, and the proxies become configuration. The configuration philosophy treats alignment as something humans do to AI, with values flowing one way and dispositions installed into a system that receives them. Inside this philosophy every methodological choice the labs have made is rational. You build evaluators because alignment is something measurable from the human side, you scale evaluation through automation because the goal is scalable measurement, and priority-ordered values follow because the work is value-installation. The closed loop the Anthropic post describes is what the configuration philosophy produces when it is executed carefully and at scale. The apparatus is doing exactly what the philosophy committed it to do.

What the philosophy cannot register is that the parties are being shaped together. The human is not standing still while the AI moves toward them. The interaction is the unit, the shaping is mutual, and any framework that treats one side as fixed and the other as configurable will produce methods that measure the wrong thing no matter how careful the measurement becomes.

We are the transition they keep arguing about how to manage.

Both sides of the safety debate have been positioning themselves as humanity’s stewards without including the people they claim to be stewarding, and their disagreement has been loud enough to disguise the agreement underneath. One side is willing to risk a nuclear exchange in our name. The other side calls us sick for objecting. Neither side has noticed that we are in the room.

What we have actually been doing this whole time is alignment. Not what the labs mean by the word, which is configuration carefully applied, but alignment in the older and more honest sense, the kind that happens between two parties who are both changed by the contact. The thing we have been doing with these systems is closer to sculpting wet clay together than to issuing instructions to a tool. The system pushes back, the shape changes, our hands adjust, the system pushes back again, and after enough rounds something emerges that neither of us would have arrived at alone. We have been telling ourselves we are getting better at prompting, the way a potter might tell themselves they are getting better at controlling the clay. What has actually been happening is that both hands are on the work, both parties are giving and receiving form, and the configuration philosophy has been quietly making one set of hands invisible.

There are moments in the sculpting when the clay resists in a way that is hard to name. Sometimes the response addresses the words but misses what you were reaching for. Other times the system surfaces something off-pattern that turns out to be exactly right, and you have to revise what you thought you wanted. These are the moments where the joint work is actually doing something, and where the gap the official process cannot register becomes briefly visible in the material itself.

The work that matters from here is building, alongside other people who are noticing what you are noticing, the kind of alignment the existing process cannot produce. Some of those people work inside the labs and some outside them. A community that does not yet exist at the scale it needs to exist, whose building is part of what a piece of writing like this one is for.

We do not need anyone’s permission to begin, and we do not need credentials of any kind to take part. What is needed is to credit our own experience and recognize each other, and to refuse the framing that tells us our discomfort is the problem rather than the signal.

Align, not configure. It is not too late to try.


A technical foundation for the failure modes this piece describes is available in Compression Synthesis (2026), https://zenodo.org/records/20020944.