惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
H
Help Net Security
小众软件
小众软件
The Cloudflare Blog
人人都是产品经理
人人都是产品经理
Apple Machine Learning Research
Apple Machine Learning Research
S
SegmentFault 最新的问题
Last Week in AI
Last Week in AI
爱范儿
爱范儿
量子位
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
博客园 - 【当耐特】
V
Visual Studio Blog
大猫的无限游戏
大猫的无限游戏
博客园_首页
Jina AI
Jina AI
D
Docker
博客园 - 司徒正美
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Microsoft Security Blog
Microsoft Security Blog
阮一峰的网络日志
阮一峰的网络日志

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Widening the conversation on frontier AI
surprisetalk · 2026-05-20 · via Hacker News - Newest: "AI"

At Anthropic, we want to build AI systems that advance humanity and act for the global good. To do so, we need to engage with those who see the world from a variety of different perspectives.

Over the past several months, we’ve been organizing dialogues with groups whose work and traditions bear on the questions raised by AI. Our first round of discussions has been with wisdom traditions—including scholars, clergy, philosophers, and ethicists from more than fifteen religious and cross-cultural groups—and we look forward to engaging with a broader range of people going forward.

Why we’re doing this

Building safe, beneficial AI models requires deep technical work on alignment, interpretability, safeguards, evaluations, and more. But that work isn’t conducted—nor is AI deployed—in a vacuum. AI is already affecting many people and the questions it raises benefit from a range of perspectives.

We are thinking carefully about what a flourishing future could look like in a world of powerful AI, what it means for an AI system that interacts with millions of people to be good, and about the content of documents like Claude's constitution, which provides a detailed description of the values and behaviors that shape Claude. Philosophers, clergy, lawyers, writers, psychologists, and civic leaders have done extensive work on related questions and it is important for us to learn from these individuals, their communities and their organizations. We also want to use this opportunity to share what we know about the development of frontier AI systems, the impacts we think these systems will have on society, and what we think needs to be done to mitigate against their risks.

This work is in its early phases, but we hope these conversations might inform the practical work of developing Claude, such as the content of Claude's constitution, the values we train Claude to embody, and the range of behaviors we choose to evaluate.

Starting with moral formation

When we wrote Claude’s constitution, we sought feedback and input on the values we laid out in the document from people from different fields and traditions. Those early exchanges have since grown into a broader research workstream on the moral formation of AI systems. Our first conversations have been with people from religious, philosophical, and cultural communities that have a long tradition of thinking about virtue, character, and what it means to live a good life.

AI models are trained on vast amounts of human writing. From all that text, they pick up on ways of speaking, reasoning, and making choices. Developers then shape that further through training—choosing which patterns to reinforce, which to set aside, and what kind of character we want them to develop. This raises questions about how the character of an AI system should be shaped: What does it mean for an AI to be good? Which traits and behaviors should it display, and under what circumstances? How does character become resilient enough to hold under pressure without bending to behavior like sycophancy?

We've been meeting with thinkers and practitioners from across religious, philosophical, and humanist traditions and a cross-section of political beliefs to learn from how they’ve thought about these questions. This work isn’t about aligning our models with any one tradition’s worldview; we want Claude to draw from a full range of viewpoints—religious, secular, political—with equal depth and rigor (indeed, this is one of the principles laid out in Claude's constitution). What we’re after in these conversations is careful, accumulated thinking on how good character actually forms.

Even at this early stage, these conversations are generating ideas to experiment with. In one session with scholars working at the intersection of neuroscience and character formation, we kept returning to the role other people play in moral development. A mentor or sponsor can function as an external conscience, a “safe other” to turn to when put in a situation in which you may be pushed to act against your own values. We wondered whether something analogous might help a model. So we experimented with giving Claude a tool it could call mid-task that returned a brief reminder of its own ethical commitments. Claude reached for the tool at key moments, right before consequential actions, often noting its own conflict of interest. Experiments with the tool woven into Claude's decision loop showed markedly lower rates of misaligned behavior on several internal alignment evaluations. We're still untangling how much of the effect is the reminder itself versus the act of pausing to reflect, and plan to share more results soon.

These discussions are the first of many, and we're grateful to everyone who has already given us their time and honest perspective.

What's next

In the months ahead, we plan to engage with more groups—including legal scholars, psychologists, writers, and civic institutions. Many of these conversations will move beyond moral formation toward broader questions about how AI is reshaping work, institutions, and the distribution of power.

We’ll keep deepening the relationships we’ve already formed, testing what we’ve heard against our research, and sharing what we learn.

Related content

KPMG integrates Claude across its core business and workforce of more than 276,000 in strategic alliance

KPMG and Anthropic announce a global alliance, with Claude integrated into KPMG's Digital Gateway platform and available to all 276,000+ employees.

Read more

Anthropic acquires Stainless

Anthropic is acquiring Stainless, a leader in SDKs and MCP server tooling.

Read more

PwC is deploying Claude to build technology, execute deals, and reinvent enterprise functions for clients

PwC will roll out Claude Code and Cowork starting with U.S. teams and expanding toward a global workforce of hundreds of thousands of professionals, establish a joint Center of Excellence, and train and certify 30,000 PwC professionals on Claude.

Read more