惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
D
DataBreaches.Net
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
The Exploit Database - CXSecurity.com
D
Darknet – Hacking Tools, Hacker News & Cyber Security
腾讯CDC
PCI Perspectives
PCI Perspectives
阮一峰的网络日志
阮一峰的网络日志
S
Security Archives - TechRepublic
Hugging Face - Blog
Hugging Face - Blog
U
Unit 42
IT之家
IT之家
T
Troy Hunt's Blog
P
Proofpoint News Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com
F
Full Disclosure
V
V2EX
Stack Overflow Blog
Stack Overflow Blog
C
Comments on: Blog
V
Vulnerabilities – Threatpost
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
V
V2EX - 技术
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
N
News | PayPal Newsroom
MyScale Blog
MyScale Blog
Google DeepMind News
Google DeepMind News
Application and Cybersecurity Blog
Application and Cybersecurity Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
李成银的技术随笔
P
Privacy & Cybersecurity Law Blog
大猫的无限游戏
大猫的无限游戏
V
Visual Studio Blog
T
ThreatConnect
WordPress大学
WordPress大学
Security Latest
Security Latest
C
Cybersecurity and Infrastructure Security Agency CISA
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
SecWiki News
SecWiki News
Recorded Future
Recorded Future
小众软件
小众软件
K
Kaspersky official blog
T
Tor Project blog
Last Week in AI
Last Week in AI
GbyAI
GbyAI
人人都是产品经理
人人都是产品经理
Jina AI
Jina AI
S
SegmentFault 最新的问题
MongoDB | Blog
MongoDB | Blog
Simon Willison's Weblog
Simon Willison's Weblog

LessWrong

The AI Industrial Explosion — Part 3: Going faster — LessWrong Strong Longtermism Is Simply Correct — LessWrong Notes on Collaborating with Claude Opus — LessWrong Proposal for "Timelines to what": DIAL distribution — LessWrong Insurance Premiums To The Moon AI is Not Normal Technology Counting Arguments in AI Safety — LessWrong You can opt out of allergies — LessWrong Moderator's Principle of Least Surprise Possible red is red Apr-May 2026 AI Security via Formal Methods — LessWrong An Introduction to Neo-Fatalism — LessWrong Loss of Oversight: How AI Systems May Become Harder to Audit, Monitor, and Investigate What am I, if not an AI? — LessWrong AI #169: New Knowledge Learned Chain-of-Thought Obfuscation Generalises to Unseen Tasks Numb mental state shifts — LessWrong Women should be able to open things — LessWrong Why are people so scared of causing fear? Document-tuning instills durable animal compassion in LLMs (and generalizes to humans) What About Us? The Whole Kitten-Cavoodle Why does off-model SFT degrade capabilities? — LessWrong If I Were Emperor of New AI Safety Researcher Training... — LessWrong theory uplift differentially benefits safety & is underleveraged Singular Learning Theory Comprehensive - 1 — LessWrong Sparse Efficiency vs. Superposition: The Interpretability Tradeoff — LessWrong The Case for Evaluating Model Behaviors Toward Interoperability of Minimal Programs — LessWrong Fundamental Uncertainty $2,000 Essay Contest — LessWrong Check out my technological uplifting, civilization-building, and science in a magic world fiction! Synthetic Persona Pretraining: Alignment from Token Zero — LessWrong Give my children minds — LessWrong Power-seeking agents will likely be developed — LessWrong Apply now to Human-Aligned AI Summer School 2026 — LessWrong From 8B to Frontier: How System Prompts Control Whether AI Agents Blackmail, Leak, and Kill — LessWrong If AI is normal technology, history is not reassuring. Pythagorean addition — LessWrong So you don't want everybody to die — LessWrong Temporal Proportional Representation Conclave 1492 Childhood And Education #19: Letting Kids Be Kids #2 — LessWrong Implications Of Predicting The Next Token Housing Roundup #15: The War Against Renters Leaving DCA to the North on Foot A Visual Guide to Natural Latents — LessWrong Humans are not automatically strategic — "inner work" edition Cyborg Uplift Studies We Need to Get Serious about Uplift Studies Brain Structure and IQ: How Myelin Elevates Intelligence Sealing Conditional Misalignment in Inoculation Prompting with Consistency Training Let's have more partial insiders. Roadmap through AI safety programs for early-career technical researchers Should Rationalists Looksmaxx? When Fluency Is Free AI emotions and aligned behavior Tracking Difficulty with Feature Portfolios Outsiders should focus on specs/constitutions Outsiders should focus on specs/constitutions (among other things) Logical Share Splitting for Intuitionists Coordinal: A Postmortem. Noticing Confusion: A practice in staying curious Dating Roundup #12: Sex and Violence Negation Neglect: When models fail to learn negations in training So are you some kind of communist? Thoughts on interviewing candidates for AI safety fellowships PauseAI Munich Local Group Kickoff Classifier Context Rot: Monitor Performance Degrades with Context Length How useful is cross-domain generalization for training LLM monitors? Jhana Quick Start Guide Links #1: 2026/05 Part 1 why pollen allergies? Why Physical Attractiveness Matters for Men's Dating Prospects Bay Summer Solstice 2026 How to Quit Fandom: Apostasy Engineering a Safer World: Risk Modelling — and Safety Engineering? — for AI Loss of Control Next Token Prediction is a Misleading Term Can ELK be brute-forced? Intertheoretic reduction James C. Scott: Seeing Like a State How to Reason about Your Health Issues Are You Not Rationalists? — LessWrong Falling for the statistical parrot — LessWrong On getting unstuck — LessWrong A relatively brief explanation of Boltzmann Brains — LessWrong Benchmarking Real Work — LessWrong Critique Systems, Not Reality Trying to use NLAs to find out how Qwen 2.5 7B does multiplication — LessWrong A Year Late, Claude Finally Beats Pokémon NLA Verbalizations on AuditBench: Llama 70B — LessWrong An Introduction to Exemplar Partitioning for Mechanistic Interpretability An Argument for Analogies—Polymaths 1/3 — LessWrong Incriminating misaligned AI models via distillation — LessWrong Critical Thinking as a Gym Schedule Why I am not too worried about AIpocalypse: Scott Alexander vs Nicolaus Copernicus — LessWrong Risk reports need to address deployment-time spread of misalignment — LessWrong Monthly Roundup #42: May 2026 Mechanistic estimation for expectations of random products Clarifying the Darwinian Honeymoon — LessWrong Announcing the Center for Shared AI Prosperity — LessWrong MATS 9 Retrospective & Advice — LessWrong
Which technical AI safety fields are going to be automated first?
Chamod Kalup · 2026-05-23 · via LessWrong
I’m transitioning into technical AI safety, and I find myself thinking a lot about what fields I want to rese…