惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News | PayPal Newsroom
IT之家
IT之家
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
大猫的无限游戏
大猫的无限游戏
GbyAI
GbyAI
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
LangChain Blog
S
SegmentFault 最新的问题
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Project Zero
Project Zero
P
Privacy & Cybersecurity Law Blog
V
Vulnerabilities – Threatpost
博客园 - 三生石上(FineUI控件)
Recorded Future
Recorded Future
The Hacker News
The Hacker News
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
宝玉的分享
宝玉的分享
aimingoo的专栏
aimingoo的专栏
T
Tor Project blog
T
The Exploit Database - CXSecurity.com
Schneier on Security
Schneier on Security
H
Help Net Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
M
MIT News - Artificial intelligence
W
WeLiveSecurity
P
Proofpoint News Feed
A
About on SuperTechFans
S
Securelist
I
InfoQ
G
Google Developers Blog
博客园 - 司徒正美
博客园 - 叶小钗
Latest news
Latest news
F
Fortinet All Blogs
G
GRAHAM CLULEY
腾讯CDC
Jina AI
Jina AI
S
Schneier on Security
I
Intezer
V
Visual Studio Blog
美团技术团队
V2EX - 技术
V2EX - 技术
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
Microsoft Security Blog
Microsoft Security Blog
Blog — PlanetScale
Blog — PlanetScale
P
Proofpoint News Feed
罗磊的独立博客
Y
Y Combinator Blog

The Decoder

Google files first joint lawsuit with FBI over Chinese AI scam network, OpenAI blocks PRC influence clusters The AI industry's platform trap is starting to look a lot like Microsoft's OpenAI buys Ona to push Codex toward long-running, autonomous coding tasks Jeff Bezos' AI startup Prometheus closes $12 billion round at a $41 billion valuation Free Deezer tool lets users on any streaming service check their playlists for AI music OpenAI vs. Anthropic: A price war over API tokens is brewing Dario Amodei's new essay reads like a Cold War playbook for the AI age Claude Fable 5: Anthropic admits "wrong tradeoff" after invisibly throttling rival AI researchers Google's new open model DiffusionGemma generates text from noise instead of word by word OpenAI's IPO slips as Altman tells staff to expect a public offering "within the next year" Anthropic study shows AI needs hours, not weeks, to build exploits from security patches OpenAI wants its biggest data center yet, and Nvidia would back the bill Claude Fable 5: The first Mythos model is powerful, expensive, and heavily filtered Germany's National Security Council greenights an AI Safety Institute modeled after the UK's AISI Google's NotebookLM now runs its own cloud computer with code execution and agent-based research Anthropic releases Claude Fable 5 and Mythos 5 with major gains in coding and science Google's Gemini 3.5 Live Translate delivers real-time voice translation across 70+ languages SpaceX wants to put data centers in orbit, and Musk says it's no big deal Landmark German ruling declares Google's AI Overviews are Google's own words and makes it liable for false answers Beijing's $295 billion AI buildout would require 80 percent domestic chips, locking out US suppliers Apple Intelligence gets a second shot with help from Google and Nvidia OpenAI now says "entirely automating everything is not the future we want" OpenAI says going public is "a complicated set of tradeoffs" and is unsure about the timing Microsoft Research's Lens proves detailed captions matter more than raw scale for training efficient image generators Intel gets a second life as Google and Nvidia explore it as a TSMC backup for AI chips Most companies are flying blind on AI spending Frontier Radar #3: How agentic AI is turning tokens into a business metric Instagram AI chatbot breach may have affected over to 20,000 accounts, Meta discloses Microsoft tightens rules for conflict zones after investigation into Israel's military use of Azure Moonshot AI targets a $30 billion valuation, more than six times its late-2025 worth Deepseek topped Ramp's trending software vendors in June 2026 as US companies chase cheaper AI OpenAI says "chat is dead" and plans to rebuild ChatGPT as a full-blown agent app Perplexity's "Search as Code" lets AI models write their own search pipelines instead of calling fixed APIs ChatGPT's new Lockdown Mode lets you disable web access and more to protect sensitive data from prompt injection Anthropic poaches OpenAI's second-ever chip engineer as both companies race toward IPOs Researchers pinpoint why larger language models pick up skills that small ones miss Sakana AI bets AI that improves itself can break the compute arms race of frontier labs Meta's Hatch AI agent could cost up to $200 a month and marks its first paid AI product Elon Musk's xAI reportedly trained its coding models on Claude outputs for months before getting cut off New open-source voice model listens nonstop and decides every 0.4 seconds whether to speak or stay silent SpaceX signs $920 million per month deal with Google for 110,000 Nvidia AI chips ahead of IPO OpenAI and the Trump administration are negotiating a government stake in the AI startup Qwen3.7-Plus is Alibaba's bid to turn multimodal AI into a full-blown autonomous agent Florida's lawsuit against OpenAI and CEO Altman treats ChatGPT as a defective product and public nuisance Satya Nadella publicly torches a VP's plan to make Microsoft's AI agent deliberately addictive Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data" Anthropic's Mythos model is reportedly powering NSA offensive cyber ops against China and Iran Anthropic says Claude now writes over 90% of its code and wants the world to have an AI pause button Cloudflare CEO says the web's future is "pay to crawl" as bots overtake human traffic ChatGPT now saves narrative dossiers about you sorted by work, hobbies, and travel preferences Bain study finds companies miss AI savings targets because humans keep getting in the way OpenAI CEO Sam Altman sees "proactive AI" as the next big phase after chatbots and agents AI can now coach amateur virologists, and top tech leaders want Congress to act on DNA security xAI updates Grok Imagine to 1.5 with image-to-video generation at 720p resolution Google Deepmind's Gemma 4 12B squeezes multimodal AI onto a laptop with just 16 GB of RAM Google lets sites opt out of AI search results, knowing most have nowhere else to go Ideogram 4.0 drops as an open-weight model with native 2K resolution and improved text rendering Trump's new executive order wants AI companies to voluntarily submit models for government safety reviews Perplexity announces hybrid AI system that decides what runs locally or in the cloud AI music startup Suno doubles its valuation to $5.4 billion while fighting major record labels in court Nous Research releases Hermes Desktop, an open-source AI agent for every platform Build 2026: Microsoft tops Google in image generation while playing catch-up on reasoning OpenAI expands Codex with role-specific plugins to build a general-purpose app for non-developers Anthropic scales Project Glasswing to 150 partners across 15 countries to hunt critical software flaws Hackers hijacked high-profile Instagram accounts by simply asking Meta's AI chatbot to change the email OpenAI turns ChatGPT into a career platform with job search and CV editor Warren Buffett's Berkshire Hathaway bets $10 billion on Alphabet's AI infrastructure buildout OpenAI models now available on Amazon Web Services Claude maker Anthropic files for IPO with the SEC Turing Award winner Richard Sutton says pure generative AI can't do real science MiniMax M3: Open-weight model with a million-token context challenges proprietary leaders Nvidia's Nemotron 3 Ultra becomes the smartest open US model, but China still leads Nvidia bets big on physical AI at GTC Taipei with a new world model, driving brain, and open humanoid robot Nvidia pitches RTX Spark as the chip that finally makes local AI agents practical on Windows devices OpenAI starts with infrastructure robots but aims for "everyone having a personal robot doing anything they need" Ask AI what goes with chicken and the answer depends on whether it learned from recipes or molecules Anthropic bans AI tools during job interviews to see how candidates actually think Anthropic study finds men use AI coding agents more than twice as often as women in social science research SoftBank plans 75 billion euro AI data center buildout in France AI search agents often confirm what they already know instead of actually researching the web Microsoft and Nvidia reportedly team up on AI PCs that run actual agents instead of Copilot Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study finds Terence Tao argues AI could bring division of labor to math for the first time in history Attackers abuse shared ChatGPT and Claude chats to spread malware OpenAI's Codex can now operate your Windows PC autonomously, hunting bugs and testing apps on its own Salesforce claims AI agents cut a 231-day migration to 13 days with fewer incidents Meta's leaked memo reveals AI pendant, supersensing glasses, and enterprise wearables strategy OpenAI gives GPT-5.5 Instant a readability upgrade while phasing out two older models Google fixes several bugs in Gemini usage limits that burned through quotas too fast One company reportedly spent $500 million on Claude in one month after failing to cap AI usage OpenAI is giving away its life sciences AI model to help governments prepare for the next pandemic Amazon kills internal AI leaderboard after employees gamed it with pointless tasks Claude company Anthropic nears a trillion-dollar valuation after raising $65 billion in Series H Anthropic ships Claude Opus 4.8 as a "modest but tangible improvement" that tops GPT-5.5 in most benchmarks Google Cloud responds to AI-accelerated cyberattacks with a platform that aims to close security gaps in minutes Google launches a tiny board that runs Gemma 3 locally Mistral rebrands LeChat as Vibe, betting its chatbot's future is as a full-blown work agent Meta One: Zuckerberg finally puts a price tag on all that AI spending Amazon builds its own AI production platform and greenlights three AI animated series for Prime Video ElevenLabs Music v2 promises opera-to-metal transitions without losing musical coherence
New review paper argues code is how AI agents think and act, not just what they produce
Jonathan Kemper · 2026-05-29 · via The Decoder

A new review paper from researchers at the University of Illinois Urbana-Champaign, Meta, and Stanford wants to change how we think about AI agents.

Their argument is that code is the foundation agents use to reason, act, and work together. So the real bottleneck for autonomous systems, they say, becomes the software layer wrapped around the model, which probably makes Gary Marcus very happy.

The authors call this layer the "harness," and it covers everything from tools and interfaces to sandboxed execution environments, memory, testing, permission boundaries, execution loops, and feedback channels. Without it, a language model is just stateless. With it, the model becomes a working agent that can grind through tasks over long stretches.

Overview graphic of the code-as-agent-harness taxonomy showing three levels - harness interface, harness mechanisms, and scaling - along with five application domains: code assistants, GUI/OS agents, scientific discovery, personalization, and embodied agents.
The paper's central overview shows how code acts as an executable, testable, and stateful layer between model and environment. | Image: Ning et al.

Why code is the right format

The authors see code as a running part of agent behavior, and they lay out several reasons why. Code is executable, so model outputs become operations you can actually check. It's traceable because intermediate calculations show up as structured traces the system can read and store. And it persists across steps because the running program logs task progress in a form the agent can pick back up later.

The paper splits long-running agent systems into three parts. There's the model's own capabilities, like reasoning and planning. Then there's the infrastructure the system provides.

And finally, the code the agent writes on the fly, everything from test scripts and throwaway helper tools to reusable skills and executable workflows. The authors say these self-generated artifacts haven't gotten nearly enough research attention.

Three layers organize the field

At the first level, code bridges the model and its environment. Methods like Program-of-Thoughts or Chain of Code offload actual computation to executable programs instead of just describing it in words. Other systems, like Code as Policies, turn natural language instructions straight into robot control code.

Diagram of the plan-execute-verify loop with four building blocks: static analysis, sandboxed execution, deterministic verification, and permissioned state transitions from read-only to full access.
Reliability comes from clearly regulated state transitions in a controlled loop around the model. | Image: Ning et al.

The second level covers what keeps an agent reliable across many steps. That means planning, memory, tool use, and a recurring cycle of plan, execute, and verify. The cycle replaces one-off troubleshooting with systematic checks. Plans spell out what the agent intends to change. Execution runs in sandboxed environments with defined permissions. A verification step then decides whether the result gets accepted, revised, or kicked to a human reviewer.

The third level is about multiple agents working together. Code collections, tests, and execution logs become a shared workspace where specialized roles like managers, planners, coders, reviewers, and testers split the work. Systems like ChatDev and MetaGPT put this into practice, and according to the researchers it's already shipping in real products. Claude Code can now farm out pull request reviews to a whole team of AI agents that scan for bugs, security flaws, and regressions in parallel without being able to approve changes themselves.

Diagram of multi-agent orchestration showing specialized roles - manager, planner, coder, reviewer, tester, executer, and verifier - with a shared code workspace and various collaboration topologies.
At the third level, specialized agents split the work through a shared code workspace and coordinate tests and execution protocols. | Image: Ning et al.

Production systems already follow this pattern

The authors point to commercial products as examples. Anthropic's Claude Code ties together the local terminal, dev environment, and browser into one workflow where the agent edits files, runs commands, and has to follow permission rules. OpenAI's Codex and GitHub Copilot's coding agents move similar workflows to managed cloud environments, bundling changes through traceable pull request outputs.

How much this layer matters became obvious by accident when Anthropic leaked roughly 500,000 lines of Claude Code's source code. Buried in there was a "dreaming" function for task consolidation and other tricks for steering models as coding agents. Anthropic later got more than 8,000 copies and forks yanked from GitHub through a copyright takedown.

Other AI labs are catching on. Deepseek plans to go head-to-head with Claude Code and Codex through its own product, Deepseek Code, and is building a dedicated "Harness" team in Beijing to handle everything beyond the model, from tool use to planning to storage. The team's core formula is that model plus harness equals AI agent.

These production systems are also turning into training data for the next round of models. Cursor's composer trains with continuous reinforcement learning on real usage traces. OpenAI's Codex-1, GPT-5-Codex, and GPT-5.1-Codex-Max are trained specifically on long, multi-step coding sessions that match the Codex workflow. The line between agent and environment is itself becoming a layer that learns.

Overview of five application domains for code as agent harness with examples, including code assistants like Claude, Codex, and OpenClaw, as well as GUI/OS agents, scientific discovery, personalization, and embodied robot agents.
The same pattern shows up across five domains, from coding assistants to GUI control and robotics. | Image: Ning et al

When the agent starts tweaking its own environment

Several research systems treat the harness itself as something to optimize. AutoHarness auto-generates code that filters out unauthorized actions, while Meta-Harness systematically hunts for better harness variants by using previous versions, their evaluations, and execution logs as a search space. Other approaches dig through telemetry data to revise individual components. Meta's hyperagents go further still, combining task resolution and self-modification in an editable program that optimizes the improvement loop itself.

But the authors flag several open problems holding the field back: more meaningful evaluations beyond raw success rates, checking the substance of results when tests alone don't cut it, harness self-improvement without regressions, shared state across multiple agents, human oversight, and extending to environments with image or sensor data like GUI agents and robots.

They're especially blunt about whether current test criteria are even good enough. Tests can be incomplete, and test programs for graphical interfaces can miss bad intermediate steps. Simulators paper over physical risks. A harness could breed false confidence precisely because it gives visible feedback, and the green checkmark doesn't mean the code is safe. The authors suggest every accepted action should come with docs that spell out which tests actually ran, which areas stayed untested, and which risks remain.

Reliability in autonomous coding agents doesn't come from better repair prompts but from tightly regulated state transitions in a controlled loop around the model, the researchers argue.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now