惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
V
Visual Studio Blog
博客园 - 司徒正美
TaoSecurity Blog
TaoSecurity Blog
博客园 - 聂微东
IT之家
IT之家
博客园_首页
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - Franky
雷峰网
雷峰网
罗磊的独立博客
S
Schneier on Security
C
Cybersecurity and Infrastructure Security Agency CISA
The Cloudflare Blog
T
Tailwind CSS Blog
B
Blog RSS Feed
H
Help Net Security
T
The Blog of Author Tim Ferriss
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
Threatpost
C
CERT Recently Published Vulnerability Notes
博客园 - 三生石上(FineUI控件)
P
Palo Alto Networks Blog
I
Intezer
G
GRAHAM CLULEY
Engineering at Meta
Engineering at Meta
S
Securelist
J
Java Code Geeks
V
V2EX
Y
Y Combinator Blog
Simon Willison's Weblog
Simon Willison's Weblog
L
LINUX DO - 热门话题
云风的 BLOG
云风的 BLOG
Spread Privacy
Spread Privacy
MongoDB | Blog
MongoDB | Blog
P
Privacy International News Feed
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
B
Blog
Forbes - Security
Forbes - Security
Google Online Security Blog
Google Online Security Blog
Help Net Security
Help Net Security
S
SegmentFault 最新的问题
N
Netflix TechBlog - Medium
Webroot Blog
Webroot Blog
Microsoft Security Blog
Microsoft Security Blog
SecWiki News
SecWiki News
Scott Helme
Scott Helme
aimingoo的专栏
aimingoo的专栏
N
News and Events Feed by Topic

Latent.Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI [AINews] The Field Guide to Fable AIEWF Daily Dispatch: The great loops debate and the state of AI engineering Vercel's Andrew Qu on why agents are a new kind of software The website of the future may assemble itself for every visitor Skill engineering and the case against one-shot AI design [AINews] not much happened today AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency Autoresearch: The feedback loop behind self-improving agents How Cursor deploys AI inside the enterprise 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI Warp CEO Zach Lloyd on why software factories are the next phase of coding AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers [AINews] Sonnet 5 today, and Fable 5 tomorrow Forward Deployed Engineers and the future of software engineering Ahmad Osman on why local AI is catching up [AINews] not much happened today [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners [AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025. [AINews] It's Meta-Harness Summer Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks [AINews] Claude Tag: Multiplayer, Proactive, Persistent Agents in Slack [AINews] SpaceX is already a $28B/yr Neocloud Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan How to AIE Good [AINews] not much happened today [AINews] GLM-5.2 is the real deal; Z.ai forecasts Open Fable by EOY The Professor of Outputmaxxing — Anjney Midha, AMP [AINews] Midjourney Medical: scan your organs like you step on a scale 🔬 The Self-Driving Lab — Joseph Krause, Radical AI [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding [AINews] Satya on Loopcraft: Building Frontier Ecosystems [AINews] Fable and Mythos officially too dangerous to release [AINews] Loopcraft: The Art of Stacking Loops [AINews] Loopcraft: The Art of Stacking Loops [AINews] Open Models, Model Labs vs Agent Labs, and What's Untrainable — Sarah Guo [AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms [AINews] FrontierCode: Benchmarking for Code Quality over Slop [AINews] not much happened today How to Stop Shipping Low-Quality RL Environments (with Examples) [AINews] not much happened today Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs [AINews] Reve 2 and Ideogram 4: Layouts in Imagegen ⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build [AINews] Microsoft Build: MAI-Thinking-1 and MAI Family models GitHub's plan for Agents — Kyle Daigle, GitHub [AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark Why Video Agent models are next — Ethan He, xAI Grok Imagine [AINews] Founders and Forward Deployed Engineers [AINews] Anthropic raises $965B Series H, releases Opus 4.8 and Dynamic Workflows/ultracode The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray [AINews] Cognition raises $1B in $26B Series D 🔬 ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub [AINews] New AI Infra decacorns: Fireworks, Baseten (with OpenRouter on the way) [AINews] All Model Labs are now Agent Labs [AINews] New AI Infra unicorns: Exa, Modal, TurboPuffer Giving Agents Computers — Ivan Burazin, Daytona [AINews] OpenAI GPT-next disproves 80 year old Erdős planar unit distance problem for under $1000 Railway: The Agent-Native Cloud — Jake Cooper [AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0 [AINews] How to land a job at a frontier lab (on Pretraining) The Autonomous Drone Tech Stack & Economics of Drones — Yaroslav Azhnyuk, The Fourth Law & Guest Host Noah Smith, Noahpinion [AINews] Cerebras' $60B IPO: Slowly, then All at Once [AINews] Everything is Conductor AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes — Janie Lee & Chai Asawa, Abridge [AINews] Codex Rises, Claude Meters Programmatic Usage [AINews] The End of Finetuning [AINews] Thinking Machines' Native Interaction Models - TML-Interaction-Small 276B-A12B - advances SOTA Realtime Voice and kills standard VAD
🔬Scaling Past Informal AI - Carina Hong, Axiom Math
RJ Honicky · 2026-06-04 · via Latent.Space

In 2025, seven-month-old startup Axiom solved all 12 of the problems Putnam exam (scoring 8/12 in the time limit) a prestigious undergraduate math exam. The 12/12 score is better than the top undergraduates (110/120) and the closest AI system that reported a result (DeepSeek 103/120), although it is unclear what the people and other systems would have scored with more time. Nonetheless, the Putnam exam is legendary for its difficulty, with the median score typically being 0 or 1 points. Taken by itself, this seems like a minor feather in the cap of AI; one of a long series of accomplishments by AI systems in elite competitions with humans, starting with Deep Blue beating Kasparov.

Fast forward to mid-2026, and Claude Code and Codex are setting the world on fire. In 2024 Anthropic’s bet on code and enterprise looked like a more pragmatic niche play vs. OpenAI’s better models and massive consume scale. Today, Amodei’s all in bet on acceleration via code (images and video be damned) seems prescient.

Despite Anthropic’s growing momentum, however, Axiom CEO Carina Hong sees coding ability as a necessary but not sufficient milestone on the path to AGI. Code arguably pushes the jagged frontier to the point of super intelligence in some domains outside of coding, but there are surprising gaps (link) that Carina believes will bottleneck AI progress. (Stats on math benchmarks).

“Verified AI” sounds like eating broccoli1 and paying taxes, but to Axiom it means something very different. “Verification to me is about scaling brilliance, compounding brilliance,” Carina told us.

It actually took a while for me to understand what she means by this (sounded like marketing-speak until it clicked). Carina brings up the legendary mathematician Srinivasa Ramanujan (“The Man who knew Infinity”) to illustrate this point. When G.H. Hardy finally persuaded Ramanujan to formally prove theorems instead of relying on his (formidable) intuition, it reportedly improved his own capabilities. This is presumably because formally proving things forced Ramanujan to articulate the details in a way that open up new lines of thinking, etc. This is how you “compound” in math — building on solid rather than shaky foundations… also known as Axioms.

But formally proving things also allowed others to benefit from his intuition: the proofs are way of communicating an intuition and persuading others that the intuition is correct. This is scaling (more people use the result) and compounding (people can learn from and build on his work).

This is the core insight that lets us understand the approach Axiom is taking.

There are two ways that Verified AI shows up: in training and in inference.

But a quick detour: to a first approximation, “Formal Verification” means using type checkers (like for TypeScript, C++ or Rust, but more capable) to verify mathematical proofs that are meticulously specified using a language like Lean2. It takes a lot of work to translate an “informal” proof (albeit one that most people would not remotely call “informal”) in to a Lean proof3. Axiom themselves have open sourced groundbreaking work with AXLE - their toolkit of interactive Lean applications for exploring, validating, and manipulating mathematical proofs.

You can imagine how this would be (very) useful during Reinforcement Learning: instead of relying on best guesses based on statistics (GRPO, RLHF, etc.), you can just verify the proof is correct using a Lean verifier. This is obviously a much stronger reward signal, akin to compiling code and testing it (which is what is typically done with RL on coding).

The catch: LLM are not (currently) very good at proving things with Lean.

Enter Axiom: While they have not officially reported benchmark numbers besides the 12/12 Putnam result, Carina reports that they have achieved a very impressive 99% (187/189) ProofGen on the Verina codegen benchmark. This benchmark is to generate code and proof of correctness for a series of problems. For context, OpenAI o3 (the last known OpenAI run) achieved 4.9% on this benchmark.

Based on the sparse benchmarking, it’s hard to say how the frontier labs are currently doing outside the annual IMO milestones, but Carina suggests that they still are not training to generate Lean proofs directly, rather relying on informal proofs.

Time will tell if the frontier labs’ current approaches will close this gap.

Carina’s Ramanujan analogy is pretty direct. Better proofs → better Lean generation → better RL. A stronger signal means higher sample efficiency and higher maximum performance. Great!

Scaling is pretty clear too: once I have proved something in Lean, the quality of the output is basically4 as high as if it came from a human, so my high quality training set has grown in a way that an informal rollout corpus cannot. I can trust my Lean proofs.

Compounding is also clear: now all of future inference and training can build upon those proofs.

On the other hand, a model trained only using statistical signals like GRPO during RL lacks the sample efficiency, maximum performance and compounding corpus that a system that uses formal verification benefits from.

Broccoli and taxes notwithstanding, verification has shown up in a lot of our conversations. In the domain of physical systems, recall Applied Intuition:

“I think [verifiability] is probably the hardest problem right now, because the as the models get better, it can be harder and harder to find the faults on the system. And so the problem of doing proper eval to find those faults, that problem also keeps getting harder as the models get better.”

Physical AI that Moves the World — Qasar Younis & Peter Ludwig, Applied Intuition

In theoretical physics, we recall Alex Lupsasca:

“…now that we’re in this regime where you can just get ChatGPT to tackle thousands of questions at the same time, it will return proofs for a significant fraction of them. Now actually the onus is back on the humans to verify all the outputs. And so, yeah, as that becomes a bottleneck, I think formalizing math and automating verification will become more valuable.”

🔬Doing Vibe Physics — Alex Lupsasca, OpenAI

Verification is, in fact, the key differences between AI for science and AI for computation: in science you to have to actually test (verify) your hypothesis by performing physical experiments. Lab in the loop systems like Radical AI and Lila build around exactly this premise (we have recorded episodes with both of these teams and will release them soon!)

And yes, formally verifying critical systems such as flight control, nuclear power plants and pacemakers is a growing focus as the software and hardware that run them becomes more complex.

Carina believes so strongly that AGI requires verified generation that she makes the unqualified claim that “We do not believe there is any other possible future.”

Lean proofs are hard generate, but they can be easily shown to be correct or incorrect. But how do you know that the proof you created maps correctly to the problem you care about? As Carina puts it: “Anything that can be specified can be proven. Humans are bad at specifying everything we want.”

Are we now in the specification business? Check out the episode to hear Carina’s take, as well as:

  • Why hardware verification is a killer app

  • Details on the AXLE open API and recently released Discovery toolkit

  • The Erdos debacle

  • The OpenAI GPT-f diaspora

Timestamps:

  • 0:00 Intro: The $200M Series A and the Math Startup Thesis

  • 4:52 Verified AI: Scaling Brilliance, Not Fixing Lousiness

  • 13:42 Axiom’s System: Lean Data, RL, and the Putnam Perfect Score

  • 22:12 Mathematical Discovery — Before the Conjecture

  • 25:12 Rice’s Theorem, Incompleteness, and Practical Limits

  • 30:42 Code With Proof — The Verina Benchmark

  • 37:57 Proof Trees, Context Windows, and Scaling Limits

  • 43:57 Markets, Moat, and the Business Case ($1.6B valuation)

  • 55:27 Personal Origin Story: Oxford, UCL Gatsby, Stanford Law

  • 1:00:57 The Erdos Controversy and the Difficulty of Search

  • 1:06:02 AlphaZero for Math, Self-Improvement

  • 1:08:47 Startup Advantage and the OpenAI GPTF Thread

  • 1:13:17 Axle API — Open Infrastructure for Lean at Scale

  • 1:20:47 Collaboration, Polymath, and Human Attention as the Bottleneck

  • 1:22:21 Founding Story — Obsession, Law School, and Julie Zhuo

  • 1:26:17 The Bigger Vision — AGI, Science, and Transfer Learning

  • 1:35:02 Bottlenecks, Fragmentation, and the Field’s Future