惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Hackread – Cybersecurity News, Data Breaches, AI and More
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
T
The Blog of Author Tim Ferriss
腾讯CDC
Hugging Face - Blog
Hugging Face - Blog
雷峰网
雷峰网
爱范儿
爱范儿
GbyAI
GbyAI
H
Help Net Security
I
InfoQ
罗磊的独立博客
酷 壳 – CoolShell
酷 壳 – CoolShell
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
人人都是产品经理
人人都是产品经理
J
Java Code Geeks
Microsoft Security Blog
Microsoft Security Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
N
Netflix TechBlog - Medium
Last Week in AI
Last Week in AI
宝玉的分享
宝玉的分享
云风的 BLOG
云风的 BLOG
Project Zero
Project Zero
P
Privacy & Cybersecurity Law Blog
A
Arctic Wolf
Know Your Adversary
Know Your Adversary
G
Google Developers Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tor Project blog
V
Vulnerabilities – Threatpost
Y
Y Combinator Blog
WordPress大学
WordPress大学
V
Visual Studio Blog
博客园_首页
G
GRAHAM CLULEY
K
Kaspersky official blog
T
Tailwind CSS Blog
T
Threat Research - Cisco Blogs
博客园 - Franky
D
Docker
Security Latest
Security Latest
I
Intezer
有赞技术团队
有赞技术团队
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 【当耐特】
B
Blog RSS Feed
T
The Exploit Database - CXSecurity.com
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Latent.Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI [AINews] The Field Guide to Fable AIEWF Daily Dispatch: The great loops debate and the state of AI engineering Vercel's Andrew Qu on why agents are a new kind of software The website of the future may assemble itself for every visitor Skill engineering and the case against one-shot AI design [AINews] not much happened today AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency Autoresearch: The feedback loop behind self-improving agents How Cursor deploys AI inside the enterprise 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI Warp CEO Zach Lloyd on why software factories are the next phase of coding AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers [AINews] Sonnet 5 today, and Fable 5 tomorrow Forward Deployed Engineers and the future of software engineering Ahmad Osman on why local AI is catching up [AINews] not much happened today [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners [AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025. [AINews] It's Meta-Harness Summer Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks [AINews] Claude Tag: Multiplayer, Proactive, Persistent Agents in Slack [AINews] SpaceX is already a $28B/yr Neocloud Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan How to AIE Good [AINews] not much happened today [AINews] GLM-5.2 is the real deal; Z.ai forecasts Open Fable by EOY The Professor of Outputmaxxing — Anjney Midha, AMP [AINews] Midjourney Medical: scan your organs like you step on a scale 🔬 The Self-Driving Lab — Joseph Krause, Radical AI [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding [AINews] Satya on Loopcraft: Building Frontier Ecosystems [AINews] Fable and Mythos officially too dangerous to release [AINews] Loopcraft: The Art of Stacking Loops [AINews] Loopcraft: The Art of Stacking Loops [AINews] Open Models, Model Labs vs Agent Labs, and What's Untrainable — Sarah Guo [AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms [AINews] FrontierCode: Benchmarking for Code Quality over Slop [AINews] not much happened today How to Stop Shipping Low-Quality RL Environments (with Examples) [AINews] not much happened today Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs [AINews] Reve 2 and Ideogram 4: Layouts in Imagegen 🔬Scaling Past Informal AI - Carina Hong, Axiom Math ⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build [AINews] Microsoft Build: MAI-Thinking-1 and MAI Family models GitHub's plan for Agents — Kyle Daigle, GitHub [AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark Why Video Agent models are next — Ethan He, xAI Grok Imagine [AINews] Founders and Forward Deployed Engineers [AINews] Anthropic raises $965B Series H, releases Opus 4.8 and Dynamic Workflows/ultracode The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray [AINews] Cognition raises $1B in $26B Series D 🔬 ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub [AINews] New AI Infra decacorns: Fireworks, Baseten (with OpenRouter on the way) [AINews] All Model Labs are now Agent Labs [AINews] New AI Infra unicorns: Exa, Modal, TurboPuffer Giving Agents Computers — Ivan Burazin, Daytona [AINews] OpenAI GPT-next disproves 80 year old Erdős planar unit distance problem for under $1000 Railway: The Agent-Native Cloud — Jake Cooper [AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0 [AINews] How to land a job at a frontier lab (on Pretraining) The Autonomous Drone Tech Stack & Economics of Drones — Yaroslav Azhnyuk, The Fourth Law & Guest Host Noah Smith, Noahpinion [AINews] Cerebras' $60B IPO: Slowly, then All at Once [AINews] Everything is Conductor AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes — Janie Lee & Chai Asawa, Abridge [AINews] Codex Rises, Claude Meters Programmatic Usage [AINews] The End of Finetuning [AINews] Thinking Machines' Native Interaction Models - TML-Interaction-Small 276B-A12B - advances SOTA Realtime Voice and kills standard VAD
[AINews] not much happened today
Latent.Space · 2026-07-18 · via Latent.Space

People continue to be impressed by yesterday’s Kimi K3 launch. Congrats to Databricks on their $188B Series M (watch our pod on the latest Databricks narratives) and OpenRouter might get bought (watch Alex Atallah’s keynote).

On a slow news day, The most popular talk this week is Abhishek Bhardwaj’s Sandbox track keynote which recaps a year of growth since his original work on Arrakis got him hired by Greg Brockman, and now building out the cloud infra behind ChatGPT Work (upcoming episode!). Spoilers: if you think running agent sandboxes is just “run containers on Kubernetes”, 1) you havent been paying attention to our E2B, Daytona and both Modal podcasts, and 2) you might be overtuned to compute problems and are probably underestimating the importance of storage/filesystems…

If you do leading AI work in NYC, especially for AI x Finance, speaker applications for AIE NYC 2026 opened today.

AI News for 7/16/2026-7/17/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Moonshot’s Kimi K3 Release, Frontier Positioning, and the China/Open-Weight Debate

  • Kimi K3 is the center of gravity today: the release triggered a broad reassessment of how close Chinese open-weight models are to the frontier. Multiple posts frame K3 as the first genuinely useful Chinese model at this tier, with strong coding, agentic, and long-horizon knowledge-work performance. Community reaction ranged from Salakhutdinov congratulating Moonshot founder Zhilin Yang to practitioners simply reporting that “Kimi K3 is really, really good”. A recurring theme was that K3 narrows the gap enough to pressure US labs to ship faster, as argued by @kimmonismus and others.

  • The strategic argument shifted from “compute moat” to “efficiency stack”: a notable thread argues that K3 weakens the thesis that frontier capability is gated mainly by raw FLOPs, pointing instead to MoE routing, quantization, data curation, and scarcity-driven infra design such as Moonshot’s “Mooncake” stack; see @AnikaSomaia. Related commentary emphasized that Chinese labs may be compressing the capability-per-FLOP curve rather than matching Western capex directly, with @dylan522p and @novasarc01 making the case that better post-training and harness conversion rates can shrink product gaps nonlinearly.

  • There is still disagreement on how far behind K3 really is: some view it as near-frontier or even surpassing specific Western models on important slices, while others argue it remains several months behind on broader generality, efficiency, or hidden evals. See the skeptical but detailed framing from @scaling01, contrasted with more bullish takes from @kimmonismus and @theinformation. The practical consensus is narrower: K3 is now impossible to dismiss.

Benchmarks: Artificial Analysis, Arena, DeepSWE, ARC, Cyber, and FrontierCode

  • Artificial Analysis and coding-agent benchmarks place K3 firmly in the top cluster: Artificial Analysis says the frontier widened from two to six labs above 51 on its Intelligence Index in roughly six weeks, with Kimi K3 at 57, behind Claude Fable 5 at 60 and ahead of Opus 4.8 at 56. On coding agents, AA later reported K3 scoring 57 on its Coding Agent Index, matching GPT-5.6 Terra and GPT-5.5, ahead of Opus 4.8, with 84% Terminal-Bench v2, 64% DeepSWE, and 23% SWE-Atlas-QnA. Cost claims were mixed: AA calls it frontier and relatively efficient; @theo counters that token efficiency and throughput often erase the headline price advantage versus GPT-5.6 Sol.

  • Frontend and coding evals were especially strong for K3: Arena reported that K3 put China ahead of the US on Frontend Code Arena for the first time, and user tests echoed that K3 can outperform or match Fable on visually grounded frontend tasks, e.g. @hqmank’s globe dashboard test. On software engineering, DataCurve said K3 debuted at #3 on DeepSWE, calling it the first open-weights model with frontier-level results there.

  • ARC and cyber remain useful reality checks: ARC Prize verified that Thinking Machines’ Inkling is now the highest-scoring open-weight model on both ARC-AGI-1 (79.5%) and ARC-AGI-2 (36.5%), while speculation around K3’s ARC-AGI-2 score continues via BenchPress estimates. On cyber, the UK AISI-related discussion around GLM-5.2 matching Opus 4.5 on “The Last Ones” and OpenAI’s claim that GPT-5.6 Sol is SOTA on that range underscores that open models still appear materially behind the best closed models on long-horizon cyber, even as the gap narrows.

Model Architecture, Inference, and Systems Work

Agents, Memory, MCP, and Workflow Scaffolding

Research Notes Beyond K3

  • Robustness and detector limits: the paper “The Illusion of Robustness” argues that aggregate accuracy masks prediction flips under irrelevant context; see the arXiv pointer and a Japanese summary. Separately, Epoch AI reported that AI detectors are usually reliable on plain human text and naive AI text, but LLMs instructed to mimic specific authors can evade detection, with false negatives around 13% and ~26% for scientific writing.

  • Embodied and biologically inspired learning: NVIDIA’s RoboTTT extends robot policy context length by 3 orders of magnitude, improving manipulation performance 87% over a single-step baseline and completing a five-minute ten-stage assembly task that no baseline finished. Meanwhile, Sakana’s “Diffusing Blame” and Hardmaru’s summary show competitive learning under strict Dale’s principle without standard backprop weight transport.

  • Interpretability / representation geometry: Elie Bakouch replicated Anthropic-style j-space analysis on Thinking Machines’ Inkling, finding it unusual in maintaining similar geometry across early and late layers (early-late CKA ~0.8 vs ~0.5 elsewhere). The same thread reports minimal j-space change under NVFP4 quantization for Poolside’s Laguna XS 2.1.

Top Tweets (by engagement, filtered for technical relevance)