惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
I
InfoQ
B
Blog
Engineering at Meta
Engineering at Meta
Y
Y Combinator Blog
GbyAI
GbyAI
T
The Blog of Author Tim Ferriss
G
Google Developers Blog
量子位
The Cloudflare Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
小众软件
小众软件
博客园 - 【当耐特】
Hugging Face - Blog
Hugging Face - Blog
阮一峰的网络日志
阮一峰的网络日志
博客园 - 司徒正美
美团技术团队
人人都是产品经理
人人都是产品经理
博客园 - 三生石上(FineUI控件)
酷 壳 – CoolShell
酷 壳 – CoolShell
大猫的无限游戏
大猫的无限游戏
罗磊的独立博客
博客园 - 聂微东
IT之家
IT之家

Latent.Space

[AINews] Zawinski's Law of MultiAgents [AINews] AMD buys Taalas [AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM??? [AINews] Megakernels are so dead and so back Unpacking ChatGPT Work: the Agent for a Billion Users [AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten [AINews] not much happened today [AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web [AINews] AI is eating Finance; AIE NYC now open [AINews] Fearing RSI: OpenAI, Anthropic, GDM, Meta, Thinky cosign letter to "Pace" AI development, as HuggingFace details Machine-Speed Offensive Cyberattack Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI [AINews] Much ado about Open Weights [AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition
[AINews] not much happened today
Latent.Space · 2026-07-18 · via Latent.Space

People continue to be impressed by yesterday’s Kimi K3 launch. Congrats to Databricks on their $188B Series M (watch our pod on the latest Databricks narratives) and OpenRouter might get bought (watch Alex Atallah’s keynote).

On a slow news day, The most popular talk this week is Abhishek Bhardwaj’s Sandbox track keynote which recaps a year of growth since his original work on Arrakis got him hired by Greg Brockman, and now building out the cloud infra behind ChatGPT Work (upcoming episode!). Spoilers: if you think running agent sandboxes is just “run containers on Kubernetes”, 1) you havent been paying attention to our E2B, Daytona and both Modal podcasts, and 2) you might be overtuned to compute problems and are probably underestimating the importance of storage/filesystems…

If you do leading AI work in NYC, especially for AI x Finance, speaker applications for AIE NYC 2026 opened today.

AI News for 7/16/2026-7/17/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Moonshot’s Kimi K3 Release, Frontier Positioning, and the China/Open-Weight Debate

  • Kimi K3 is the center of gravity today: the release triggered a broad reassessment of how close Chinese open-weight models are to the frontier. Multiple posts frame K3 as the first genuinely useful Chinese model at this tier, with strong coding, agentic, and long-horizon knowledge-work performance. Community reaction ranged from Salakhutdinov congratulating Moonshot founder Zhilin Yang to practitioners simply reporting that “Kimi K3 is really, really good”. A recurring theme was that K3 narrows the gap enough to pressure US labs to ship faster, as argued by @kimmonismus and others.

  • The strategic argument shifted from “compute moat” to “efficiency stack”: a notable thread argues that K3 weakens the thesis that frontier capability is gated mainly by raw FLOPs, pointing instead to MoE routing, quantization, data curation, and scarcity-driven infra design such as Moonshot’s “Mooncake” stack; see @AnikaSomaia. Related commentary emphasized that Chinese labs may be compressing the capability-per-FLOP curve rather than matching Western capex directly, with @dylan522p and @novasarc01 making the case that better post-training and harness conversion rates can shrink product gaps nonlinearly.

  • There is still disagreement on how far behind K3 really is: some view it as near-frontier or even surpassing specific Western models on important slices, while others argue it remains several months behind on broader generality, efficiency, or hidden evals. See the skeptical but detailed framing from @scaling01, contrasted with more bullish takes from @kimmonismus and @theinformation. The practical consensus is narrower: K3 is now impossible to dismiss.

Benchmarks: Artificial Analysis, Arena, DeepSWE, ARC, Cyber, and FrontierCode

  • Artificial Analysis and coding-agent benchmarks place K3 firmly in the top cluster: Artificial Analysis says the frontier widened from two to six labs above 51 on its Intelligence Index in roughly six weeks, with Kimi K3 at 57, behind Claude Fable 5 at 60 and ahead of Opus 4.8 at 56. On coding agents, AA later reported K3 scoring 57 on its Coding Agent Index, matching GPT-5.6 Terra and GPT-5.5, ahead of Opus 4.8, with 84% Terminal-Bench v2, 64% DeepSWE, and 23% SWE-Atlas-QnA. Cost claims were mixed: AA calls it frontier and relatively efficient; @theo counters that token efficiency and throughput often erase the headline price advantage versus GPT-5.6 Sol.

  • Frontend and coding evals were especially strong for K3: Arena reported that K3 put China ahead of the US on Frontend Code Arena for the first time, and user tests echoed that K3 can outperform or match Fable on visually grounded frontend tasks, e.g. @hqmank’s globe dashboard test. On software engineering, DataCurve said K3 debuted at #3 on DeepSWE, calling it the first open-weights model with frontier-level results there.

  • ARC and cyber remain useful reality checks: ARC Prize verified that Thinking Machines’ Inkling is now the highest-scoring open-weight model on both ARC-AGI-1 (79.5%) and ARC-AGI-2 (36.5%), while speculation around K3’s ARC-AGI-2 score continues via BenchPress estimates. On cyber, the UK AISI-related discussion around GLM-5.2 matching Opus 4.5 on “The Last Ones” and OpenAI’s claim that GPT-5.6 Sol is SOTA on that range underscores that open models still appear materially behind the best closed models on long-horizon cyber, even as the gap narrows.

Model Architecture, Inference, and Systems Work

Agents, Memory, MCP, and Workflow Scaffolding

Research Notes Beyond K3

  • Robustness and detector limits: the paper “The Illusion of Robustness” argues that aggregate accuracy masks prediction flips under irrelevant context; see the arXiv pointer and a Japanese summary. Separately, Epoch AI reported that AI detectors are usually reliable on plain human text and naive AI text, but LLMs instructed to mimic specific authors can evade detection, with false negatives around 13% and ~26% for scientific writing.

  • Embodied and biologically inspired learning: NVIDIA’s RoboTTT extends robot policy context length by 3 orders of magnitude, improving manipulation performance 87% over a single-step baseline and completing a five-minute ten-stage assembly task that no baseline finished. Meanwhile, Sakana’s “Diffusing Blame” and Hardmaru’s summary show competitive learning under strict Dale’s principle without standard backprop weight transport.

  • Interpretability / representation geometry: Elie Bakouch replicated Anthropic-style j-space analysis on Thinking Machines’ Inkling, finding it unusual in maintaining similar geometry across early and late layers (early-late CKA ~0.8 vs ~0.5 elsewhere). The same thread reports minimal j-space change under NVFP4 quantization for Poolside’s Laguna XS 2.1.

Top Tweets (by engagement, filtered for technical relevance)