惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Palo Alto Networks Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
C
CERT Recently Published Vulnerability Notes
C
Cybersecurity and Infrastructure Security Agency CISA
S
Schneier on Security
S
Securelist
酷 壳 – CoolShell
酷 壳 – CoolShell
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cyberwarzone
Cyberwarzone
Apple Machine Learning Research
Apple Machine Learning Research
S
SegmentFault 最新的问题
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
GbyAI
GbyAI
Security Latest
Security Latest
Last Week in AI
Last Week in AI
Microsoft Security Blog
Microsoft Security Blog
云风的 BLOG
云风的 BLOG
Recorded Future
Recorded Future
Webroot Blog
Webroot Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
TaoSecurity Blog
TaoSecurity Blog
C
Cisco Blogs
博客园 - 【当耐特】
Blog — PlanetScale
Blog — PlanetScale
Hugging Face - Blog
Hugging Face - Blog
B
Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Attack and Defense Labs
Attack and Defense Labs
The Last Watchdog
The Last Watchdog
U
Unit 42
阮一峰的网络日志
阮一峰的网络日志
Project Zero
Project Zero
WordPress大学
WordPress大学
L
LINUX DO - 最新话题
F
Fortinet All Blogs
L
LINUX DO - 热门话题
PCI Perspectives
PCI Perspectives
Simon Willison's Weblog
Simon Willison's Weblog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
MongoDB | Blog
MongoDB | Blog
Latest news
Latest news
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
The Hacker News
The Hacker News
爱范儿
爱范儿
O
OpenAI News
J
Java Code Geeks
T
The Exploit Database - CXSecurity.com
H
Hackread – Cybersecurity News, Data Breaches, AI and More

Latent.Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI [AINews] The Field Guide to Fable AIEWF Daily Dispatch: The great loops debate and the state of AI engineering Vercel's Andrew Qu on why agents are a new kind of software The website of the future may assemble itself for every visitor Skill engineering and the case against one-shot AI design [AINews] not much happened today AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency Autoresearch: The feedback loop behind self-improving agents How Cursor deploys AI inside the enterprise 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI Warp CEO Zach Lloyd on why software factories are the next phase of coding AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers [AINews] Sonnet 5 today, and Fable 5 tomorrow Forward Deployed Engineers and the future of software engineering Ahmad Osman on why local AI is catching up [AINews] not much happened today [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners [AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025. [AINews] It's Meta-Harness Summer Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks [AINews] Claude Tag: Multiplayer, Proactive, Persistent Agents in Slack [AINews] SpaceX is already a $28B/yr Neocloud Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan How to AIE Good [AINews] not much happened today [AINews] GLM-5.2 is the real deal; Z.ai forecasts Open Fable by EOY The Professor of Outputmaxxing — Anjney Midha, AMP [AINews] Midjourney Medical: scan your organs like you step on a scale 🔬 The Self-Driving Lab — Joseph Krause, Radical AI [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding [AINews] Satya on Loopcraft: Building Frontier Ecosystems [AINews] Fable and Mythos officially too dangerous to release [AINews] Loopcraft: The Art of Stacking Loops [AINews] Loopcraft: The Art of Stacking Loops [AINews] Open Models, Model Labs vs Agent Labs, and What's Untrainable — Sarah Guo [AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms [AINews] FrontierCode: Benchmarking for Code Quality over Slop [AINews] not much happened today How to Stop Shipping Low-Quality RL Environments (with Examples) [AINews] not much happened today Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs [AINews] Reve 2 and Ideogram 4: Layouts in Imagegen 🔬Scaling Past Informal AI - Carina Hong, Axiom Math ⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build [AINews] Microsoft Build: MAI-Thinking-1 and MAI Family models GitHub's plan for Agents — Kyle Daigle, GitHub [AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark Why Video Agent models are next — Ethan He, xAI Grok Imagine [AINews] Founders and Forward Deployed Engineers [AINews] Anthropic raises $965B Series H, releases Opus 4.8 and Dynamic Workflows/ultracode The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray [AINews] Cognition raises $1B in $26B Series D 🔬 ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub [AINews] New AI Infra decacorns: Fireworks, Baseten (with OpenRouter on the way) [AINews] All Model Labs are now Agent Labs [AINews] New AI Infra unicorns: Exa, Modal, TurboPuffer Giving Agents Computers — Ivan Burazin, Daytona [AINews] OpenAI GPT-next disproves 80 year old Erdős planar unit distance problem for under $1000 Railway: The Agent-Native Cloud — Jake Cooper [AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0 [AINews] How to land a job at a frontier lab (on Pretraining) The Autonomous Drone Tech Stack & Economics of Drones — Yaroslav Azhnyuk, The Fourth Law & Guest Host Noah Smith, Noahpinion [AINews] Cerebras' $60B IPO: Slowly, then All at Once [AINews] Everything is Conductor AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes — Janie Lee & Chai Asawa, Abridge [AINews] Codex Rises, Claude Meters Programmatic Usage [AINews] The End of Finetuning [AINews] Thinking Machines' Native Interaction Models - TML-Interaction-Small 276B-A12B - advances SOTA Realtime Voice and kills standard VAD
[AINews] not much happened today
Latent.Space · 2026-07-15 · via Latent.Space

Yesterday’s headline story became even more true, with Superapp usage adding yet another 1M users since we last wrote:

X avatar for @swyx

swyx@swyx

uhm this gpt 5.6 launch might be the openai's most successful model ever since... since chatgpt? this is IPO altering stuff going on here

X avatar for @latentspacepod

Latent.Space @latentspacepod

Did... Codex just overtake Claude Code? 24.5 hours ago Tibo announced 6M active users. this means Codex usage jumped 1M in ~ONE DAY. the last user number we heard from Claude Code was 2M in Feb: https://t.co/jghFZlpjEq more analysis within, but this is very big if true. https://t.co/cMw1QUyj9C

10:43 PM · Jul 14, 2026 · 10.4K Views

20 Replies · 4 Reposts · 101 Likes

In other news, Richard MacManus published his final AIEWF26 recap of recaps:

Including coverage of Addy Osmani’s excellent keynote covering what AI engineers should continue doing even when the cost of code generation trends to zero:

AI News for 7/13/2026-7/14/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Coding Agents, Harnesses, and the Shift From Chat to Execution

Open Models, Quantization, and Local Inference Compression

  • Aggressive compression is bringing frontier-adjacent models onto consumer devices: PrismML released Bonsai 27B, based on Qwen 3.6 27B, in two compact variants: Ternary Bonsai 27B at 5.9 GB / 1.71 effective bits and 1-bit Bonsai 27B at 3.9 GB / 1.125 effective bits, both under Apache 2.0. The claim is notable not just for size, but for preserving multimodal, tool-using, long-context agentic workflows locally; a demo shows Hermes running it on an RTX 5090, while Locally AI highlighted phone deployment. In parallel, Tencent Hunyuan released 1-bit and 4-bit Hy3, describing a 295B flagship-scale model that can be served on a single GPU via llama.cpp with MTP enabled.

  • Quantization and edge deployment continue to broaden the open-model operating envelope: @danielhanchen announced NVFP4 dynamic quants across the Gemma-4 family and additional large models including Qwen3.5-122B-A10B and GLM-4.7-Flash. @MiaAI_lab’s DGX Spark thread sketched practical multi-node local deployments, including 1M-context DeepSeek v4 Flash and MiMo-V2.5 on 2× DGX Sparks, and GLM 5.2 NVFP4 across four. The common theme across these posts is that local inference is no longer just a toy path: it is becoming viable for serious agentic workflows, especially when paired with low-bit weight formats and optimized harnesses.

Multimodal and World-Model Systems: Video, Realtime VLMs, and Motion

  • Realtime multimodal interaction is moving from “watch then answer” to continuous perception: OpenMOSS released MOSS-VL-Realtime, an 11B vision-language family under Apache 2.0 with 256K context, designed for continuous video streams. Its key systems property is that it can keep watching while generating, revise or interrupt answers as scenes change, and remain silent when evidence is insufficient. A companion technical thread from @Open_MOSS emphasizes a cross-attention architecture, XRoPE for unified temporal-spatial positioning, and unified templates across offline/streaming/realtime settings.

  • Long-video understanding is increasingly framed as active evidence search, not passive frame ingestion: a dense summary from @ZhihuFrontier described OmniAgent, built on Qwen2.5-Omni-7B, which uses an Observation–Thought–Action loop to request only the frames/audio it needs. On LVBench, OmniAgent-7B reportedly scored 50.5, beating Qwen2.5-VL-72B at 47.3, while consuming only ~203 frames vs 768. The training recipe is also notable: passive SFT hurt performance, while 58K agentic trajectories and entropy-weighted RL via TAURA improved it. The larger research pattern here aligns with Andrew Carr’s note that motion is a fundamentally novel data type requiring dedicated collection, infra, and model treatment rather than being reduced to images-with-time.

  • Open world models are inching toward interactive, longer-horizon simulation: @RekaAILabs outlined the data stack behind omni world models, stressing petabytes of video, 6 pipeline stages, and the doubled payoff from data-quality improvements when models both generate and understand video. @omarsar0 summarized LingBot-World 2.0 as one of the first open releases claiming hour-scale, 720p/60fps interactive generation, though still without long-term memory. On the application side, PixVerse Game was highlighted as pursuing the harder problem of real-time interactive video response rather than canned game-like clips.

Research Infrastructure, Benchmarks, and Evaluation Methodology

  • Perplexity open-sourced WANDR, a benchmark for wide-and-deep agentic research: @perplexity_ai described WANDR as a 500-task benchmark built from de-identified production research tasks, requiring 170,495 source-backed records across multiple difficulty tiers. Rather than grading against a static gold set, WANDR re-fetches cited pages and checks claims against underlying evidence, which better matches dynamic web research. @AravSrinivas framed this as the internal benchmark behind Perplexity Computer’s deep-and-wide research harness, while @denisyarats emphasized its additional role as an RL environment synthesized from production traces.

  • Eval design is getting more adversarial and more realistic: Agent Arena highlighted work cutting system costs by 89% while matching the best static config’s accuracy, arguing that full system config > LLM routing alone. Relatedly, Google DeepMind work on model routing argued that routers should be judged not just by accuracy/cost but by behavioral differentiation among experts and stability under paraphrase; otherwise routing may be functionally meaningless. @HamelHusain’s automated evals post landed in a similar place: these systems can spot issues humans miss, but still lack enough domain taste and feedback loops to replace experts.

  • Benchmarks are expanding beyond one-shot SWE tasks toward degradation and search realism: mini-swe-agent marked one year while now powering multiple software benchmarks; SlopCodeBench was cited as measuring how agents erode codebases over sequential tasks rather than just solving one isolated issue. This broadens the benchmark surface from “can it solve a task?” to “can it avoid making the repository worse over time?”

Physical AI, Collective Intelligence, and Robotics

  • Sakana AI pushed collective intelligence from software into physical self-repairing systems: across multiple posts, Sakana introduced “Smart Cellular Bricks”, published in Nature Communications. The system consists of many identical cubes, each running a small neural network and communicating only with physical neighbors, yet able to infer global shape and detect damage without centralized control. A follow-up detail is especially notable: the cells can detect missing neighbors across six spatial directions with 95% accuracy and regrow target structures; in simulation, the method scaled to 18,000+ cubes (detail thread).

  • Physical autonomy is also showing up in much smaller form factors: @alextoussss posted a striking demo of an autonomous micro-drone achieving an air-to-air kill of a flying moth, framed as a step toward mosquito eradication. Separately, @fchollet highlighted Airtap, which turns SMS into a headless agentic execution layer for mobile apps, using text as the control plane and intervening only for authentication. These are different ends of the autonomy spectrum, but both point to interfaces where humans specify goals while systems handle embodied or semi-embodied execution.

Top tweets (by engagement)