惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Security @ Cisco Blogs
罗磊的独立博客
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
美团技术团队
T
Tailwind CSS Blog
博客园 - 三生石上(FineUI控件)
博客园 - Franky
G
Google Developers Blog
Jina AI
Jina AI
Stack Overflow Blog
Stack Overflow Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
Visual Studio Blog
腾讯CDC
S
SegmentFault 最新的问题
Recent Announcements
Recent Announcements
博客园 - 叶小钗
Microsoft Security Blog
Microsoft Security Blog
雷峰网
雷峰网
L
LangChain Blog
Vercel News
Vercel News
Forbes - Security
Forbes - Security
PCI Perspectives
PCI Perspectives
N
News | PayPal Newsroom
S
Security Affairs
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 司徒正美
J
Java Code Geeks
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Hacker News: Ask HN
Hacker News: Ask HN
Schneier on Security
Schneier on Security
A
About on SuperTechFans
Attack and Defense Labs
Attack and Defense Labs
Google Online Security Blog
Google Online Security Blog
aimingoo的专栏
aimingoo的专栏
MongoDB | Blog
MongoDB | Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
Cloudbric
Cloudbric
B
Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
P
Proofpoint News Feed
D
DataBreaches.Net
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
B
Blog RSS Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
N
News and Events Feed by Topic

Latent.Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO [AINews] The Field Guide to Fable AIEWF Daily Dispatch: The great loops debate and the state of AI engineering Vercel's Andrew Qu on why agents are a new kind of software The website of the future may assemble itself for every visitor Skill engineering and the case against one-shot AI design [AINews] not much happened today AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency Autoresearch: The feedback loop behind self-improving agents How Cursor deploys AI inside the enterprise 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI Warp CEO Zach Lloyd on why software factories are the next phase of coding AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers [AINews] Sonnet 5 today, and Fable 5 tomorrow Forward Deployed Engineers and the future of software engineering Ahmad Osman on why local AI is catching up [AINews] not much happened today [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners [AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025. [AINews] It's Meta-Harness Summer Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks [AINews] Claude Tag: Multiplayer, Proactive, Persistent Agents in Slack [AINews] SpaceX is already a $28B/yr Neocloud Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan How to AIE Good [AINews] not much happened today [AINews] GLM-5.2 is the real deal; Z.ai forecasts Open Fable by EOY The Professor of Outputmaxxing — Anjney Midha, AMP [AINews] Midjourney Medical: scan your organs like you step on a scale 🔬 The Self-Driving Lab — Joseph Krause, Radical AI [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding [AINews] Satya on Loopcraft: Building Frontier Ecosystems [AINews] Fable and Mythos officially too dangerous to release [AINews] Loopcraft: The Art of Stacking Loops [AINews] Loopcraft: The Art of Stacking Loops [AINews] Open Models, Model Labs vs Agent Labs, and What's Untrainable — Sarah Guo [AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms [AINews] FrontierCode: Benchmarking for Code Quality over Slop [AINews] not much happened today How to Stop Shipping Low-Quality RL Environments (with Examples) [AINews] not much happened today Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs [AINews] Reve 2 and Ideogram 4: Layouts in Imagegen 🔬Scaling Past Informal AI - Carina Hong, Axiom Math ⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build [AINews] Microsoft Build: MAI-Thinking-1 and MAI Family models GitHub's plan for Agents — Kyle Daigle, GitHub [AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark Why Video Agent models are next — Ethan He, xAI Grok Imagine [AINews] Founders and Forward Deployed Engineers [AINews] Anthropic raises $965B Series H, releases Opus 4.8 and Dynamic Workflows/ultracode The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray [AINews] Cognition raises $1B in $26B Series D 🔬 ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub [AINews] New AI Infra decacorns: Fireworks, Baseten (with OpenRouter on the way) [AINews] All Model Labs are now Agent Labs [AINews] New AI Infra unicorns: Exa, Modal, TurboPuffer Giving Agents Computers — Ivan Burazin, Daytona [AINews] OpenAI GPT-next disproves 80 year old Erdős planar unit distance problem for under $1000 Railway: The Agent-Native Cloud — Jake Cooper [AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0 [AINews] How to land a job at a frontier lab (on Pretraining) The Autonomous Drone Tech Stack & Economics of Drones — Yaroslav Azhnyuk, The Fourth Law & Guest Host Noah Smith, Noahpinion [AINews] Cerebras' $60B IPO: Slowly, then All at Once [AINews] Everything is Conductor AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes — Janie Lee & Chai Asawa, Abridge [AINews] Codex Rises, Claude Meters Programmatic Usage [AINews] The End of Finetuning [AINews] Thinking Machines' Native Interaction Models - TML-Interaction-Small 276B-A12B - advances SOTA Realtime Voice and kills standard VAD
[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI
Latent.Space · 2026-07-08 · via Latent.Space

Congrats to Meta Superintelligence on having the top 2/3 image/video models in the world! This would’ve been a candidate for a title story, but unfortunately that is pretty much all the detail we have about Muse Image/Video - no paper, no technical detail whatsoever. Still, this beats the Microsoft MAI models from last month which is nice.

We are noted Lilian Weng fans, so we take notice whenever she drops another research recap, especially rare now that she is a cofounder at Thinky. Today she is thinking about the relationship of harnesses to RSI:

X avatar for @lilianweng

Lilian Weng@lilianweng

new post on harness engineering for AI self-improvement: lilianweng.github.io/posts/2026-07-… It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter

lilianweng.github.io

Harness Engineering for Self-Improvement

5:58 AM · Jul 7, 2026 · 405K Views

72 Replies · 534 Reposts · 3.89K Likes

While we have written before about how even Greg Brockman is now quietly endorsing agent/harness engineering, it is refreshing for a respected thinker and neolab cofounder like Lilian to also agree that “Even when many harness improvement[s] get eventually internalized into core model, the need to specify goals and context will not disappear.”

Her post breaks out the main proven design trends in harnesses that everyone should know, and then recaps the harness optimization literature, most notably from the well known ACE paper to even more recent trends like Meta-Harnesses, which we have covered anecdotally on AINews.

It surely also provides a hint as to what Thinky is Thinking, beyond just Interaction Models.

AI News for 7/06/2026-7/07/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Agent Products, Harnesses, and Long-Running Workflows

  • Anthropic expands “background agent” UX on top of Claude: The biggest product launch by engagement was Claude Cowork coming to mobile and web, positioning Claude as a task-running background teammate rather than a foreground chat UI. Related posts show the product convergence around a shared home tab and tighter Chat/Cowork integration from @mikeyk. Separately, Anthropic extended access to Claude Fable 5 on paid plans through July 12 in a highly engaged announcement from @claudeai, though many users noted the awkward timing relative to weekly limits in reactions from @kimmonismus and others.

  • Harness engineering is increasingly the center of agent design: Lilian Weng’s new post was widely referenced as reframing recursive self-improvement around the harness, not direct weight self-modification; Sakana’s summary connects this to The AI Scientist, ShinkaEvolve, and Darwin Gödel Machine in their thread. LangChain echoed the same shift with a new Deep Agents course and an open-source harness project in posts from @LangChain and @hwchase17. Google is also productizing this direction: Gemini API Managed Agents added background execution, remote MCP servers, custom function calling, and credential refresh in posts from @_philschmid and @OfficialLoganK.

  • Practical agent infra keeps getting more opinionated: There were several notable operator-facing updates: Codex Mobile iOS added task management, filtered diffs, SSH key login, branch comparison, and attachment flows in posts from @Dimillian and @reach_vb; Hermes Agent added pluggable secrets managers plus native 1Password integration and export of sessions/datasets to formats including private Hugging Face repos in @Teknium’s threads; Weaviate 1.38 made its MCP server GA with runtime-gated write access, notably allowing MCP_SERVER_WRITE_ACCESS_ENABLED to be flipped live without restart in @victorialslocum’s post. A more experimental pattern came from @omarsar0, using a Dial MCP server so agents can escalate decisions via phone call/SMS/iMessage for human-in-the-loop control.

Model and Modality Releases: Audio, Speech, Robotics, and Media Generation

  • Meta’s Muse Image/Muse Video push agentic generation into media: Meta Superintelligence Labs launched Muse Image and previewed Muse Video in announcements from @AIatMeta, @alexandr_wang, and @_tim_brooks. The notable technical angle is not just image quality, but an explicitly agentic generation loop: planning, web search, tool use, code execution, and self-refinement before rendering. Meta also says performance improves with scaled test-time compute, and that self-refinement behavior emerged during RL rather than being hand-scripted in this follow-up. On public evals, Muse Image quickly reached #2 on Image Arena behind GPT Image 2 in Arena’s ranking, while Muse Video debuted at #3 on Video Arena in another Arena post.

  • NVIDIA and Cohere both shipped strong audio releases: NVIDIA released Audex, a 30B parameter / 3B active MoE with 1M context for unified text+audio work, summarized by @HuggingPapers and described in more detail by @_weiping. The model’s core claim is preserving text intelligence while adding broad audio generation and understanding via a single MoE backbone. Cohere launched Cohere Transcribe Arabic, described as the most accurate open-source Arabic ASR model, under Apache 2.0, with emphasis on dialects, code-switching, and Arabic-accented English in posts from @cohere and @JayAlammar.

  • Open robotics keeps consolidating around Hugging Face + NVIDIA: NVIDIA expanded its robotics stack into the HF ecosystem by bringing GR00T 1.7 and Isaac Teleop into LeRobot, aimed at open humanoid robotics workflows, in @NVIDIARobotics’s announcement and integration guide. On the embodied side, UMA showed a strong full-stack robotics narrative: @RemiCadene described a prototype built by a small team in 9 months, while the Northstar reveal and @psermanet’s safety note emphasized vertically integrated hardware/software for trustworthy robots.

Training, Inference, and Post-Training Techniques

  • Liquid AI’s “Antidoom” directly targets reasoning-loop failure modes: One of the clearest technical releases of the day was Liquid AI’s Antidoom, an open-source training method to reduce doom loops where small reasoning models repeat tokens until context exhaustion. The reported reductions are substantial: LFM2.5-2.6B from 10.2% → 1.4% and Qwen3.5-4B from 22.9% → 1% under greedy sampling, with downstream eval gains. The method, FTPO (Final Token Preference Optimization), relabels the loop-triggering token and redistributes probability toward alternatives, summarized well by @helloiamleonie and @LiorOnAI. This is a good example of the field’s recent pattern: removing specific failure modes rather than only scaling parameters.

  • Inference efficiency and compression remain a major frontier: NVIDIA’s Puzzle-75B-A9B compression work got strong attention via @omarsar0: compressing a hybrid MoE parent model while preserving reasoning, coding, long-context, and agentic quality, with roughly 2x server throughput and 1M-context concurrency on H100 rising from 1 request to 8. On the tooling side, Nsight Python 1.0 launched in @HagedornBastian’s post, making GPU perf analysis scriptable in Python. Unsloth also shipped GGUFs for DeepSeek-V4-Flash, plus export to NVFP4/FP8 and speedups for GRPO and MoEs in @danielhanchen’s update.

  • Agent RL and verification are getting more specialized: @cwolferesearch highlighted how GRPO-style normalization is being adapted for agentic RL at the task or environment level to handle higher reward variance in multi-turn environments. Separately, @omarsar0 flagged a training-free verifier paper from Stanford/NVIDIA/Berkeley that reads calibrated continuous scores off scoring-token logits, posting strong numbers across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench and suggesting verification is becoming an independent scaling axis.

Interpretability, Model Internals, and the “J-Space” Debate

  • Anthropic’s J-space work dominated interpretability discussion, but also drew sharp criticism: The community split between seeing the work as useful mechanistic analysis and objecting to the consciousness framing. Strong critiques came from @danburonline, @paul_cal, and @scaling01, who argued the vectors are causal largely by construction under the Jacobian-lens definition. A useful historical reference came from @jacobandreas, pointing readers back to the original Jacobian lenses paper.

  • The stronger technical takeaway is cross-model structure, not consciousness rhetoric: @eliebakouch computed CKA similarity on J-lens geometry across 38 open models and found surprisingly universal layer/depth organization, even across unrelated families like Llama and OLMo. Anthropic and Neuronpedia also released J-lens weights for open models, noted in this follow-up. In parallel, Goodfire introduced Block-Sparse Featurizers for multidimensional concepts in activations, arguing many vision concepts are inherently 2–4 dimensional blocks rather than single directions, in their thread.

Benchmarks, Evaluations, and Domain-Specific Systems

  • Agent and legal benchmarks continue to expose the gap between “passes many criteria” and “fully solves real work”: Agent Arena placed Claude Sonnet 5 (Thinking) at #6, with strongest signals in confirmed task success and bash usage, but still with uncertainty around steerability. Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark over 120 private legal tasks across 24 practice areas, where Claude Fable 5 led at 14.2% all-pass rate; Claude Opus 4.8 and GLM-5.2 tied at 7.5%, with GLM hitting that at roughly ~6% of Fable’s cost per task in their release. The big message is that models can satisfy many individual rubric items yet still fail to produce acceptable end-to-end deliverables.

  • Research automation and specialized domain systems are broadening: Google promoted Experience AI Scientist, a multi-agent system for end-to-end scientific workflows, in this ICML post. DeepMind also launched Predicting the Past, grounding Gemini in Aeneas and Ithaca for Greek/Latin historical analysis via plain-English interactions, in their thread. On legal AI commercialization, Norm Ai announced a $120M Series C at $1.2B valuation and described a full-stack “agentic law” setup spanning software plus an AI-native law firm in @johnjnay’s post.

Top tweets (by engagement)

  • New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0) (Activity: 653): Tencent released the non-preview Hy3 open model collection on Hugging Face, described as a 295B-parameter MoE with 21B active parameters, now under Apache 2.0 rather than the prior restrictive community license. The post highlights that the earlier license reportedly excluded use in regions including South Korea, the UK, and the EU, while top comments point to claimed benchmark gains over HY3-Preview and frame this as potentially relevant for high-end local/home inference setups. Commenters viewed the Apache 2.0 relicensing as the most important change, especially given Tencent’s recent translation models also using Apache licensing. There was cautious optimism that the reported benchmark improvements may translate to real-world usefulness, but with implicit skepticism until tested outside vendor charts.

    • Commenters highlighted that Hunyuan/HY3 is now listed as Apache 2.0, contrasting it with the prior “community” license that reportedly restricted usage in regions such as South Korea, the UK, and the EU. This was viewed as technically important for deployment because Apache 2.0 removes many commercial and geographic usage barriers.

    • Several users focused on whether Tencent’s claimed benchmark improvements over HY3-Preview will translate into real-world workloads. Given the reported 295B total / 21B active MoE-style configuration, commenters suggested it could be relevant for “high-end home setups” if inference formats such as GGUF become available.

    • There was early speculation that HY3 could become an alternative to Qwen and MiniMax models in local/open-weight workflows, but commenters were waiting for quantized releases and independent testing before drawing conclusions.