惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
N
Netflix TechBlog - Medium
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园_首页
S
SegmentFault 最新的问题
H
Help Net Security
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
T
The Blog of Author Tim Ferriss
量子位
GbyAI
GbyAI
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
P
Privacy & Cybersecurity Law Blog
B
Blog
月光博客
月光博客
博客园 - 聂微东
Vercel News
Vercel News
罗磊的独立博客
腾讯CDC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
Blog — PlanetScale
Blog — PlanetScale
L
Lohrmann on Cybersecurity
I
Intezer
小众软件
小众软件
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
C
Check Point Blog
AWS News Blog
AWS News Blog
C
Cisco Blogs
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
宝玉的分享
宝玉的分享
S
Security Affairs
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
雷峰网
雷峰网
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Hacker News: Front Page
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Security Blog
Microsoft Security Blog

Latent.Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI [AINews] The Field Guide to Fable AIEWF Daily Dispatch: The great loops debate and the state of AI engineering Vercel's Andrew Qu on why agents are a new kind of software The website of the future may assemble itself for every visitor Skill engineering and the case against one-shot AI design [AINews] not much happened today AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency Autoresearch: The feedback loop behind self-improving agents How Cursor deploys AI inside the enterprise 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI Warp CEO Zach Lloyd on why software factories are the next phase of coding AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers [AINews] Sonnet 5 today, and Fable 5 tomorrow Forward Deployed Engineers and the future of software engineering Ahmad Osman on why local AI is catching up [AINews] not much happened today [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners [AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025. [AINews] It's Meta-Harness Summer Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks [AINews] Claude Tag: Multiplayer, Proactive, Persistent Agents in Slack [AINews] SpaceX is already a $28B/yr Neocloud Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan How to AIE Good [AINews] not much happened today [AINews] GLM-5.2 is the real deal; Z.ai forecasts Open Fable by EOY The Professor of Outputmaxxing — Anjney Midha, AMP [AINews] Midjourney Medical: scan your organs like you step on a scale 🔬 The Self-Driving Lab — Joseph Krause, Radical AI [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding [AINews] Satya on Loopcraft: Building Frontier Ecosystems [AINews] Fable and Mythos officially too dangerous to release [AINews] Loopcraft: The Art of Stacking Loops [AINews] Loopcraft: The Art of Stacking Loops [AINews] Open Models, Model Labs vs Agent Labs, and What's Untrainable — Sarah Guo [AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms [AINews] FrontierCode: Benchmarking for Code Quality over Slop [AINews] not much happened today How to Stop Shipping Low-Quality RL Environments (with Examples) [AINews] not much happened today Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs [AINews] Reve 2 and Ideogram 4: Layouts in Imagegen 🔬Scaling Past Informal AI - Carina Hong, Axiom Math ⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build [AINews] Microsoft Build: MAI-Thinking-1 and MAI Family models GitHub's plan for Agents — Kyle Daigle, GitHub [AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark Why Video Agent models are next — Ethan He, xAI Grok Imagine [AINews] Founders and Forward Deployed Engineers [AINews] Anthropic raises $965B Series H, releases Opus 4.8 and Dynamic Workflows/ultracode The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray [AINews] Cognition raises $1B in $26B Series D 🔬 ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub [AINews] New AI Infra decacorns: Fireworks, Baseten (with OpenRouter on the way) [AINews] All Model Labs are now Agent Labs [AINews] New AI Infra unicorns: Exa, Modal, TurboPuffer Giving Agents Computers — Ivan Burazin, Daytona [AINews] OpenAI GPT-next disproves 80 year old Erdős planar unit distance problem for under $1000 Railway: The Agent-Native Cloud — Jake Cooper [AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0 [AINews] How to land a job at a frontier lab (on Pretraining) The Autonomous Drone Tech Stack & Economics of Drones — Yaroslav Azhnyuk, The Fourth Law & Guest Host Noah Smith, Noahpinion [AINews] Cerebras' $60B IPO: Slowly, then All at Once [AINews] Everything is Conductor AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes — Janie Lee & Chai Asawa, Abridge [AINews] Codex Rises, Claude Meters Programmatic Usage [AINews] The End of Finetuning [AINews] Thinking Machines' Native Interaction Models - TML-Interaction-Small 276B-A12B - advances SOTA Realtime Voice and kills standard VAD
[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)
Latent.Space · 2026-07-16 · via Latent.Space

Thinky only seems to come up for air once every few months; most recently with Interaction models - but each time they do they impress, showing both taste and depth. Today they introduced Inkling — not a SOTA model, but a very solid new family for a baseline American open model:

  • Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active.

  • It supports a context window of up to 1M tokens.

  • It was pretrained on 45 trillion tokens of text, images, audio and video.

  • It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency.

  • Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort

The Huggingface breakdown covers some interesting technical highlights:

AI News for 7/14/2026-7/15/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Thinking Machines Lab launched Inkling, its first fully released open-weights foundation model family entry, positioning it as a customizable multimodal base model rather than a benchmark-maxed flagship.

  • Thinking Machines announced Inkling as an open-weights model that “reasons efficiently across text, image, and audio modalities,” with full weights available and immediate support on its Tinker platform and Playground @thinkymachines.

  • Mira Murati described Inkling as the company’s “first model,” “trained from scratch,” with open weights and same-day fine-tuning on Tinker @miramurati.

  • Soumith Chintala framed it as Thinking Machines’ “first general model,” stressing open weights, 975B parameters, native multimodality, and availability on Tinker, Hugging Face, and partners @soumithchintala.

  • John Schulman added timeline context: pretraining began last winter, and from mid-January a small team built coding, reasoning, and agentic training on top @johnschulman2.

  • Lilian Weng characterized Inkling as a foundation model aimed at “solid performance across a broad categories of capabilities” and intended for practical use plus customization @lilianweng.

  • TML staff repeatedly emphasized that this is a day-1 release and a foundation for future iterations rather than their final frontier push @soumithchintala, @cHHillee, @keirp1.

  • The release landed with unusually broad day-0 ecosystem support across vLLM, SGLang, Modal, Baseten, Databricks, Hugging Face, and quantization/community tooling @vllm_project, @lmsysorg, @modal, @baseten, @Yuchenj_UW, @huggingface, @danielhanchen.

  • Independent commentators immediately tagged it as the strongest U.S.-based open-weight release so far, though generally still behind the top Chinese open-weight and best closed models on some benchmarks @natolambert, @ArtificialAnlys, @scaling01.

Several technically literate reactions extracted architectural choices from the release:

  • Hybrid/sliding-window attention with a 5:1 local-to-global layer ratio and window size 512 @eliebakouch, @ariG23498.

  • Relative positional encoding / relative attention bias instead of RoPE; multiple posters called this one of the most novel large-scale choices @stochasticchasm, @eliebakouch, @rasbt, @arohan, @ChangJonathanC.

  • Short convolution layers added around attention/FFN streams; commenters flagged this as unusually scaled-up usage of short convs @eliebakouch, @stochasticchasm, @rasbt, @SonglinYang4.

  • MoE with shared expert sinks / 2 shared experts, noted as atypical since many recent MoEs use 1 shared expert @eliebakouch, @ariG23498.

  • DeepSeek-style auxiliary-loss-free load balancing was cited in community readings of the architecture @eliebakouch.

  • muP and Muon/weight decay variants were inferred from the writeup and confirmed by optimizer expert reaction: Aaron Defazio said they are using his corrected weight decay approach, “MuonC/AdamC” @aaron_defazio, while community readers also pointed out muP @stochasticchasm, @Laz4rz.

  • 8 MTP heads for speculative decoding were highlighted by vLLM @vllm_project.

  • Inkling-Small is repeatedly referenced as an upcoming or separately discussed smaller model @LiorOnAI, @teortaxesTex.

  • Community summaries describe Inkling-Small as 276B total / 12B active and unexpectedly competitive versus the larger model on several evaluations @eliebakouch, @nrehiew_.

  • Artificial Analysis said Inkling debuts at 41 on the Intelligence Index, making it the leading U.S. open-weights release and ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29), and gpt-oss-120b (24) @ArtificialAnlys.

  • Artificial Analysis also said Inkling averages 25K output tokens per Intelligence Index task, vs 43K for GLM-5.2 max, 38K for Kimi K2.6, and 37K for DeepSeek v4 Pro max, framing it as relatively token-efficient @ArtificialAnlys.

  • Natolambert called it a “clear step up from Nemotron Ultra” and “new best American model,” but still “a bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal” @natolambert.

  • Design Arena said Inkling entered Agentic Web App Arena at #9 overall, Elo 1257, in the same band as Claude Opus 4.6 and Gemini 3.5 Flash, and called it the highest-ranking U.S.-based open-weight model for agentic workloads @DesignArena.

  • Arena added Inkling to Agent Arena / Text / Vision / Code Arena on launch day @arena.

From Artificial Analysis:

  • GDPval-AA v2 Elo 1238, higher than Kimi K2.6 (1190) and DeepSeek v4 Flash max (1189) @ArtificialAnlys.

  • τ³-Banking 24%, above Kimi K2.6 (21%) and slightly above DeepSeek v4 Flash max (23%) @ArtificialAnlys.

Positive:

  • “Sharp and concise” reasoning, not rambly @MichaelElabd.

  • Strong tool calling and good long-horizon error recovery on agentic tasks @MichaelElabd.

  • Good “quality of mind” / unsycophantic flavor @skirano, @tinkerapi.

  • Alex Kirillov claimed Inkling avoids the common “audio in = intelligence penalty” seen in many omni models, though another user asked for stronger supporting evidence and benchmarks @alex_kirillov, @giffmana, @alex_kirillov.

More mixed / critical:

  • Scaling01 argued the benchmarks are “not that great,” describing it as roughly “another Kimi-K2.6” and behind all closed models and GLM-5.2, speculating the release may have been timed ahead of Kimi-K3 and DeepSeek-V4-GA @scaling01.

  • Stochasticchasm said it seems “very strong for multimodal” but “not super strong for terminal bench etc.” @stochasticchasm.

  • JJitsev pushed back on hype around “only open-weight model trained without distilling,” saying Inkling uses distillation from open weights and underperforms GLM 5.2 on TerminalBench-style evals @JJitsev.

  • TeortaxesTex offered a contrarian positive spin: mediocre benchmark-maxing may actually suggest less corner-cutting/distillation contamination and a more independent data pipeline @teortaxesTex.

  • NVIDIA said Inkling was trained on GB300 NVL72 and that an NVFP4 checkpoint was available on Hugging Face on day 0 @NVIDIAAI.

  • vLLM said day-0 support includes NVFP4 and BF16, optimized for Blackwell and Hopper, reaching up to 380 tok/s/user on 4× GB200 with MTP @vllm_project.

  • Inferact detailed system work: sconv-aware tensor-parallel sharding, low-latency fused collectives (5× faster at bs=1), and direct integration of TML’s FA4 sheared-bias kernel @inferact.

  • LMSYS/SGLang said Inkling architecture support was implemented natively, including ShortConv, relative positional attention, shared expert sink MoE, prefill full CUDA graph, MXFP8 KV cache, full parameter and LoRA RL in customized Megatron backend, routing replay, cross-runtime parameter sync, and DFlash speculative decoding from Modal @lmsysorg.

  • Modal said Inkling on Modal uses a custom DFlash speculator for 67% higher throughput and interactivity @modal.

  • Soumith Chintala separately amplified that Modal’s DFlash speculator is “much faster than MTP” @soumithchintala.

  • Lysandre reported replacing TML’s causal Conv1D with causal-conv1d yielded +4% tok/s, and replacing attention with FlashAttention-4 yielded another +11%, for ~15% total throughput gain without retraining @LysandreJik.

  • Unsloth released 1-bit GGUF quants said to be 86% smaller (270GB vs 1.9TB) while retaining 74.2% of top-1% accuracy, with vision and audio support @danielhanchen.

  • Artificial Analysis listed Tinker pricing as:

    • 64K context: $1.87 / 1M input, $0.374 cached, $4.68 output

    • 256K context: $3.74 / 1M input, $0.748 cached, $9.36 output
      @ArtificialAnlys

  • Available on Tinker, Hugging Face, and via launch partners including Databricks, Baseten, Modal, vLLM/SGLang stacks @soumithchintala, @Yuchenj_UW, @baseten, @modal.

  • “Best American open model” / “saved American open-source frontier” are judgments, albeit repeated by several respected observers @natolambert, @karinanguyen, @saranormous.

  • Claims that Inkling is especially important because it is not distilled from OpenAI/Anthropic are disputed. Jxmnop called it “the ONLY open-weight model” without such distillation @jxmnop, then partially walked it back: “apparently they did distill lol. but only a tiny bit” @jxmnop. Andrew Carr also contested the purity framing, noting use of Kimi 2.5 for SFT traces @andrew_n_carr.

  • Claims that Inkling was “rushed” ahead of Chinese releases are speculation from critics, not evidenced by the launch materials @scaling01.

  • Claims that relative attention gives TML a finetuning moat because backward is hard are speculative @typedfemale.

  • Claims that Inkling avoids multimodal intelligence loss are promising but not yet benchmark-complete in the tweet set @alex_kirillov.

  • Open-weight and permissive license as strategic win: Many saw the Apache-2.0 release as a major boost to the U.S./Western open ecosystem @latkins, @saranormous, @brexton, @hyperindexed.

  • Customization over leaderboard chasing: Researchers and builders praised the explicit framing that Inkling is a broad, tunable foundation rather than a benchmark-maxed point solution @gneubig, @ben_burtenshaw, @thealexker.

  • Strong release quality: Several users praised the transparency, grounded tone, and comprehensive technical documentation @lvwerra, @saranormous, @rasbt.

  • Architecture interest: The non-RoPE positional choice and scaled short-conv usage drew positive attention as evidence TML is willing to make meaningful architecture bets @stochasticchasm, @rasbt, @ChangJonathanC.

  • Strong but not top overall: The most balanced reads place Inkling as the new U.S. open-weight leader, but behind GLM/Kimi/DeepSeek or top closed models on some fronts @natolambert, @ArtificialAnlys, @stochasticchasm.

  • Good base model thesis: Multiple analysts read the release as a systems/business move: ship a solid, efficient, post-trainable base and let Tinker plus downstream RL/fine-tuning create differentiation @ben_burtenshaw, @kimmonismus, @tinkerapi.

  • Not frontier overall: Critics argued it is still clearly behind top Chinese open-weight models and the strongest closed models @scaling01, @JJitsev.

  • Purity claims overstated: Some pushback focused on exaggerated claims that it is uniquely “pure” or non-distilled; the thread set includes both hype and corrections @jxmnop, @jxmnop, @andrew_n_carr, @JJitsev.

  • Benchmark middlingness as concern: Some readers saw the moderate benchmark profile as evidence it may simply lag current Chinese open frontier rather than inaugurate a new frontier @scaling01.

  • First major TML public model: This is the first true external model release from Thinking Machines after months of anticipation around a lab staffed by ex-OpenAI leaders and researchers. That made the choice of open weights itself notable @Hesamation, @TechCrunch.

  • A U.S. open-weight answer to Chinese momentum: Many reactions explicitly compare Inkling to GLM, Kimi, DeepSeek, and Qwen. The release lands amid concern that Western open-weight models have trailed Chinese ones on capability and release cadence @scaling01, @teortaxesTex, @sriramk.

  • Open base + post-training stack thesis: TML’s messaging strongly suggests a strategy similar to “ship a competent open substrate, then differentiate via customization/fine-tuning/RL infrastructure.” That aligns with Tinker distribution and with user reactions centering controllable reasoning, concise outputs, and adaptation rather than raw leaderboard supremacy @thinkymachines, @MichaelElabd, @ben_burtenshaw.

  • Inference ecosystem maturity: The release also showcases how far open inference stacks have come. Day-0 support for a 1T-class multimodal MoE with new architectural components and multiple kernel-level optimizations would have been far less plausible a year earlier @vllm_project, @inferact, @LysandreJik.

  • Architectural experimentation at scale: Relative positional bias instead of RoPE and large-scale short-conv usage are the kind of choices researchers watch closely because they may indicate future architecture trends if they prove robust under scaling and post-training @stochasticchasm, @rasbt, @ChangJonathanC.

  • Release style as signal: Several commentators praised the unusually restrained release language, explicit admission that it is not the strongest overall model, and detailed technical notes. For expert audiences, that improved credibility relative to more benchmark-maxed launches @eliebakouch, @lvwerra, @thealexker.