惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
Vercel News
Vercel News
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
MyScale Blog
MyScale Blog
P
Proofpoint News Feed
量子位
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
博客园 - 【当耐特】
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
D
DataBreaches.Net
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
小众软件
小众软件
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
T
The Blog of Author Tim Ferriss
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Jina AI
Jina AI
博客园 - 三生石上(FineUI控件)

Latent.Space

[AINews] Zawinski's Law of MultiAgents [AINews] AMD buys Taalas [AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM??? [AINews] Megakernels are so dead and so back Unpacking ChatGPT Work: the Agent for a Billion Users [AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten [AINews] not much happened today [AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web [AINews] AI is eating Finance; AIE NYC now open [AINews] Fearing RSI: OpenAI, Anthropic, GDM, Meta, Thinky cosign letter to "Pace" AI development, as HuggingFace details Machine-Speed Offensive Cyberattack Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI [AINews] Much ado about Open Weights [AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" Inside the Model Factory — Eiso Kant, Poolside AI [AINews] AI Cybersecurity becomes top of mind 🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist) [AINews] not much happened today [AINews] not much happened today [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) [AINews] not much happened today 5 Trends That Defined AI Engineering at World’s Fair 2026 [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code?? [AINews] not much happened today [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp
[AINews] SpaceXAI launches Grok 4.5, first Opus-class mod...
Latent.Space · 2026-07-09 · via Latent.Space

As GPT 5.6 is confirmed to launch tomorrow, today is pretty much the last day anyone will be excited about a GPT 5.5 equivalent model launch, and that is exactly what SpaceXAI did:

X avatar for @cursor_ai

Cursor@cursor_ai

We've partnered with SpaceXAI to train Grok 4.5. It’s our most powerful model yet and the first we've built for more than software engineering.

5:57 PM · Jul 8, 2026 · 3.24M Views

605 Replies · 1.26K Reposts · 14.4K Likes

The new Grok 4.5 is a different weight class than the Composer series (1.5T) and despite the solid evals still performs very comparably to the current workhorse Opus and GPTs, although per OpenAI’s evals team even the mighty SWE-Bench Pro is now saturated/terminally flawed - leaving presumably a small list of successors including FrontierCode.

As for training and data disclosures, this is all the information we have.

AI News for 7/07/2026-7/08/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

Top Story: Grok 4.5 release

xAI/“SpaceXAI” publicly launched Grok 4.5 as a new coding-and-agents-focused frontier model, positioned on capability-per-dollar rather than absolute benchmark supremacy.

  • Elon Musk first said Grok 4.5 would be made public “tomorrow” based on strong beta feedback, calling it “Opus-class,” but faster, more token-efficient, and lower cost @elonmusk.

  • Musk later framed Grok 4.5 internally as “roughly comparable to Opus 4.7, but much faster,” emphasizing usefulness to Tesla and SpaceX engineers over benchmark chasing @elonmusk.

  • The official launch came from xAI’s account, describing Grok 4.5 as “our first model trained specifically for coding and agents,” trained with Cursor, and offering “frontier intelligence at leading speeds and cost efficiency” @SpaceXAI.

  • Cursor said it partnered with xAI to train Grok 4.5, called it “our most powerful model yet,” and stressed that it was “the first we’ve built for more than software engineering” @cursor_ai.

  • Cursor also announced in-product availability with “double usage for the first week” @cursor_ai.

  • Cursor clarified that “Grok 4.5 and Composer are two different model weight classes,” and that Composer 2.5 would remain available with future models in that smaller class @cursor_ai.

  • Early ecosystem support appeared immediately: Grok 4.5 became available in Grok Build/API/Cursor @milichab, day-0 support was announced for Hermes Agent @Teknium, and later live availability in Hermes Agent/Portal/OpenRouter/Grok subscriptions was confirmed @Teknium.

  • Musk said the context window would likely move from 500k back to 1M “by next week” @elonmusk.

Officially, xAI’s message was not “best overall model,” but near-Opus quality with materially better economics and speed:

  • “Opus-class model, but faster, more token-efficient and lower cost” @elonmusk

  • “First model trained specifically for coding and agents” @SpaceXAI

  • “Frontier intelligence at leading speeds and cost efficiency” @SpaceXAI

  • “Most powerful model yet” and “first we’ve built for more than software engineering” @cursor_ai

This framing matters: xAI is explicitly targeting the coding-agent workflow market that has recently been dominated by Anthropic/OpenAI/Cursor-style tool-using systems, not just general chat.

The concrete numbers that surfaced:

  • Official pricing: $2 / 1M input tokens, $6 / 1M output tokens @scaling01

  • Artificial Analysis repeated the same price point and added:

    • cache hits discounted by 75% to $0.5 / 1M tokens

    • long inputs over 200k tokens cost double

    • 500k context window, down from Grok 4.3’s 1M

    • vision input retained

    • configurable reasoning retained @ArtificialAnlys

  • Musk later said the context window would probably upgrade back to 1M soon @elonmusk.

Relative pricing comparisons cited by users:

  • Grok 4.5: $2 in / $6 out

  • GPT-5.6: $5 in / $30 out

  • Opus 4.8: $5 in / $25 out @kimmonismus

One important spec surfaced via third-party reporting of Musk’s disclosure:

That is a notable jump, and likely central to why multiple observers interpreted 4.5 as xAI’s first entry into the true flagship coding-agent tier rather than an iterative refresh.

Artificial Analysis provided the most substantive external evaluation in the tweet set.

Key results:

  • #4 on Artificial Analysis Intelligence Index, score 54, behind only Fable 5, GPT-5.5, and Opus 4.8 @ArtificialAnlys

  • +16 points vs Grok 4.3 on the same index @ArtificialAnlys

  • GDPval-AA v2 Elo 1543, also ranking #4, behind Anthropic’s latest Claude releases @ArtificialAnlys

  • Top score on τ³-Banking: 33%, above 31% for GPT-5.5 (xhigh) @ArtificialAnlys

  • Artificial Analysis Coding Agent Index score 76 in Grok Build, “on par with GPT-5.5 in Codex” and below Fable 5 in Claude Code @ArtificialAnlys

  • Cost per Intelligence Index task: $0.31 @ArtificialAnlys

  • Cost per GDPval task: $0.49 @ArtificialAnlys

  • Cost per Coding Agent Index task: $2.59 @ArtificialAnlys

  • Average output tokens per Intelligence Index task: ~14k, over 60% lower than Opus 4.8 @ArtificialAnlys

  • Average total tokens per Coding Agent Index task: 1.9M, versus 7.2M for Fable 5 in Claude Code and 6.2M for GPT-5.5 in Codex @ArtificialAnlys

Artificial Analysis’ interpretation was clear: Grok 4.5 is near-frontier on capability, but unusually strong on efficiency, making it sit on the Pareto frontier for cost/performance.

Musk explicitly amplified the Artificial Analysis assessment @elonmusk.