惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
H
Hackread – Cybersecurity News, Data Breaches, AI and More
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
小众软件
小众软件
月光博客
月光博客
D
DataBreaches.Net
L
LangChain Blog
美团技术团队
S
SegmentFault 最新的问题
MyScale Blog
MyScale Blog
大猫的无限游戏
大猫的无限游戏
博客园 - 司徒正美
aimingoo的专栏
aimingoo的专栏
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Help Net Security
阮一峰的网络日志
阮一峰的网络日志
Y
Y Combinator Blog
I
InfoQ
U
Unit 42
Microsoft Azure Blog
Microsoft Azure Blog
J
Java Code Geeks
博客园 - 三生石上(FineUI控件)
腾讯CDC
Martin Fowler
Martin Fowler

OfficeChai

These Are The 10 Cheapest AI Models In The World [June 2026] 18 Best AI Tools For English Speaking (With Examples) [2026] AI Impact? Vacancy Rates For US Office Properties Are Now Highest Since The 2008 Crisis KPMG Pulls Report Praising AI After It Was Found To Have Fake AI-Generated Citations India's Sarvam Raises $234 Million At $1.5 Billion Valuation After SpaceX Stock Pops 20%, Musk Has Made More Money In The Last 24 Hours Than Warren Buffett Made In His Entire Career OfficeChai Nobody Is Using AI Better Than Meta: NVIDIA CEO Jensen Huang 21 Best AI Tools For Animation (With Examples) [2026] 22 Best AI Tools For Architecture (With Examples) [2026] Datacenter Construction Spending Has Eclipsed Public Transportation Spending In The US China Scraps 12,000 Degree Courses, Mainly In Arts And Humanities, To Prepare For AI Age OfficeChai There Is No Job Loss With AI: David Friedberg Loop Between Human Capital And "Token Capital" Will Be The New IP For Firms, Says Satya Nadella How to Reduce Dependency on Key Employees 8 Google Index Checker Use Cases Beyond New Blog Posts Memory Squeeze? Smartphone Purchases Are Down Globally 21 Best AI Tools For Accounting (With Examples) [2026] AI For Voice Generation: 22 Best Options (With Examples) [2026] These Are The Most Popular Image Generation Models On OpenRouter [June 2026] Search Traffic For Websites Is Down 25% Over The Last Year Because Of AI: a16z Data Agentic Coding Has Led To A 50% Increase In Number Of Apps, But Most Are Finding Very Few Users: SimilarWeb Data OpenRouter Launches Fusion API, Which Uses A Combination Of Models To Achieve Fable-Like Performance At Half The Price Dario Amodei Refused To De-Deploy Or Fix Vulnerabilities In Fable Before US Export Controls, Says David Sacks 23 Best AI Tools For Notes Making (With Examples) [2026] 16 Best AI Tools For Astrology (With Examples) [2026] How Jensen Huang Once Had To Ask SEGA's CEO To Pay NVIDIA For A Technology That Didn't Work ChatGPT Already Has 11% Of The Search Market: OpenAI CFO Sarah Friar SpaceX Has Now Launched More Satellites Than Rest Of Humanity Combined Across History
China's GLM 5.2 Beats All OpenAI, Google Models On GDPval...
OfficeChai Team · 2026-06-23 · via OfficeChai

Chinese models aren’t merely 6 months behind US labs — they now seem to be better than anything produced by top US labs.

GLM-5.2, the latest model from Beijing-based Knowledge Atlas (Z.AI), has placed third on GDPval-AA v2, Artificial Analysis’s benchmark for real-world, economically valuable knowledge work. The model scored 1524 Elo on the leaderboard, which is anchored to a human baseline of 1,000. Above it are only Claude Fable 5 (1783) and Claude Opus 4.8 (1615) — two Anthropic models. Every OpenAI and Google model sits below it.

GPT-5.5 at its highest reasoning setting scores 1509. Gemini 3.5 Flash, the best Google model on this leaderboard, lands at 1357. GLM-5.2 is ahead of both, and by a meaningful margin in a field where the gaps between top models have been narrowing.

What makes the GDPval-AA result more interesting than a standard benchmark finish is what the benchmark actually measures. These aren’t reasoning puzzles or code challenges in isolation — the tasks are agentic, multi-turn, and designed to mirror actual paid knowledge work. GLM-5.2 averaged roughly 31 turns per task across 1,999 matches. That’s not a model being asked to answer a question; it’s a model being asked to do a job, repeatedly, over a long horizon.

The open weights dimension adds another layer. MiniMax-M3, the next open model on the leaderboard, scores 1408 — 116 points behind GLM-5.2. In a space where open models have historically trailed proprietary systems by a wide margin on capability, that gap from open to open is almost as striking as the gap from GLM-5.2 to the proprietary models below it.

The pattern holds on AA-Briefcase, Artificial Analysis’s separate agentic knowledge work benchmark which combines rubric pass rate, analytical quality, and presentation into a single Elo score. GLM-5.2 again takes the top spot among open models with 1266, behind Claude Fable 5 (1587) and Claude Opus 4.8 (1356), but ahead of GPT-5.5 at xhigh reasoning (1159). On a benchmark built specifically around the kind of work people are paid to do — research, analysis, structured deliverables — GLM-5.2 is outperforming OpenAI’s best publicly available model.

Artificial Analysis put GLM-5.2 and three frontier models through the same set of real professional briefs: a daily task list for a retail supervisor, an IEC emergency-stop circuit schematic, and a moodboard for an orchestral ballad music video. Each deliverable was rendered exactly as the model produced it. GLM-5.2 held its own against Claude Fable 5, GPT-5.5, and Gemini 3.5 Flash across all three.

This isn’t the first time GLM-5.2 has shown up on a major leaderboard. The model placed fourth on the Artificial Analysis Intelligence Index, scoring 51 behind only Claude Fable 5 (60), Claude Opus 4.8 (56), and GPT-5.5 (55). It also leads open weights on the Agentic Index, the same category where GDPval-AA sits. These results aren’t isolated — they’re consistent across every evaluation that Artificial Analysis runs on GLM-5.2.

Knowledge Atlas has been releasing at a pace that most Western labs haven’t matched. GLM-5 launched in February, GLM-5.1 followed in late March, and GLM-5.2 dropped in June — roughly one significant model release every six weeks. GLM-5.1 had already topped SWE-Bench Pro ahead of GPT-5.4 and Claude Opus 4.6, making it the first Chinese model to lead that leaderboard. GLM-5.2 extended that streak on a different and arguably more practically relevant benchmark.

All of this is happening on Huawei Ascend chips. Knowledge Atlas has been on the US Entity List since January 2025, meaning it has no access to Nvidia hardware. The export control thesis — that restricting chip access would slow Chinese AI development — keeps running into GLM releases.

The price point adds to the picture. GLM-5.2 is priced at $1.40 per million input tokens and $4.40 per million output tokens. Claude Opus 4.8 costs $15 per million input and $75 per million output. For an open model at that price to rank alongside proprietary frontier systems on agentic work benchmarks built around real professional tasks, the gap between Chinese and American AI labs looks considerably different than the conventional narrative suggested even twelve months ago.