惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Troy Hunt's Blog
Blog — PlanetScale
Blog — PlanetScale
Engineering at Meta
Engineering at Meta
F
Full Disclosure
Recorded Future
Recorded Future
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
博客园_首页
博客园 - 叶小钗
MongoDB | Blog
MongoDB | Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hacker News: Front Page
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
博客园 - 司徒正美
Webroot Blog
Webroot Blog
Google DeepMind News
Google DeepMind News
Help Net Security
Help Net Security
Cloudbric
Cloudbric
PCI Perspectives
PCI Perspectives
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
量子位
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tailwind CSS Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
P
Proofpoint News Feed
N
News and Events Feed by Topic
罗磊的独立博客
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
T
Tor Project blog
IT之家
IT之家
M
MIT News - Artificial intelligence
S
Security @ Cisco Blogs
O
OpenAI News
AI
AI
S
Securelist
Simon Willison's Weblog
Simon Willison's Weblog
The Last Watchdog
The Last Watchdog
月光博客
月光博客
Security Archives - TechRepublic
Security Archives - TechRepublic
L
LINUX DO - 热门话题

HN's home page

More than 6 out of 10 people turn to AI for psychological support databow: a Rust CLI to query any database with an ADBC driver Pluto.jl 1.0 release – reactive notebook for Julia Use your Nvidia GPU's VRAM as swap space on Linux Show HN: Paseo – Beautiful open-source coding agent interface 4K years ago, Mohenjo-daro grew more equal over time Gleam v1.17.0 Released I'm skeptical about efforts to revolutionize schooling CT scans of BYD car parts Branchless Quicksort faster than std:sort and pdqsort with C and C++ API My thoughts after using Clojure for about a month The advertising cartel coming to your web browser Open Repair Data Standard – Open Repair Alliance JLink JTAG Access on the Pinecil Gmail Thinks I'm Stupid, So I Left HP re-releases classic computer science calculator: The HP-16C Show HN: Edsger – A handwritten Clojure REPL for the reMarkable 2 Microsoft's MAI-Code-1-Flash Scores 51% SWE-Bench Pro with Just 5B Active Params MAI-Thinking-1 Microsoft Announces AI Autopilot | Hacker News Morningstar values SpaceX at $780B, half its IPO target GitHub Copilot App | Hacker News Bringing Up DeepSeek-V4-Flash on AMD MI300X U.S. Army Corps of Engineers Bay Model Anthropic scales Claude Mythos to critical infrastructure in 15 countries QBE – Compiler Backend – 1.3 Larry Ellison: "Citizens will be on their best behavior because we’re recording" (2024) Three Ways to Get Paid (2018) Coreutils for Windows | Hacker News Trump signs executive order granting oversight of AI models Rethinking Search as Code Generation How we index images for RAG 1-Click GitHub Token Stealing via a VSCode Bug thunderbolt-ibverbs: We have InfiniBand at home WiFi Time | Hacker News Preparing for KDE Plasma's Last X11-Supported Release Please don't spam people looking for employment. It's just cruel Fidonet: Technology, Use, Tools, and History (1993) A walking tour of surveillance infrastructure in Seattle Expanding Project Glasswing Apple rejected my dictation app for using the accessibility API CSS-Native Parallax Effect | Hacker News Adafruit receives demand letter from Fenwick legal counsel on behalf of Flux.ai Stop Ruining It Why Janet? (2023) | Hacker News You Don't Love Systemd Timers Enough Show HN: Eyeball | Hacker News Strace-ui, Bonsai_term, and the TUI renaissance macOS needs its grid back How is Groq raising more money? Can the stockmarket swallow Anthropic, SpaceX and OpenAI? Age verification for social media, the beginning of the end for a free internet? Chipotlai Max | Hacker News OpenAI frontier models and Codex are now available on AWS Debug Project | Hacker News Should you normalize RGB values by 255 or 256? AI Agent Guidelines for CS336 at Stanford The newest Instagram “exploit” is the goofiest I've seen Anthropic confidentially submits draft S-1 to the SEC The Dirt That Refused to Die KDE at 30 The Pirate Bay Remains Resilient, 20 Years After the Raid CS336: Language Modeling from Scratch Sysadmining Like It's 2009 | Hacker News Nvidia Cosmos 3 Malicious npm packages detected across Red Hat Cloud Services Windows GOG DOS Games on M-Series Macs Flipper Zero Zig Template | Hacker News Linux Basics for Hackers (2019) Launch HN: Expanse (YC P26) – Unlock Wasted GPU Capacity Microsoft builds MacBook Pro rival with NVIDIA-powered Surface Laptop Ultra Now is the best time to be a duct tape engineer Go Experiments Explained | Hacker News Using Git's rerere feature to escape recurring conflict hell How turkey hacked the hair-transplant industry A 10 year old Xeon is all you need Sum-product, unit distances, and number fields Chuwi Minibook X | Hacker News Cloudflare Turnstile requiring fingerprintable WebGL Dav2d | Hacker News Squillions: How money laundering won London's Free Roof Terraces | Hacker News The Website Specification | Hacker News Why Custom Attributes in .NET Give Me Nightmares Muxcard, a dyi credit card size computer Webcam head tracking, webcam to control in‑game FOV CQL: Categorical Databases | Hacker News Decades of Effort Restore Steelhead and Salmon Passage on Alameda Creek Reviving Teletext for Ham Radio Unix in East Germany (GDR) (1990) Benchmarking SurrealDB 3.x vs. Postgres, Mongo, Neo4j and Redis (With Fsync) Key chemistry question answered, no quantum computer required New Beam Spring Keyboards | Hacker News Finding success in industry as a chip designer Linux/M68k | Hacker News Fooling around with encrypted reasoning blobs The Genius of the Barn Owl's Feathers Having your insulin pump die while you're on vacation Tracing HTTP Requests with Go's net/HTTP/httptrace Only 17% of all 64-bit Integers are products of two 32-bit integers
GLM 5.2 Performance Benchmarks | Hacker News
theanonymous · 2026-06-17 · via HN's home page
GLM 5.2 Performance Benchmarks (artificialanalysis.ai)
97 points by theanonymousone 8 hours ago | hide | past | favorite | 35 comments
 help


It does really well on "AA-Omniscience Non-Hallucination Rate", far higher than DeepSeek, GPT 5.5 or Fable. I really like that benchmark because it's one of the few benchmarks that allows LLMs to elect not to answer if they are unsure and punishes them for trying to bullshit their way through the benchmark


This implies that other benchmarks (for which every AI provider is optimizing?) are actively encouraging bullshitting?


A lot of benchmarks are setup to not punish false positives (irrelevant answers or extra text) and punish false negatives (missing the snippet being looked for).

This leads to answer bloat and/or hallucination if you benchmaxx on those


There is a tradeoff where as factual accuracy increases, creativity decreases, and the model becomes more "rigid" and less general. Unfortunately it seems that creativity is a good quality for reasoning and ultimately problem solving.

So we have a situation where models that can solve challenging problems, also tend to have problems with hallucinating, but those hallucinations seem be the breeding ground for the solutions that got them high "Wow" factor intelligence.


Yes. Most benchmarks just measure how many answers are correct. The best way to optimize that is to confidently state something, in hopes it's correct. Which is exactly how most LLMs behave, despite plenty of evidence that they do know whether they "know" something


if this is the case, then GLM 5.2 model seems better than gpt 5.5 or maybe even "Fable" depending upon what you are trying to achieve.

Fable model being removed from Anthropic because of security concerns by the US government (or well, also partially because of the personal vendetta between US govt and Anthropic)


Bullshitting is how LLMs work. It doesn't require active encouragement. All it takes is a machine without consciousness or physical access to the world and an actually-lived life. A training set that contains lots of confident answers and few to no refusals doesn't help either.


It's simpler than that.

An LLM outputs tokens, one-by-one. It stops the loop if it outputs the end-of-text token. Which is, of course, statistically much rarer than any other kind of token.

(This is why you cannot, in general, prompt an LLM with something like "don't answer if the result is correct". It has to output something, by design.)


They are, especially multiple choice questions. The same happens with humans exams:

Let's say there are 100 questions, with 4 answers each. A good answer is worth 1 point. By just guessing you get an average of 25/100, way more than 0/100 by not replying.

If instead a wrong answer is -1 point, by just guessing you get on average -75/100, way worse than 0/100.


Where do you see that? I see they have GPT-5.5 (xhigh) at 55, GPT-5.5 (high) at 53, and Muse Spark at 43. Muse Spark does beat GPT-5.4 mini (xhigh) which scores 40, but the key there is "mini".

In the coding index, GPT-5.5 gets 59.1, 58.5, 56.2, and 52.1 for xhigh, high, medium, and low while Muse Spark is behind at 47.5. For agentic, GPT-5.5 gets 74.1, 72.0, 69.4, and 59.7 (xhigh, high, medium, low) while Muse Spark gets 62.0 (beating only GPT-5.5 low).

GPT-5.5 only gets beaten by Opus 4.8 in their general index, is the top spot for coding, and is #3 behind Opus 4.8 and GLM-5.2 for agentic (excluding Fable 5 which takes the top spot, but is unavailable).


It's always nice to see how open source models growing, hope we will have good performance with lower tier hardware some day.


Local models are already useful today. The next milestone is getting this level of performance onto truly affordable hardware.


NVidia has less than zero reason to ship cards ideal for this at low prices.

AMD’s stock price reflects a hope they launch a CUDA alternative. But this is unlikely for the near future.

There is a lot of interest in preventing China coming in with cheap AI hardware.

So I expect the direction to be good local models that few can run effectively.


The Chinese will flood the market with cheap AI chips just like they did with EV cars. As consumers we can’t thank them enough.


I think it will eventually result in regulation and a potential grey market, and/or implosion of the centralized LLM services — I doubt they can keep hardware from becoming cheaper forever, and diminishing returns will make consumer hardware suitable for all but the hardest problems. At that point, the hardware “moat” will be completely gone and have become an extreme unrecoverable sunk cost.


Well, most people were not liking Fable when it was available anyway, because it refused to answer questions very often.


On some it does yes, also in real usage.

It avoided answering 2/21 tests in this specific benchmark mark, that's already 90% max score already.


I'm glad those tests apparently work out for you but a benchmark where three of the top 5 models are different flavors of Gemini Flash and zero are anything by Anthropic, is just so wildly divergent from my personal experience with the models that it's not useful to me.

Whatever it is you're measuring, it's not anything related to what I use models for.


Thanks for the feedback!

What are you using Claude models for? Coding only? Computer use? Which harness?


Not only coding but also general knowledge work, anything from learning about how some things work (e.g. walking me through PNP vs NPN transistors) to summarizing texts, doing web research, and occasionally some OCR.

I've experimented with a few models for all this and have found Gemini the best at OCR but quite a bit worse at the rest. Claude is worse than GPT at web research-shaped things, but Opus 4.8 wins my anecdote benchmark for the other tasks besides those two.

But really, for code or knowlege stuff Gemini is markedly worse than the others, while Opus and GPT 5.5 are very very close.


There’s no way Anthropic can keep jacking up the prices like this for every marginally better model. I think even tokenmaxxing companies are going to soon balk at $50/million output tokens.


Anthropic wants to ban the alternatives through regulation and ideally provide differential access with differential pricing.