惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
P
Privacy International News Feed
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog
S
Schneier on Security
Project Zero
Project Zero
T
Threatpost
Spread Privacy
Spread Privacy
阮一峰的网络日志
阮一峰的网络日志
C
Cybersecurity and Infrastructure Security Agency CISA
AWS News Blog
AWS News Blog
H
Heimdal Security Blog
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
P
Privacy & Cybersecurity Law Blog
J
Java Code Geeks
罗磊的独立博客
博客园 - Franky
博客园 - 叶小钗
S
Security Affairs
月光博客
月光博客
Application and Cybersecurity Blog
Application and Cybersecurity Blog
The Last Watchdog
The Last Watchdog
WordPress大学
WordPress大学
人人都是产品经理
人人都是产品经理
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
A
Arctic Wolf
Cloudbric
Cloudbric
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V2EX - 技术
V2EX - 技术
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LINUX DO - 最新话题
Y
Y Combinator Blog
宝玉的分享
宝玉的分享
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
N
News | PayPal Newsroom
Hugging Face - Blog
Hugging Face - Blog
美团技术团队
W
WeLiveSecurity
云风的 BLOG
云风的 BLOG
The Register - Security
The Register - Security
I
InfoQ
F
Fortinet All Blogs
T
The Exploit Database - CXSecurity.com
S
SegmentFault 最新的问题
Recent Announcements
Recent Announcements
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
L
Lohrmann on Cybersecurity

Maggie Appleton

The Dark Forest and Generative AI One Developer, Two Dozen Agents, Zero Alignment Gas Town’s Agent Patterns, Design Bottlenecks, and Vibecoding at Scale January 2026 | Maggie Appleton A Treatise on AI Chatbots Undermining the Enlightenment A Brief History & Ethos of the Digital Garden Vibe Code is Legacy Code May 2025 | Maggie Appleton Home-Cooked Software and Barefoot Developers Statistically, When Will My Baby Be Born? Speculative Calendar Events ChatGPT Would be a Decent Policy Advisor March 2025 | Maggie Appleton The Expanding Dark Forest and Generative AI Squish Meets Structure Common Misconceptions in AI Undetected AI Exam Answers Unbaited Smidgeons Growing a Human: The First 30 Weeks How to Import Academic Papers from Zotero into Tana December 2024 | Maggie Appleton Aesthetic Command Lines with Hyper, Spaceship, and Oh My Zsh Leaving Elicit July 2024 | Maggie Appleton A Short History of Bi-Directional Links The Pattern Language of Project Xanadu Assumed Audiences Ambient Co-presence On Opening Essays, Conference Talks, and Jam Jars Spinning Worlds, Seasickness, and Dealing with Vestibular Neuritis A Collection of Design Engineers Gathering Structures Daily Notes Pages Historical Trails December 2023 | Maggie Appleton September 2023 | Maggie Appleton Digital Gardening for Non-Technical Folks Language Model Sketchbook, or Why I Hate Chatbots June 2023 | Maggie Appleton Computational Notebooks Folk Interfaces Reverse Outlining with Language Models Command K Bars Spatial Web Browsing A Picture Worth a Thousand Programmes Programmable Notes Programming Portals Teenage Skeuomorphic Desktop Designs Tending Evergreen Notes in Roam Research Growing the Evergreens Why You Own an iPad and Still Can't Draw A Brief Introduction to Digital Anthropology Transclusion and Transcopyright Dreams The Block-Paved Path to Structured Data Empty Pointers and Constellations of AI Metaphors We Web By The Gift Economy Epistemic Disclosure November 2022 | Maggie Appleton Joining Ought July 2022 | Maggie Appleton The Linear Oppression of Note-taking Apps Paleolithic Nostalgia Interoperable Personal Libraries and Ad Hoc Reading Groups The Finest Narrative Non-Fiction Essays Algorithmic Transparency October 2021 | Maggie Appleton Plebeian Programming with Keyboard Maestro The Cultural Anthropology of React August 2021 | Maggie Appleton Natureculture, Moral Purity, and Cultural Boundaries The Echo & Narcissus Writing Club Pink, Soft, Glittering Developers Fetishism & Mechanical Keyboards Making Programming Visual, Spatial, and Learnable Organic, Local, Artisan Data Storage Positioning Elements & Scrollytelling in CSS Painting Roam Research with Custom CSS A Digital Anthropology Reading List The Eponymous Laws of Programming A History of Cyborgs Neologisms GreenSock Animations with React Hooks The Bare Essentials of Greensock September 2020 | Maggie Appleton Illustrating Gatsby's Key Concepts Problematic Proteins New Harvest & Illustrating the Cultivated Meat Podcast Synecdoche: Drawing the Part for the Whole A Meta-Tour of This Site Douglas, Dirt, and Matter Out of Place The Knowledge Hydrant A Naïve Exploration of Computer-Supported Collaborative Learning Silent Synchronous Reading Sessions What the Fork is React Suspense? Visually Workshopping the AWS Cloud Are Data Unions the Future of Data? Pattern Languages in Programming and Interface Design A Metaphorical Reading Collection
Humanity's Last Exam
Center for AI Safety (CAIS) and Scale AI · 2025-02-20 · via Maggie Appleton

We have a new(ish) Okay, it’s not that new – created in September 2024 – but we’ve only recently seen companies using when they announce new models. benchmark, cutely named “Humanity’s Last Exam.”

If you’re not familiar with benchmarks, they’re how we measure the capabilities of particular AI models like o1 or Claude Sonnet 3.5. Each one is a standardised test designed to check a specific skill set.

For example:

  • MMLU (Massive Multitask Language Understanding) measures understanding across 57 academic subjects including STEM, social science, and the humanities.
  • HumanEval measures code generation skills.
  • GPQA (Graduate-Level Google-Proof Q&A Benchmark) measures correctness on a set of questions written by PhD students and domain experts in biology, physics, and chemistry.

When you run a model on a benchmark it gets a score, which allows us to create leaderboards showing which model is currently the best for that test. To make scoring easy, the answers are usually formatted as multiple choice, true/false, or unit tests for programming tasks.

Among the many problems with using benchmarks as a stand-in for “intelligence” (other than the fact they’re multiple choice standardised tests – do you think that’s a reasonable measure of human capabilities in the real world?), is that our current benchmarks aren’t hard enough.

New models routinely achieve 90%+ on the best ones we have. So there’s a clear need for harder benchmarks to measure model performance against.

Hence, Humanity’s Last Exam .

Made by ScaleAI and the Center for AI Safety, they’ve crowdsourced “the hardest and broadest set of questions ever” by experts across domains. 2,700 questions at the moment, some of which they’re keeping private to prevent future models training on the dataset and memorising answers ahead of time. Questions like this:

Samples of the diverse and challenging questions submitted to Humanity's Last Exam.
Samples of the diverse and challenging questions submitted to Humanity's Last Exam.
Samples of the diverse and challenging questions submitted to Humanity's Last Exam.

So far, it’s doing it’s job well – the highest scoring model is OpenAI’s Deep Research at 26.6%, with other common models like GPT-4o, Grok, and Claude only getting 3-4% correct. Maybe it’ll last a year before we have to design the next “last exam.”

A quick note on benchmarks and sweeping generalisations

When people make sweeping statements like “language models are bullshit machines” or “ChatGPT lies,” it usually tells me they’re not seriously engaged in any kind of AI/ML work or productive discourse in this space.

First, because saying a machine “lies” or “bullshits” implies motivated intent in a social context, which language models don’t have. Models doing statistical pattern matching aren’t purposefully trying to deceive or manipulate their users.

And second, broad generalisations about “AI”‘s correctness, truthfulness, or usefulness is meaningless outside of a specific context. Or rather, a specific model measured on a specific benchmark or reproducible test.

So, next time you hear someone making grand statements about AI capabilities (both critical and overhyped), ask: which model are they talking about? On what benchmark? With what prompting techniques? With what supporting infrastructure around the model? Everything is in the details, and the only way to be a sensible thinker in this space is to learn about the details.