惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
云风的 BLOG
云风的 BLOG
人人都是产品经理
人人都是产品经理
T
The Blog of Author Tim Ferriss
阮一峰的网络日志
阮一峰的网络日志
罗磊的独立博客
J
Java Code Geeks
博客园 - 聂微东
B
Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
WordPress大学
WordPress大学
腾讯CDC
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
Microsoft Azure Blog
Microsoft Azure Blog
D
DataBreaches.Net
The GitHub Blog
The GitHub Blog
美团技术团队
博客园 - Franky
Google DeepMind News
Google DeepMind News
V
V2EX
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
月光博客
月光博客
The Cloudflare Blog

Latest from TechRadar in Chatgpt

‘He needed to have total control over it’ — Altman testifies Musk never trusted shared leadership… I tried a viral 'backwards calendar' ChatGPT prompt — and it completely changed how I plan my week OpenAI snaps up consulting company to help spread the word about AI I asked ChatGPT and Gemini how to make French Toast as good as my mother used to make — one nailed the… OpenAI's ‘Trusted Contact’ feature for ChatGPT users in crisis is a sign AI is becoming something much… ChatGPT now lets you nominate a Trusted Contact who gets alerted if your interaction with AI 'indicates a serious… OpenAI has 3 new AI voice models that the ChatGPT maker says will ‘unlock a new class of voice apps for… AI may kill the app grid, but I still think complex tasks need apps — and that I asked ChatGPT to ruin my photos with ugly 90s-style MS Paint art — and the results are weirdly brilliant Google is bringing Gemini to Mac in a bid to help organize your files and much more ChatGPT just gave me a glimpse of my life in three years ChatGPT just got a major personality overhaul — fewer hallucinations, fewer emojis, much shorter answers, and I can already notice the difference ‘Without me, OpenAI wouldn’t exist,’ says Elon Musk as courtroom clash with Sam Altman turns personal — and exposes a deeper fight over who really built the company behind ChatGPT Sam Altman says some companies are ‘AI washing’ by blaming unrelated layoffs on the technology — but… I used ChatGPT Images 2.0 to meet my childhood self — and this nostalgic photo prompt is going viral for a reason ‘The worst-case situation is where it is a Terminator situation’ — Elon Musk invokes killer robots in… Gen Z hate AI? The Musk vs Altman trial heats up, OpenAI phone rumors buzz and more of the week’s most surprising… OpenAI is making ChatGPT accounts much more secure – including some literal physical security keys I asked ChatGPT to reimagine The Devil Wears Prada 2 ending based on the shocking Runway magazine AI twist in the sequel — and the results aren't as dreadful as you'd think Everyone’s switching from ChatGPT to Claude — but new tests say neither is the smartest free AI, and the… ChatGPT just made it easier to pick the right model, just like Gemini does — here’s when to use Instant,… I noticed ChatGPT slowly drifting off topic in long chats — this tiny prompt forces it to reset itself every few messages and keeps the conversation surprisingly on track Sam Altman just dropped a big hint that GPT-6 is coming soon — ‘with extra goblins’ 'I won’t provide instructions, tactics, or advice that could help someone commit a crime': ChatGPT claims it won't assist would-be felons, despite claims to the contrary from Florida AG ChatGPT just announced it can finally pass the simple ‘how many “r”s in strawberry’ test, but users are still tripping it up by switching to ‘cranberry’ Musk vs Altman heads to trial in a battle that could reshape the future of AI for everyone I swapped my iPhone for a pair of Ray-Ban Meta (2nd gen) on vacation, and I felt liberated — but these smart glasses have a way to go yet I stopped asking AI for answers and started asking for frameworks — and suddenly it all clicked Would you buy a ChatGPT-powered iPhone rival? OpenAI is reportedly developing a smartphone chip, which teases the… I compared ChatGPT Images 2.0 and Google’s Nano Banana 2 using real-world prompts — from portraits to product shots — and the AI image generator that came out on top genuinely surprised me
AIs like ChatGPT fall apart in classic 'Stroop' psycholog...
Darren Allan · 2026-06-05 · via Latest from TechRadar in Chatgpt
A robot standing thoughtfully in front of a giant digital display with code on it
(Image credit: Getty Images)

  • A new study tasked AIs with tackling the 'Stroop' test
  • GPT and Claude performed very poorly compared to humans
  • There are nuances here, but broadly, the researchers argue that improving this side of AIs is crucial for achieving artificial general intelligence

A freshly published study has pointed out a limitation of big-name AI models such as ChatGPT, albeit causing some controversy as the primary piece of research uses now outdated versions of those models – but there are nuances therein, and this doesn't make the findings irrelevant.

I'll go into that more shortly, but first, let's look at the study itself, which was highlighted on Reddit ('New study reveals top AI models completely fail the classic 'Stroop' psychological attention test') and published via the Oxford University Press in the journal PNAS Nexus.

The research consists of testing the so-called 'Stroop effect' with GPT-4o and Claude 3.5 Sonnet. As noted, these aren't the cutting-edge versions of those AIs (Large Language Models, or LLMs) – but they were at the time the initial study was carried out.

The Stroop effect refers to the phenomenon whereby the human brain gets confused when asked to name the color of the ink used to write a word, when that word can be the written version of another (incongruent) color in some cases. So, if the word 'red' is written in blue ink, that'll cause a slower response – or possibly a wrong response, where the viewer will accidentally say "red" rather than the actual color of the ink, which is blue.

This is because the brain is trying to juggle two different tasks – reading comprehension and color recognition – and so cognitive interference arises. Overriding the compulsion to read the word and say the color instead requires "executive control of attention," and this is what the authors were testing in the AI models. Both color-naming and word-reading were tested in shorter and longer lists of words (5, 10, 20, and 40 words).

The study observes: "Like humans, both LLMs [GPT and Claude] showed relatively high accuracy on the word-reading task and performed worse in the incongruent condition [where the word doesn't match the color] than in the congruent and neutral conditions for the color-naming task."

For color naming, humans maintain around 95% accuracy even in very long tests (up to an hour), but LLMs' accuracy declined very swiftly with longer word lists under the incongruent condition (mismatched color and word name). GPT-4o was 91% accurate in a five-word test, but dropped off to 57% with 10 words, and fell away completely to 22% with 20 words (and was only 15% accurate at 40 words).

Sign up for breaking news, reviews, opinion, top tech deals, and more.

Claude 3.5 Sonnet did better, staying 76% accurate at 20 words, but again fell hopelessly to 24% in the longest test of 40 words.

The authors conclude: "The significant degradation pattern of the two LLMs suggests fundamental limitations compared with human attention."


Analysis: another necessary step on the path to AGI?

An AI face in profile against a digital background.

(Image credit: Shutterstock / Ryzhi)

If you've scanned through the Reddit thread, you doubtless noticed that, as mentioned at the outset, there's a lot of flak fired at this study by commenters due to the usage of outdated models of GPT and Claude.

Indeed, these older LLMs are called "state of the art" at one point by the authors – and of course, as already noted, they were cutting-edge when the main study was conducted. Still, this is unfortunate phrasing that should've been updated and tweaked now that the paper has just been published (after peer review and so forth).

However, the researchers did conduct tests on GPT-5, Claude Opus 4.1, and Gemini 2.5 Pro in September 2025, although this is somewhat buried in the paper. That more recent testing found that these models offered only "slight" improvements on their predecessors, and that they still exhibited "ongoing executive attention deficiencies, consistent with our comprehensive analysis of earlier transformer models" (as did Gemini 2.5 Pro, which was a new introduction here).

Granted, a smaller sample size was used, but the researchers still argue that overall, their study reflects a fundamental limitation which is "inherent to the architectural constraints of transformer-based LLMs".

The authors note that a caveat is that GPT-5 in 'Thinking' mode can write and then run code to ensure it performs the Stroop test flawlessly – and similar functionality can be utilized by other LLMs – but this is essentially the AI (cleverly) fudging around its inadequacies. It isn't changing the way it works or reasons more broadly, of course.

The researchers note that transformer architecture innovations for LLMs are focused on enhancing memory capabilities, which fail to address the "core limitations of attention mechanisms, specifically the need for sophisticated alerting, orienting, and executive control networks to enable cognitive flexibility."

The ultimate aim is effective goal-directed behavior, and the study observes: "Future [LLM] development might benefit from implementing more sophisticated executive control systems that can handle decision conflicts through structured, goal-directed processing rather than relying solely on enhanced memory capabilities."

The authors argue that "incorporating executive control mechanisms akin to those in biological attention is crucial for achieving artificial general intelligence [AGI]."


Google logo on a black background next to text reading 'Click to follow TechRadar'

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.


An Apple MacBook Air against a white background

Darren is a freelancer writing news and features for TechRadar (and occasionally T3) across a broad range of computing topics including CPUs, GPUs, various other hardware, VPNs, antivirus and more. He has written about tech for the best part of three decades, and writes books in his spare time (his debut novel - 'I Know What You Did Last Supper' - was published by Hachette UK in 2013).