惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

小众软件
小众软件
V
Visual Studio Blog
博客园 - 三生石上(FineUI控件)
Last Week in AI
Last Week in AI
Blog — PlanetScale
Blog — PlanetScale
爱范儿
爱范儿
J
Java Code Geeks
A
About on SuperTechFans
F
Fortinet All Blogs
B
Blog
aimingoo的专栏
aimingoo的专栏
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Engineering at Meta
Engineering at Meta
Y
Y Combinator Blog
有赞技术团队
有赞技术团队
G
Google Developers Blog
Apple Machine Learning Research
Apple Machine Learning Research
V
V2EX
博客园_首页
博客园 - 叶小钗
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
D
Docker
云风的 BLOG
云风的 BLOG

Mashable

AdultFriendFinder 2016 data breach: Security improvements 5 AdultFriendFinder scams to avoid The best hookup apps of 2026: I swiped until my thumb hurt How to delete your AdultFriendFinder account Tax Day 2026 deals: Score free food from Burger King, Krispy Kreme, Popeyes, Wendy's, and more XChat to launch on iPhone and iPad The 9 best headphones and earbuds for working out in 2026 Health chatbots could pave the way for 'AI privilege' in court UFC 2026 livestream: How to watch UFC for free 'Mexodus' review: This live-looped musical is a theatrical miracle 'Zelda: Ocarina of Time' remake: 4 things I really, really want Boston Bruins vs. Tampa Bay Lightning 2026 livestream: How to watch NHL for free The DJI Mini 5 Pro drone is down to its record-low price at Amazon — save over $500 Best Hulu deals and bundles: Best streaming deals in April 2026 NYT Connections Sports Edition hints and answers for April 11: Tips to solve Connections #565 NYT Strands hints, answers for April 11, 2026 Today's Hurdle hints and answers for April 11, 2026 NYT Pips hints, answers for April 11, 2026 NYT Connections hints and answers for April 11. Tips to solve 'Connections' #1035. Wordle today: The answer and hints for April 11, 2026 Artemis 2 splashdown: Photos, videos of the astronauts' return Artemis II crew return to Earth with perfect splashdown All the streaming apps that raised prices in 2026 so far Artemis II: All the Apple, GoPro, and Microsoft gadgets on Orion 'Moon joy' takes off as NASA embraces a new space-age catchphrase The pros and cons of switching from Kindle to Kobo e-readers Apple will close its first unionized retail store 'The AI Doc' director: Cynicism is the only wrong answer to AI Artemis II return: How to livestream reentry and splashdown BTS 'Arirang' World Tour: How to watch it live in cinemas
OpenAI’s GPT-5.5 vs Claude Opus 4.7: Which is better?
2026-04-25 · via Mashable

How do the latest models from these AI heavy hitters compare? We take a look at benchmarks, leaderboards, and overall feature set.

 

By

Timothy Beck Werth

 

on 

Share on Facebook Share on Twitter Share on Flipboard

illustration of claude opus 4.7

Credit: Anthropic

OpenAI released its latest model, GPT-5.5, on April 23, just a week after Anthropic introduced Claude Opus 4.7.

As the two leading models from the two leading AI labs, we wanted to see how the new models compare.

Spoiler alert: We think Claude Opus 4.7 has an edge on advanced and agentic coding, but GPT-5.5 performs better on most benchmarks.


You May Also Like



Want to learn more about getting the best out of your tech? Sign up for Mashable's Top Stories and Deals newsletters today.

GPT-5.5 and Opus 4.7: Leaderboards

GPT-5.5 isn't yet ranked on all AI leaderboards yet, but it should be very competitive with Claude Opus 4.7. On the leaderboards of verified benchmark tests such as Arc Prize, GPT-5.5 beats Opus 4.7 (more on this below).

On the popular Arena leaderboard, which is based on user testing, Claude Opus 4.7 Thinking has the top overall spot. Interestingly, Opus 4.7 is currently ranked below Opus 4.6, though that will likely change in time. Currently, new Anthropic models hold the top four overall spots. What's more, Anthropic's unreleased Claude Mythos isn't ranked, and Anthropic says it performs even better than Opus 4.7.

On the Epoch Capabilities Index (ECI) leaderboard, GPT-5.4 Pro has the top score for now. (ECI combines several benchmarks into a single score.) You'll find Gemini 3.1 Pro and GPT-5.4 in the second and third positions.

GPT-5.5 and Opus 4.7: Benchmarks

How do the new models perform on the most common benchmark tests? We have to rely primarily on OpenAI and Anthropic's self-reported scores for these tests. They both achieve high marks, as you'd expect, but GPT-5.5 definitely has the edge.

Here's how they compare on some top AI benchmark tests:

Mashable Light Speed

  • SWE-Bench Pro: GPT-5.5 scored 58.6; Opus 4.7 scored 64.3 percent

  • Terminal-Bench 2.0: GPT-5.5 scored 82.7 percent; Opus 4.7 scored 69.4 percent

  • Humanity's Last Exam: GPT-5.5 scored 40.6 percent; Opus 4.7 scored 31.2 percent*

  • Humanity's Last Exam (with tools): GPT-5.5 scored 52.2 percent; Opus 4.7 scored 54.7 percent

  • BrowseComp: GPT-5.5 scored 84.4 percent; Opus 4.7 scored 79.3 percent

  • GPQA Diamond: GPT-5.5 scored 93.6 percent; Opus 4.7 scored 94.2 percent

  • ARC-AGI-1 (Verified): GPT-5.5 (High) scored 94.5 percent; Claude 4.7 (High) scored 92 percent**

  • ARC-AGI-2 (Verified): GPT-5.5 (High) scored 83.3 percent; Claude 4.7 (High) scored 68.3 percent**

*For Humanity's Last Exam, we're citing Artificial Analysis's verified HLE results. Notably, Anthropic reports that Opus 4.7 scored 46.9 percent on this test.

**See the full results at the Arc Prize website.

GPT 5.5 and Opus 4.7: Availability and pricing

OpenAI says GPT 5.5 is "our smartest and most intuitive to use model yet." Claude Opus 4.7 is Anthropic's most advanced model available to Claude users, though Anthropic says the unreleased Claude Mythos Preview is the more capable model overall.

As such, only paid subscribers can access these frontier models.

GPT 5.5 is only available to OpenAI Plus, Pro, Business, and Enterprise users in ChatGPT and Codex (sorry, ChatGPT Go users). Pro, Business, and Enterprise users can also access GPT-5.5 Pro, while Plus, Pro, Business, and Enterprise customers can access GPT-5.5 Thinking.

OpenAI is raising prices for GPT-5.5 in its API, though the company says it's more token-efficient. API pricing starts at "$5 per 1M input tokens and $30 per 1M output tokens, with a 1M context window."

Opus 4.7 is available to Pro and Max customers; via the API, it's available for "$5 per million input tokens and $25 per million output tokens."

GPT-5.5 and Opus 4.7: Feature set

OpenAI says that GPT-5.5 makes noticeable improvements in "agentic coding, computer use, knowledge work, and early scientific research." Anthropic says Claude Opus 4.7 improves in advanced coding, visual intelligence, and document analysis.

ChatGPT and Claude have similar overall feature sets, though there are some exceptions. Broadly speaking, you can use both of these AI chatbots for research, coding, creative projects, and everyday professional work. You can also use both of the new models in OpenAI and Anthropic's coding platforms, Codex and Claude Code.

It's easier to talk about the differences than the similarities. While GPT-5.5 is not an image model, within ChatGPT you can use the new ChatGPT Images 2.0 model. Anthropic recently rolled out Claude Design, but it only offers data visualizations, graphics, and slides, not full image generation. So, if you need to generate images or interactive graphics for a project, GPT-5.5 will have more tools available to call.

interactive visualization of orion, moon, and sun orbits created by chatgpt

GPT-5.5 can be used to create complex and interactive data visualizations. Credit: OpenAI

ChatGPT has more app and shopping integrations, though thanks to its recent acquisition of OpenClaw, Anthropic has the edge on agentic capabilities.

TL;DR: If we had to pick one of these models for everyday professional work, GPT-5.5 would have the edge thanks to ChatGPT's broader overall feature set. However, for advanced and agentic coding, we'd go with Claude Opus 4.7.


Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.

headshot of timothy beck werth, a handsome journalist with great hair

Tech Editor

Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men’s grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he’s also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl.

Tim studied print journalism at the University of Southern California. He currently splits his time between Brooklyn, NY and Charleston, SC. He's currently working on his second novel, a science-fiction book.

Mashable Potato

These newsletters may contain advertising, deals, or affiliate links. By clicking Subscribe, you confirm you are 16+ and agree to our Terms of Use and Privacy Policy.