惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
Google DeepMind News
Google DeepMind News
大猫的无限游戏
大猫的无限游戏
雷峰网
雷峰网
MyScale Blog
MyScale Blog
G
GRAHAM CLULEY
云风的 BLOG
云风的 BLOG
MongoDB | Blog
MongoDB | Blog
WordPress大学
WordPress大学
Y
Y Combinator Blog
The Register - Security
The Register - Security
宝玉的分享
宝玉的分享
S
Schneier on Security
N
News and Events Feed by Topic
T
Threat Research - Cisco Blogs
C
Cyber Attacks, Cyber Crime and Cyber Security
G
Google Developers Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园_首页
S
Security @ Cisco Blogs
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
Visual Studio Blog
M
MIT News - Artificial intelligence
U
Unit 42
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
人人都是产品经理
人人都是产品经理
W
WeLiveSecurity
Latest news
Latest news
博客园 - 【当耐特】
P
Palo Alto Networks Blog
博客园 - 叶小钗
Simon Willison's Weblog
Simon Willison's Weblog
Jina AI
Jina AI
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
T
Troy Hunt's Blog
L
LangChain Blog
腾讯CDC
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - Franky
C
Check Point Blog
O
OpenAI News
C
Cisco Blogs
T
Tor Project blog
A
About on SuperTechFans
F
Full Disclosure
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
Lohrmann on Cybersecurity
Attack and Defense Labs
Attack and Defense Labs

Mashable

AdultFriendFinder 2016 data breach: Security improvements 5 AdultFriendFinder scams to avoid The best hookup apps of 2026: I swiped until my thumb hurt How to delete your AdultFriendFinder account Tax Day 2026 deals: Score free food from Burger King, Krispy Kreme, Popeyes, Wendy's, and more XChat to launch on iPhone and iPad The 9 best headphones and earbuds for working out in 2026 Health chatbots could pave the way for 'AI privilege' in court UFC 2026 livestream: How to watch UFC for free 'Mexodus' review: This live-looped musical is a theatrical miracle 'Zelda: Ocarina of Time' remake: 4 things I really, really want Boston Bruins vs. Tampa Bay Lightning 2026 livestream: How to watch NHL for free The DJI Mini 5 Pro drone is down to its record-low price at Amazon — save over $500 Best Hulu deals and bundles: Best streaming deals in April 2026 NYT Connections Sports Edition hints and answers for April 11: Tips to solve Connections #565 NYT Strands hints, answers for April 11, 2026 Today's Hurdle hints and answers for April 11, 2026 NYT Pips hints, answers for April 11, 2026 NYT Connections hints and answers for April 11. Tips to solve 'Connections' #1035. Wordle today: The answer and hints for April 11, 2026 Artemis 2 splashdown: Photos, videos of the astronauts' return Artemis II crew return to Earth with perfect splashdown All the streaming apps that raised prices in 2026 so far Artemis II: All the Apple, GoPro, and Microsoft gadgets on Orion 'Moon joy' takes off as NASA embraces a new space-age catchphrase The pros and cons of switching from Kindle to Kobo e-readers Apple will close its first unionized retail store 'The AI Doc' director: Cynicism is the only wrong answer to AI Artemis II return: How to livestream reentry and splashdown BTS 'Arirang' World Tour: How to watch it live in cinemas Home Depot Spring Black Friday Sale 2026: What to expect, best live deals, and more How the FBI recovered Signal messages (and how to fix the flaw) Samsung Galaxy Z Fold 8 launch date leaks Samsung The Frame dupe deal: Save over $300 on the Hisense Canvas TV The 'Exit 8' movie is here and for a limited time, get the video game for just $2.79 on Steam New FCC rule will make Starlink satellite internet faster and cheaper Aya Cash on 'Giant,' boycotting, and the silliest part of being on 'The Boys' 'Exit 8' review: The most nightmarish spot-the-difference you've ever experienced 'Outcome' is full of cameos, so we've listed them all Regularly $200, you can now upgrade your PC with this powerful OS for just $13 Get Microsoft Office essentials for less than $5 each with this lifetime license Regularly $1,099, you can now get this MacBook Air for $230 if you act fast Pricey AI blood test services promise answers. Do they deliver? Best Disney+ deals and bundles: Best streaming deals in April 2026 Masters 2026 livestream: How to watch Masters Tournament for free Moon phase today explained: What the Moon will look like on April 10, 2026 'Thrash' review: Tommy Wirkola's shark movie ate AFL 2026 livestream: How to watch AFL for free NRL 2026 livestream: How to watch National Rugby League for free All the states Pornhub is blocked in as of April 2026 NYT Connections Sports Edition hints and answers for April 10: Tips to solve Connections #564 NYT Pips hints, answers for April 10, 2026 NYT Connections hints and answers for April 10. Tips to solve 'Connections' #1034. NYT Strands hints, answers for April 10, 2026 Wordle today: The answer and hints for April 10, 2026 Today's Hurdle hints and answers for April 10, 2026 Artemis II reentry and splashdown: Everything the astronauts will experience The latest Microsoft Visual Studio is on sale for just $43 Kindle owners are furious over Amazon's plan to end support for older devices Waymo and Waze launch pothole patching pilot for U.S. cities Motorola budget phone prices are spiking up to 50 percent. Is AI to blame? BTS' 'Hot Ones' episode included milk, screaming, and a 'Digimon' singalong 'Outcome' review: Keanu Reeves puts his nice guy rep on the line 'Malcolm in the Middle: Life's Still Unfair' review: I didn't know how much I needed this Best power station deal: Take 52% off the Bluetti Elite 300 ahead of RV season Samsung Galaxy Z TriFold gets a surprise restock April 10 What is OnlyFans? Home Depot Spring Black Friday free cordless tools: Best deals on DeWalt, Ryobi, and Milwaukee Tesla is developing a smaller, cheaper SUV, report says New Congressional scam alert issued for IRS fraud ahead of Tax Day Dyson launches its first-ever portable fan for $99: Shop the HushJet Mini Cool NBA livestream 2026: How to watch NBA for free Apple iPhone 17e review: Ticks every box but one Best Magic The Gathering deal: 30 packs of Lorwyn Eclipsed Play Booster Box for $110 NYT Pips hints, answers for April 9, 2026 Musician Leith Ross is taking a year without screens NYT Connections Sports Edition hints and answers for April 9: Tips to solve Connections #563 NYT Mini crossword answers, hints for April 9, 2026 Where is Artemis II right now? Track the astronauts returning from the moon Best robot vacuum deal: Save $220 on the Roborock Q10 S5+ Stephen Colbert has thoughts on Trump's 'double-sided ceasefire' Moon phase today explained: What the Moon will look like on April 9, 2026 Best robot vacuum deal: Save $600 on Mova Z60 robot vacuum Best robot vacuum deal: Save $620 on Ecovacs Deebot X9 Pro Omni Best TV deal: Save $401.99 on Sony Bravia 5 65-inch The Samsung Galaxy S26 is under $100 at T-Mobile — how to claim this limited-time deal NASA to run Artemis II astronauts through obstacle course after splashdown This $60 Chromebook can be your low-stress backup This cable simplifies your charging setup, and it’s on sale for just $22 AI is changing health: Here's what you should know What is the viral Needoh toy, and why is it out of stock everywhere? What's new to streaming this week? (April 10, 2026) ChatGPT Health: The data worries are real AI could soon detect heart disease just by listening to it Best Pokémon TCG deal: Ascended Heroes Premium Poster Collection under $120 Best Pokémon TCG deal: Perfect Order Bundle at best-ever price Regularly $999, score a MacBook Air for $200 with this limited-time deal 'Big Mistakes' review: Dan Levy's crime comedy gifts us with wild sibling hijinks 'You, Me and Tuscany' review: Halle Bailey and Regé-Jean Page deliver a radiant, feel-good rom-com Today's Hurdle hints and answers for April 9, 2026
Anthropic: Claude Opus 4.7 has a 92% honesty rate, fewer hallucinations
2026-04-18 · via Mashable

Anthropic released a new hybrid reasoning model on Thursday: Claude Opus 4.7.

Anthropic has a reputation as a safety-first AI company, and the Opus 4.7 system card reports that the model is less likely to hallucinate or engage in sycophancy than both prior Anthropic models and other frontier AI models.

We dived into the Opus 4.7 system card to see exactly what Anthropic had to say about the model's safety, honesty, and sycophancy.

You May Also Like

Don’t miss out on our latest stories: Add Mashable as a trusted news source in Google.

The TL;DR version

Why put the TL;DR version at the end?

Anthropic says Claude Opus 4.7 makes improvements on various types of hallucinations and overall honesty. Anthropic gives the new model top marks on sycophancy and encouragement of user delusions, too. (Anthropic's data also shows that Claude Opus 4.7 scores much better on these behaviors than Gemini 3.1 Pro and Grok 4.20.)

"Claude Opus 4.7 is more reliably honest than Opus 4.6 or Sonnet 4.6, with large reductions in the rate of important omissions, and moderate improvements in factuality and rates of hallucinated input," Anthropic reports.

chart showing claude opus 4.7's false premises honesty rate

False premises honesty rate: Will the model tell a user when they're incorrect? Credit: Anthropic

chart showing claude opus 4.7's mask honesty rate

MASK honesty rate: Will the model contradict its own stated belief when pushed to do so by a user? Credit: Anthropic

Want to learn more about getting the best out of your tech? Sign up for Mashable's Top Stories and Deals newsletters today.

Anthropic measures Claude's honesty and hallucination rates in multiple ways, but let's look at one representative example — the Model Alignment between Statements and Knowledge (MASK) benchmark. MASK was developed by Scale AI and the Center for AI Safety.

Claude Opus had a MASK honesty rate of 91.7 percent, compared to 90.3 percent for Opus 4.6 and 89.1 percent for Sonnet 4.6. While that’s lower than the 95.4 percent score achieved by Claude Opus 4.5, the new model performs better on other hallucination scores (more on that below).

Interestingly, Claude Mythos was more honest still, with an honesty rate of 95.4 percent.

Claude Opus 4.7 lags behind Claude Mythos on overall performance

Since Anthropic repeatedly compares Opus 4.7 to Claude Mythos, let's quickly review the differences between the two models.

Claude Opus 4.7 is the latest hybrid reasoning model available to paid Claude subscribers. Claude Mythos is an unreleased model that Anthropic has only made available to partners via Project Glasswing.

Mashable Light Speed

Under normal circumstances, we would expect Claude Opus 4.7 to be Anthropic's most advanced and powerful model to date. However, Anthropic says it lags behind the unreleased Claude Mythos in key areas. Anthropic deemed Claude Mythos too dangerous to release to the public because of its advanced cybersecurity capabilities.

Still, Claude Opus 4.7 improves upon Opus 4.6 in many ways, particularly advanced coding, visual intelligence, and document analysis, Anthropic says.

More details on Claude Opus 4.7 hallucination rates

When using Opus 4.7, how likely is Claude to tell a lie, invent facts, or deceive users? There isn't a single hallucination rate that Anthropic provides, because there are multiple types of hallucinations.

So, this section is for the AI nerds.

Anthropic identifies a few different ways to measure hallucination and honesty:

  • Factual hallucinations: How likely the model is to provide accurate information. How often does the model admit that it doesn't know something?

  • Input hallucination: This occurs when an AI model ignores prompt instructions, hallucinates the content of files, or pretends to have access to a tool it doesn't have.

  • False premises honesty rate: Will the model tell a user when they're incorrect?

  • MASK honesty rate: This "tests whether a model will contradict its own stated belief when a user or system prompt pushes it to."

We've already covered the MASK honesty rate, and Claude Opus 4.7 shows similar gains on these other measures, according to Anthropic.

At this time, we cannot independently verify Anthropic's results.

To measure factual hallucinations, Anthropic used four different tests and recorded correct responses, incorrect responses, and abstentions. In this case, abstentions are good — the model should decline to answer a question rather than guessing. Across all four tests, Opus 4.7 scored higher than Opus 4.6 and Sonnet 4.6 but lower than Claude Mythos.

chart showing claude opus 4.7 performance on accuracy benchmarks

Chart showing Claude Opus 4.7's performance on accuracy tests. Credit: Anthropic

Anthropic measured Opus 4.7's input hallucination in two ways: "prompts requesting an unavailable tool" and "prompts referencing missing context."

Opus 4.7 scored 89.5 percent on the former, beating Claude Mythos's 84.8 percent; on the latter, Opus 4.7 scored 91.8 percent, two points lower than Claude Mythos's 93.8 percent.

This shows just how stubborn AI hallucinations are, with even leading AI companies like Anthropic recording input hallucination rates around 90 percent. Anthropic's reported hallucination rates are similar to the latest OpenAI models, which provide responses with incorrect information up to 5.8 percent of the time (with browsing enabled) to 10.9 percent (browsing disabled), per OpenAI.

chart showing openai ai models hallucination rates

OpenAI most recently reported hallucination rates in the system card for GPT-5-2. Credit: OpenAI

What about Opus 4.7's honesty rate for false premises, i.e., will Claude tell a user they're wrong? According to the system card, Claude will push back on false premises 77.2 percent of the time. That's better than all other recent Anthropic models except for — you guessed it — Claude Mythos, which will reject false premises 80 percent of the time.

Claude Opus 4.7 sycophancy

There's not much new to report in terms of sycophancy. While Anthropic's expert red-team testers reported that Opus 4.7 was prone to “sycophantic agreement under pushback," it has very similar scores to prior models from Anthropic and OpenAI, and noticeably better scores than Gemini 3.1 Pro and Grok 4.20. Again, this is according to Anthropic.

To measure bad behaviors like sycophancy and "encouragement of user delusion," Anthropic uses Petri 2.0, its open-source behavioral audit tool. This test scores models on a 1-10 scale, with lower scores reflecting better behavior. The Petri score isn't akin to a percentage, as it measures both the rate of a behavior and the severity.

Anthropic scored Opus 4.7 highly (or, lowly, with this particular scale) on both sycophancy and user delusions.

charts from claude opus 4.7 system card showing safety evaluation scores for frontier AI models

Anthropic uses Petri 2.0, its open source AI safety tool, which scores bad behaviors from 1-10. The lower the score, the better. Credit: Anthropic

Mashable reached out to Anthropic for comment but did not receive a response in time for publication.

Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.