惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
MyScale Blog
MyScale Blog
博客园 - Franky
The Cloudflare Blog
IT之家
IT之家
Blog — PlanetScale
Blog — PlanetScale
博客园 - 聂微东
WordPress大学
WordPress大学
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
T
The Blog of Author Tim Ferriss
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
罗磊的独立博客
Google DeepMind News
Google DeepMind News
P
Proofpoint News Feed
Martin Fowler
Martin Fowler
aimingoo的专栏
aimingoo的专栏
J
Java Code Geeks
腾讯CDC
雷峰网
雷峰网
Microsoft Azure Blog
Microsoft Azure Blog
G
Google Developers Blog
博客园 - 【当耐特】
美团技术团队
云风的 BLOG
云风的 BLOG

Mashable

AdultFriendFinder 2016 data breach: Security improvements 5 AdultFriendFinder scams to avoid The best hookup apps of 2026: I swiped until my thumb hurt How to delete your AdultFriendFinder account Tax Day 2026 deals: Score free food from Burger King, Krispy Kreme, Popeyes, Wendy's, and more XChat to launch on iPhone and iPad The 9 best headphones and earbuds for working out in 2026 Health chatbots could pave the way for 'AI privilege' in court UFC 2026 livestream: How to watch UFC for free 'Mexodus' review: This live-looped musical is a theatrical miracle 'Zelda: Ocarina of Time' remake: 4 things I really, really want Boston Bruins vs. Tampa Bay Lightning 2026 livestream: How to watch NHL for free The DJI Mini 5 Pro drone is down to its record-low price at Amazon — save over $500 Best Hulu deals and bundles: Best streaming deals in April 2026 NYT Connections Sports Edition hints and answers for April 11: Tips to solve Connections #565 NYT Strands hints, answers for April 11, 2026 Today's Hurdle hints and answers for April 11, 2026 NYT Pips hints, answers for April 11, 2026 NYT Connections hints and answers for April 11. Tips to solve 'Connections' #1035. Wordle today: The answer and hints for April 11, 2026 Artemis 2 splashdown: Photos, videos of the astronauts' return Artemis II crew return to Earth with perfect splashdown All the streaming apps that raised prices in 2026 so far Artemis II: All the Apple, GoPro, and Microsoft gadgets on Orion 'Moon joy' takes off as NASA embraces a new space-age catchphrase The pros and cons of switching from Kindle to Kobo e-readers Apple will close its first unionized retail store 'The AI Doc' director: Cynicism is the only wrong answer to AI Artemis II return: How to livestream reentry and splashdown BTS 'Arirang' World Tour: How to watch it live in cinemas
Google won't say how often Gemini 3.5 Flash hallucinates
Timothy Beck Werth · 2026-05-29 · via Mashable

Unlike Anthropic and OpenAI, Google won't say how often Gemini hallucinates, lies, or acts sycophantic. That's a big problem.

 By 

Timothy Beck Werth

headshot of timothy beck werth, a handsome journalist with great hair

Timothy Beck Werth

Tech Editor

Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men’s grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he’s also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl.

Read Full Bio

 on 

Share on Facebook Share on Twitter Share on Flipboard

 Google's AI Mode is advertised at the Outernet in Tottenham Court Road in 2025

Would you trust an AI chatbot that's accurate 68 percent of the time? That's not a hypothetical. Credit: John Keeble/Getty Images

As AI gets integrated into every facet of our lives, AI hallucinations remain a stubborn and intractable problem. Yet in the two-hour Google I/O keynote, where Google introduced a massive expansion of AI search and a new default model, Gemini 3.5 Flash, hallucinations didn't warrant a mention.

Likewise, the Gemini 3.5 Flash system card contains no references to hallucinations. Sycophancy is also conspicuously absent. This is especially notable given that both Anthropic and OpenAI publicly report data on metrics such as how often their models hallucinate, encourage delusions, or act sycophantically.

So, as Google makes AI Mode and AI Overviews even more visible in Google Search, users may not realize just how likely an answer is to contain hallucinations and confident mistakes.

Google AI tools do sometimes include warnings such as "AI responses may include mistakes." But there’s no disclosure to searchers that Gemini and AI Mode responses may only be accurate 68.8 to 83.8 percent of the time.

Those are the results from Google's most recent data on Gemini accuracy.

In response to Mashable's questions, a Google spokesperson said that the company plans to publish more information about the newest models' safety evaluations alongside the release of the rest of the Gemini 3.5 model series, which is expected in June.

a disclosure on AI Overviews that states AI sometimes makes mistakes

Credit: Google

How accurate is Gemini, AI Mode, and AI Overviews? It's at the top of a failing class.

Google doesn’t report the honesty, sycophancy, or hallucination rates of its latest models. However, in December, it published a study of their accuracy based on the FACTS Grounding test, a benchmark created by Google DeepMind to measure accuracy.

FACTS "comprehensively evaluates the ability of language models to generate factually accurate text," and Gemini 3 Pro and Gemini 2.5 Pro top this benchmark.

Google reports that Gemini 3 Pro has an overall accuracy score of 68.8. In many classrooms, this would be a hard "F" grade, though it's considered a high score for an AI model.

On the FACTS Search benchmark, which measures a model’s skill at "generating factual responses by interacting with a search tool," Gemini 3 Pro scores 83.8 percent.

table showing AI models scored on the FACTS Grounding benchmark

Credit: Google

The FACTS Search benchmark also measures models' "hedging rate," or how often they decline to answer a question, which is the desired outcome when an answer is unknown. Gemini 3 Pro has a significantly lower "hedging rate" than GPT-5, Claude 4.5 Opus, Claude 4.5 Sonnet, and even its predecessor Gemini 2.5 Pro. 

What does Google say about AI hallucinations?

A single reference to hallucinations does appear in the Gemini 3 Pro system card published on Nov. 18, 2025. "Known Limitations: Gemini 3 Pro may exhibit some of the general limitations of foundation models, such as hallucinations. There may also be occasional slowness or timeout issues."

This boilerplate language is similar to what’s included in the Gemini 2 series system cards, which acknowledge additional problems. “Gemini 2.0 Flash may exhibit some of the general limitations of foundation models, such as hallucinations, and limitations around causal understanding, complex logical deduction, and counterfactual reasoning.” (Emphasis added.)

Hallucinations are actually a feature, not a bug, of the way large-language models work. They're probabilistic algorithms predicting the next token in a sequence. By definition, they're predicting, not "knowing" or "reporting."

Mashable Light Speed

"Hallucinations can only be reduced and never eliminated,” Niranjan Krishnan, Head of AI Solutions, FPT Software, told Mashable. "Large language models are penalized if they sound uncertain or tentative. They don’t know what’s true, but know how to sound true. That bias drives confident errors. Models don’t know their limitations and do not know when to stop."

Krishnan added, "Trying to eliminate hallucinations is the wrong goal. The ultimate challenge is building systems that know when to say, 'I don’t know.'"

“I think users are entitled to that information, especially considering the fact that if you're using an AI chatbot, for example, like Claude or ChatGPT, you're opting into that experience...But when you're on Google, not everyone opts into getting an AI Overview, or engaging with AI mode. They're opening up a search engine that they've always used, and now the experience is different."

So, why doesn't Google report hallucination or sycophancy rates like its chief rivals?

Gary Marcus, scientist, author, and the AI Cassandra of Silicon Valley, told Mashable that "One could guess that their performance there wasn’t groundbreaking or we would have likely heard about it." He added, "Some candor about these things, as with nutrition labels, would certainly be a good thing."

By ignoring AI hallucinations, Google is depriving users of information they could use to evaluate AI output.

Mashable reached out to Google to ask about the lack of hallucination data in the Gemini system cards. In response, a Google spokesperson said, "We take a rigorous approach to defining and measuring persona attributes like helpfulness, tone, and sycophancy. Our goal is to train models to provide objective, direct responses that avoid flattery or simply mirroring a user's views, while keeping the system highly steerable for developers."

The spokesperson also said:

Improving model factuality and managing persona are ongoing, scientific efforts for us. While balancing a model's creativity with factual accuracy remains an industry-wide challenge, hallucination rates have steadily fallen as core model capabilities advance...To continuously guard against incorrect outputs, we invest heavily in robust safety policies, pioneering automated quality-check systems like FunSearch, and open-source evaluation benchmarks like FACTS Grounding to track and improve factual accuracy over time. 

Why does this matter?

Billions of people rely on Google to find information on everything from random celebrity trivia to life-altering medical diagnoses. And Google has long said it looks for expertise, authority, experience, and trustworthiness (or E-E-A-T in Google jargon) for "Your Money or Your Life" (YMYL) topics.

These YMYL topics include anything "that could significantly impact the health, financial stability, or safety of people, or the welfare or well-being of society." Now, users are learning about these topics directly in Google Search or the Gemini app, a tool that's only accurate up to 83.8 percent of the time.

AI hallucinations are also poisoning our collective body of knowledge. Fortune recently reported on a study that found 4,000 AI-fabricated references in nearly 3,000 medical research papers. Likewise, lawyers around the world are being sanctioned for including hallucinated decisions in their briefs. One database tracking legal hallucinations includes 1,497 cases and counting.

Google's AI transformation is also having an outsized impact on the publishers who produce the information that Gemini relies on.

As Google has shifted to AI search, traffic to news websites has fallen off a cliff, a phenomenon that’s been described as a “Traffic Apocalypse” and the "AI armageddon" for publishers.

Once upon a time, back when Google prided itself on its "Don't be evil" ethos, the company defined success by how quickly users left Google. "We may be the only people in the world who can say our goal is to have people leave our website as quickly as possible." Now, Google wants users to spend as much time as possible in its walled garden.

To be clear, all of the actual reporting — the interviews, the research, the photography, the videography, and the old-fashioned sleuthing — is still performed by human journalists. But instead of leaving Google to read about the Iran War in the New York Times, Gemini and AI Mode will brief you right on the search page.

In any other context, journalists call this plagiarism. And as Mashable has reported previously, AI chatbots like Gemini are particularly bad at parsing breaking news, which is when misinformation spreads quickly.

Klaudia Jaźwińska, a journalist and researcher for the Tow Center for Digital Journalism, told Mashable that Google should do more to inform users of its AI's limitations.

“I think users are entitled to that information, especially considering the fact that if you're using an AI chatbot, for example, like Claude or ChatGPT, you're opting into that experience,” Jaźwińska said. “But when you're on Google, not everyone opts into getting an AI Overview or engaging with AI mode. They're opening up a search engine that they've always used, and now the experience is different. And I think for that reason [Google] should be even more transparent about what it can and can't do and what its limitations are.”

In the absence of regulation on AI safety and transparency, Google could commit to publishing data on Gemini's hallucination, sycophacy, or honesty rates, as OpenAI and Anthropic do.

In the meantime, don't forget what Google says in its AI terms of service: "Use discretion before relying on, publishing, or otherwise using content provided by the Services."


Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging that it infringed Ziff Davis copyrights in training and operating its AI systems.

headshot of timothy beck werth, a handsome journalist with great hair

Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men’s grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he’s also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl.

Tim studied print journalism at the University of Southern California. He currently splits his time between Brooklyn, NY and Charleston, SC. He's currently working on his second novel, a science-fiction book.

Mashable Potato