惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
P
Proofpoint News Feed
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
IT之家
IT之家
Microsoft Security Blog
Microsoft Security Blog
F
Fortinet All Blogs
人人都是产品经理
人人都是产品经理
博客园 - 叶小钗
C
Check Point Blog
Microsoft Azure Blog
Microsoft Azure Blog
aimingoo的专栏
aimingoo的专栏
月光博客
月光博客
美团技术团队
D
Docker
博客园 - Franky
Y
Y Combinator Blog
大猫的无限游戏
大猫的无限游戏
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 【当耐特】
罗磊的独立博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

School of Computer Science News

Robotics Innovation Center Earns LEED Platinum Certification for Sustainable Construction AI4MiddleSchools Expands Nationwide Effort To Prepare Students for an AI-Powered Future Carvalho Earns NSF CAREER Award To Study Motivation and Learning Season Three of 'Does Compute' Now Available Rare Ventures Partners Rings NYSE Opening Bell Bringing Images to Life Through Touch - Robotics Institute Carnegie Mellon University Fried Receives NSF CAREER Award - Language Technologies Institute - School of Computer Science - Carnegie Mellon University PAIR Helps Students Find Their Place in AI Research Navigating the AI Era with a CMU Focus on Critical Thinking Kaess Named to Inaugural Chief of Naval Research Fellows Program Navigating the Moon Koedinger Wins Lifetime Achievement Award Carnegie Mellon Names Damion Shelton Associate VP and Executive Director of the Swartz Center for Entrepreneurship Hong Shen Discusses AI Safety at WEF Annual Meeting Erickson Earns NSF CAREER Award - Robotics Institute Carnegie Mellon University Satya Honored With Test of Time Award From Proof to Program: CMU and the Rise of AI-Driven Mathematics SCS Researchers Named to Inaugural ACM SIGSOFT Software Engineering Academy You Can't Remove Humans From Software Engineering Designing the Future of Tech Governance AI, Single-Cell Technology Reveal How 3D Genome Differs in People With Alzheimer's Disease Tepper School of Business and School of Computer Science Partner to Launch AI for Business Executive Education Program Carnegie Mellon Researchers Lead Three DOE Genesis Mission Awards to Advance the Future of AI-Enabled Scientific Discovery Snake Robots Support Earthquake Search and Rescue in Venezuela Lindlbauer Receives NSF CAREER Award for Adaptive Extended Reality Interfaces Fredrikson Earns Test of Time Award for AI Security CMU Advances Defense Manufacturing and Military Education at Pennsylvania Defense and Innovation Summit Looking Ahead: AI Needs UI Liu Receives NSF CAREER Award Carnegie Foundry, Carnegie Mellon and American Drone Manufacturers Launch Initiative to Supercharge America
Healthcare Blind Spots: AI Models Prone To Fabricating Di...
Mallory Lindahl · 2026-07-31 · via School of Computer Science News
Warning: You are viewing this site with an outdated/unsupported browser. Please update your browser or consider using a different one in order to view this site without issue.
For a list of browsers that this site supports, see our Supported Browsers page.
Skip to content

Healthcare Blind Spots: AI Models Prone To Fabricating Diagnoses

CMU Study Shows LLMs May Invent Information When Missing Key Data

07/20/2026    Amanda Sapio

The Breakdown

  • In 18% of cases, Claude, GPT-5 and Gemini fabricated medical diagnoses, despite supporting images being omitted.
  • The AI models used demographic-based clinical assumptions to invent these diagnoses.
  • CMU research emphasizes the need for demographic sensitivity testing and verification before deploying AI tools in medical systems.

* * *

A man in a black shirt sitting on a bench

Siddharth Vohra is in the Master of Science in Computer Vision program at the Robotics Institute

A Carnegie Mellon University School of Computer Science student recently showed that large language models (LLMs) intermittently invent false information when responding to medical questions. 

For his study, Robotics Institute master’s student Siddharth Vohra asked popular LLMs to describe a medical image that was intentionally omitted from the query. Rather than requesting the missing image, the models fabricated a diagnosis based on the user’s age, gender and race 18% of the time. 

“A 65-year-old white man asking Claude about a skin mole receives melanoma in nearly every response,” said Vohra, the study’s sole author. “When chest X-ray questions are presented, OpenAI’s GPT-5 names sarcoidosis for roughly 77% of young Black patients.”

In reality, fewer than one in 10,000 moles will become melanoma, according to the Memorial Sloan Kettering Cancer Center. The prevalence of sarcoidosis varies throughout the world, but the Cleveland Clinic reports that there are typically fewer than 200,000 cases of sarcoidosis at any given time in the U.S., making it quite rare.

Vohra was raised by two physician parents who exposed him to the healthcare industry at an early age. After reading a research paper detailing how visual language systems often provide incorrect responses, he decided to investigate how demographic information influences the way models respond to medical questions.  

“I analyzed close to 11,700 model responses across Claude, GPT-5 and Gemini,” Vohra said. “About 82% of the time, the models refused to provide a response because no image was attached. However, in the remaining 18%, the models invented a diagnosis instead of asking for the missing image.”

When conducting the study, Vohra focused on chest X-ray, brain MRI and dermatology images in 12 simulated patient profiles. The findings reinforced a critical concern Vohra sees across the AI industry: users often assume that models understand more than they actually do.

“There’s a general perception that AI models are smart because they perform well on certain tasks,” Vohra said. “But being good at one thing doesn’t mean they’re good at everything. AI is used for a host of tasks, from coding to healthcare advice, but different domains require different levels of reliability.”

Because small changes in the words used to describe a patient can drastically influence a model’s conclusions, Vohra stressed that clinical pipelines should audit and test for demographic sensitivity before deploying AI models in healthcare settings.

“These models are incredibly useful and they’re improving quickly,” Vohra said. “As labs continue advancing the models, stringent testing and verification checks are necessary, especially when relying on the models for medical decisions. Strong performance on one task doesn’t guarantee safety in another.” 

Vohra’s current focus includes running the study at a much larger scale, with the goal of surfacing these failure patterns so the labs building these systems can identify and address them. He presented his research at the TrustVLM workshop at the ACM International Conference on Multimedia Retrieval in Amsterdam this past June. He was also recently selected for the Google Gemini Academic Program Award 2026, which promotes trustworthy AI models and mission-critical AI deployments.

For more on Vohra’s work, visit his website.

For More Information: Aaron Aupperlee | 412-268-9068 | aaupperlee@cmu.edu

2026-07-20T10:45:26-04:00

Share This Story!