惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
M
MIT News - Artificial intelligence
Recent Announcements
Recent Announcements
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Stack Overflow Blog
Stack Overflow Blog
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
罗磊的独立博客
博客园 - 【当耐特】
H
Help Net Security
腾讯CDC
大猫的无限游戏
大猫的无限游戏
GbyAI
GbyAI
Last Week in AI
Last Week in AI
Jina AI
Jina AI
博客园 - 聂微东
Blog — PlanetScale
Blog — PlanetScale
A
About on SuperTechFans
Apple Machine Learning Research
Apple Machine Learning Research
P
Proofpoint News Feed
Y
Y Combinator Blog
C
Check Point Blog
博客园 - 司徒正美
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
Did AI Really Beat ER Doctors At Diagnosis? No, Here’s Wh...
Jesse Pines, · 2026-05-23 · via Forbes - Innovation
Generative AI Apps

AI tools like Claude and ChatGPT are increasingly being tested in clinical settings—but a viral study raised questions about what that really means for diagnosis.

getty

A study published April 30 in the journal Science found that AI was more accurate than doctors in diagnosing cases in the ER.

Within hours of the study’s publication, headlines highlighting the story ricocheted across social media, cable news, and the inboxes of hospital administrators. OpenAI’s o1 model, the coverage incorrectly proclaimed, outperformed the reasoning of emergency physicians to diagnose triage complaints.

For example, the headline published on the National Public Radio website read: In real-world test, an AI model did better than doctors at diagnosing patients.

Many ER physicians took issue with how the findings were characterized by the media. As an emergency physician, I too read the study. To me, what this study actually means is quite interesting but also nuanced.

One of the study’s authors has also since offered some insightful clarification on the study.

MORE FOR YOU

Here’s The Study And What It Actually Found

The experiment presented OpenAI’s o1 and 4o models with the electronic medical records of 76 real patients who had come through the Beth Israel Deaconess emergency department and were admitted to the hospital.

Two internal medicine attending physicians reviewed the same cases. Then two separate internal medicine physicians, blinded to whether the diagnosis came from a human or an AI, evaluated the results.

OpenAI’s o1 model identified the exact or closely related diagnosis in 67% of triage cases, compared to 55% and 50% for the two physicians. AI’s advantage was largest at the first touchpoint, initial triage, where the least information is available. The researchers were careful to note that the AI was given the same raw, unprocessed electronic health record data available at the time of each diagnostic decision.

Yet, the headlines largely missed that the emergency department was just one of six experiments in the paper. The other five drew on more established benchmarks used to evaluate AI diagnostic systems.

Across all six experiments, the results were impressive. But none should be mistaken for proof that AI is ready to diagnose patients independently. Nevertheless, since publication, ER physicians have raised concerns about the study on emergency medicine diagnoses.

First, the doctors in the study weren’t ER doctors. They were internal medicine doctors, who have different training and focus. In addition, the primary goal of emergency medicine is not always about landing on the precise diagnosis. It’s about ruling out life threats, managing uncertainty and moving patients safely through a high-volume, high-stakes environment.

Spend a shift in a busy ER and you will quickly understand why a text-based diagnostic exercise, however well designed, doesn’t capture how real-life emergency medicine works. In the study, the AI read notes. It did not see the patient who appeared ill (or not) in ways that might change the differential diagnosis. It didn’t see the subtle neurological exam finding or notice that the patient’s story shifted between triage and the exam room.

The AI was not practicing emergency medicine. It was offering a written opinion based on selected information.

A Study Author Responds To Critics

In response, one of the paper’s own authors, an emergency physician himself, sees it differently. Dr. Adrian Haimovich, an assistant professor of emergency medicine at Harvard Medical School and an attending physician at Beth Israel Deaconess Medical Center, has offered a different framing.

“Even the toughest cases published in medical journals are now regularly solved by LLMs,” he wrote. “When a patient is admitted to the hospital, they will typically be seen and stabilized by ER doctors who then pass the patient to the internal medicine doctors for the hospital stay. This experiment compares how well LLMs and internal medicine doctors do at guessing the diagnosis of patients admitted to the hospital using only the information that was available in the ER.

Indeed, ERs are messy, real-world clinical environments where reasoning under pressure matters most. He went on to explain, "We restricted the data to the ER because it reflects when the diagnosis is most uncertain and so represents the toughest challenge.”

To Haimovich, the study wasn’t meant to be a head-to-head contest between doctors and machines. The primary finding in his view is that OpenAI’s o1, one of the first true “reasoning” models, can actually perform clinical reasoning across domains.

How We Should Interpret The Study’s Findings

In my view, the study results are quite important. This is why the editors of Science one of the most prestigious peer-reviewed journals, chose to publish it.

The most important finding is not the comparative accuracy. But rather it’s the fact that AI performed so well on messy, real-world, unprocessed clinical data. Prior comparisons of doctors to AI rely on polished case presentations that bear little resemblance to actual emergency care.

The fact that o1 held its own with all the uncertainty is a meaningful signal. Another important consideration: the study data at this point are old, by AI standards. New models have since eclipsed o1, so whatever benchmark o1 set in these experiments, the ceiling has since moved.

The study’s authors were also cautious about what they thought the next step should be: prospective trials. Not deployment. Not replacement of physicians.

How AI Could (Eventually) Play A Role In Real-Life Diagnoses

At this point in mid-2026, the debate over whether AI will play a role in clinical diagnosis is settled. It absolutely will.

Today, ER doctors and other specialists use AI to get second opinions on real cases. In some cases, the AI’s insights prove quite helpful. Given this is true, the more consequential questions surround governance, accountability and integration.

There is currently no formal accountability framework for AI-generated diagnoses. If a patient is harmed based on an AI recommendation that a physician acted on, or failed to act on who is responsible? The physician who acted incorrectly? The hospital who purchased the software? The vendor who created the AI model?

These are questions that will determine whether AI diagnostic tools get adopted thoughtfully, imposed recklessly or at some point get entirely shut down because healthcare as a field is so risk averse. When an incorrect AI diagnosis proves demonstrably lethal to a patient, the system could overreact and hit the kill switch.

Is It An All-Hands-On-Deck Moment?

Haimovich frames the current moment correctly: it’s an all-hands-on-deck in emergency medicine. The question now isn’t whether models are capable. They are. It’s how to make them work in ways that help physicians care for patients and improve the physician experience.

The research pipeline being assembled around this study reflects the kinds of questions that matter. Can AI systems help with reduce medical errors? How accurately can AI navigate disposition decisions? Can AI help double-check that subtle diagnostic findings aren’t missed or help read an equivocal electrocardiogram to make the decision about whether an urgent heart catheterization is needed?

Groups are actively working within specialty organizations like the American College of Emergency Medicine Physicians and the Society for Academic Emergency Medicine to address these questions.

Ultimately, we should interpret the headlines on AI beating doctors skeptically but take the underlying science seriously. So is AI better than ER doctors at diagnosis? The study didn’t ask that question, but it did signal where this technology is heading.