惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Jina AI
Jina AI
Hugging Face - Blog
Hugging Face - Blog
博客园 - 三生石上(FineUI控件)
博客园 - 【当耐特】
大猫的无限游戏
大猫的无限游戏
IT之家
IT之家
宝玉的分享
宝玉的分享
WordPress大学
WordPress大学
有赞技术团队
有赞技术团队
Apple Machine Learning Research
Apple Machine Learning Research
酷 壳 – CoolShell
酷 壳 – CoolShell
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
爱范儿
爱范儿
小众软件
小众软件
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
S
SegmentFault 最新的问题
博客园 - Franky
博客园_首页
T
Tailwind CSS Blog
雷峰网
雷峰网
罗磊的独立博客

WhatIs

Hims & Hers launches AI agent for lab results Twilio revamps, updates customer engagement platform CISA launches critical infrastructure cyber resilience initiative Most patients find appointment scheduling, billing overly complex Teradata's latest targets putting agentic AI into production AHA, Joint Commission launch cyber resilience program Tableau in transition as AI forces BI vendors to evolve California hospitals sue Elevance over out-of-network penalty CMS Health Tech Ecosystem adds electronic prior auth pledge Atlassian MCP updates take aim at AI token usage Leapfrog: Hospitals improved in 17 patient safety measures United promises another 30% cut to prior auths in 2026 ServiceNow's Autonomous CRM takes aim at Salesforce ServiceNow reintroduces itself as an AI 'security company' New Tableau leader talks vendor's evolution in era of AI Deloitte warns of a "bubble effect" caused by the GLP-1 boom Tableau repositions for AI, unveils new knowledge layer IBM Bob AI coding agent ships, HashiCorp AIOps previewed DOJ forms West Coast Strike Force to stop healthcare fraud Most people benefit from the ACA's free preventive services SAP acquisitions of Dremio, Prior Labs target AI development Bridging the gap: Legacy tools gain enterprise AI support Amazon Connect Talent: AWS enters AI interviewing market AHA, West Health launch health tech adoption initiative How are states preparing for Medicaid work requirements? Medical device security improves, but cyberattacks remain pervasive Weekly news roundup: Musk vs. Altman, Google’s Pentagon AI deal, China and EU hit Meta Skin substitute spending driven by patients, products, prices Clinical AI company Aidoc snags $150M in new funding Qlik's Capone departs after eight years as CEO
AI outperforms docs on clinical reasoning, but not ready ...
2026-05-05 · via WhatIs

Anuja Vaidya

By

Published: 05 May 2026

New research shows that a large language model outperformed physicians in various clinical reasoning tasks; however, the study's authors cautioned that the findings do not mean that AI tools are ready to autonomously practice medicine.

The question of whether AI tools can accurately perform clinical reasoning tasks has been top of mind since LLMs exploded onto the healthcare scene in late 2022. Generally, research shows that LLMs' clinical reasoning abilities are improving, but the models still struggle with certain tasks and should remain under human supervision.

However, few studies have compared the clinical reasoning capabilities of advanced LLMs with the baseline performance of human physicians. Thus, researchers from Harvard Medical School and Beth Israel Deaconess Medical Center sought to establish these baselines and assess an LLM's performance against them in a new study published in Science.

The researchers evaluated the clinical reasoning capabilities of the OpenAI o1 series. They compared the AI model's performance against hundreds of physicians across various experiments, including published patient vignettes, evaluations of new emergency room patients, and clinical tasks involving diagnoses and clinical management planning.

Overall, the AI model outperformed physicians across the experiments, including those using real, unstructured clinical data from the EHR in an emergency department. In the ER experiment, the model was presented with patients at various points in their diagnostic journey. They provided the model with information at each stage of the journey, from triage to admission decisions, and asked it to generate likely diagnoses and a treatment plan. Overall, o1 outperformed both ChatGPT-4o and two expert attending physicians, as assessed by two other attending physicians.

In another experiment, researchers used five clinical vignettes to test the AI model's ability to provide next steps in clinical management. Using the mixed-effects model, they found that the o1-preview model scored 41 percentage points higher than GPT-4 alone, 41.9 percentage points higher than physicians using GPT-4 and 48.4 percentage points higher than physicians with conventional resources.  

"Our findings suggest that LLMs have now eclipsed most benchmarks of clinical reasoning," the researchers concluded.

Are humans-in-the-loop still necessary for clinical AI?

In short, the answer is yes.

The researchers noted the study's limitations, including that it examined only six aspects of clinical reasoning, whereas researchers have identified dozens of other tasks that may have a greater impact on actual clinical care and need to be studied.

They also emphasized that the study only assessed text-based performance for both humans and AI. But clinical medicine is multifaceted and involves various non-text inputs, including auditory and visual information. 

"A model might get the top diagnosis right but also suggest unnecessary testing that could expose a patient to harm," Peter Brodeur, a Harvard Medical School clinical fellow in medicine at Beth Israel Deaconess and the study's co-first author, said in a press release. "Humans should be the ultimate baseline when it comes to evaluating performance and safety."

The researchers noted that new testing and research approaches are required as AI models evolve, including new benchmarks, human-computer interaction studies and prospective clinical trials.

"Models are increasingly capable," Brodeur said in the press release. "We used to evaluate models with multiple-choice tests; now they are consistently scoring close to 100%, and we can't track progress anymore because we're already at the ceiling."

Anuja Vaidya has covered the healthcare industry since 2012. She currently covers the virtual healthcare landscape, including telehealth, remote patient monitoring and digital therapeutics.

Dig Deeper on Artificial intelligence in healthcare