惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Recent Announcements
Recent Announcements
雷峰网
雷峰网
The GitHub Blog
The GitHub Blog
罗磊的独立博客
月光博客
月光博客
J
Java Code Geeks
A
About on SuperTechFans
Microsoft Security Blog
Microsoft Security Blog
D
Docker
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
F
Fortinet All Blogs
U
Unit 42
C
Check Point Blog
Martin Fowler
Martin Fowler
有赞技术团队
有赞技术团队
博客园 - 叶小钗
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
酷 壳 – CoolShell
酷 壳 – CoolShell
Blog — PlanetScale
Blog — PlanetScale
大猫的无限游戏
大猫的无限游戏
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
阮一峰的网络日志
阮一峰的网络日志
MyScale Blog
MyScale Blog

WhatIs

Hims & Hers launches AI agent for lab results Twilio revamps, updates customer engagement platform Most patients find appointment scheduling, billing overly complex Teradata's latest targets putting agentic AI into production AHA, Joint Commission launch cyber resilience program Tableau in transition as AI forces BI vendors to evolve California hospitals sue Elevance over out-of-network penalty CMS Health Tech Ecosystem adds electronic prior auth pledge Atlassian MCP updates take aim at AI token usage Leapfrog: Hospitals improved in 17 patient safety measures United promises another 30% cut to prior auths in 2026 AI outperforms docs on clinical reasoning, but not ready for solo work ServiceNow's Autonomous CRM takes aim at Salesforce ServiceNow reintroduces itself as an AI 'security company' New Tableau leader talks vendor's evolution in era of AI Deloitte warns of a "bubble effect" caused by the GLP-1 boom Tableau repositions for AI, unveils new knowledge layer IBM Bob AI coding agent ships, HashiCorp AIOps previewed DOJ forms West Coast Strike Force to stop healthcare fraud Most people benefit from the ACA's free preventive services SAP acquisitions of Dremio, Prior Labs target AI development Bridging the gap: Legacy tools gain enterprise AI support Amazon Connect Talent: AWS enters AI interviewing market AHA, West Health launch health tech adoption initiative How are states preparing for Medicaid work requirements? Medical device security improves, but cyberattacks remain pervasive Weekly news roundup: Musk vs. Altman, Google’s Pentagon AI deal, China and EU hit Meta Skin substitute spending driven by patients, products, prices Clinical AI company Aidoc snags $150M in new funding Qlik's Capone departs after eight years as CEO
General-purpose AI beats out specialized clinical AI in s...
Anuja Vaidya · 2026-06-15 · via WhatIs

A new study challenges the value proposition of specialized clinical AI tools, showing they underperformed compared to general-purpose AI models across medical benchmarks.

After large language models exploded on the scene in late 2022, developers rushed to explore their use in healthcare, creating clinical AI tools for healthcare-specific use cases. But now, a new study reveals that general-purpose AI can outperform specialized clinical AI on several medical benchmarks.

The study, published in nature medicine, tested two specialized LLM-based clinical AI tools, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. The results call into question the industry's focus on designing LLMs specifically for healthcare use cases.

Investment in specialized clinical AI is growing. Earlier this year, OpenEvidence raised $250 million in a closed series D funding round, sending its valuation skyrocketing to $12 billion. Since then, the company has expanded rapidly, releasing audio telehealth, AI coding, prescription and prioritization features.

The study authors noted that, though proprietary clinical AI tools claim to provide enhanced clinical performance over general-purpose AI, their architectures, base models and training pipelines are not publicly available. As a result, providers must assess their value and safety without independent evidence, making it harder for them to challenge the results of clinical AI compared with general-purpose tools.

Thus, researchers from NYU Langone Health and the University of Texas at Austin set out to evaluate the tools against three medical benchmarks.

The evaluation included testing the AI models using three types of assessments: 500 US Medical Licensing Examination-style MedQA questions assessing medical knowledge, 500 HealthBench items evaluating agreement with expert clinicians and 100 real clinical queries drawn from physicians' LLM queries. Twelve clinicians conducted a randomized, blinded review of the RCQ stage.

Model performance varied, with general-purpose AI coming out on top

The general-purpose frontier AI tools outperformed the specialized clinical AI tools in all three evaluations, the study revealed.

In the MedQA questions assessment, Gemini achieved the highest accuracy at 97.4%, followed by GPT at 94.2% and Claude at 90.2%. Meanwhile, OpenEvidence achieved an accuracy of 89.6% and UpToDate achieved 88.4%.

Similarly, GPT scored highest in the HealthBench assessment, receiving a score of 88 on a 100-point scale, followed by Gemini at 79.3 and Claude at 77. Both specialized clinical AI tools scored lower: OpenEvidence at 62.6 and UpToDate at 61.3.

In the RCQ benchmark evaluation, two performance tiers emerged. The first tier, which comprised the general-purpose tools, outperformed the second tier of clinical AI tools on most individual questions, not just on average. The researchers also included Google Search AI Overview in the RCQ evaluation because it is routinely encountered by clinicians. The clinical AI tools performed comparably to the Google Search AI Overview on the RCQ. 

"Clinical AI tools may carry institutional legitimacy and are likely safe for routine use, but our results show that they are not superior to frontier models on knowledge, communication or clinical alignment," the researchers wrote.

However, the researchers are not necessarily arguing that providers only use general-purpose AI tools. Rather, they suggest that providers develop hospital-specific LLMs that leverage institutional data and use them alongside general-purpose models for less-sensitive tasks.

Anuja Vaidya has covered the healthcare industry since 2012. She currently covers healthcare IT and innovation, including artificial intelligence, digital healthcare, EHRs and interoperability.

Dig Deeper on Artificial intelligence in healthcare