惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
F
Fortinet All Blogs
小众软件
小众软件
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Microsoft Azure Blog
Microsoft Azure Blog
J
Java Code Geeks
有赞技术团队
有赞技术团队
D
DataBreaches.Net
Hugging Face - Blog
Hugging Face - Blog
V
Visual Studio Blog
A
About on SuperTechFans
I
InfoQ
The GitHub Blog
The GitHub Blog
Engineering at Meta
Engineering at Meta
雷峰网
雷峰网
H
Hackread – Cybersecurity News, Data Breaches, AI and More
罗磊的独立博客
C
Check Point Blog
大猫的无限游戏
大猫的无限游戏
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
酷 壳 – CoolShell
酷 壳 – CoolShell
MyScale Blog
MyScale Blog

WhatIs

Hims & Hers launches AI agent for lab results Twilio revamps, updates customer engagement platform Most patients find appointment scheduling, billing overly complex Teradata's latest targets putting agentic AI into production AHA, Joint Commission launch cyber resilience program Tableau in transition as AI forces BI vendors to evolve California hospitals sue Elevance over out-of-network penalty CMS Health Tech Ecosystem adds electronic prior auth pledge Atlassian MCP updates take aim at AI token usage Leapfrog: Hospitals improved in 17 patient safety measures United promises another 30% cut to prior auths in 2026 AI outperforms docs on clinical reasoning, but not ready for solo work ServiceNow's Autonomous CRM takes aim at Salesforce ServiceNow reintroduces itself as an AI 'security company' New Tableau leader talks vendor's evolution in era of AI Deloitte warns of a "bubble effect" caused by the GLP-1 boom Tableau repositions for AI, unveils new knowledge layer IBM Bob AI coding agent ships, HashiCorp AIOps previewed DOJ forms West Coast Strike Force to stop healthcare fraud Most people benefit from the ACA's free preventive services SAP acquisitions of Dremio, Prior Labs target AI development Bridging the gap: Legacy tools gain enterprise AI support Amazon Connect Talent: AWS enters AI interviewing market AHA, West Health launch health tech adoption initiative How are states preparing for Medicaid work requirements? Medical device security improves, but cyberattacks remain pervasive Weekly news roundup: Musk vs. Altman, Google’s Pentagon AI deal, China and EU hit Meta Skin substitute spending driven by patients, products, prices Clinical AI company Aidoc snags $150M in new funding Qlik's Capone departs after eight years as CEO
General-purpose AI beats out specialized clinical AI in s...
Anuja Vaidya · 2026-06-15 · via WhatIs

A new study challenges the value proposition of specialized clinical AI tools, showing they underperformed compared to general-purpose AI models across medical benchmarks.

After large language models exploded on the scene in late 2022, developers rushed to explore their use in healthcare, creating clinical AI tools for healthcare-specific use cases. But now, a new study reveals that general-purpose AI can outperform specialized clinical AI on several medical benchmarks.

The study, published in nature medicine, tested two specialized LLM-based clinical AI tools, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. The results call into question the industry's focus on designing LLMs specifically for healthcare use cases.

Investment in specialized clinical AI is growing. Earlier this year, OpenEvidence raised $250 million in a closed series D funding round, sending its valuation skyrocketing to $12 billion. Since then, the company has expanded rapidly, releasing audio telehealth, AI coding, prescription and prioritization features.

The study authors noted that, though proprietary clinical AI tools claim to provide enhanced clinical performance over general-purpose AI, their architectures, base models and training pipelines are not publicly available. As a result, providers must assess their value and safety without independent evidence, making it harder for them to challenge the results of clinical AI compared with general-purpose tools.

Thus, researchers from NYU Langone Health and the University of Texas at Austin set out to evaluate the tools against three medical benchmarks.

The evaluation included testing the AI models using three types of assessments: 500 US Medical Licensing Examination-style MedQA questions assessing medical knowledge, 500 HealthBench items evaluating agreement with expert clinicians and 100 real clinical queries drawn from physicians' LLM queries. Twelve clinicians conducted a randomized, blinded review of the RCQ stage.

Model performance varied, with general-purpose AI coming out on top

The general-purpose frontier AI tools outperformed the specialized clinical AI tools in all three evaluations, the study revealed.

In the MedQA questions assessment, Gemini achieved the highest accuracy at 97.4%, followed by GPT at 94.2% and Claude at 90.2%. Meanwhile, OpenEvidence achieved an accuracy of 89.6% and UpToDate achieved 88.4%.

Similarly, GPT scored highest in the HealthBench assessment, receiving a score of 88 on a 100-point scale, followed by Gemini at 79.3 and Claude at 77. Both specialized clinical AI tools scored lower: OpenEvidence at 62.6 and UpToDate at 61.3.

In the RCQ benchmark evaluation, two performance tiers emerged. The first tier, which comprised the general-purpose tools, outperformed the second tier of clinical AI tools on most individual questions, not just on average. The researchers also included Google Search AI Overview in the RCQ evaluation because it is routinely encountered by clinicians. The clinical AI tools performed comparably to the Google Search AI Overview on the RCQ. 

"Clinical AI tools may carry institutional legitimacy and are likely safe for routine use, but our results show that they are not superior to frontier models on knowledge, communication or clinical alignment," the researchers wrote.

However, the researchers are not necessarily arguing that providers only use general-purpose AI tools. Rather, they suggest that providers develop hospital-specific LLMs that leverage institutional data and use them alongside general-purpose models for less-sensitive tasks.

Anuja Vaidya has covered the healthcare industry since 2012. She currently covers healthcare IT and innovation, including artificial intelligence, digital healthcare, EHRs and interoperability.

Dig Deeper on Artificial intelligence in healthcare