惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 最新话题
A
Arctic Wolf
I
Intezer
V
Vulnerabilities – Threatpost
C
Cisco Blogs
MyScale Blog
MyScale Blog
NISL@THU
NISL@THU
Y
Y Combinator Blog
C
CERT Recently Published Vulnerability Notes
P
Privacy International News Feed
H
Hackread – Cybersecurity News, Data Breaches, AI and More
酷 壳 – CoolShell
酷 壳 – CoolShell
Recorded Future
Recorded Future
云风的 BLOG
云风的 BLOG
S
SegmentFault 最新的问题
Microsoft Security Blog
Microsoft Security Blog
L
LangChain Blog
博客园 - 聂微东
博客园 - 叶小钗
F
Fortinet All Blogs
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Recent Announcements
Recent Announcements
C
Cyber Attacks, Cyber Crime and Cyber Security
Latest news
Latest news
Simon Willison's Weblog
Simon Willison's Weblog
P
Palo Alto Networks Blog
S
Schneier on Security
C
Cybersecurity and Infrastructure Security Agency CISA
V
V2EX
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Hacker News
The Hacker News
博客园 - 司徒正美
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
LINUX DO - 热门话题
罗磊的独立博客
K
Kaspersky official blog
Last Week in AI
Last Week in AI
Know Your Adversary
Know Your Adversary
小众软件
小众软件
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
D
DataBreaches.Net
Scott Helme
Scott Helme
P
Proofpoint News Feed
P
Privacy & Cybersecurity Law Blog
P
Proofpoint News Feed
博客园 - 三生石上(FineUI控件)
Hugging Face - Blog
Hugging Face - Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
F
Full Disclosure

Forbes - Healthcare

How to Prevent Domestic Violence Deaths UK Smoking Ban Highlights Debate Over The Proper Function Of Government What To Do When Someone You Love Has Cancer Psychedelic Medicine Goes Mainstream: Breakthrough or Bubble? Humana Profits Eclipse $1 Billion As Medicare Costs Ease Slightly What Are Peptides And Why Is Everyone Talking About Them? Tonsillectomy Doesn’t Lead To Illness, But Tonsillitis Just Might Does Retail Pharmacy Have A Tower Records Problem? Precision Radiation Therapy Could Offer New Hope For Hard-To-Treat Cancers Centene’s Obamacare Enrollment Drops By 2 Million After Congress Strips Subsidies RFK Jr.’s Messaging Could Be Impacting Food And Pharmaceutical Choices Over A Million Road Crash Deaths Annually Prompt $350 Million Investment Breast Cancer Screening Tool Avoids Radiation, Compression, Contrast Large Study Finds Benefits Of Doula Care On Postpartum Outcomes TrumpRx Has Signed Deals With Nearly Every Major Drugmaker. Are Prices Actually Falling? America Can’t Lower Healthcare Costs Without A Moonshot Trump’s Orders Elevate The Medical Status Of Psychedelics And Cannabis Mark Cuban’s Cost Plus Drugs, Humana Partner To Take On Employer Drug Costs Cell, Gene And Specialty Drug Costs Intensify For Health Plans U.S. Tennis Participation Continues Growth. Up 54 Percent Since 2019 New AMA Study Finds Burnout Is Decreasing Among Medical Residents And Fellows Daytime Naps May Be A Sign Of Serious Health Problems, Study Reveals New Antibody Drugs Target Disease From Within Concierge Medicine Was Built For The Few. Here’s How To Open It To The Many Burnout in Medicine Is Still Prevalent, With Emergency Medicine Leading Who Is Actually Qualified To Give Advice On Peptides And Who Isn’t What the 49ers Can Teach Leaders About Handling False And Misleading Narratives Do Older Adults Need Routine Colonoscopies Or Low Thyroid Drugs? Your Period, Your Proteins, Your Health Doctors Say Hegseth’s Flu Vaccine Decision Will Weaken Military Readiness Where Bullets Fly, Malaria Kills Using AI To Personalize Healthcare–Without Losing Patient Trust Progress For Preeclampsia Allowing Our Military To Refuse Flu Vaccination Is A Bad Idea. Here’s Why Can Vaccine Development Weather Political Storms? A Virus From Farmed Seafood Is Causing A New Eye Disease In People Elevance Health Profits Eclipse $1.7 Billion Despite Elevated Costs The UK Passes A Lifetime Smoking Ban. Could America Be Next? There's No Such Thing As Brain Honey UnitedHealth Group Profits Eclipse $6 Billion As Medical Costs Ease AI Is Already Here. The Real Risk In Public Health Is Sitting It Out UnitedHealthcare Reduces Need For Prior Approvals For Patients In Rural America Why No Child Should Have To Sacrifice School To Care For Their Family Oscar Health Launches Consumer Marketplace For Insurance Beyond Its Own Calling The Iconic 867-5309 Now Goes To A Cancer Helpline FDA Lists Xanax Recall. Here’s What You Need To Know What Trump’s Ibogaine Executive Order Means For Veterans With PTSD 20 Years Of Priority Review Vouchers, A Tool For Spurring Needed Drugs Rotavirus Is Surging Across The US — Here’s What Parents Need To Know Leadership Dysfunctional In Healthcare: “Split The Baby” Thinking ‘Bedtime Stacking’ Trends On TikTok. Here Are The Risks Why Do Weight Loss Drugs Work For Some And Not Others? It’s In The Genes Hospital Safety: How to Avoid Medical Errors and Protect Yourself Medicare Can Save $4 Billion On Four Cancer Drugs — Can You Guess Which Ones? After 25 Years Of Consumer-Directed Healthcare, What’s Missing? This Sam Altman-Backed $1.8 Billion Startup Bets AI Can Get Drugs Through Clinical Trials Faster RFK Jr. Pushes To Expand Access To Peptides. A Doctor Explains The Risks How The Trump Administration Is Blocking Access To Home Care Genome Sequencing Solves Rare Disease Mysteries Breakthrough HIV Drug Is Out Of Reach For Many Who Need It Most New Drug Protects Against Life-Threatening Pancreatitis This Pill May Help Pancreatic Cancer Patients Live Longer What Should We Do When The Patient Is Racist? Attention Turns To UnitedHealth Earnings For Signs Of Insurer Rebound New Pancreatic Cancer Drug Nearly Doubles Survival. Here’s What Patients Should Know Why Sex Exists A Novel Approach To The Treatment Of Antibiotic Resistant Infections Democrat-Leaning Plan Takes Aim At Health Plans With New Regulations Trump Administration Weighs Default Medicare Advantage Plans For Seniors An AI System Passed Peer Review. The Scientific Community Isn’t Ready Prior Authorization Reform Is Here — And It Could Change How Millions Get Care The More We Add To U.S. Healthcare, The Worse It Gets How Two Sisters Built A $1 Billion HealthTech Unicorn CDC Delays Reporting Of COVID-19 Vaccine Benefits—Here’s What To Know This Startup Wants To Use AI To Help Digitize History Are Nicotine Pouches Like Zyn And VELO Safe To Use? A Doctor Answers America’s Healthcare Innovation Problem GLP-1 Weight Loss Drugs Are Easy To Get—But Are They Safe? Why Cleveland Clinic Chose This AI Startup To Rewire Key Healthcare Operations Upset About The High Price Of Your Hospital Stay? Medicaid Cuts Might Be To Blame Trump’s New Pharmaceutical Tariffs Will Hit Small Drugmakers Hardest A New Way To Target Metastatic Cancer What A Florida Birth Case Reveals About Post-Dobbs Maternal Healthcare 5 Reasons Why the Medicare Program Can’t Go Broke Lowering Healthcare Costs Without A Disastrous Government-Run Model Promising Study Links Coffee Consumption To Reduced Dementia Risk Gene Regulation May Control How Long We Live Health Insurers Get 2.5% Medicare Rate Hike They Feared Would Be Flat Engineered Antibodies Pry Apart The Most Difficult Viruses Centene Latest Health Insurer To Shakeup Management Ranks 1.6 Million Teens Are Vaping. Health Risks Are Worse Than You Think Increasing Burdens Medical Debt And Bankruptcy Are Uniquely American Medicaid Work Requirements Go Live Soon. Here’s How Many Could Lose Coverage What SpaceX’s IPO Means For The Space Economy Thus Far, Most Favored Nation Drug Prices Have Had Little Impact FDA Approves New Oral Weight Loss Pill Foundayo — Here’s What To Know ‘Medicare By Choice’ Plans Could Work, But More Details Needed Criticism of NFL's Rooney Rule Misses How Hiring Actually Works Navigating Health In The Age Of Misinformation NASA Artemis II astronaut health risks explained
How The ARISE Network Is Rethinking Clinical AI
Spencer Dorn · 2026-05-20 · via Forbes - Healthcare
Merge arrows infographic

The Arise Network aims to understand and explain what AI can do in healthcare.

getty

You’ve seen the headlines: AI aces the medical boards. AI outperforms expert physicians. But what does this actually mean? And how do we evaluate technology that’s advancing faster than we can fully make sense of it?

The AI Research and Science Evaluation (ARISE) Healthcare Network was formed to help answer these questions. Spanning multiple medical centers and led by physicians at Harvard and Stanford with diverse and complementary backgrounds, ARISE is trying to understand what AI systems can do in medicine and how we can evaluate and explain their performance.

They are working to define what holds up in real-world medicine, what we mean by clinical reasoning, how clinicians and AI should work together, when either may perform better alone, and how we might recognize if AI approaches “medical superintelligence.”

The Physician Data Scientist And AI Magic Tricks

Physician, magician, and data scientist Jonathan H. Chen is working to demystify clinical AI.

ARISE Network

Arthur C. Clarke famously wrote, “Any sufficiently advanced technology is indistinguishable from magic.”

Decades later, many people see A.I. as magic. So, who better to spot a magic trick than Jonathan H. Chen, a physician, data scientist, and performing magician?

Chen’s path is not typical. He started college at 13 and worked as a software engineer before returning to school to earn an MD and PhD in computer science and then training in internal medicine. Since joining the Stanford faculty in 2017, he’s been evaluating how AI applies to medical problems.

He points out that the first rule of magic is (mis)directing the audience to look where you want them to look. So, when LLMs like ChatGPT arrived, he knew to look in the other direction to understand what they’re doing, where they fail, and how clinicians might use them.

A core theme of his research is understanding how physicians and AI can best work together.

In late 2024, his team made headlines after finding that, on diagnostic reasoning tasks, LLMs alone outperformed both physicians using AI and physicians working alone. This ran counter to the long-held “fundamental theorem” of informatics that physicians plus AI will outperform either alone.

Part of the explanation was timing. It was still early, and many physicians used LLMs like search engines.

So, in a follow-up trial, the team tested a customized LLM tailored for clinical collaboration that taught clinicians in real time how to use it. This time, physician-plus-AI outperformed physicians alone, while matching—but still not surpassing—AI alone in diagnostic reasoning.

The group later reported similar results in another study on management reasoning tasks. Through a new ARPA-H grant, ARISE is now building a “flight simulator” for medicine to study and improve how clinicians and AI work together.

Taken together, these findings raise a deeper question: if AI alone sometimes outperforms physicians working with AI on reasoning tasks, what exactly are we measuring when we talk about “clinical reasoning” in the first place?

The Physician Historian And The Nature of Reasoning

A medical historian, Dr. Adam Rodman draws on the past to understand the present and predict the future.

Danielle Duffey

Adam Rodman, a fast-talking and even faster-thinking Harvard internist, medical historian, and clinical educator, has spent the past two decades studying clinical reasoning and decision-making.

The first thing he will tell you is that none of this is new. Technology has always changed what it means to be a doctor. Think of the stethoscope, anesthesia, penicillin, MRI, and the electronic health record.

What’s different this time is that AI moves up the cognitive stack, shifting knowledge and even thinking to machines. Yet there’s also a long history behind this work, and clinical reasoning may not be what medicine portrays it to be.

Rodman points out that modern ideas about both clinical reasoning and AI surprisingly share common roots in World War II-era signal detection theory, which gave rise to frameworks such as sensitivity, specificity, and ROC curves.

Building on this tradition, pioneers such as Robert Ledley and Lee Lusted argued in a landmark 1959 article that medical decision-making could be understood through logic, probability, and value theory. Their work laid the groundwork for a series of computerized diagnostic tools like INTERNIST-1 and Isabel that sought to model the clinical reasoning of expert physicians.

Rodman believes that computer science, in turn, shaped how medical schools teach clinical reasoning, using frameworks such as Fagan’s nomogram, pretest probabilities, and rule-based heuristics. While these approaches are useful for teaching and assessing trainees, he believes they may not fully capture how experts actually practice.

In his words, “We train doctors in ways that reflect the appearance of expertise, based on cognitive models and computer-era abstractions, rather than how real experts behave, which is often fast, intuitive, and non-linear. And we are now building AI systems that mimic that same abstraction.”

These same traditions shaped how medical AI systems came to be evaluated, often using complex clinical case vignettes drawn from the New England Journal of Medicine clinicopathological case conference series.

Following this tradition, when LLMs emerged, Rodman and colleagues were the first to report that GPT-4 provided the correct diagnosis in its differential in two-thirds of these challenging cases. Still, he was quick to admit that studies like his are limited by saturated benchmarks and a lack of physician comparators.

So, Rodman and the ARISE team went a few steps further in a set of experiments recently published in Science. They found that OpenAI’s o1 reasoning model outperformed physicians across multiple historical clinical reasoning tasks.

More notably, o1 performed as well or better than two Harvard internists in generating differential diagnoses based on EHR data for 76 real-world emergency cases.

While the study captured widespread attention, Rodman sees this as an incremental step on a much longer journey.

“What we need now,” he told me, “are prospective clinical trials in real-world patient care settings.”

Of course, diagnosis is just one aspect of clinical reasoning. And clinical reasoning is just one domain in which medical AI is being developed. As AI systems begin to perform differently across tasks—and sometimes outperform physicians—how should we evaluate them in ways that actually matter?

The Physician Bridge Builder And Communicating Science

Dr. Ethan Goh aims to clearly explain what AI can do in healthcare.

ARISE Network

Ethan Goh is a thoughtful, mild-mannered physician with a diverse range of experience beyond his years. After starting his career as a hospitalist in Europe and Asia, he served as a policymaker in Singapore, an advisor to the UK National Health Service, and an executive at a digital health startup before joining Stanford as a postdoctoral fellow in informatics.

Now serving as ARISE’s Executive Director, he draws on his diverse background to connect AI development, academic research, and real-world clinical care.

Goh sees ARISE’s main role as understanding and clearly explaining what AI can do in healthcare.

Traditionally, AI was evaluated using medical exam questions, such as those on the USMLE. Yet while these exams assess knowledge, they do not reflect real-world practice, where clinicians iteratively gather information from patients who often present differently than textbook descriptions. And a machine that performs well on a standardized knowledge test will not necessarily provide good clinical care.

The field is now moving toward simulations that more closely mirror clinical practice, often using rubrics rather than single correct answers.

Still, even these newer benchmarks typically lack context and focus on isolated cognition rather than actual clinical work.

Because medicine is not a single task and “doctoring” is not a single function, Goh argues that benchmarks must be framed around precise tasks such as triage, diagnosis, treatment, and communication, each with different thresholds for AI readiness.

Accordingly, ARISE introduced the Medical AI Superintelligence Test (MAST), which combines multiple domains of clinical competence—including diagnosis, management, reasoning, safety, and agentic workflow use—all benchmarked against realistic physician baselines and incorporating physicians working with AI, not just models alone.

As Goh explained, “Our goal is to open-source benchmarks and constantly index the latest models to find out where they are strong or weak on key clinical tasks, rather than the industry relying on its own limited benchmarks.”

One MAST component benchmark is NOHARM, which quantifies how often an LLM makes potentially harmful recommendations. Recently, the ARISE team reported that even top models generate potentially harmful advice in up to 22% of cases, typically due to errors of omission. Still, the best models outperformed generalist physicians on safety, and ensembles of models made fewer errors than individual models.

Another MAST component is MedAgentBench, which assesses models’ ability to independently perform 300 patient-specific, clinically relevant tasks—like ordering medications and aggregating test results—in a realistic FHIR-based EHR setting.

In mid-2025, the ARISE team reported that the best-performing model achieved a 70% success rate, with most failures clustering around tasks requiring three or more steps. However, just six months later, Anthropic announced that its Opus 4.6 model achieved a 92% success rate, underscoring how quickly these capabilities are advancing and how quickly benchmarks themselves may become outdated.

In response, ARISE developed PhysicianBench, a new benchmark designed to evaluate how well AI agents complete multi-step medical consultation and execution tasks in realistic EHR settings.

Yet AI may quickly outpace this benchmark, too. If these trends continue and AI approaches superintelligence—which ARISE defines as outperforming top clinicians across a range of clinically meaningful tasks under real-world conditions—evaluating AI based on concordance with physician experts will break down, just as experts in the game of Go were confounded by AlphaGo’s winning Move 37.

This will force a shift to real-world randomized controlled trials with hard clinical outcomes.

Where Are We Going?

Chen, Rodman, and Goh each believe we are on the verge of a fundamental shift in what it means to be a doctor. Whether they are right or wrong, AI is forcing medicine to reconsider some of its deepest assumptions about clinical reasoning and expertise, human-machine interaction, and how we define good care.

In the process, AI is pushing us to think more carefully about what physicians do, where AI helps, and how and when the two should work together. These questions are no longer theoretical.