惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Azure Blog
Microsoft Azure Blog
Engineering at Meta
Engineering at Meta
A
About on SuperTechFans
T
The Blog of Author Tim Ferriss
I
InfoQ
博客园_首页
G
Google Developers Blog
爱范儿
爱范儿
Last Week in AI
Last Week in AI
量子位
阮一峰的网络日志
阮一峰的网络日志
雷峰网
雷峰网
酷 壳 – CoolShell
酷 壳 – CoolShell
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
GbyAI
GbyAI
月光博客
月光博客
The GitHub Blog
The GitHub Blog
V
Visual Studio Blog
N
Netflix TechBlog - Medium
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 司徒正美
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 聂微东

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Towards Conversational AI for Disease Management - Nature
Schaekermann, Mike · 2026-06-17 · via Hacker News - Newest: "AI"
  • Article
  • Published:

Nature (2026) Cite this article

We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.

Subjects

Abstract

While large language models (LLMs) have shown promise in diagnostic dialogue1, their capabilities for effective management reasoning—including disease progression, therapeutic response, and safe medication prescription—remain under-explored. We advance the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMIE)1−3 through a new LLM-based agentic system optimized for multi-visit clinical management and dialogue. To ground its reasoning in authoritative clinical knowledge, AMIE leverages Gemini’s long-context capabilities4, combining in-context retrieval with structured reasoning to align its output with up-to-date clinical practice guidelines and drug formularies. In a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) study, AMIE was compared to 21 primary care physicians (PCPs) across 100 multi-visit case scenarios designed to reflect UK NICE Guidance and BMJ Best Practice guidelines. AMIE was non-inferior to PCPs in management reasoning as assessed by specialists and scored better in both preciseness of treatments and investigations, and in its alignment with and grounding in clinical guidelines. To benchmark medication reasoning, we developed RxQA, a multiple-choice question benchmark derived from two national drug formularies (US, UK) and validated by board-certified pharmacists. Though AMIE and PCPs both benefited from the ability to access external drug information, AMIE outperformed PCPs on higher difficulty questions. While further research would be needed before real-world translation, AMIE’s strong performance across evaluations marks a significant step towards conversational AI as a tool in disease management.

This is a preview of subscription content, access via your institution

Access options

Access Nature and 54 other Nature Portfolio journals

Get Nature+, our best-value online-access subscription

27,99 € / 30 days

cancel any time

Subscribe to this journal

Receive 52 print issues and online access

185,98 € per year

only 3,58 € per issue

Rent or buy this article

Prices vary by article type

from$1.95

to$39.95

Prices may be subject to local taxes which are calculated during checkout

Author information

Author notes

  1. These authors jointly supervised this work: Alan Karthikesalingam, Mike Schaekermann

  2. These authors contributed equally: Valentin Liévin, Anil Palepu

Authors and Affiliations

  1. Google DeepMind, Mountain View, California, USA

    Valentin Liévin, Khaled Saab, David Stutz, Yong Cheng, S. Sara Mahdavi, Joëlle Barral, Ryutaro Tanno & Tao Tu

  2. Google Research, Mountain View, California, USA

    Anil Palepu, Wei-Hung Weng, Kavita Kulkarni, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam & Mike Schaekermann

Authors

  1. Valentin Liévin
  2. Anil Palepu
  3. Wei-Hung Weng
  4. Khaled Saab
  5. David Stutz
  6. Yong Cheng
  7. Kavita Kulkarni
  8. S. Sara Mahdavi
  9. Joëlle Barral
  10. Dale R. Webster
  11. Katherine Chou
  12. Avinatan Hassidim
  13. Yossi Matias
  14. James Manyika
  15. Ryutaro Tanno
  16. Vivek Natarajan
  17. Adam Rodman
  18. Tao Tu
  19. Alan Karthikesalingam
  20. Mike Schaekermann

Corresponding authors

Correspondence to Valentin Liévin, Anil Palepu, Alan Karthikesalingam or Mike Schaekermann.

Supplementary information

Supplementary Information (download PDF )

Supplementary discussion, methods and results (Sections 1-16). Contains related work, details on the system design for the Mx agent and Dialogue agent, details on the OSCE evaluation study (inter-rater reliability analysis, clinician metadata, scenario metadata, ablation analysis), and methods details and further results for the RxQA medication reasoning benchmark.

Reporting Summary (download PDF )

Supplementary Data 1 (download PDF )

Detailed view of two sample scenarios with AMIE and PCP output and evaluation gradings. Full details for two sample scenarios used in the OSCE evaluation study, including scenario information, AMIE-patient-actor conversations, PCP-patient-actor conversations, specialist physician gradings and patient actor gradings for all three visits per scenario.

Supplementary Data 2 (download PDF )

Details for all 120 OSCE scenarios with AMIE output (PDF). Scenario details and AMIE output for all 120 scenarios used either in the OSCE evaluation study (100) or for validation purposes (20), in human-readable PDF format.

Supplementary Data 3 (download CSV )

Details for all 120 OSCE scenarios with AMIE output (CSV). Scenario details and AMIE output for all 120 scenarios used either in the OSCE evaluation study (100) or for validation purposes (20), in machine-readable CSV format.

Peer Review File (download PDF )

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Liévin, V., Palepu, A., Weng, WH. et al. Towards Conversational AI for Disease Management. Nature (2026). https://doi.org/10.1038/s41586-026-10764-5

Download citation

  • Received:

  • Accepted:

  • Published:

  • DOI: https://doi.org/10.1038/s41586-026-10764-5

Associated content