惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
A
About on SuperTechFans
The Last Watchdog
The Last Watchdog
A
Arctic Wolf
S
Schneier on Security
Cisco Talos Blog
Cisco Talos Blog
K
Kaspersky official blog
Spread Privacy
Spread Privacy
The Hacker News
The Hacker News
P
Proofpoint News Feed
Attack and Defense Labs
Attack and Defense Labs
NISL@THU
NISL@THU
AWS News Blog
AWS News Blog
Schneier on Security
Schneier on Security
TaoSecurity Blog
TaoSecurity Blog
H
Hacker News: Front Page
L
LangChain Blog
Y
Y Combinator Blog
T
Tenable Blog
Microsoft Security Blog
Microsoft Security Blog
L
Lohrmann on Cybersecurity
量子位
N
News and Events Feed by Topic
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
The GitHub Blog
The GitHub Blog
云风的 BLOG
云风的 BLOG
W
WeLiveSecurity
Martin Fowler
Martin Fowler
Cloudbric
Cloudbric
S
SegmentFault 最新的问题
Project Zero
Project Zero
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园 - 叶小钗
V
Vulnerabilities – Threatpost
小众软件
小众软件
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
腾讯CDC
博客园 - 聂微东
F
Full Disclosure
WordPress大学
WordPress大学
PCI Perspectives
PCI Perspectives
P
Privacy International News Feed
M
MIT News - Artificial intelligence
Forbes - Security
Forbes - Security
Blog — PlanetScale
Blog — PlanetScale
T
The Blog of Author Tim Ferriss
Webroot Blog
Webroot Blog
S
Security @ Cisco Blogs
Last Week in AI
Last Week in AI

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Agentic AI Comes to Medicine
Eric Topol · 2026-06-18 · via Hacker News - Newest: "AI"

It was just a matter of time. Agentic autonomous AI has already been applied to life science and many other domains, and today there were 2 notable publications in Nature that move this concept forward for healthcare. One is called MIRA from Jacob Kather and colleagues from Germany, and the other is called AIME, from Mike Schaekermann and colleagues at Google (acronyms defined below). This work is getting well beyond AI support for narrow applications, such as help in making diagnoses, to full management, end-to-end care plans. They are both very complicated papers with a lot to unpack, including tens of pages of supplementary information to fully describe what they assessed. In this issue of Ground Truths, I’m going to get to the core results and implications. First, a summary Table that compares the 2 systems.

Note the Towards in the title of the 2 papers:

This was designed to be embedded in a health system EHR to provide reasoning and action steps. There were 2 agents, the patient and the AI physician (MIRA). MIRA queried the patient’s history, the physical exam results, and could order labs, blood cultures, scans, medications, procedures, surgery, and triage for hospital admission. This was done in 500 emergency department established real cases with the MIRA results directly compared to 4 board-certified (BC) physicians, and also to a hybrid group of 2 BC physicians and 2 residents. (I won’t review the results of the hybrid group further since their performance in all tasks was lower than the 4 BC physicians.) It simulated the sequential way a patient’s data would be interrogated and processed. MIRA was enabled with 11 different tools and choices from >85,000 action options, operating in a standards-compliant framework for multi-step reasoning (using FHIR, ICD-10, RxNORM, ATC, LOINC, and SNOWMED-CT). The system was built on OpenAI’s GPT-4o.

The overall diagnostic accuracy for MIRA was 87.8% compared with BC physicians at 78.1%. That increase was especially notable for specific diagnoses like pancreatitis (95.2% vs 78.6%), and appendicitis (100% for MIRA, 88% for BC physicians). While MIRA ordered more blood tests (51% vs 28%), resource consumption was countered this by ordering substantially less scans. For therapy, MIRA surpassed the BC physicians 53.5% vs 38.3% for correctly ordering procedures such as laparoscopic appendectomy or cholecystectomy (Figure below). Other advantages for MIRA therapy included better IV fluid management and analgesic adherence to guidelines, and an overall 35% increased alignment of clinical guidelines compared with the BC physicians. Of 468 medications ordered by MIRA, 99.8% were correct for indication and safety (such as allergy, interactions, and kidney dosing). MIRA triaged more cases than physicians for hospital admission, which reflects lack of being economically driven.

Because the design of the system would allow leak of the case data to the AI, considerable effort was made to avoid premature information flow. That worked well, with 0 of 933 cases exhibiting any leak. 880 adversarial prompts were tested and the system held up well, as it did for stress-testing attempts at hacking, medico-legal threats, and other patient agent trickery. MIRA also assessed multiple patient perturbations including high anxiety, non-English speaking, paranoia, and diagnostic denial, without affecting its performance.

This system had a very different design, with its focus on longitudinal assessment of outpatients, with the primary goal of developing first-rate management plans. Like MIRA, there were 2 agents used. The Dialogue Agent was conversational, interacting with the patient, representing fast, System 1 thinking (À la Danny Kahneman, using Gemini 1.5 Flash), and asynchronous to the Mx management agent, System 2, slow thinking that used long, context processing (even though it was quick). AMIE assessed 100 patients with 3 visits (each separated by ~2 days) spanning 5 different specialties. The results were compared with 21 BC primary care physicians. A noteworthy feature was the Ensemble Refinement which took 4 different treatment plans developed and came up with a consensus, mimicking a real medical treatment board, as is typically seen with cancer management. The massive >600 clinical guidelines were fully tokenized (not just parts of them) to provide the grounding and citations for management. This ensemble only took about 80 seconds to produce. Like MIRA, this was all text based, which was defended as necessary in AIME for the intent to maintain blinding. 30 physicians rated the performance of AIME outputs vs the 21 BC PCPs. Unlike MIRA, patient actors were used. The performance wheel below shows the differences for AIME vs the PCPs for 6 metrics, all leaning to AIME for being slightly or substantially better.

Overall, for management reasoning, AIME was non-inferior to the 21 PCPs. By the 3rd outpatient visit, the rating of AMIE’s management plan was 98% vs 81% of the PCPs. Preciseness of treatment was 95% vs 67%. Expert guidelines alignment was 100% vs 86%, respectively. So while the overall management was non-inferior, there were several ratings that favored AIME’s management.

A new benchmark for medication management was developed, called RxQA, built on 600 questions to board-certified pharmacists and the content from the UK and US formularies. AIME’s medication management outperfomed the PCPs for the correct medication, dose, duration, side-effects, and follow-up evaluation. This was found with the more difficult cases even when the PCPs were tested open-book (58% vs 48% in favor of AI).

Share Ground Truths

The new studies raise the level of capabilities for medical AI from what has been previously studied, which were relatively narrow, mainly for support in making a diagnosis or answering questions on a medical exam or from a patient. MIRA took on autonomous agent end-to-end assessment and action plans for each patient presenting to an emergency department. AIME was geared to provide longitudinal assessment across 3 outpatient visits. Both used only 2 agents, one interacting with the patient and the other to do the AI work.

There are many reasons to think these are preliminary findings that do not reflect the real world of medicine. Both MIRA and AIME are text-only AI, meaning all the other things that are part of medicine, from the patient’s non-verbal communication and tone of voice to the review of actual medical images, were not included. The cases used in both studies were “clean,” complete data from established datasets. MIRA interactions were capped at 20 conversation turns. AIME used patient actors. There was nothing done that truly represents the practice of medicine, which is typically characterized by incomplete and conflicting data. The 3 outpatients visits in AIME were set up 2 days apart, hardly simulating how hard it is to get appointments with doctors! Likewise, only 5 specialties of medicine were addressed.

But there were some findings that can be viewed as filling in gaps of current medical care. Improving diagnostic accuracy in MIRA was clear, although the comparator group was very small, with only 4 board-certified physicians. Getting 100% correct diagnosis and ordering a laparoscopic procedure for appendicitis is impressive, as was the accuracy of the medications selected. The breadth of conditions in MIRA were limited (n=8) even though there was an array of diagnostic tests, procedures and surgeries as part of the management option mix. The lack of economic incentives in the practice of MIRA management is desirable and even resulted in less expansive scan tests ordered.

For AIME, a big advantage was the longitudinal context across 3 clinic visits. Unlike real patients today, the Patient Dialogue Agent didn’t have to fill out forms. There was memory and efficiency (not what we see in US healthcare). That agent was set up to be empathetic and attentive which showed up in the patient preferences, a pretty big gradient favoring the AI.

Both of these medical AI models were remarkably precise with respect to sticking to guidelines or providing specific plans. For example, the PCP in the AIME study wrote “give an antibiotic whereas the AI prescribed amoxicillin 500 mg orally, 3 tablets daily for 7 days, and wrote to check for a penicillin allergy.

But here is the rub. Guidelines in medicine are meant to apply to the vast majority of patients, without specific regards for needs, fears, prior patient experiences, cost, and many other factors. Much of present day guidelines are provided by “experts,” that is they represent opinions that are eminence-based, not evidence-based medicine. So this very tight, superior alignment of the AI with guidelines is not necessarily a good thing, with the claims of great “precision.” Frankly, over-adherence to guidelines may presage the loss of the art of medicine, not taking in the human factor of each patient, and the human-to-human bond that would be the foundation of the patient-doctor relationship.

Where is this headed? The large language models (LLMs) will keep getting better. In fact, the ones used in these 2 reports are already obsolete. In the agentic AI era, there could literally be hundreds of specialized agents, such as one for labs, one for scans, one for sensors, one for environmental exposures, one for genetics/genomics, and so on. Today, with these new papers, we are only seeing the rudimentary use of agents.

You can think of MIRA and AIME as a major step forward within the constraints of a simulation, not real medicine. But the improvements in AI’s capabilities are coming fast, and it would not be surprising to see some of the benefits here extended to the actual practice of medicine. To prove that, we ideally will need randomized trials of 3 strategies to assess outcomes: (1) end-to-end medical AI ; (2) human clinicians only; and (3) combining both. That won’t happen soon since are still on a steep slope of model improvement and such a large trial would not be easy to get funded, no less executed. In the meantime, we’ve got new evidence for potential ways that generative AI will improve medical communication, diagnosis, and management.

NB This post was written by me, no AI. I have no COI related to the content of the post.

A big thanks to Ground Truths subscribers from every US state and 212 countries. Your subscription to these free essays and podcasts makes my work in putting them together worthwhile. If you’re not a subscriber, please join!

If you found this interesting PLEASE share it!

Share Ground Truths

Paid subscriptions are voluntary and all proceeds from them go to support Scripps Research. They do allow for posting comments and questions, which I do my best to respond to. Please don’t hesitate to post comments and give me feedback. Let me know topics that you would like to see covered.

Leave a comment

Many thanks to those who have contributed—they have greatly helped fund our summer internship programs for the past two years. It enabled us to accept and support a record number of 51 summer interns coming in 2026! These are high school, college and medical students selected from thousands of applicants. We couldn’t do this expanded program without the funds coming in through Ground Truths.

Discussion about this post

Ready for more?