惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Vulnerabilities – Threatpost
月光博客
月光博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
美团技术团队
Last Week in AI
Last Week in AI
Jina AI
Jina AI
P
Privacy International News Feed
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享
T
Tenable Blog
阮一峰的网络日志
阮一峰的网络日志
P
Proofpoint News Feed
T
Tailwind CSS Blog
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy & Cybersecurity Law Blog
人人都是产品经理
人人都是产品经理
S
Schneier on Security
Google DeepMind News
Google DeepMind News
爱范儿
爱范儿
C
Cisco Blogs
K
Kaspersky official blog
C
Cybersecurity and Infrastructure Security Agency CISA
Hugging Face - Blog
Hugging Face - Blog
博客园_首页
O
OpenAI News
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
IT之家
IT之家
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
酷 壳 – CoolShell
酷 壳 – CoolShell
T
Tor Project blog
博客园 - 【当耐特】
腾讯CDC
V
V2EX
A
Arctic Wolf
Webroot Blog
Webroot Blog
S
Securelist
小众软件
小众软件
大猫的无限游戏
大猫的无限游戏
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园 - 三生石上(FineUI控件)
The GitHub Blog
The GitHub Blog
量子位
J
Java Code Geeks
博客园 - 叶小钗
S
SegmentFault 最新的问题
Project Zero
Project Zero
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Scott Helme
Scott Helme
Cyberwarzone
Cyberwarzone

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Built a Pneumonia Detection AI on My MacBook — Here's Exactly How It Works
nandhu_sauce · 2026-04-24 · via DEV Community

I just finished building a deep learning system that identifies pneumonia from chest X-rays with 96% accuracy and an AUC-ROC of 0.99. I ran the entire training process on my MacBook Pro M5 using the GPU acceleration provided by Apple's Metal Performance Shaders (MPS). This project wasn't about complex math; it was about using transfer learning to turn a consumer laptop into a medical diagnostic tool. Understanding how to handle messy data and verify what an AI is actually "seeing" is more important than having a massive server room.

WHAT I ACTUALLY BUILT

At its core, I built a computer program that acts like a specialized set of eyes for doctors. You feed it a digital chest X-ray image, and in less than a second, it tells you whether it sees signs of pneumonia or a normal, healthy lung. It doesn't just guess; it provides a confidence percentage for its decision. The system is designed to catch cases that might be subtle to the human eye, acting as a second pair of eyes to reduce diagnostic errors.

THE DATASET PROBLEM NOBODY MENTIONS

When I first looked at the data, I found a massive problem: class imbalance. The training set had 1,341 "Normal" images but a whopping 3,875 "Pneumonia" images. If I had trained the model as-is, it would have quickly learned that "guessing pneumonia every time" results in 74% accuracy without actually learning a single thing about lungs. This is a common trap in AI where a high accuracy score hides a completely useless model.

To fix this, I used two specific techniques to level the playing field. First, I implemented a WeightedRandomSampler, which you can think of as "photocopying rare examples." It ensures that during training, the model sees the "Normal" images more frequently so it doesn't forget what a healthy lung looks like. Second, I added class weights to the loss function—essentially a "bigger penalty for missing rare cases." If the model misclassifies a "Normal" lung as pneumonia, the error signal it receives is mathematically amplified, forcing it to pay closer attention to those specific patterns.

WHY I DIDN'T TRAIN FROM SCRATCH

I didn't start with a blank slate. Instead, I used a technique called transfer learning with an architecture known as ResNet-18. Think of it like this: a radiologist who already knows what edges, textures, and shapes look like doesn't need to relearn basic vision from scratch. They already have the "visual foundation" from years of looking at the world; they just need to learn what sick lungs look like specifically. ResNet-18 comes pre-trained on millions of everyday images (like dogs, cars, and trees), so it already understands how to detect lines and textures.

I used a two-phase training strategy to refine this pre-existing knowledge. In Phase 1 (Epochs 1-5), I froze most of the model and only trained the very last layer. This allowed the model to get a "feel" for the new medical data without overwriting its basic visual skills. In Phase 2 (Epochs 6-20), I unfroze everything and used a tiny learning rate to fine-tune the entire network. At the start of Phase 2, I saw a temporary spike in the loss curve. This is normal and expected—it's the mathematical equivalent of the model being slightly "confused" as it starts adjusting its deep-seated visual patterns to the nuances of X-ray tissue.

CAN YOU TRUST IT? GRAD-CAM EXPLAINABILITY

A 96% accurate AI is useless if it's a "black box" that you can't verify. In medical AI, you have to know why a decision was made. I implemented Grad-CAM (Gradient-weighted Class Activation Mapping), which is a tool that "shows you which pixels the model was looking at when it made its decision." Without this, a model might achieve 99% accuracy just by learning that images with a certain hospital's patient ID label in the corner are usually the pneumonia cases.

When I ran Grad-CAM on my results, the heatmaps were revealing. For pneumonia cases, the "heat" (red and yellow zones) was concentrated directly on the lung tissue where opacities usually appear. For normal cases, the focus was much more diffuse across the entire chest cavity. This gave me the confidence that the model was actually learning medical features, not just memorizing background noise or image artifacts.

THE RESULTS

After 20 epochs and about 45 minutes of training on my MacBook, here is how the system performed on the 624 images in the test set:

Metric Score
Test Accuracy 96%
AUC-ROC 0.99
NORMAL F1 0.94
PNEUMONIA F1 0.96
False Negatives 15/390
False Positives 13/234

Let's break these down into plain English. Test Accuracy means the model was right 96 times out of 100 overall. AUC-ROC measures how well the model can distinguish between the two classes across different confidence levels; 0.99 is nearly perfect separation. F1 Score is a balanced average of precision (not flagging healthy people as sick) and recall (not missing sick people).

Most importantly, we have to look at the 15 missed pneumonia cases (False Negatives). In a clinical context, missing a sick patient is far worse than accidentally flagging a healthy one. This is why "Recall" matters more than "Accuracy" in medical AI. While 15 misses out of 390 is low, it highlights that this system is a diagnostic assistant, not a replacement for a human doctor who would catch those edge cases.

WHAT I LEARNED

This project taught me a few genuine technical lessons that go beyond the usual tutorials:

  1. The Validation Set Trap: The original dataset only had 16 images in the validation folder. This made the validation accuracy bounce around wildly and become meaningless during training. You need a representative validation set to know if your model is actually improving.
  2. Watch the Loss, Not the Accuracy: Accuracy is a "lagging indicator." The loss curve tells you the "quality" of the model's learning. If the loss is still going down but accuracy is flat, you're still making progress.
  3. Grad-CAM is Mandatory: For medical AI, explainability isn't a "nice to have." It's the difference between a useful tool and a legal liability. If you can't see the heatmaps, you shouldn't trust the predictions.
  4. Apple Silicon is Ready: Training this on MPS (Metal Performance Shaders) was surprisingly fast. For this size of workload, you don't need a dedicated Linux server with a massive GPU; a modern MacBook Pro handles it in under an hour.

WHAT'S NEXT

I'm not finished with this system yet. My next steps are specific:

  • Fix the Validation Set: I'm going to move about 500 images from the training set into the validation set to get more reliable feedback during training.
  • Try DenseNet-121: This architecture is the current gold standard in chest X-ray research papers because of how it handles feature reuse.
  • Build a Web UI: I want to use Streamlit to create a simple drag-and-drop interface so anyone can test the model without looking at code.
  • Kaggle Notebook: I've already published a self-contained version of this project as a Kaggle notebook for the community to play with.

CLOSING

This project demonstrated that you don't need a supercomputer to build high-performing medical AI. By using transfer learning and being smart about how you handle imbalanced data, you can achieve professional-grade results on consumer hardware. It's a testament to how accessible deep learning has become, provided you focus on the data and the "why" behind the predictions.


Disclaimer: This model is for educational purposes only and is not intended for clinical use.

View the full notebook on Kaggle
View the code on GitHub