惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Last Week in AI
Last Week in AI
阮一峰的网络日志
阮一峰的网络日志
P
Proofpoint News Feed
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
MongoDB | Blog
MongoDB | Blog
云风的 BLOG
云风的 BLOG
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
J
Java Code Geeks
WordPress大学
WordPress大学
T
The Blog of Author Tim Ferriss
V
Visual Studio Blog
小众软件
小众软件
Microsoft Azure Blog
Microsoft Azure Blog
博客园_首页
IT之家
IT之家
Vercel News
Vercel News
C
Check Point Blog
Google DeepMind News
Google DeepMind News
月光博客
月光博客
D
DataBreaches.Net
酷 壳 – CoolShell
酷 壳 – CoolShell
美团技术团队
Y
Y Combinator Blog
Hugging Face - Blog
Hugging Face - Blog

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More

Learn Command Line Interface (CLI) Development with Dart: From Zero to a Fully Published Developer Tool How to Bypass Cloud SMTP Restrictions Using Brevo and HTTP APIs How to Build a Live Options Database in Python – A Complete Guide How to Migrate to S3 Native State Locking in Terraform How to Use SCons to Build Software Projects [Full Handbook] How to Run Open Source LLMs Locally and in the Cloud QuRT: The Real-Time OS Inside Your Phone's Processor [Full Handbook] The Real Infrastructure Behind Remote Work (It’s Not Just Wi-Fi) The Lithography Handbook: Machines, Markets, and the Next Wave of Semiconductor Startups ITCM vs DTCM vs DDR: Embedded Memory Types Explained [Full Handbook] AI Paper Review: Improving Language Understanding by Generative Pre-Training (GPT-1) How to Build a Market Research Copilot with MCP and Python [Full Handbook] How to Build a Scoped Note-Taking API with Django Rest Framework and SimpleJWT The Complete SOC 2 Type II Implementation Handbook for Engineers: A Month-by-Month Roadmap with Real Commands Mastering the JavaScript Event Loop Data Science Insights: Why the Mean Lies When Handling Messy Retail Data How to Build High-Ranking SEO Landing Page How to Query Data in DynamoDB Using .Net How to Unblock Your AI PR Review Bottleneck: A Tech Lead’s Guide to Building a Codebase-Aware Reviewer How to Navigate Microservices as a Frontend Engineer How to Compress PDF Files in the Browser Using JavaScript (Step-by-Step) Stanford's youngest instructor talks InfoSec, AI, and catching cheaters - Rachel Fernandez interview [Podcast #217] Product Experimentation with Propensity Scores: Causal Inference for LLM-Based Features in Python How to Build a Multi-Agent AI System with LangGraph, MCP, and A2A [Full Book] How to Land Your First Cloud or DevOps Role: What Hiring Managers Actually Look For How to Deploy a Serverless Spam Classifier Using Scikit-Learn, AWS Lambda, & API Gateway How to Dockerize a Go Application – Full Step-by-Step Walkthrough Learn Hardware, Cloud, DevOps, Networking, Security, Databases, DNS, Git, and Linux Inside TreeHacks 2026, Stanford’s Elite Student Hakc Inside Stanford’s Elite Student Hackathon [Full Documentary]
CUDA Programming for NVIDIA H100s
2026-04-09 · via freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
CUDA Programming for NVIDIA H100s

Learn CUDA programming for NVIDIA Hopper GPUs.

We just posted a course on the freeCodeCamp.org YouTube channel that will teach you to build efficient WGMMA pipelines and leverage Cutlass optimizations to perform the massive matrix multiplications that power modern AI.

Beyond single-chip performance, the curriculum covers multi-GPU scaling and NCCL primitives necessary for training trillion-parameter models. To get the most out of these lessons, you should have a foundational grasp of C++ syntax and linear algebra, particularly how matrices are tiled and multiplied.

Here are all the sections in this massive course:

  • Course Introduction

  • Table of Contents & Course Overview

  • LESSON 1 — H100 Hopper GPU Architecture

  • H100 Specifications: HBM3, Bandwidth & Power

  • Tensor Cores Overview

  • Tensor Memory Accelerator (TMA)

  • Transformer Engine

  • L2 Cache Architecture

  • GPCs, TPCs & SM Layout

  • Thread Block Clusters

  • Distributed Shared Memory

  • SM Sub-Partitions (SMSPs)

  • Warp Schedulers & Dispatch Units

  • Shared Memory & Data Movement

  • Occupancy

  • LESSON 2 — Clusters, Data Types, Inline PTX & Pointers

  • Thread Block Clusters Programming

  • Configuring Cluster Dimensions

  • Inline PTX Assembly

  • State Spaces

  • Data Types in PTX

  • Generic Pointers

  • Address Space Conversion

  • LESSON 3 — Asynchronicity & Barriers

  • Introduction to Async Operations

  • Proxies

  • Fences & Memory Ordering

  • Fence Ordering & Visibility

  • Fence Scopes

  • Acquire & Release Fences

  • Expected Count & Thread Arrival

  • M-Barrier Arrive Operations

  • M-Barrier PTX Instructions

  • Barrier Wait Operations

  • Phase & Parity

  • Commit Operations

  • LESSON 4 — CuTensorMap Descriptors

  • Tensor Shape, Stride & Data Type

  • Element Stride & Dimensions

  • Box Dimensions (Tile Size)

  • Bank Conflicts

  • Swizzling

  • Swizzle Formula Deep Dive

  • Interleave Layouts

  • Out-of-Bounds Fill (OOB)

  • LESSON 5 — cp.async.bulk (Async Bulk Copies via TMA)

  • Bulk Tensor Operations (1D–5D)

  • Multicast Operations

  • Prefetch

  • LESSON 6 — WGMMA Part 1 (Warp Group Matrix Multiply Accumulate)

  • Warp Groups & Matrix Multiplication

  • WGMMA Descriptors

  • Accumulators & Register Reuse

  • Scale Factors (Scale D, Scale A, Scale B)

  • Core Matrices & 16×16 Tiles

  • LESSON 7 — WGMMA Part 2

  • Commit Groups & Wait Groups

  • WGMMA with FP8 Data Types

  • LESSON 8 — Kernel Design

  • Compute-Bound vs. Memory-Bound Kernels

  • Warp Specialization

  • Cooperative vs. Ping-Pong Pipelines

  • Pipelining Fundamentals

  • Circular Buffering

  • Ping-Pong Pipeline Deep Dive

  • Epilogue Handling in Pipelines

  • Persistent Scheduling

  • Split-K & Stream-K Strategies

  • Data-Parallel Tile Scheduling

  • Epilogue Fusion (Bias, Activation, Scaling)

  • Epilogue Operations Overview

  • CUTLASS SOURCE CODE WALKTHROUGH

  • Main Loop & Scheduling Policies

  • Dispatch Policy

  • SM90 Tile Scheduler

  • SM90 Epilogue (TMA Warp Specialized)

  • SM90 Builder

  • Collective Builder

  • FAST.CU KERNEL WALKTHROUGH

  • Main Loop Implementation

  • Producer Warp Group (Dependence Wall)

  • Consumer Warp Group

  • Prologue

  • MULTI-GPU PROGRAMMING — Part 1

  • NVSwitch

  • Topology & System Architecture

  • NVSwitch, BlueField DPUs & Storage Fabrics

  • CUDA Peer-to-Peer Communication

  • MPI (Message Passing Interface)

  • P2P Limitations & Trade-offs

  • MULTI-GPU PROGRAMMING — Part 2

  • SLURM Resource Allocation

  • PMIx Process Management

  • NCCL (NVIDIA Collective Communications Library)

  • NCCL Internals & Ring Algorithm

  • AllReduce Operations

  • NCCL Collectives: Broadcast, AllGather, ReduceScatter

  • Parallelism Strategies: Data, Tensor, Pipeline & Expert Parallelism

  • Course Conclusion & Next Steps

Watch the course on the freeCodeCamp.org YouTube channel (24-hour watch).



Learn to code for free. freeCodeCamp's open source curriculum has helped more than 40,000 people get jobs as developers. Get started