惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
Stack Overflow Blog
Stack Overflow Blog
Latest news
Latest news
T
The Blog of Author Tim Ferriss
D
DataBreaches.Net
C
Check Point Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Security Latest
Security Latest
宝玉的分享
宝玉的分享
S
Schneier on Security
Blog — PlanetScale
Blog — PlanetScale
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Cisco Talos Blog
Cisco Talos Blog
MyScale Blog
MyScale Blog
B
Blog RSS Feed
N
Netflix TechBlog - Medium
P
Privacy & Cybersecurity Law Blog
L
LINUX DO - 热门话题
Apple Machine Learning Research
Apple Machine Learning Research
T
Tenable Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
NISL@THU
NISL@THU
Google DeepMind News
Google DeepMind News
Hacker News: Ask HN
Hacker News: Ask HN
Schneier on Security
Schneier on Security
博客园 - Franky
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
S
Secure Thoughts
T
Threat Research - Cisco Blogs
D
Docker
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
人人都是产品经理
人人都是产品经理
G
GRAHAM CLULEY
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 【当耐特】
PCI Perspectives
PCI Perspectives
GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
Cyberwarzone
Cyberwarzone
V
Vulnerabilities – Threatpost
F
Fortinet All Blogs
罗磊的独立博客
Engineering at Meta
Engineering at Meta
Y
Y Combinator Blog
SecWiki News
SecWiki News
A
Arctic Wolf
小众软件
小众软件
T
Troy Hunt's Blog
博客园 - 三生石上(FineUI控件)
Know Your Adversary
Know Your Adversary

Lei Mao's Log Book

2026 FIFA World Cup 备受嘲讽的会徽 Python Debugging Via VS Code In Docker Container Vargas Plateau Regional Park 徒步 Vargas Plateau Regional Park Apex 2026 FIFA World Cup 小组赛赛程 Retaining EXIF Metadata In GIMP San Francisquito Creek Joint Powers Authority 2025 Calendar Photo 麦当劳 The FIFA World Cup 套餐 Ardenwood Historic Farm 徒步 Ardenwood Historic Farm Synchronizations With TorchRec KeyedJaggedTensor Pacific Commons Linear Park 徒步 Pacific Commons Linear Park 2026 San Jose Half Marathon 竞赛 目标 Mountain View Shoreline Park 徒步 Mountain View Shoreline Park PyTorch AOTInductor Hybrid Lowering Carquinez Strait Regional Shoreline 徒步 Carquinez Strait Regional Shoreline PyTorch Triton Kernel Transparent Tracing and Compilation 脸庞 PyTorch Fake Export 2026 BRAIN Foundation 10K 竞赛 2026 Wild and Scenic Film Festival 参观 2026 Wild and Scenic Film Festival 系统工程程序员修 Bug FIFA 官方网站的语言 PyTorch Custom Operation 汉堡王 The Mandalorian and Grogu 套餐 Tilden Regional Parks Botanic Garden 参观 Tilden Regional Park Tilden Regional Parks Botanic Garden Tilden Regional Park 徒步 《寻秦记》电影版 ICML 2026 Area Chair Experience 2026 Foster City 5K Fun Run 竞赛 2026 年 3 月和 4 月该入手的模型手办 Docker Container GUI Display Using Wayland 马拉松破二 2026 Heart & Soles Run 5K 竞赛 How Is FARS, The Fully Automated Research System? 算计: 七天的死亡游戏 Lake Chabot Regional Park 徒步 Lake Chabot Regional Park 2023 年恐怖电影《感恩节》 2026 Airport Runway Run at San Carlos Airport 5K 竞赛 Page Table for Page-Locked Host Memory Don Edwards San Francisco Bay National Wildlife Refuge - Ravenswood 徒步 Don Edwards San Francisco Bay National Wildlife Refuge - Ravenswood 法外风云 PyTorch Graph Symbolic Integer Contra Costa Canal Regional Trail 徒步 Contra Costa Canal Regional Trail 娑婆诃 PyTorch Export 2026 Western Pacific 5K 竞赛 浮躁的科研和胡扯的自媒体 Connecting Logitech Devices On Linux 2026 Oakland Half Marathon 竞赛 Del Valle Regional Park 徒步 Del Valle Regional Park 踏切时间 Wildcat Canyon Regional Park 徒步 Wildcat Canyon Regional Park Credit Card Unauthorized Transaction While In Possession 莎拉的真伪人生 2026 Union City Superhero Fun Run 5K 竞赛 McLaughlin Eastshore State Park 徒步 McLaughlin Eastshore State Park Cloudflare Worker Proxy R2 Bucket Access Zorro 和 Batman Fix MacBook Pro Space Key Stuck Problem 2026 年 1 月和 2 月该入手的模型手办 2026 Brazen Victory 10K 竞赛 儿时的玩伴李峰 Marsh Creek Regional Trail 徒步 Marsh Creek Regional Trail Perfetto GPU Flow Artifacts 百万人推理 System Performance Optimizations QQ 幻想 2026 Brazen Bay Breeze 5K 竞赛 CUDA Shared Memory Bank Conflict-Free Vectorized Access Dota 闪电站出售 Mountain View Downtown 徒步 Mountain View Downtown C++ Latch and Barrier 2025 年跑步总结 2026 Rotary Mission Ten Half Marathon 竞赛 狗的素质等于人的素质 CUDA Rendezvous Stream Pleasanton Ridge Regional Park 徒步 Pleasanton Ridge Regional Park Xfinity Internet 多年来的使用感受 Randomized SVD Don Castro Regional Recreation Area 徒步 Don Castro Regional Recreation Area 拖车公司的大汉们
CUDA_LAUNCH_BLOCKING=1
Lei Mao · 2026-03-20 · via Lei Mao's Log Book

Introduction

When we run CUDA programs on GPUs, we would sometimes encounter asynchronous errors which are only reported after synchronization. To localize where exactly the error occurred, usually we could see two approaches:

  1. In a relatively simple application or system where there are not many kernel launches, we could put cudaDeviceSynchronize() or cudaStreamSynchronize(stream) after every kernel launch to force synchronization and error checking. We will rebuild the application or system and rerun it to see where exactly the error occurred.
  2. In a more complex application, we will just set the environment variable CUDA_LAUNCH_BLOCKING=1 to force all kernel launches to be synchronous. There is no need to rebuild the application or system. We will just rerun it to see where exactly the error occurred.

In this blog post, I would like to quickly discuss why the second approach CUDA_LAUNCH_BLOCKING=1 is favored over the first approach for debugging CUDA programs.

CUDA_LAUNCH_BLOCKING=1 Is Favored

There are some analogies of CUDA_LAUNCH_BLOCKING=1, saying that it is equivalent as if you put a cudaDeviceSynchronize() or cudaStreamSynchronize(stream) after every kernel launch, i.e., the two approaches mentioned above are equivalent. However, this is incorrect. There are scenarios where the latter approach could reveal the correct error location, while the former approach could not.

Suppose we have two CPU threads. On CPU thread 1, we launch kernel A on the CUDA stream_1 followed by a cudaStreamSynchronize(stream_1). On CPU thread 2, we launch kernel B on the CUDA stream_2 followed by a cudaStreamSynchronize(stream_2). This approach cannot guarantee that the execution of kernel A has no overlap with the execution of kernel B. For example, if the CPU instruction issue order is launch A on stream_1, launch B on stream_2, synchronize stream_1, synchronize stream_2, then the execution of kernel A and kernel B could overlap. For some simple asynchronous errors, such as illegal memory access, this approach might still reveal the correct error location since the synchronization is stream specific. However, for debugging more complex problems, such as a racing condition due to kernel A and kernel B writing to the same memory location, this approach will likely fail to help us root cause the problem. I once got a CUDA kernel which runs perfectly fine in a single-CPU-thread and single-CUDA-stream environment. However, when I run it in a multi-CPU-thread and multi-CUDA-stream environment, it will sometimes produce incorrect results. It turns out that the kernel has a global __device__ variable which will be overwritten in the CUDA kernel. Launching the kernel in multiple CPU threads and multiple CUDA streams will cause racing conditions on the global __device__ variable, hence producing incorrect results. In this case, if I could make CUDA kernel launches truly synchronous, I would be able to observe that the program produces correct results, confirming my debugging hypothesis.

Using CUDA_LAUNCH_BLOCKING=1, the execution of kernel A and kernel B will never overlap. The mental model of CUDA_LAUNCH_BLOCKING=1 is that GPU will only execute one kernel at a time, i.e., there is only one CUDA stream available, and the kernel launch call on the CPU thread will not return until the kernel finishes execution. Because of this, no matter how many CPU threads and CUDA streams we have, CUDA kernel launches will always be synchronous.

Therefore, using CUDA_LAUNCH_BLOCKING=1 should be the preferred approach to debug asynchronous errors in CUDA applications and it seems to be easier to use than the other approach as well.

References