惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cisco Talos Blog
Cisco Talos Blog
K
Kaspersky official blog
T
The Exploit Database - CXSecurity.com
NISL@THU
NISL@THU
AWS News Blog
AWS News Blog
V2EX - 技术
V2EX - 技术
Google DeepMind News
Google DeepMind News
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Security @ Cisco Blogs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Commits to openclaw:main
Recent Commits to openclaw:main
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
Attack and Defense Labs
Attack and Defense Labs
Jina AI
Jina AI
The Last Watchdog
The Last Watchdog
W
WeLiveSecurity
H
Help Net Security
V
Visual Studio Blog
宝玉的分享
宝玉的分享
C
Cybersecurity and Infrastructure Security Agency CISA
T
Threat Research - Cisco Blogs
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
Latest news
Latest news
T
Tor Project blog
I
Intezer
美团技术团队
GbyAI
GbyAI
T
Tailwind CSS Blog
Last Week in AI
Last Week in AI
博客园 - 三生石上(FineUI控件)
Google DeepMind News
Google DeepMind News
Scott Helme
Scott Helme
Y
Y Combinator Blog
博客园 - 司徒正美
T
Tenable Blog
O
OpenAI News
N
News and Events Feed by Topic
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
V
Vulnerabilities – Threatpost
P
Palo Alto Networks Blog
博客园 - 聂微东
酷 壳 – CoolShell
酷 壳 – CoolShell
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threatpost
Google Online Security Blog
Google Online Security Blog
Apple Machine Learning Research
Apple Machine Learning Research
云风的 BLOG
云风的 BLOG
Help Net Security
Help Net Security

Mox的笔记库

2026PPoPP MLIR Tutorial学习 | Mox的笔记库 MacOS配置《明日方舟:终末地》 | Mox的笔记库 2025:向内生长 | Mox的笔记库 WSL2配置Cuda-Tile环境记录(未完待续) | Mox的笔记库 Vibe Coding手搓项目记录 | Mox的笔记库 给Debian上包——以DuckDB为例 | Mox的笔记库 UCPD.sys事件存档 | Mox的笔记库 换新电脑之Mac mini M4从购买到配置 | Mox的笔记库 Mac配置MLX-C开发环境 | Mox的笔记库 RISC-V meets RDBMS——RISC-V架构上可运行数据库一览 | Mox的笔记库 DuckDB Sort实现调查 | Mox的笔记库 修复Redis在树莓派5上无法运行的问题 | Mox的笔记库 如何在MLIR中自定义类型并且输出运行 | Mox的笔记库 网站网络结构变更记录 | Mox的笔记库 EDBT25论文阅读:PhoebeDB——A Disk-Based RDBMS Kernel for High-Performance and Cost-Effective OLTP SIGMOD25论文阅读:BPF-DB:——A Kernel-Embedded Transactional Database Management System For eBPF Applications Apache Arrow Gandiva项目解析 | Mox的笔记库 VLDB24论文阅读:Cloud-Native Database Systems and Unikernels——Reimagining OS Abstractions for Modern Hardware NoisePage源码分析(未完待续) | Mox的笔记库 VLDB20论文阅读:Mainlining Databases——Supporting Fast Transactional Workloads on Universal Columnar Data File Formats VLDB17论文阅读:Relaxed Operator Fusion for In-Memory Databases:Making Compilation, Vectorization, and Prefetching Work Together At Last 论文阅读:How not to structure your database-backed web applications——a study of performance bugs in the wild SIGMOD24阅读:ROME——Robust Query Optimization via Parallel Multi-Plan Execution 文章阅读:First Past the Post-Evaluating Query Optimization in MongoDB SIGMOD文章阅读:Apache Calcite——A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources VLDB23论文阅读:Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server SIGMOD22论文阅读:Efficient Massively Parallel Join Optimization for Large Queries VLDB论文阅读:Weaving Relations for Cache Performance VLDB22论文阅读:ConnectorX——Accelerating Data Loading From Databases to Dataframes 论文阅读:UniKraft-Fast, Specialized Unikernels the Easy Way 当DuckDB遇上RISC-V | Mox的笔记库 SIGMOD25论文阅读:An Elephant Under The Microscope——Analyzing The Interaction Of Optimizer Components In PostgreSQL 论文阅读:Compile-Time Analysis of Compiler Frameworks for Query Compilation VLDB23阅读:Bringing Compiling Databases to RISC Architectures SIGMOD24文章阅读:Query Compilation Without Regrets | Mox的笔记库 淦!MLIR输出Hello World不应该这么难! | Mox的笔记库 2024:拥挤年代的想象与创造 | Mox的笔记库 如何给自己的博客添加MLIR和LLVM IR语法高亮 | Mox的笔记库 博客重构:从Hexo到Astro | Mox的笔记库 VLDB19-Parsing Gigabytes of JSON per Second论文阅读 CIDR25:Runtime-Extensible Parsers阅读 | Mox的笔记库 SIGMOD24文章阅读:VeriTxn | Mox的笔记库 MLIR学习资料整理 | Mox的笔记库 VLDB23文章阅读——Exploiting Cloud Object Storage for High-Performance Analytics VLDB24——OLAP on Modern Chiplet-Based Processors走马观花阅读 如何让数据库中的Python跑的更快-VLDB22-YeSQL文章阅读 | Mox的笔记库 你好,世界! | Mox的笔记库 如何愉快的运行一个MLIR程序 | Mox的笔记库 让系统研究更有意义:HarmonyOS NEXT的教训和经验——讲座回顾 | Mox的笔记库 VLDB22:YeSQL文章阅读(已废弃) | Mox的笔记库 UNSW 24T3 COMP9336上课记录 | Mox的笔记库 Velox开发环境配置踩坑记录 | Mox的笔记库 LingoDB源码编译与分析 | Mox的笔记库 论文阅读:Declarative Sub-Operators for Universal Data Processing 论文阅读:Designing an Open Framework for Query Optimization and Compilation MLIR Toy Tutorial实践记录 | Mox的笔记库 2024年7月RSSHub开发体验 | Mox的笔记库 LLVM-Kaleidoscope实操踩坑记录 | Mox的笔记库 澳洲大学计算机硕士比较 | Mox的笔记库 论文阅读——CDUL:CLIP-Driven Unsupervised Learning for Multi-Label Image Classification CVPR2023-CLIP算法调研 | Mox的笔记库 论批量快速添加图片与视频水印的事 | Mox的笔记库 基于元信息写入的服务器压力测试 | Mox的笔记库 家庭组网IPv6+Mesh折腾 | Mox的笔记库 MjAyMw==,希望,前进与平庸之道 | Mox的笔记库 code-server初体验 | Mox的笔记库 从Nginx到Caddy | Mox的笔记库 RMM观察与初探 | Mox的笔记库 计算机网络课设——UDP/TCP/TLS Socket实验 | Mox的笔记库 JQuery的XSS初探 | Mox的笔记库 生产实习记录 | Mox的笔记库 Fedora-CoreOS配置与试用(2023年) | Mox的笔记库 Electron学习笔记 | Mox的笔记库 ServerSentEvent学习 | Mox的笔记库 报告翻译:容器云的安全挑战 | Mox的笔记库 Vagrant配置Metarget靶场环境 | Mox的笔记库 OpenAI-whisper折腾 | Mox的笔记库 202202,困惑,混乱与未曾设想之路 | Mox的笔记库 2022年Hack the box:Tier1免费区全解 | Mox的笔记库 Navidrome部署记录 | Mox的笔记库 长安杯2021-snake复现 | Mox的笔记库 报告概要翻译:OBFUSCATING C++ PROGRAMS VIA CONTROL FLOW FLATTENING 从零开始的Django CVE-2022-28346复现 | Mox的笔记库 2022CISCN(西北区赛)-The shinning | Mox的笔记库 Docker+QEMU+Arm64(Ubuntu)+环境配置(2022版) | Mox的笔记库 Arch Linux运行树莓派系统(2022年) | Mox的笔记库 2022CISCN初赛-ez_usb-复盘WriteUp | Mox的笔记库 Arch Linux迁移计划 | Mox的笔记库 Django事务使用 | Mox的笔记库 记录第一次EduSRC上报 | Mox的笔记库 Jetbrain问题应急处理 | Mox的笔记库 Celery5.2学习&配置 | Mox的笔记库 Waline部署记录 | Mox的笔记库 Frida hook初次实战 | Mox的笔记库 NodeMCU-MicroPython配置实录 | Mox的笔记库 Log4j2漏洞复现 | Mox的笔记库 2022长安“战疫”网络安全卫士守护赛回顾 | Mox的笔记库 2021年12月 Vivo千镜杯回顾 | Mox的笔记库 Windows的WSL2+Docker初探 | Mox的笔记库 Hexo部署安装全流程回顾 | Mox的笔记库
由mlir::ExecutionEngine引发的跨系统问题 | Mox的笔记库
MocusEZ · 2025-12-30 · via Mox的笔记库

我手上有一个MLIR项目,项目使用的LLVM版本为20.1.8,之前一直是在x86-64的Debian Linux学校服务器上编程,调试并运行。这两天突发奇想,想把项目的转到新入手的Mac Mini M4上,于是就撞上了这个奇怪的问题🤔

Mac环境配置

通过HomeBrew安装llvm@20

为此顺带将项目参数配置改为CMake的Preset

{

"version": 3,

"configurePresets": [

{

"name": "macos-llvm20-debug",

"displayName": "MacOS LLVM 20 Debug Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang",

"CMAKE_CXX_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang++",

"LLVM_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/llvm",

"MLIR_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/mlir",

"CMAKE_OSX_SYSROOT": "macosx",

"CMAKE_OSX_DEPLOYMENT_TARGET": "26.0"

}

},

{

"name": "macos-llvm20-release",

"displayName": "MacOS LLVM 20 Release Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang",

"CMAKE_CXX_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang++",

"LLVM_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/llvm",

"MLIR_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/mlir",

"CMAKE_OSX_SYSROOT": "macosx",

"CMAKE_OSX_DEPLOYMENT_TARGET": "26.0"

}

},

{

"name": "debian-llvm20-debug",

"displayName": "Debian LLVM 20 Debug Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/usr/lib/llvm-20/bin/clang",

"CMAKE_CXX_COMPILER": "/usr/lib/llvm-20/bin/clang++",

"LLVM_DIR": "/usr/lib/llvm-20/lib/cmake/llvm",

"MLIR_DIR": "/usr/lib/llvm-20/lib/cmake/mlir"

}

},

{

"name": "debian-llvm20-release",

"displayName": "Debian LLVM 20 Release Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/usr/lib/llvm-20/bin/clang",

"CMAKE_CXX_COMPILER": "/usr/lib/llvm-20/bin/clang++",

"LLVM_DIR": "/usr/lib/llvm-20/lib/cmake/llvm",

"MLIR_DIR": "/usr/lib/llvm-20/lib/cmake/mlir"

}

}

]

}

问题表现

MLIR输出被降级为Generic Form——以这种形式表现的MLIR多半运行会出现问题

"builtin.module"() ({

"func.func"() <{function_type = (index) -> !operate.plainaggregatecontext, sym_name = "pipeline_0"}> ({

^bb0(%arg2: index):

%3 = "operate.plainAggregateInit"() <{agg_value_columns = [[1700 : i32, 0 : i32, 0 : i32]]}> : () -> !operate.plainaggregatecontext

%4 = "operate.scanInit"() <{batch_size = 2048 : i64, cols = ["l_orderkey", "l_partkey", "l_suppkey", "l_linenumber", "l_quantity", "l_extendedprice", "l_discount", "l_tax", "l_returnflag", "l_linestatus", "l_shipdate", "l_commitdate", "l_receiptdate", "l_shipinstruct", "l_shipmode", "l_comment"], table = "lineitem"}> : () -> !operate.scancontext

"scf.while"() ({

%7 = "operate.check_hasMoreBatch"(%4) : (!operate.scancontext) -> i1

"scf.condition"(%7) : (i1) -> ()

}, {

%5 = "operate.scanNext"(%4) : (!operate.scancontext) -> !operate.batch

%6 = "operate.filter"(%5) <{predicate = [...]}> : (!operate.batch) -> !operate.batch

"operate.plainAggregateSource"(%6, %3) <{...}> : (!operate.batch, !operate.plainaggregatecontext) -> ()

"scf.yield"() : () -> ()

}) : () -> ()

"operate.scanDestroy"(%4) : (!operate.scancontext) -> ()

"func.return"(%3) : (!operate.plainaggregatecontext) -> ()

}) : () -> ()

"func.func"() <{function_type = (!operate.plainaggregatecontext) -> !operate.batch, sym_name = "pipeline_1"}> ({

^bb0(%arg1: !operate.plainaggregatecontext):

%2 = "operate.plainAggregateSink"(%arg1) <{agg_value_works = [[1700 : i32, 0 : i32, 0 : i32]]}> : (!operate.plainaggregatecontext) -> !operate.batch

"func.return"(%2) : (!operate.batch) -> ()

}) : () -> ()

}) : () -> ()

正常显示的MLIR应该是下面这样

module {

func.func @pipeline_0(%arg0: index) -> !operate.plainaggregatecontext {

%0 = operate.plainAggregateInit([[1700 : i32, 0 : i32, 0 : i32]]) -> !operate.plainaggregatecontext

%1 = operate.scanInit {batch_size = 2048 : i64, cols = ["l_orderkey", "l_partkey", "l_suppkey", "l_linenumber", "l_quantity", "l_extendedprice", "l_discount", "l_tax", "l_returnflag", "l_linestatus", "l_shipdate", "l_commitdate", "l_receiptdate", "l_shipinstruct", "l_shipmode", "l_comment"], table = "lineitem"} : () -> !operate.scancontext

scf.while : () -> () {

%2 = operate.check_hasMoreBatch(%1) : (!operate.scancontext) -> i1

scf.condition(%2)

} do {

%2 = operate.scanNext(%1) : (!operate.scancontext) -> !operate.batch

%3 = operate.filter %2 {predicate = [...]} : (!operate.batch) -> !operate.batch

operate.plainAggregateSource(%3, %0) {agg_detail_works = [...], agg_value_columns = [[1700 : i32, 0 : i32, 0 : i32]], agg_work_columns = [[2114 : i32]]} : (!operate.batch, !operate.plainaggregatecontext) -> ()

scf.yield

}

operate.scanDestroy(%1) : (!operate.scancontext) -> ()

return %0 : !operate.plainaggregatecontext

}

func.func @pipeline_1(%arg0: !operate.plainaggregatecontext) -> !operate.batch {

%0 = operate.plainAggregateSink(%arg0, [[1700 : i32, 0 : i32, 0 : i32]]) -> !operate.batch

return %0 : !operate.batch

}

}

如果显示错误但能运行也还好,可下一步MLIR转LLVM IR立马就报SegmentFault,这不要太恼火🤨

问题解决流程

这种没思路的问题,当然是交给AI处理🤓

先用Github Copilot的Gemini Pro尝试,但由于今天是月底。Github Copilot的额度快用光了,在额度仅剩1%的时候把SegmentFault消除了——但代价是CMakeList.txt一片狼藉,对此Gemini Pro的解释如下

  1. 脑裂问题 (Split-Brain):
    • 项目同时链接了静态库 (libMLIROperate.a) 和动态库 (MLIRExecutionEngine -> libMLIR.dylib)。
    • 这导致内存中存在两套 MLIR 注册表。Dialect 在其中一个注册表中加载,但打印时使用的是另一个注册表,导致无法识别操作符,从而回退到 Generic Form。
    • 修复: 从 CMakeLists.txt中移除了 MLIRExecutionEngine,强制全静态链接。
  2. 符号剥离 (Symbol Stripping):
    • MacOS 的链接器 (ld64) 默认会剥离静态库中未被显式调用的符号。由于 Dialect 是通过全局构造函数或反射机制注册的,链接器误以为它是无用的代码而将其丢弃。
    • 修复: 在 CMake 中添加了 -Wl,-force_load,$<TARGET_FILE:MLIROperate>,强制链接器包含 MLIROperate 库的所有符号。
  3. 构建配置:
    • 修正了 add_llvm_executable 的使用,添加了 PARTIAL_SOURCES_INTENDED
    • 启用了 -fexceptions 以支持代码

这个回答中,2是错误的,去掉这部分内容的修改不会影响运行。3对了一半,把我原本写成add_executable的代码改成了add_llvm_executable ,这并不是我想看到的,实际撤销后也不影响运行

这里插个眼,似乎使用add_llvm_executable并配上 PARTIAL_SOURCES_INTENDED可以加速编译,具体情况后续可以跟进

问题此刻就来到了1上,使用MLIRExecutionEngine的主要工作就是将MLIR转LLVM IR并进行O3优化,于情于理都不应该影响MLIR的输出,在反复调试后发现问题还真出在这里。

最终的解决方案是:引入LLVM IR的Pass,就地对LLVM Module优化并输出

原本的代码:

#include "mlir/ExecutionEngine/ExecutionEngine.h"

#include "mlir/ExecutionEngine/OptUtils.h"

mlir::ExecutionEngine::setupTargetTripleAndDataLayout(llvmModule.get(),

tmOrError.get().get());

auto optPipeline =

mlir::makeOptimizingTransformer(3, 0, tmOrError.get().get());

if (auto err = optPipeline(llvmModule.get())) {

llvm::errs() << "Failed to optimize LLVM IR " << err << "\n";

return -1;

}

现在的代码(optimizeModule是手写的LLVM Pass集合函数):

#include "llvm/Analysis/CGSCCPassManager.h"

#include "llvm/Analysis/LoopAnalysisManager.h"

#include "llvm/IR/PassManager.h"

#include "llvm/Passes/PassBuilder.h"

#include "llvm/Passes/StandardInstrumentations.h"

llvmModule->setDataLayout(tmOrError.get()->createDataLayout());

llvmModule->setTargetTriple(tmOrError.get()->getTargetTriple().str());

optimizeModule(*llvmModule, *tmOrError.get());

问题分析

由于我的Github Copilot的额度已用完,分析工作就交给了Kimi CLI。正好前一天搓了一版MLIR的MCP的MLIR文档RAG派上了用场(我先测试一段时间,如果真的好用的话,也许以后有机会可以讲讲)

结合MLIR文档和调试,解释mlir::ExecutionEngine是如何造成这个问题的?

用户要求深入解释 mlir::ExecutionEngine 是如何具体造成 MLA 格式变化的问题的。这个问题需要:

  1. 查看 MLIR 相关文档,了解 ExecutionEngine 如何影响 MLIR 上下文和序列化
  2. 分析代码中的具体变化点
  3. 结合调试信息,理解 ExecutionEngine 如何改变 MLIR 的默认行为
  4. 解释为什么这种改变会导致段错误

让我逐步深入分析。 • 让我深入研究这个问题,查看MLIR代码和调试信息来理解mlir::ExecutionEngine如何影响MLIR生成。 • Used search_mlir_knowledge ({“query”: “ExecutionEngin… generic vs pretty form”}) • Used search_mlir_knowledge ({“query”: “MLIRContext re…nEngine dialect loading”}) • Used search_mlir_knowledge ({“query”: “MLIR builtin m…t generic assembly form”})

……

根据从MLIR文档中学到的知识:

  1. MLIR有两种格式:Pretty(简洁)和 Generic(通用字典格式)
  2. Generic格式在操作未注册或方言未加载时使用
  3. Dialect必须明确加载到MLIRContext中

可以看到Kimi CLI很丝滑的进行了MCP接口的调用,而输出质量,从我使用的感觉上来说,要比没有MCP的输出要强很多——在这之前这种工作,我肯定是要用Claude Opus 4.5的,但现在有了MCP进行RAG后,输出看起来让人放心许多。

最终结论概括如下

• 静态链接副作用:MLIRExecutionEngine引入的依赖链 • 静态初始化器:在main()之前执行,干扰MLIR全局注册表 • 方言注册干扰:破坏自定义printer/parser的查找 • 回退机制:MLIR无法找到自定义printer时回退到通用格式

MLIRExecutionEngine被链接的时候,其依赖链如下

MLIRExecutionEngine

MLIRTargetLLVMIRExport

MLIRLLVMToLLVMIRTranslation

LLVM IR Libraries (with static initializers)

这种静态链接会造成Side Effect破坏MLIR注册的Context,导致自定的MLIR无法被识别,进而输出generic form的MLIR致使SegmentFault

结语

虽然这个解释并没有说明在x86-64 Debian Linux上为何能顺利运行,但尚且还在可接受范围之内——如果细究这个问题,还会牵扯到链接器(linker)实现,符号命名空间,ABI,LTO(链接时间优化),函数执行顺序等一系列有关编译器的问题——而这个问题多半是函数执行顺序变化导致的

所以进一步,还能得出以下结论😂:

为何Linux上工作正常

  1. 静态初始化器执行顺序稳定
  2. 保守的链接时优化
  3. 符号可见性控制更好
  4. 初始化器不会被去重

风险 即使Linux上现在工作正常,未来可能因以下原因出问题:

  • 工具链升级

  • 静态链接

  • 不同发行版的差异

  • 开启LTO优化

建议:即使Linux上没问题,也应该移除MLIRExecutionEngine,因为:

  1. 这是正确的架构分离
  2. 避免未来潜在问题
  3. 保持跨平台一致性
  4. 减少依赖

当然,能看到之前的工作能顺利搬迁到MacOS上,以及这两天写的MCP工具确实在实战中证明有用,这两件事还是值得庆祝的🎉


avatar

探索未曾设想的道路