惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
The Last Watchdog
The Last Watchdog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
G
Google Developers Blog
博客园 - 叶小钗
雷峰网
雷峰网
人人都是产品经理
人人都是产品经理
博客园_首页
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 三生石上(FineUI控件)
Help Net Security
Help Net Security
Cloudbric
Cloudbric
AI
AI
N
News | PayPal Newsroom
博客园 - 聂微东
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 【当耐特】
Forbes - Security
Forbes - Security
美团技术团队
Stack Overflow Blog
Stack Overflow Blog
SecWiki News
SecWiki News
H
Heimdal Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
MyScale Blog
MyScale Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
P
Proofpoint News Feed
S
Security @ Cisco Blogs
Google DeepMind News
Google DeepMind News
V
V2EX
大猫的无限游戏
大猫的无限游戏
阮一峰的网络日志
阮一峰的网络日志
S
Security Affairs
L
LangChain Blog
The Hacker News
The Hacker News
F
Full Disclosure
aimingoo的专栏
aimingoo的专栏
Hacker News - Newest:
Hacker News - Newest: "LLM"
腾讯CDC
Webroot Blog
Webroot Blog
A
About on SuperTechFans
H
Hacker News: Front Page
Cyberwarzone
Cyberwarzone
WordPress大学
WordPress大学
L
LINUX DO - 热门话题
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Attack and Defense Labs
Attack and Defense Labs
M
MIT News - Artificial intelligence

Mox的笔记库

2026PPoPP MLIR Tutorial学习 | Mox的笔记库 MacOS配置《明日方舟:终末地》 | Mox的笔记库 2025:向内生长 | Mox的笔记库 WSL2配置Cuda-Tile环境记录(未完待续) | Mox的笔记库 Vibe Coding手搓项目记录 | Mox的笔记库 给Debian上包——以DuckDB为例 | Mox的笔记库 UCPD.sys事件存档 | Mox的笔记库 换新电脑之Mac mini M4从购买到配置 | Mox的笔记库 Mac配置MLX-C开发环境 | Mox的笔记库 RISC-V meets RDBMS——RISC-V架构上可运行数据库一览 | Mox的笔记库 DuckDB Sort实现调查 | Mox的笔记库 修复Redis在树莓派5上无法运行的问题 | Mox的笔记库 如何在MLIR中自定义类型并且输出运行 | Mox的笔记库 网站网络结构变更记录 | Mox的笔记库 EDBT25论文阅读:PhoebeDB——A Disk-Based RDBMS Kernel for High-Performance and Cost-Effective OLTP SIGMOD25论文阅读:BPF-DB:——A Kernel-Embedded Transactional Database Management System For eBPF Applications Apache Arrow Gandiva项目解析 | Mox的笔记库 VLDB24论文阅读:Cloud-Native Database Systems and Unikernels——Reimagining OS Abstractions for Modern Hardware NoisePage源码分析(未完待续) | Mox的笔记库 VLDB20论文阅读:Mainlining Databases——Supporting Fast Transactional Workloads on Universal Columnar Data File Formats VLDB17论文阅读:Relaxed Operator Fusion for In-Memory Databases:Making Compilation, Vectorization, and Prefetching Work Together At Last 论文阅读:How not to structure your database-backed web applications——a study of performance bugs in the wild SIGMOD24阅读:ROME——Robust Query Optimization via Parallel Multi-Plan Execution 文章阅读:First Past the Post-Evaluating Query Optimization in MongoDB SIGMOD文章阅读:Apache Calcite——A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources VLDB23论文阅读:Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server SIGMOD22论文阅读:Efficient Massively Parallel Join Optimization for Large Queries VLDB论文阅读:Weaving Relations for Cache Performance VLDB22论文阅读:ConnectorX——Accelerating Data Loading From Databases to Dataframes 论文阅读:UniKraft-Fast, Specialized Unikernels the Easy Way 当DuckDB遇上RISC-V | Mox的笔记库 SIGMOD25论文阅读:An Elephant Under The Microscope——Analyzing The Interaction Of Optimizer Components In PostgreSQL 论文阅读:Compile-Time Analysis of Compiler Frameworks for Query Compilation VLDB23阅读:Bringing Compiling Databases to RISC Architectures SIGMOD24文章阅读:Query Compilation Without Regrets | Mox的笔记库 淦!MLIR输出Hello World不应该这么难! | Mox的笔记库 2024:拥挤年代的想象与创造 | Mox的笔记库 如何给自己的博客添加MLIR和LLVM IR语法高亮 | Mox的笔记库 博客重构:从Hexo到Astro | Mox的笔记库 VLDB19-Parsing Gigabytes of JSON per Second论文阅读 CIDR25:Runtime-Extensible Parsers阅读 | Mox的笔记库 SIGMOD24文章阅读:VeriTxn | Mox的笔记库 MLIR学习资料整理 | Mox的笔记库 VLDB23文章阅读——Exploiting Cloud Object Storage for High-Performance Analytics VLDB24——OLAP on Modern Chiplet-Based Processors走马观花阅读 如何让数据库中的Python跑的更快-VLDB22-YeSQL文章阅读 | Mox的笔记库 你好,世界! | Mox的笔记库 如何愉快的运行一个MLIR程序 | Mox的笔记库 让系统研究更有意义:HarmonyOS NEXT的教训和经验——讲座回顾 | Mox的笔记库 VLDB22:YeSQL文章阅读(已废弃) | Mox的笔记库 UNSW 24T3 COMP9336上课记录 | Mox的笔记库 Velox开发环境配置踩坑记录 | Mox的笔记库 LingoDB源码编译与分析 | Mox的笔记库 论文阅读:Declarative Sub-Operators for Universal Data Processing 论文阅读:Designing an Open Framework for Query Optimization and Compilation MLIR Toy Tutorial实践记录 | Mox的笔记库 2024年7月RSSHub开发体验 | Mox的笔记库 LLVM-Kaleidoscope实操踩坑记录 | Mox的笔记库 澳洲大学计算机硕士比较 | Mox的笔记库 论文阅读——CDUL:CLIP-Driven Unsupervised Learning for Multi-Label Image Classification CVPR2023-CLIP算法调研 | Mox的笔记库 论批量快速添加图片与视频水印的事 | Mox的笔记库 基于元信息写入的服务器压力测试 | Mox的笔记库 家庭组网IPv6+Mesh折腾 | Mox的笔记库 MjAyMw==,希望,前进与平庸之道 | Mox的笔记库 code-server初体验 | Mox的笔记库 从Nginx到Caddy | Mox的笔记库 RMM观察与初探 | Mox的笔记库 计算机网络课设——UDP/TCP/TLS Socket实验 | Mox的笔记库 JQuery的XSS初探 | Mox的笔记库 生产实习记录 | Mox的笔记库 Fedora-CoreOS配置与试用(2023年) | Mox的笔记库 Electron学习笔记 | Mox的笔记库 ServerSentEvent学习 | Mox的笔记库 报告翻译:容器云的安全挑战 | Mox的笔记库 Vagrant配置Metarget靶场环境 | Mox的笔记库 OpenAI-whisper折腾 | Mox的笔记库 202202,困惑,混乱与未曾设想之路 | Mox的笔记库 2022年Hack the box:Tier1免费区全解 | Mox的笔记库 Navidrome部署记录 | Mox的笔记库 长安杯2021-snake复现 | Mox的笔记库 报告概要翻译:OBFUSCATING C++ PROGRAMS VIA CONTROL FLOW FLATTENING 从零开始的Django CVE-2022-28346复现 | Mox的笔记库 2022CISCN(西北区赛)-The shinning | Mox的笔记库 Docker+QEMU+Arm64(Ubuntu)+环境配置(2022版) | Mox的笔记库 Arch Linux运行树莓派系统(2022年) | Mox的笔记库 2022CISCN初赛-ez_usb-复盘WriteUp | Mox的笔记库 Arch Linux迁移计划 | Mox的笔记库 Django事务使用 | Mox的笔记库 记录第一次EduSRC上报 | Mox的笔记库 Jetbrain问题应急处理 | Mox的笔记库 Celery5.2学习&配置 | Mox的笔记库 Waline部署记录 | Mox的笔记库 Frida hook初次实战 | Mox的笔记库 NodeMCU-MicroPython配置实录 | Mox的笔记库 Log4j2漏洞复现 | Mox的笔记库 2022长安“战疫”网络安全卫士守护赛回顾 | Mox的笔记库 2021年12月 Vivo千镜杯回顾 | Mox的笔记库 Windows的WSL2+Docker初探 | Mox的笔记库 Hexo部署安装全流程回顾 | Mox的笔记库
由mlir::ExecutionEngine引发的跨系统问题 | Mox的笔记库
MocusEZ · 2025-12-30 · via Mox的笔记库

我手上有一个MLIR项目,项目使用的LLVM版本为20.1.8,之前一直是在x86-64的Debian Linux学校服务器上编程,调试并运行。这两天突发奇想,想把项目的转到新入手的Mac Mini M4上,于是就撞上了这个奇怪的问题🤔

Mac环境配置

通过HomeBrew安装llvm@20

为此顺带将项目参数配置改为CMake的Preset

{

"version": 3,

"configurePresets": [

{

"name": "macos-llvm20-debug",

"displayName": "MacOS LLVM 20 Debug Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang",

"CMAKE_CXX_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang++",

"LLVM_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/llvm",

"MLIR_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/mlir",

"CMAKE_OSX_SYSROOT": "macosx",

"CMAKE_OSX_DEPLOYMENT_TARGET": "26.0"

}

},

{

"name": "macos-llvm20-release",

"displayName": "MacOS LLVM 20 Release Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang",

"CMAKE_CXX_COMPILER": "/opt/homebrew/opt/llvm@20/bin/clang++",

"LLVM_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/llvm",

"MLIR_DIR": "/opt/homebrew/opt/llvm@20/lib/cmake/mlir",

"CMAKE_OSX_SYSROOT": "macosx",

"CMAKE_OSX_DEPLOYMENT_TARGET": "26.0"

}

},

{

"name": "debian-llvm20-debug",

"displayName": "Debian LLVM 20 Debug Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/usr/lib/llvm-20/bin/clang",

"CMAKE_CXX_COMPILER": "/usr/lib/llvm-20/bin/clang++",

"LLVM_DIR": "/usr/lib/llvm-20/lib/cmake/llvm",

"MLIR_DIR": "/usr/lib/llvm-20/lib/cmake/mlir"

}

},

{

"name": "debian-llvm20-release",

"displayName": "Debian LLVM 20 Release Config",

"binaryDir": "${sourceDir}/build",

"generator": "Ninja",

"cacheVariables": {

"CMAKE_C_COMPILER": "/usr/lib/llvm-20/bin/clang",

"CMAKE_CXX_COMPILER": "/usr/lib/llvm-20/bin/clang++",

"LLVM_DIR": "/usr/lib/llvm-20/lib/cmake/llvm",

"MLIR_DIR": "/usr/lib/llvm-20/lib/cmake/mlir"

}

}

]

}

问题表现

MLIR输出被降级为Generic Form——以这种形式表现的MLIR多半运行会出现问题

"builtin.module"() ({

"func.func"() <{function_type = (index) -> !operate.plainaggregatecontext, sym_name = "pipeline_0"}> ({

^bb0(%arg2: index):

%3 = "operate.plainAggregateInit"() <{agg_value_columns = [[1700 : i32, 0 : i32, 0 : i32]]}> : () -> !operate.plainaggregatecontext

%4 = "operate.scanInit"() <{batch_size = 2048 : i64, cols = ["l_orderkey", "l_partkey", "l_suppkey", "l_linenumber", "l_quantity", "l_extendedprice", "l_discount", "l_tax", "l_returnflag", "l_linestatus", "l_shipdate", "l_commitdate", "l_receiptdate", "l_shipinstruct", "l_shipmode", "l_comment"], table = "lineitem"}> : () -> !operate.scancontext

"scf.while"() ({

%7 = "operate.check_hasMoreBatch"(%4) : (!operate.scancontext) -> i1

"scf.condition"(%7) : (i1) -> ()

}, {

%5 = "operate.scanNext"(%4) : (!operate.scancontext) -> !operate.batch

%6 = "operate.filter"(%5) <{predicate = [...]}> : (!operate.batch) -> !operate.batch

"operate.plainAggregateSource"(%6, %3) <{...}> : (!operate.batch, !operate.plainaggregatecontext) -> ()

"scf.yield"() : () -> ()

}) : () -> ()

"operate.scanDestroy"(%4) : (!operate.scancontext) -> ()

"func.return"(%3) : (!operate.plainaggregatecontext) -> ()

}) : () -> ()

"func.func"() <{function_type = (!operate.plainaggregatecontext) -> !operate.batch, sym_name = "pipeline_1"}> ({

^bb0(%arg1: !operate.plainaggregatecontext):

%2 = "operate.plainAggregateSink"(%arg1) <{agg_value_works = [[1700 : i32, 0 : i32, 0 : i32]]}> : (!operate.plainaggregatecontext) -> !operate.batch

"func.return"(%2) : (!operate.batch) -> ()

}) : () -> ()

}) : () -> ()

正常显示的MLIR应该是下面这样

module {

func.func @pipeline_0(%arg0: index) -> !operate.plainaggregatecontext {

%0 = operate.plainAggregateInit([[1700 : i32, 0 : i32, 0 : i32]]) -> !operate.plainaggregatecontext

%1 = operate.scanInit {batch_size = 2048 : i64, cols = ["l_orderkey", "l_partkey", "l_suppkey", "l_linenumber", "l_quantity", "l_extendedprice", "l_discount", "l_tax", "l_returnflag", "l_linestatus", "l_shipdate", "l_commitdate", "l_receiptdate", "l_shipinstruct", "l_shipmode", "l_comment"], table = "lineitem"} : () -> !operate.scancontext

scf.while : () -> () {

%2 = operate.check_hasMoreBatch(%1) : (!operate.scancontext) -> i1

scf.condition(%2)

} do {

%2 = operate.scanNext(%1) : (!operate.scancontext) -> !operate.batch

%3 = operate.filter %2 {predicate = [...]} : (!operate.batch) -> !operate.batch

operate.plainAggregateSource(%3, %0) {agg_detail_works = [...], agg_value_columns = [[1700 : i32, 0 : i32, 0 : i32]], agg_work_columns = [[2114 : i32]]} : (!operate.batch, !operate.plainaggregatecontext) -> ()

scf.yield

}

operate.scanDestroy(%1) : (!operate.scancontext) -> ()

return %0 : !operate.plainaggregatecontext

}

func.func @pipeline_1(%arg0: !operate.plainaggregatecontext) -> !operate.batch {

%0 = operate.plainAggregateSink(%arg0, [[1700 : i32, 0 : i32, 0 : i32]]) -> !operate.batch

return %0 : !operate.batch

}

}

如果显示错误但能运行也还好,可下一步MLIR转LLVM IR立马就报SegmentFault,这不要太恼火🤨

问题解决流程

这种没思路的问题,当然是交给AI处理🤓

先用Github Copilot的Gemini Pro尝试,但由于今天是月底。Github Copilot的额度快用光了,在额度仅剩1%的时候把SegmentFault消除了——但代价是CMakeList.txt一片狼藉,对此Gemini Pro的解释如下

  1. 脑裂问题 (Split-Brain):
    • 项目同时链接了静态库 (libMLIROperate.a) 和动态库 (MLIRExecutionEngine -> libMLIR.dylib)。
    • 这导致内存中存在两套 MLIR 注册表。Dialect 在其中一个注册表中加载,但打印时使用的是另一个注册表,导致无法识别操作符,从而回退到 Generic Form。
    • 修复: 从 CMakeLists.txt中移除了 MLIRExecutionEngine,强制全静态链接。
  2. 符号剥离 (Symbol Stripping):
    • MacOS 的链接器 (ld64) 默认会剥离静态库中未被显式调用的符号。由于 Dialect 是通过全局构造函数或反射机制注册的,链接器误以为它是无用的代码而将其丢弃。
    • 修复: 在 CMake 中添加了 -Wl,-force_load,$<TARGET_FILE:MLIROperate>,强制链接器包含 MLIROperate 库的所有符号。
  3. 构建配置:
    • 修正了 add_llvm_executable 的使用,添加了 PARTIAL_SOURCES_INTENDED
    • 启用了 -fexceptions 以支持代码

这个回答中,2是错误的,去掉这部分内容的修改不会影响运行。3对了一半,把我原本写成add_executable的代码改成了add_llvm_executable ,这并不是我想看到的,实际撤销后也不影响运行

这里插个眼,似乎使用add_llvm_executable并配上 PARTIAL_SOURCES_INTENDED可以加速编译,具体情况后续可以跟进

问题此刻就来到了1上,使用MLIRExecutionEngine的主要工作就是将MLIR转LLVM IR并进行O3优化,于情于理都不应该影响MLIR的输出,在反复调试后发现问题还真出在这里。

最终的解决方案是:引入LLVM IR的Pass,就地对LLVM Module优化并输出

原本的代码:

#include "mlir/ExecutionEngine/ExecutionEngine.h"

#include "mlir/ExecutionEngine/OptUtils.h"

mlir::ExecutionEngine::setupTargetTripleAndDataLayout(llvmModule.get(),

tmOrError.get().get());

auto optPipeline =

mlir::makeOptimizingTransformer(3, 0, tmOrError.get().get());

if (auto err = optPipeline(llvmModule.get())) {

llvm::errs() << "Failed to optimize LLVM IR " << err << "\n";

return -1;

}

现在的代码(optimizeModule是手写的LLVM Pass集合函数):

#include "llvm/Analysis/CGSCCPassManager.h"

#include "llvm/Analysis/LoopAnalysisManager.h"

#include "llvm/IR/PassManager.h"

#include "llvm/Passes/PassBuilder.h"

#include "llvm/Passes/StandardInstrumentations.h"

llvmModule->setDataLayout(tmOrError.get()->createDataLayout());

llvmModule->setTargetTriple(tmOrError.get()->getTargetTriple().str());

optimizeModule(*llvmModule, *tmOrError.get());

问题分析

由于我的Github Copilot的额度已用完,分析工作就交给了Kimi CLI。正好前一天搓了一版MLIR的MCP的MLIR文档RAG派上了用场(我先测试一段时间,如果真的好用的话,也许以后有机会可以讲讲)

结合MLIR文档和调试,解释mlir::ExecutionEngine是如何造成这个问题的?

用户要求深入解释 mlir::ExecutionEngine 是如何具体造成 MLA 格式变化的问题的。这个问题需要:

  1. 查看 MLIR 相关文档,了解 ExecutionEngine 如何影响 MLIR 上下文和序列化
  2. 分析代码中的具体变化点
  3. 结合调试信息,理解 ExecutionEngine 如何改变 MLIR 的默认行为
  4. 解释为什么这种改变会导致段错误

让我逐步深入分析。 • 让我深入研究这个问题,查看MLIR代码和调试信息来理解mlir::ExecutionEngine如何影响MLIR生成。 • Used search_mlir_knowledge ({“query”: “ExecutionEngin… generic vs pretty form”}) • Used search_mlir_knowledge ({“query”: “MLIRContext re…nEngine dialect loading”}) • Used search_mlir_knowledge ({“query”: “MLIR builtin m…t generic assembly form”})

……

根据从MLIR文档中学到的知识:

  1. MLIR有两种格式:Pretty(简洁)和 Generic(通用字典格式)
  2. Generic格式在操作未注册或方言未加载时使用
  3. Dialect必须明确加载到MLIRContext中

可以看到Kimi CLI很丝滑的进行了MCP接口的调用,而输出质量,从我使用的感觉上来说,要比没有MCP的输出要强很多——在这之前这种工作,我肯定是要用Claude Opus 4.5的,但现在有了MCP进行RAG后,输出看起来让人放心许多。

最终结论概括如下

• 静态链接副作用:MLIRExecutionEngine引入的依赖链 • 静态初始化器:在main()之前执行,干扰MLIR全局注册表 • 方言注册干扰:破坏自定义printer/parser的查找 • 回退机制:MLIR无法找到自定义printer时回退到通用格式

MLIRExecutionEngine被链接的时候,其依赖链如下

MLIRExecutionEngine

MLIRTargetLLVMIRExport

MLIRLLVMToLLVMIRTranslation

LLVM IR Libraries (with static initializers)

这种静态链接会造成Side Effect破坏MLIR注册的Context,导致自定的MLIR无法被识别,进而输出generic form的MLIR致使SegmentFault

结语

虽然这个解释并没有说明在x86-64 Debian Linux上为何能顺利运行,但尚且还在可接受范围之内——如果细究这个问题,还会牵扯到链接器(linker)实现,符号命名空间,ABI,LTO(链接时间优化),函数执行顺序等一系列有关编译器的问题——而这个问题多半是函数执行顺序变化导致的

所以进一步,还能得出以下结论😂:

为何Linux上工作正常

  1. 静态初始化器执行顺序稳定
  2. 保守的链接时优化
  3. 符号可见性控制更好
  4. 初始化器不会被去重

风险 即使Linux上现在工作正常,未来可能因以下原因出问题:

  • 工具链升级

  • 静态链接

  • 不同发行版的差异

  • 开启LTO优化

建议:即使Linux上没问题,也应该移除MLIRExecutionEngine,因为:

  1. 这是正确的架构分离
  2. 避免未来潜在问题
  3. 保持跨平台一致性
  4. 减少依赖

当然,能看到之前的工作能顺利搬迁到MacOS上,以及这两天写的MCP工具确实在实战中证明有用,这两件事还是值得庆祝的🎉


avatar

探索未曾设想的道路