惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

AI
AI
T
Tailwind CSS Blog
雷峰网
雷峰网
人人都是产品经理
人人都是产品经理
Microsoft Azure Blog
Microsoft Azure Blog
爱范儿
爱范儿
P
Proofpoint News Feed
D
Docker
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 【当耐特】
云风的 BLOG
云风的 BLOG
有赞技术团队
有赞技术团队
The Cloudflare Blog
Engineering at Meta
Engineering at Meta
酷 壳 – CoolShell
酷 壳 – CoolShell
U
Unit 42
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 司徒正美
博客园 - Franky
The GitHub Blog
The GitHub Blog
月光博客
月光博客
小众软件
小众软件
Martin Fowler
Martin Fowler
P
Palo Alto Networks Blog
Google DeepMind News
Google DeepMind News
Hugging Face - Blog
Hugging Face - Blog
T
Threat Research - Cisco Blogs
阮一峰的网络日志
阮一峰的网络日志
IT之家
IT之家
K
Kaspersky official blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Securelist
Spread Privacy
Spread Privacy
博客园 - 聂微东
B
Blog
The Hacker News
The Hacker News
Simon Willison's Weblog
Simon Willison's Weblog
T
The Exploit Database - CXSecurity.com
S
Schneier on Security
P
Privacy International News Feed
腾讯CDC
C
Cyber Attacks, Cyber Crime and Cyber Security
T
Threatpost
T
Tenable Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
Stack Overflow Blog
Stack Overflow Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Heimdal Security Blog

Mox的笔记库

2026PPoPP MLIR Tutorial学习 | Mox的笔记库 MacOS配置《明日方舟:终末地》 | Mox的笔记库 2025:向内生长 | Mox的笔记库 由mlir::ExecutionEngine引发的跨系统问题 | Mox的笔记库 WSL2配置Cuda-Tile环境记录(未完待续) | Mox的笔记库 Vibe Coding手搓项目记录 | Mox的笔记库 给Debian上包——以DuckDB为例 | Mox的笔记库 UCPD.sys事件存档 | Mox的笔记库 换新电脑之Mac mini M4从购买到配置 | Mox的笔记库 Mac配置MLX-C开发环境 | Mox的笔记库 RISC-V meets RDBMS——RISC-V架构上可运行数据库一览 | Mox的笔记库 DuckDB Sort实现调查 | Mox的笔记库 修复Redis在树莓派5上无法运行的问题 | Mox的笔记库 如何在MLIR中自定义类型并且输出运行 | Mox的笔记库 网站网络结构变更记录 | Mox的笔记库 EDBT25论文阅读:PhoebeDB——A Disk-Based RDBMS Kernel for High-Performance and Cost-Effective OLTP SIGMOD25论文阅读:BPF-DB:——A Kernel-Embedded Transactional Database Management System For eBPF Applications Apache Arrow Gandiva项目解析 | Mox的笔记库 VLDB24论文阅读:Cloud-Native Database Systems and Unikernels——Reimagining OS Abstractions for Modern Hardware NoisePage源码分析(未完待续) | Mox的笔记库 VLDB20论文阅读:Mainlining Databases——Supporting Fast Transactional Workloads on Universal Columnar Data File Formats VLDB17论文阅读:Relaxed Operator Fusion for In-Memory Databases:Making Compilation, Vectorization, and Prefetching Work Together At Last 论文阅读:How not to structure your database-backed web applications——a study of performance bugs in the wild SIGMOD24阅读:ROME——Robust Query Optimization via Parallel Multi-Plan Execution 文章阅读:First Past the Post-Evaluating Query Optimization in MongoDB SIGMOD文章阅读:Apache Calcite——A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources VLDB23论文阅读:Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server SIGMOD22论文阅读:Efficient Massively Parallel Join Optimization for Large Queries VLDB论文阅读:Weaving Relations for Cache Performance VLDB22论文阅读:ConnectorX——Accelerating Data Loading From Databases to Dataframes 论文阅读:UniKraft-Fast, Specialized Unikernels the Easy Way 当DuckDB遇上RISC-V | Mox的笔记库 SIGMOD25论文阅读:An Elephant Under The Microscope——Analyzing The Interaction Of Optimizer Components In PostgreSQL 论文阅读:Compile-Time Analysis of Compiler Frameworks for Query Compilation VLDB23阅读:Bringing Compiling Databases to RISC Architectures SIGMOD24文章阅读:Query Compilation Without Regrets | Mox的笔记库 淦!MLIR输出Hello World不应该这么难! | Mox的笔记库 2024:拥挤年代的想象与创造 | Mox的笔记库 如何给自己的博客添加MLIR和LLVM IR语法高亮 | Mox的笔记库 博客重构:从Hexo到Astro | Mox的笔记库 VLDB19-Parsing Gigabytes of JSON per Second论文阅读 CIDR25:Runtime-Extensible Parsers阅读 | Mox的笔记库 SIGMOD24文章阅读:VeriTxn | Mox的笔记库 MLIR学习资料整理 | Mox的笔记库 VLDB24——OLAP on Modern Chiplet-Based Processors走马观花阅读 如何让数据库中的Python跑的更快-VLDB22-YeSQL文章阅读 | Mox的笔记库 你好,世界! | Mox的笔记库 如何愉快的运行一个MLIR程序 | Mox的笔记库 让系统研究更有意义:HarmonyOS NEXT的教训和经验——讲座回顾 | Mox的笔记库 VLDB22:YeSQL文章阅读(已废弃) | Mox的笔记库 UNSW 24T3 COMP9336上课记录 | Mox的笔记库 Velox开发环境配置踩坑记录 | Mox的笔记库 LingoDB源码编译与分析 | Mox的笔记库 论文阅读:Declarative Sub-Operators for Universal Data Processing 论文阅读:Designing an Open Framework for Query Optimization and Compilation MLIR Toy Tutorial实践记录 | Mox的笔记库 2024年7月RSSHub开发体验 | Mox的笔记库 LLVM-Kaleidoscope实操踩坑记录 | Mox的笔记库 澳洲大学计算机硕士比较 | Mox的笔记库 论文阅读——CDUL:CLIP-Driven Unsupervised Learning for Multi-Label Image Classification CVPR2023-CLIP算法调研 | Mox的笔记库 论批量快速添加图片与视频水印的事 | Mox的笔记库 基于元信息写入的服务器压力测试 | Mox的笔记库 家庭组网IPv6+Mesh折腾 | Mox的笔记库 MjAyMw==,希望,前进与平庸之道 | Mox的笔记库 code-server初体验 | Mox的笔记库 从Nginx到Caddy | Mox的笔记库 RMM观察与初探 | Mox的笔记库 计算机网络课设——UDP/TCP/TLS Socket实验 | Mox的笔记库 JQuery的XSS初探 | Mox的笔记库 生产实习记录 | Mox的笔记库 Fedora-CoreOS配置与试用(2023年) | Mox的笔记库 Electron学习笔记 | Mox的笔记库 ServerSentEvent学习 | Mox的笔记库 报告翻译:容器云的安全挑战 | Mox的笔记库 Vagrant配置Metarget靶场环境 | Mox的笔记库 OpenAI-whisper折腾 | Mox的笔记库 202202,困惑,混乱与未曾设想之路 | Mox的笔记库 2022年Hack the box:Tier1免费区全解 | Mox的笔记库 Navidrome部署记录 | Mox的笔记库 长安杯2021-snake复现 | Mox的笔记库 报告概要翻译:OBFUSCATING C++ PROGRAMS VIA CONTROL FLOW FLATTENING 从零开始的Django CVE-2022-28346复现 | Mox的笔记库 2022CISCN(西北区赛)-The shinning | Mox的笔记库 Docker+QEMU+Arm64(Ubuntu)+环境配置(2022版) | Mox的笔记库 Arch Linux运行树莓派系统(2022年) | Mox的笔记库 2022CISCN初赛-ez_usb-复盘WriteUp | Mox的笔记库 Arch Linux迁移计划 | Mox的笔记库 Django事务使用 | Mox的笔记库 记录第一次EduSRC上报 | Mox的笔记库 Jetbrain问题应急处理 | Mox的笔记库 Celery5.2学习&配置 | Mox的笔记库 Waline部署记录 | Mox的笔记库 Frida hook初次实战 | Mox的笔记库 NodeMCU-MicroPython配置实录 | Mox的笔记库 Log4j2漏洞复现 | Mox的笔记库 2022长安“战疫”网络安全卫士守护赛回顾 | Mox的笔记库 2021年12月 Vivo千镜杯回顾 | Mox的笔记库 Windows的WSL2+Docker初探 | Mox的笔记库 Hexo部署安装全流程回顾 | Mox的笔记库
VLDB23文章阅读——Exploiting Cloud Object Storage for High-Performance Analytics
MocusEZ · 2024-11-24 · via Mox的笔记库

来自TUM的Database的论文

关键字:云数据仓库

下载链接

相关仓库

摘要选取

“远程网络和本地 NVMe 带宽之间的差距正在缩小,这使得云存储更具吸引力”

emmm,但你们想怎么解决高延迟的差距😅放到之前我也许会相信,但是今年黑5的云存储也不怎么诱人,价格并不吸引人

we present AnyBlob, a novel download manager for query engines that optimizes throughput while minimizing CPU usage.

期待(搓手手😋)

We discuss the integration of high-performance data retrieval in query engines and demonstrate it by incorporating AnyBlob in our database system Umbra

与Umbra结合,Sounds Good

Our experiments show that even without caching, Umbra with integrated AnyBlob achieves similar performance to state-of-the-art cloud data warehouses that cache data on local SSDs while improving resource elasticity.

他们想怎么比较?

正文选取

1 Introduction

In future data centers, database systems may even run on hardware that separates memory and compute.

俗称存算分离

In 2018, AWS introduced instances with 100 Gbit/s (≈12 GB/s) networking – resulting in a fourfold increase in per-instance bandwidth [22, 26].

In contrast to Infiniband, 100 Gbit/s Ethernet has not only become widely-available but also affordable

有点好奇AWS是如何实现100Gbit/s的网络的,自定义FPGA网卡必然是用上的,但不运行在Infiniband,那要怎么做?

Surprisingly, no empirical study for general-purpose analytics (OLAP) on cloud object stores has been conducted.

论文的Motivation:在云上进行OLAP还没怎么出过论文

Challenge 1: Achieving instance bandwidth. Because the latency of each object request is high, saturating high-bandwidth networks requires many concurrently outstanding requests. Therefore, a careful network integration into the DBMS is crucial to achieve the complete bandwidth available on network-optimized instances. Challenge 2: Network CPU overhead. In contrast to fetching

将网络集成到DMBS里解决高延迟问题?有意思😍

In contrast to the desire for multi-cloud systems, each cloud vendor provides its own networking library.

不同厂商的云技术很多是割裂的,但我认为K8s可以解决这个问题

Contribution 1: Experimental study of cloud object stores.

这种实验是动态的,2024年的实验不一定对2024有效

Contribution 2: AnyBlob, a low overhead multi-cloud library.

In contrast to existing solutions, our approach does not need to start new threads for parallel requests because it uses io_uring [21],

现在的数据库Paper基本都用io_uring吧😂

2 Cloud Storage Charaterstic

Methodlogy: 测AWS,Google,IBM,Azure,OCI三家公司的S3存储

image-20241123150655439

恕我直言:这个测试我感觉不好😅

既没有测试Minio这个最大的开源S3系统,也没有测Cloudflare D2这个廉价且有CDN的S3系统

选择的都是这几家的普通服务,也没有花钱看看大家云厂商是怎么优化的

2.1 Object Storage Architecture

问就是S3😆说到对象存储基本都是S3

2.2 Object Storage Cost

That‘s is a good question

Cost structure

每家都有自己的规则,都最终都会归结于存储成本和带宽成本

2.3 latency

Different request sizes

分布式存储会增大延迟,热数据会减少延迟

Noisy neighbors

高峰时间段,即使你没有高负载,你访问的延迟也会增大

Latency variations between cloud vendors.

差异不大

第二章先看到这里

2.4-2.8充斥着大量实验细节,感兴趣直接看原文

3 AnyBolb

使用异步请求减少线程带来的上下文切换(AWS的做法是每请求就通过Curl开一个线程)

3.1 AnyBlob Design

看图说话时间,但我怎么看都像是重新讲了一遍IO_Uring的工作原理

image-20241123153410899

3.2 Authentication & Security

Transparent authentication

透明化各家的认证系统

AnyBlob enables encryption-at-rest

推荐encryption-at-rest方案

静态数据加密的工作原理

  1. 数据加密
  • 数据在写入磁盘之前被加密。
  • 常用的加密算法包括 AES(Advanced Encryption Standard)。
  • 加密密钥由专门的密钥管理系统(KMS,Key Management System)管理。
  1. 数据解密
  • 数据在从磁盘读取到内存时被解密。
  • 只有持有正确密钥的系统或用户才能访问解密后的数据。
  1. 密钥管理
  • 密钥的保护至关重要。如果密钥被泄露,攻击者可以解密数据。
  • 通常,密钥存储在专用的硬件安全模块(HSM)或云服务的密钥管理系统中。

实现静态数据加密的方式

静态数据加密可以在以下层级中实现:

1) 文件级别加密

  • 特定文件或目录加密。
  • 适用于操作系统或应用程序的文件系统,例如 eCryptfsEncFS
  • 适合小范围数据保护,但对大规模存储性能可能有影响。

2) 数据库加密

  • 数据库本身支持静态数据加密,例如 MySQL 的 Transparent Data Encryption (TDE) 和 PostgreSQL 的加密扩展。
  • 数据在写入数据库存储文件前被加密。

3) 磁盘/卷加密

  • 整个存储设备或卷被加密。
  • 常用工具:dm-crypt (Linux)、BitLocker (Windows)、FileVault (macOS)。
  • 优点:易于管理,对操作系统透明。
  • 缺点:加密层对所有数据一视同仁,缺乏细粒度控制。

4) 云存储加密

  • 云服务商提供的静态数据加密。
  • AWS S3 提供服务端加密(SSE),支持 KMS 集成。
  • Google Cloud Storage 和 Azure 也提供类似功能。

与其他加密技术的对比

特性静态数据加密 (At Rest)传输数据加密 (In Transit)使用中数据加密 (In Use)
数据状态存储在磁盘或数据库上网络中传输的数据在内存中被处理的数据
主要目标防止物理设备泄露风险防止中间人攻击和窃听防止运行时内存数据泄露
典型工具BitLocker, TDETLS, HTTPS, VPNSGX, Homomorphic Encryption

3.3 Domain Name Resolver Strategies

Resolution overhead.

提前缓存好不同存储节点的IP,从而避免域名解析带来的延迟

Throughput-based resolver.

启用负载均衡

MTU-based resolver.

减少或消除巨型帧,从而降低CPU拆包开销

统一各个节点的MTU量

MTU discovery.

寻找支持高MTU的路径,尽可能提高MTU量

3.4 Performance Evaluation

4 Cloud Storage Integregation

4.1 Database Engine Design

列式存储

下面这张图主要关于元信息存储

image-20241123163936013

4.2 Table Scan Operator

很多操作都还是Umbra原样

Scan design preliminaries.

Morsel picking.

Worker jobs.

image-20241123164419410

4.3 Object Scheduler

还是Umbra

4.4 Relation & Storage Format

还是Umbra

4.5 Encryption & Compressio

lz4和前面的encryption-at-rest

5 EXPERIMENTAL EVALUATION

Cloud DBMS: AWS Athena和RedShift等

Processing in the cloud: 各大云厂商的Spark和Hadoop

Spot instances(竞价实例): 研发的AnyBolb很适合这种场景

Serverless computing: 不在同一个竞争赛道

Cloud storage for DBMS: 各类数据湖(Apache Iceberg)

Memory disaggregation: 未来可期

Networking and kernel APIs:使用更快的RDMA以及Nvme SSD

4.3 Object Scheduler

总结

一个基于io_uring的Umbra数据库S3读取方案,Over