惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CXSECURITY Database RSS Feed - CXSecurity.com
A
About on SuperTechFans
H
Help Net Security
Engineering at Meta
Engineering at Meta
G
Google Developers Blog
aimingoo的专栏
aimingoo的专栏
The Register - Security
The Register - Security
WordPress大学
WordPress大学
MongoDB | Blog
MongoDB | Blog
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
C
Check Point Blog
人人都是产品经理
人人都是产品经理
云风的 BLOG
云风的 BLOG
N
Netflix TechBlog - Medium
酷 壳 – CoolShell
酷 壳 – CoolShell
雷峰网
雷峰网
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - Franky
I
InfoQ
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The GitHub Blog
The GitHub Blog
Last Week in AI
Last Week in AI
D
Docker
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 叶小钗
Jina AI
Jina AI
F
Fortinet All Blogs
宝玉的分享
宝玉的分享
小众软件
小众软件
有赞技术团队
有赞技术团队
F
Full Disclosure
月光博客
月光博客
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Recorded Future
Recorded Future
Apple Machine Learning Research
Apple Machine Learning Research
P
Proofpoint News Feed
量子位
U
Unit 42
T
The Blog of Author Tim Ferriss
Microsoft Azure Blog
Microsoft Azure Blog
V
V2EX
O
OpenAI News
S
Secure Thoughts
罗磊的独立博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Google Online Security Blog
Google Online Security Blog
Cloudbric
Cloudbric
W
WeLiveSecurity
IT之家
IT之家

Jiajun的技术笔记

你好,2026! TiDB 源码阅读(六):TiDB Coprocessor 源码解析 性能优化的核心思想 TiDB 源码阅读(五):索引 TiDB 源码阅读(四):AST、逻辑计划、物理计划 CockroachDB Serverless Architecture podman 无故退出 Cursor Control-L (CTRL-L) Keyboard Shortcuts in Terminal Replace docker with podman Using xmonad with xfce4 A RC script for freebsd frpc 自己动手写一个k8s controller AI 会取代你的(编程)岗位吗? 自建DERP服务器提升Tailscale连接速度(使用Nginx转发) 自动升级Docker容器 再读《程序员修炼之道-从小工到专家》 让浏览器下载文件 再读《软件随想录》/《黑客与画家》/《软技能》 HTTP 压力测试中的 Coordinated Omission 2的补码 编程语言中的 context 是什么? flutter macOS 构建出错 Flatpak 使用小记 Golang CAS 操作是怎么实现的 PostgreSQL 当MQ来使用 Clash 结合 工作VPN 的网络设计 使用 PostgreSQL 搭建 JuiceFS PostgreSQL 配置优化和日志分析 有GitHub Copilot?那就可以搭建你的ChatGPT4服务 窗口函数的使用(以PG为例) 读《为什么学生不喜欢上学》 OpenAI Prompt Engineering 摘录和总结 读《打造真正的新产品》 VueJS 总结 Linux 自动挂载 alist 提供的webdav FreeBSD 使用 vm-bhyve 安装Debian虚拟机 FreeBSD 和 Linux 网卡聚合实现提速 GPT 帮我搞定了时区转换问题 长任务系统如何处理? macOS/Linux 编译 InputLeap 使用开源软KVM - synergy-core 解决 macOS 终端hostname一直变化问题 KVM 共享 Intel 集成显卡 PromQL 备忘 读《格鲁夫给经理人的第一课》 读《打开心智》 为什么要把复杂的联表操作拆成多个单表查询? 红包系统的设计 MySQL Index Condition Pushdown Optimization Go mod 简明教程 OpenWRT 使用 Android/iOS USB 网络 搭建旁路由 Golang gRPC 错误处理 编写可维护的单元测试代码 OAuth 2 详解(六):Authorization Code Flow with PKCE OAuth 2 详解(五):Device Authorization Flow OAuth 2 详解(三):Resource Owner Password Credentials Grant OAuth 2 详解(四):Client Credentials Flow OAuth 2 详解(二):Implict Grant Flow OAuth 2 详解(一):简介及 Authorization Code 模式 ElasticSearch 学习笔记 三种git流程以及发版模型 错误处理实践 权限模型(RBAC/ABAC) OIDC(OpenID Connect) 简介 任务队列简介 PostgreSQL 操作笔记 使用Drone CI构建CI/CD系统 Golang migrate 做数据库变更管理 使用PostgreSQL做搜索引擎 Nginx 源码阅读(三): 连接池、内存池 Nginx 源码阅读(二): 请求处理 Nginx 源码阅读(一): 启动流程 Go 泛型简明教程 KVM 显卡穿透给 Windows 使用 HTTP Router 处理 Telegram Bot 按钮回调 使用反射(reflect)对结构体赋值 GIN 是如何绑定参数的 你好 2022(2021 年终总结) 用Go导入大型CSV到PostgreSQL 使用 OpenWRT 搭建软路由 使用软KVM切换器 barrier 共享键鼠 SQL 防注入及原理 使用 gomock 测试 Go 代码 gevent不是黑魔法(二): gevent 实现 gevent不是黑魔法(一): greenlet 实现 用 entgo 替代 gorm 应用内使用crontab不是那么方便 单测时要不要 mock 数据库? Sentry 自建指南 用selenium完成自动化任务 用闲置的安卓手机做垃圾电话短信过滤 推荐三个时间管理工具 一次事故反思 当JS遇到uint64:JS整数溢出问题 SQLite3 存储以及ACID原理 Redis源码阅读:pub/sub实现 Redis源码阅读:zset实现 Redis源码阅读:bitmap 位图的运算 Redis源码阅读:set是怎么做交并集运算的?
WAL(Write-ahead logging)的套路
Jiajun Huang · 2021-05-15 · via Jiajun的技术笔记

WAL的全称是 Write-ahead logging, 是一种常见的用于持久化数据的方式, 通常性能都很不错, 利用了磁盘连续写的性能高于随机读写的这一个特性. 一般来说, WAL都是用于这样一种场景, 记录操作日志, 数据库收到合法请求之后, 首先在WAL里写入一条记录, 然后再开始进行内存 操作以及需要更长时间的操作, 假如此时应用崩溃, 那么应用可以读取WAL来进行重建/修复, 也就是说, 只要提交到WAL里并且落盘的数据, 就可以认为是一定被持久化了的.

我去阅读了一下 Redis的AOF , 以及 RocksDB 的实现, 结合之前看NSQ的经验, 总结了一下WAL的一些实现套路. 接下来我们分别来看看.

Redis

在Redis中有一种持久化方式就叫做 AOF , 全称Append-Only-File, 简单来说, 就是如下代码:

    /* Open the AOF file if needed. */
    if (server.aof_state == AOF_ON) {
        server.aof_fd = open(server.aof_filename,
                               O_WRONLY|O_APPEND|O_CREAT,0644);
        if (server.aof_fd == -1) {
            serverLog(LL_WARNING, "Can't open the append-only file: %s",
                strerror(errno));
            exit(1);
        }
    }

也就是说, 我们以追加的方式写入文件, 操作系统来保证每一次调用都总是写在文件的末尾. 我们来看看Redis是 怎么把操作日志写入的:

void flushAppendOnlyFile(int force) {
    ssize_t nwritten;
    int sync_in_progress = 0;
    mstime_t latency;
    ...
    latencyStartMonitor(latency);
    nwritten = aofWrite(server.aof_fd,server.aof_buf,sdslen(server.aof_buf));
    latencyEndMonitor(latency);
    ...
}

/* This is a wrapper to the write syscall in order to retry on short writes
 * or if the syscall gets interrupted. It could look strange that we retry
 * on short writes given that we are writing to a block device: normally if
 * the first call is short, there is a end-of-space condition, so the next
 * is likely to fail. However apparently in modern systems this is no longer
 * true, and in general it looks just more resilient to retry the write. If
 * there is an actual error condition we'll get it at the next try. */
ssize_t aofWrite(int fd, const char *buf, size_t len) {
    ssize_t nwritten = 0, totwritten = 0;

    while(len) {
        nwritten = write(fd, buf, len);

        if (nwritten < 0) {
            if (errno == EINTR) continue;
            return totwritten ? totwritten : -1;
        }

        len -= nwritten;
        buf += nwritten;
        totwritten += nwritten;
    }

    return totwritten;
}

对于 write 这个系统调用, 操作系统是可能会阻塞线程的, 至少在把内容提交到内核的缓冲区这段时间是 阻塞的, 所以Redis的主线程可能被阻塞, 不过正常情况下, 当内核的缓冲区够用时, 这个系统调用可以很快 的返回, 也正是因此, 这个系统调用完成后, 并不等于数据真的写入到磁盘了, 因此Redis会根据 appendfsync 这个参数, 如果设置的是 everysec, 那么每秒会执行一次 fsync(Linux上使用的 fdatasync), 这个 系统调用就会确保缓冲区内的内容真的落到磁盘上了.

$ man 2 fdatasync
SYNOPSIS
       #include <unistd.h>

       int fdatasync(int fildes);

DESCRIPTION
       The fdatasync() function shall force all currently queued I/O operations associated with the file indicated by file descriptor fildes to the synchronized I/O com‐
       pletion state.

$ man 2 write
       A successful return from write() does not make any guarantee that data has been committed to disk.  On some filesystems, including NFS, it does not even guarantee
       that space has successfully been reserved for the data.  In this case, some errors might be delayed until a future write(), fsync(2), or even close(2).  The  only
       way to be sure is to call fsync(2) after you are done writing all your data.

对于常见的应用, 我们最常见的配置就是 AOF + 每秒刷盘, 这样是一个数据安全+写入性能好好的权衡, 但同时 可以见得, 如果Redis服务器在执行 fdatasync 之后的一秒之内断电了, 那么这部分数据就会丢失.

综上所述, Redis的AOF策略就是内存保存 aof_buf, 然后写入到操作系统的缓存里, 加上定时刷盘. 至于写入的 内容, 其实就是 Redis Protocol 的内容.

NSQ

策略和Redis差不多, 参考: https://jiajunhuang.com/articles/2020_08_16-nsq_source_code.md.html

RocksDB

RocksDB的设计比较复杂一些, 如下:

首先日志文件有如下格式:

       +-----+-------------+--+----+----------+------+-- ... ----+
 File  | r0  |        r1   |P | r2 |    r3    |  r4  |           |
       +-----+-------------+--+----+----------+------+-- ... ----+
       <--- kBlockSize ------>|<-- kBlockSize ------>|

  rn = variable size records
  P = Padding

r0, r1…r4 是每一条记录, 也就是一个record. 可以看到, 日志文件被 分为多个块, 每一个块大小为 kBlockSize, 也就是32KB. 如果一个block里 的剩余空间可以放得下一个record, 那么就放下去, 比如r0, 如果放不下了, 那么剩余的空间就会被置空, 比如 r1和r2中间的P.

Record格式:

  • The Legacy Record Format

    +---------+-----------+-----------+--- ... ---+
    |CRC (4B) | Size (2B) | Type (1B) | Payload   |
    +---------+-----------+-----------+--- ... ---+
    
    CRC = 32bit hash computed over the payload using CRC
    Size = Length of the payload data
    Type = Type of record
       (kZeroType, kFullType, kFirstType, kLastType, kMiddleType )
       The type is used to group a bunch of records together to represent
       blocks that are larger than kBlockSize
    Payload = Byte stream as long as specified by the payload size
    
    
  • The Recyclable Record Format

    +---------+-----------+-----------+----------------+--- ... ---+
    |CRC (4B) | Size (2B) | Type (1B) | Log number (4B)| Payload   |
    +---------+-----------+-----------+----------------+--- ... ---+
    Same as above, with the addition of
    Log number = 32bit log file number, so that we can distinguish between
    records written by the most recent log writer vs a previous one.
    
    

上述字段中的Type有如下选项:

  • FULL 说明所有数据都在这一个record里
  • FIRST 说明record太大, 一个block装不下, 这是第一个
  • MIDDLE 说明这是record切分后, 中间的
  • LAST 说明这是record切分后, 该record最后一个内容

RocksDB处理WAL的逻辑是, 每当打开一个DB, 或者column family刷盘了, 就创建一个新的WAL. 因此可以看出, RocksDB对于WAL的处理也是写磁盘, 然后策略触发flush.

总结

这篇博客里, 看了一下几种WAL的设计和实现, 可以总结出来, 大部分的WAL 实现, 都是 write + 定期 fdatasync 这种模式, 这样可以保证 数据的最大可用性, 至于WAL文件的格式本身, 这取决于应用将要如何使用 WAL, 比如NSQ和Redis只需要从头读取里面的内容, 因此基本上就是直接写入, Redis为了防止AOF太大, 还会进行AOF重写. 而RocksDB则实现的跟复杂一些, 保存了payload大小等内容到内容的前面, 这样就可以快速的跳过一些内容进行查询.


Ref: