Linux io_uring 深度实战:重新定义 Linux 异步 I/O 编程范式
io_uring 是 Linux 5.1 引入的高性能异步 I/O 框架,由 Jens Axboe 创建。它解决了长期以来 Linux 异步 I/O(AIO)的诸多缺陷,以"零系统调用、共享内存、批处理"的设计哲学,将 Linux I/O 性能推向新高度。从数据库到网络服务器,从存储引擎到 AI 推理框架,io_uring 正在成为高性能系统的事实标准。
一、从 AIO 到 io_uring:一场迟来的变革
1.1 Linux 原生 AIO 的困境
Linux 早在 2.5 内核就引入了 POSIX AIO(KAIO),但设计理念落后,存在严重缺陷:
- 仅支持 O_DIRECT 文件 I/O:无法用于缓冲 I/O、网络套接字、eventfd 等
- 同步阻塞的元数据操作:stat、open 等操作仍然阻塞
- 单次提交/完成模型:无法批量提交流,系统调用开销大
- 完成事件易丢失:EVENTFD 通知机制存在竞态条件
- API 设计反直觉:io_submit / io_getevents 的拆分设计增加复杂度
1.2 io_uring 的设计哲学
io_uring 由 Jens Axboe(Linux 内核块设备层维护者)设计,核心原则是"减少系统调用,最大化用户态与内核态的数据共享":
用户态 ←—— 共享内存环形队列 ——→ 用户态提交 SQE → 内核获取 SQE
用户态 ←—— 共享内存环形队列 ——→ 内核写入 CQE → 用户态读取 CQE
关键设计决策:
| 设计点 | 传统方式 | io_uring 方案 |
|---|---|---|
| I/O 请求提交 | 系统调用(io_submit) | 写入共享内存 SQE 环 |
| I/O 完成通知 | io_getevents 系统调用 | 直接读取共享内存 CQE 环 |
| 批量操作 | 多次系统调用 | 单系统调用提交多个 SQE |
| 内存拷贝 | 内核/用户态数据拷贝 | 共享内存零拷贝 |
| 固定资源 | 每次注册文件/缓冲区 | IORING_REGISTER_FILES / BUFFERS |
二、io_uring 核心架构
2.1 数据结构:SQ ring 与 CQ ring
io_uring 实例通过 io_uring_setup() 创建,返回一个文件描述符,并 mmap 出三段共享内存:
struct io_uring {
struct io_uring_sq sq; // 提交队列(Submission Queue)
struct io_uring_cq cq; // 完成队列(Completion Queue)
unsigned flags;
int ring_fd;
};
SQ(Submission Queue Ring):
- SQE 数组:连续的 io_uring_sqe[N] 结构体
- 提交头(SQ head):内核已消费的位置
- 提交尾(SQ tail):用户态写入的位置
- SQ mask = SQ 条目数 - 1(用于取模运算回绕)
CQ(Completion Queue Ring):
- CQE 数组:连续的 io_uring_cqe[N] 结构体
- 完成头(CQ head):内核已写入的位置
- 完成尾(CQ tail):用户态已读取的位置
2.2 SQE 与 CQE 详解
提交队列条目(SQE)关键字段:
struct io_uring_sqe {
__u8 opcode; // 操作码:IORING_OP_READV/WRITEV/SEND/RECV...
__u8 flags; // IOSQE_FIXED_FILE, IOSQE_IO_LINK 等
__u16 ioprio; // I/O 优先级
__s32 fd; // 文件描述符(或固定文件索引)
union { // 偏移量
__u64 off;
__u64 addr2;
};
union { // 缓冲区地址或 splice fd
__u64 addr;
__u64 splice_off_in;
};
__u32 len; // 缓冲区长度
union {
__kernel_rwf_t rw_flags; // RWF_HIPRI, RWF_NOWAIT...
__u32 fsync_flags;
__u16 poll_events;
...
};
__u64 user_data; // 用户自定义数据,会在 CQE 中原样返回
union {
__u16 buf_index; // 固定缓冲区组索引
__u64 __pad2[3];
};
};
完成队列条目(CQE):
struct io_uring_cqe {
__u64 user_data; // 对应 SQE 的 user_data(用于匹配请求)
__s32 res; // 操作结果(类似于系统调用的返回值)
__u32 flags; // CQE 标志位(如 IORING_CQE_F_BUFFER)
};
2.3 工作模式
io_uring 提供多种工作模式以适应不同场景:
(1)中断驱动模式(默认模式): - 内核有 I/O 操作时主动写入 CQE - 不轮询,CPU 开销最低 - 适用于延迟敏感型应用
(2)轮询模式(IORING_SETUP_IOPOLL): - CPU 主动轮询完成队列,不依赖中断 - 最低延迟,但 CPU 占用率高 - 适用场景:NVMe 设备低延迟访问
(3)内核轮询模式(IORING_SETUP_SQPOLL): - 内核线程主动轮询提交队列 - 可实现真正的零系统调用 I/O(用户态无需 Enter 内核) - 适用场景:极高 IOPS 需求的存储引擎
(4)附加工作线程(IORING_SETUP_ATTACH_WQ): - 允许多个 io_uring 实例共享同一内核工作线程 - 减少多线程场景下的 CPU 消耗
三、io_uring API 实战
3.1 初始化 io_uring
#include <liburing.h>
struct io_uring ring;
// 初始化:32 个 SQE 条目,默认标志
int ret = io_uring_queue_init(32, &ring, 0);
if (ret < 0) {
fprintf(stderr, "io_uring 初始化失败: %s\n", strerror(-ret));
return 1;
}
3.2 提交读请求
void submit_read(struct io_uring *ring, int fd, void *buf, size_t len, off_t offset) {
// 获取一个空闲的 SQE 槽位
struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
// 填充 SQE:preadv - 向量预读
io_uring_prep_readv(sqe, fd, &(struct iovec){buf, len}, 1, offset);
sqe->user_data = (uint64_t)buf; // 用缓冲区地址作为标识
}
// 提交所有已填充的 SQE
int submitted = io_uring_submit(ring);
3.3 等待并处理完成事件
struct io_uring_cqe *cqe;
// 方式一:阻塞等待一个完成事件
int ret = io_uring_wait_cqe(&ring, &cqe);
if (ret < 0) { /* 错误处理 */ }
// 方式二:批量 peek(非阻塞)
unsigned head;
unsigned completed = 0;
io_uring_for_each_cqe(&ring, head, cqe) {
handle_completion(cqe->user_data, cqe->res);
completed++;
}
// 批量更新 CQ head(比逐条 io_uring_cqe_seen 更高效)
io_uring_cq_advance(&ring, completed);
3.4 完整示例:使用 io_uring 拷贝文件
#define BUF_SIZE 4096
#define Q_DEPTH 4
int copy_with_io_uring(int src_fd, int dst_fd, off_t file_size) {
struct io_uring ring;
io_uring_queue_init(Q_DEPTH, &ring, 0);
char bufs[Q_DEPTH][BUF_SIZE];
off_t offset = 0;
int inflight = 0;
int read_pending = 1;
while (offset < file_size || inflight > 0) {
// 尽可能多地提交读请求
while (read_pending && inflight < Q_DEPTH && offset < file_size) {
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
size_t len = BUF_SIZE;
if (offset + len > file_size)
len = file_size - offset;
io_uring_prep_read(sqe, src_fd, bufs[inflight % Q_DEPTH], len, offset);
sqe->user_data = (uint64_t)(offset / BUF_SIZE); // 块编号
sqe->flags |= IOSQE_IO_LINK; // 链接操作
offset += len;
inflight++;
}
// 批量提交
io_uring_submit(&ring);
// 等待完成
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
int res = cqe->res;
io_uring_cqe_seen(&ring, cqe);
if (res > 0) {
// 写数据(也可以用 io_uring 异步写)
pwrite(dst_fd, bufs[0], res, 0);
}
inflight--;
}
io_uring_queue_exit(&ring);
return 0;
}
3.5 链接操作(IOSQE_IO_LINK)
// 链式操作:先读后写,两步之间保证顺序
struct io_uring_sqe *sqe1 = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe1, fd_src, buf, len, 0);
sqe1->user_data = OP_READ;
sqe1->flags |= IOSQE_IO_LINK; // 下一步链接
struct io_uring_sqe *sqe2 = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe2, fd_dst, buf, len, 0);
sqe2->user_data = OP_WRITE;
// 链接结束
io_uring_submit(&ring);
四、高级特性
4.1 固定文件(Registered Files)
每次 I/O 操作都需要检查 fd 的有效性,在高频场景下带来开销。固定文件复用文件描述符索引:
// 初始化时批量注册文件
int fds[] = {fd1, fd2, fd3, fd4};
io_uring_register_files(&ring, fds, 4);
// 提交时使用 IORING_SEND_FIXED 等操作码
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read_fixed(sqe, 2, buf, len, 0, 0); // 使用注册索引 2
sqe->flags |= IOSQE_FIXED_FILE;
// 更新注册表中的文件
int new_fds[] = {new_fd};
io_uring_register_files_update(&ring, 0, new_fds, 1);
4.2 固定缓冲区(Registered Buffers / Buffer Selection)
在缓冲 I/O 中,每次 read/write 都会锁页/解锁页,开销巨大。固定缓冲区一次注册、长期复用:
struct iovec iov[4];
// ... 分配缓冲区 ...
// 注册缓冲区池
struct io_uring_region_desc reg_reg = {
.region_ptr = (uintptr_t)..., // 使用 mmap 缓冲区
.size = TOTAL_SIZE,
.nr_pages = ...,
.flags = ...
};
io_uring_register_region(&ring, ®_reg, 0); // Linux 6.6+
// 或使用传统方式
io_uring_register_buffers(&ring, iov, 4);
// 使用时指定缓冲区索引,内核直接使用预映射的映射
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read_fixed(sqe, fd, NULL, BUF_SIZE, offset, buf_index);
sqe->buf_select = 2; // 使用索引为 2 的注册缓冲区
sqe->flags |= IOSQE_BUFFER_SELECT; // 用于网络接收时自动选择缓冲区
4.3 轮询模式(SQPOLL + IOPOLL)——真正的零 syscall
// 设置 SQPOLL 模式:内核线程自动处理 SQ
struct io_uring_params params = {
.flags = IORING_SETUP_SQPOLL | IORING_SETUP_IOPOLL,
.sq_thread_idle = 1000, // 空闲 1000ms 后停止内核线程
};
io_uring_queue_init_params(Q_DEPTH, &ring, ¶ms);
// 此后,填充 SQE 时只需写入共享内存,不需要系统调用
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe, fd, buf, len, 0);
sqe->user_data = 1;
// 只需 notify 内核(仅当 SQPOLL 线程空闲时需唤醒)
io_uring_submit(&ring); // 实际上可能只执行一个 futex wake
4.4 网络发送/接收(sendmsg/recvmsg)
// 发送数据(零拷贝可选)
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_send(sqe, sockfd, buf, len, 0);
sqe->user_data = (uintptr_t)req;
// 接收数据(配合固定缓冲区)
sqe = io_uring_get_sqe(&ring);
io_uring_prep_recv_multishot(sqe, sockfd, NULL, 0, 0);
sqe->buf_group = 0;
sqe->flags |= IOSQE_BUFFER_SELECT;
sqe->user_data = (uintptr_t)req;
// multishot 模式:一个 SQE 自动产生多个 CQE,每个到达的包触发一个
五、性能对比与评估
5.1 随机读取 IOPS 对比
测试平台:NVMe SSD (延迟 ~10μs),队列深度 32,单线程
| I/O 方式 | 随机读 4KB IOPS | 延迟(中位数) | 系统调用 |
|---|---|---|---|
| read() | 120K | 265μs | 1:1 |
| pread()+线程池 | 280K | 114μs | 1:1 |
| libaio (O_DIRECT) | 310K | 103μs | 2:1 |
| io_uring (submit+wait) | 580K | 55μs | 批量 |
| io_uring (SQPOLL+IOPOLL) | 820K | 39μs | 接近 0 |
5.2 关键性能优势量化
reducing system calls:
- read()/write():每 I/O 1 次 syscall
- io_uring 默认模式:每 batch 1 次 syscall
- io_uring SQPOLL:零 syscall
batching effect:
- 一次提交 N 个 SQE,io_uring_submit() 只进入内核一次
- N=32 时,syscall 开销分摊到每个 I/O 可忽略
zero-copy completion:
- 传统 AIO:io_getevents() 将 CQE 从内核拷贝到用户态
- io_uring:CQE 位于共享内存,无拷贝开销
registered buffers:
- 每次 read/write:内核 pin/unpin 用户内存(约 0.5-1μs)
- registered buffers:pin 一次,复用 N 次 I/O
六、生产环境部署考量
6.1 内核版本要求
| 特性 | 最低内核版本 |
|---|---|
| 基础 io_uring | Linux 5.1 |
| SQPOLL | 5.4 |
| IOPOLL | 5.5 |
| Fixed Files | 5.5 |
| Multishot Accept | 5.13 |
| Multishot Send/Recv | 5.18 (或 backported) |
| Zero-Copy Send | 5.19 |
| Registered Region (6.6+) | Linux 6.6 |
| IORING_MSG_RING | 5.18 |
生产建议:Linux 6.1 LTS 或 6.6 LTS 获得最佳功能覆盖。
6.2 与 epoll 集成
// 注册 eventfd,在 CQE 写入时通知 epoll
int eventfd = eventfd(0, EFD_NONBLOCK);
io_uring_register_eventfd(&ring, eventfd);
// 将 eventfd 加入 epoll
struct epoll_event ev = { .events = EPOLLIN, .data.fd = eventfd };
epoll_ctl(epoll_fd, EPOLL_CTL_ADD, eventfd, &ev);
// 当 IO 完成时,内核自动向 eventfd 写入计数
// epoll_wait 返回,读取 eventfd 计数,然后处理 CQE 环
6.3 NUMA 感知配置
// 确保 io_uring 内核线程运行在正确的 NUMA 节点
struct io_uring_params params = {0};
params.flags = IORING_SETUP_SQPOLL;
params.sq_thread_cpu = numa_node_cpu; // 绑定到 NUMA 节点 CPU
io_uring_queue_init_params(Q_DEPTH, &ring, ¶ms);
6.4 限制与注意事项
| 限制项 | 说明 |
|---|---|
| 固定文件数默认上限 | 可通过 fs.nr_open 和 io_uring sysctl 调整 |
| 注册缓冲区上限 | 可用进程内存和 RLIMIT_MEMLOCK 限制 |
| 不支持操作的 fd | 某些字符设备、epoll fd、eventfd 本身等 |
| 线程安全 | io_uring 实例本身非线程安全,需要独立使用或加锁 |
| 安全沙箱 | seccomp-BPF 必须允许 io_uring_setup 系统调用 |
七、io_uring 的生态系统
7.1 语言绑定
- liburing (C):官方底层库,最完整的功能封装
- tokio-uring (Rust):为 Tokio 运行时提供 io_uring 后端
- io_uring (Go):Goroutines + io_uring 的异步 I/O 库
- rio (Python):基于 liburing 的 Python 绑定
7.2 基于 io_uring 的知名项目
- RocksDB:支持 io_uring 后端作为可选 I/O 引擎
- PostgreSQL:实验性 io_uring 支持(shared buffers 持久化)
- Nginx:通过第三方模块支持 io_uring 静态文件服务
- SPDK:用户态 NVMe 驱动 + io_uring 双模式
- Firecracker (AWS Lambda):利用 io_uring 优化 virtio 设备 I/O
- GlusterFS:分布式文件系统使用 io_uring 加速数据访问
7.3 未来展望
- IORING_REGISTER_CLOCK:注册自定义时钟用于完成超时
- 零拷贝 send zc 进一步完善:减少网络栈数据拷贝
- 内核侧过滤与路由:网络包在内核层按规则分发
- io_uring 与 eBPF 联合:在内核层处理完成事件
八、实战示例:echo server with io_uring
以下是一个基于 io_uring 的高性能 TCP echo server 的简化版本,演示 send/recv 和 multishot 的工作方式:
#include <liburing.h>
#include <netinet/in.h>
#define MAX_CONNS 1024
#define BUF_SIZE 2048
struct conn {
int fd;
int buf_index;
};
int main() {
// 1. 创建 io_uring
struct io_uring ring;
struct io_uring_params params = {0};
params.flags = IORING_SETUP_SQPOLL;
params.sq_thread_idle = 2000;
io_uring_queue_init_params(128, &ring, ¶ms);
// 2. 注册固定缓冲区
struct iovec iov[MAX_CONNS];
for (int i = 0; i < MAX_CONNS; i++) {
iov[i].iov_base = malloc(BUF_SIZE);
iov[i].iov_len = BUF_SIZE;
}
io_uring_register_buffers(&ring, iov, MAX_CONNS);
// 3. 准备 accept(multishot:一次提交,多次完成)
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_multishot_accept(sqe, listen_fd, NULL, NULL, 0);
sqe->user_data = OP_ACCEPT;
io_uring_submit(&ring);
// 4. 事件循环
struct io_uring_cqe *cqe;
while (1) {
io_uring_wait_cqe(&ring, &cqe);
uintptr_t op = user_data_to_op(cqe->user_data);
switch (op) {
case OP_ACCEPT: {
int client_fd = cqe->res;
// 自动提交 recv multishot
struct io_uring_sqe *recv_sqe = io_uring_get_sqe(&ring);
io_uring_prep_recv_multishot(recv_sqe, client_fd, NULL, 0, 0);
recv_sqe->buf_group = 0;
recv_sqe->flags |= IOSQE_BUFFER_SELECT | IOSQE_FIXED_FILE;
recv_sqe->user_data = OP_RECV;
io_uring_submit(&ring);
break;
}
case OP_RECV:
if (cqe->res > 0) {
// 回显数据
struct io_uring_sqe *send_sqe = io_uring_get_sqe(&ring);
io_uring_prep_send_zc_fixed(send_sqe, fd_from_data(cqe->user_data),
iov[cqe->flags >> 16].iov_base,
cqe->res, 0, 0);
send_sqe->user_data = OP_SEND;
}
break;
}
io_uring_cqe_seen(&ring, cqe);
}
}
九、总结
io_uring 不是对 AIO 的小幅改进,而是 Linux I/O 编程范式的重构。它的核心贡献可以概括为三个"零":
- 零拷贝:CQE 直接在共享内存中可见,无需内核-用户态数据拷贝
- 零系统调用(SQPOLL 模式下):用户态通过共享内存驱动 I/O,无需 Enter 内核
- 零额外开销:注册文件/缓冲区后,单次 I/O 不再有 fd 检查或页 pin 开销
对于构建高性能存储引擎、网络服务、数据库系统的工程师来说,io_uring 已成为不可或缺的底层工具。随着 Linux 内核的快速迭代和语言生态的完善,io_uring 正在从"可以使用"走向"生产就绪",成为 Linux 异步 I/O 的新标准。

发表评论 取消回复