引言:Linux I/O 模型演进的必然选择
在 Linux 内核 5.1 发布(2019 年 5 月)之际,Jens Axboe 提交了一个改变高性能 I/O 编程范式的全新框架——io_uring。彼时,epoll 已是网络编程的事实标准,AIO(POSIX async I/O)在实际使用中却因其实现限制(仅支持 O_DIRECT 文件、不支持 socket)而名存实亡。五年已过,io_uring 已成为存储引擎(RocksDB、SPDK)、网络代理(Nginx 通过第三方模块)、数据库系统(PostgreSQL 14+ 实验性支持)和高性能计算领域的首选异步 I/O 方案。
本文将从零开始,深度剖析 io_uring 的设计哲学、核心数据结构、liburing API 使用方法、完整的高性能实战案例,以及生产环境部署要点。读者在阅读完毕后,应能独立使用 io_uring 构建异步 I/O 密集型应用。
一、io_uring 要解决的三个核心痛点
1.1 减少系统调用开销
传统 Linux I/O 模型中,每次 read/write 至少涉及两次系统调用(提交+等待完成)。在高 IOPS 场景(如 NVMe 单盘可达百万 IOPS),系统调用本身就成了瓶颈。epoll 虽然解决了"等事件"的问题,但真正的读写仍然需要单独的 syscall,且每次调用都要经历用户态→内核态→用户态的完整上下文切换。
io_uring 的解法:用户态与内核态共享环形缓冲区(Ring Buffer),批量提交和收割 I/O 请求,理想情况下可以做到整个生命周期仅 1-2 次系统调用。
1.2 消除数据拷贝
POSIX AIO 和 epoll 都要求用户态维护独立的请求数据结构,内核在处理时需要将这些数据拷贝到内核空间。io_uring 通过共享内存 + 寄存器文件(Fixed Files)+ 预注册缓冲区(Fixed Buffers),实现了真正的零拷贝提交。
1.3 支持所有 I/O 类型
与 POSIX AIO 只能用于 O_DIRECT 文件不同,io_uring 对以下所有操作提供统一异步接口:
- 文件读写(read/write/preadv/pwritev)
- 网络 I/O(send/recv/sendmsg/recvmsg)
- 文件同步(fsync/fdatasync/fallocate)
- 文件事件(poll_add/mod/del)
- 其他 syscall(connect/accept/open/close/statx 等)
二、核心架构:三个环形缓冲区
2.1 Submission Queue (SQ) — 提交队列
SQ 是一个生产者-消费者模型的环形缓冲区。用户态程序将 I/O 请求封装为 Submission Queue Entry (SQE),写入 SQ 尾部。内核从 SQ 头部取出 SQE 执行。
SQE 的数据结构(C 语言):
struct io_uring_sqe {
__u8 opcode; // 操作码:IORING_OP_READV, IORING_OP_WRITEV 等
__u8 flags; // 标志位:IOSQE_FIXED_FILE, IOSQE_IO_LINK 等
__u16 ioprio; // I/O 优先级
__s32 fd; // 目标文件描述符/固定文件索引
union { __u64 off; unsigned long long addr2; };
union { __u64 addr; unsigned long long splice_off; };
__u32 len; // 缓冲区长度或 iovec 数量
union { __kernel_rwf_t rw_flags; __u32 fsync_flags; ... };
__u64 user_data; // 用户自定义标识,原样回传到 CQE
union { __u16 buf_index; __u16 buf_group; };
__u16 personality;
union { __s32 splice_fd_in; __u32 file_index; };
__u64 __pad2[2];
};
2.2 Completion Queue (CQ) — 完成队列
CQ 是内核向用户态报告 I/O 完成结果的通道。每个完成的 I/O 请求以 Completion Queue Entry (CQE) 形式写入 CQ 尾部。
CQE 的数据结构:
struct io_uring_cqe {
__u64 user_data; // 与 SQE 中的 user_data 对应
__s32 res; // 操作结果:>=0 成功(字节数),<0 错误码
__u32 flags; // 完成标志:IORING_CQE_F_BUFFER 等
};
2.3 Submission Queue Entry Array (SQEArray)
SQEArray 是一个独立的数组。SQ 环形缓冲区中存储的不是 SQE 本身,而是 SQEArray 的索引。用户先向 SQEArray 中写入完整 SQE,再将索引写入 SQ。这种间接设计使得用户可以预先准备大量 SQE 后再批量提交。
三、liburing 实战编程
3.1 环境准备
io_uring 需要 Linux 内核 ≥ 5.1,推荐 5.10+(含重要特性如 Fixed Buffers、Multi-shot accept)。安装 liburing:
# Debian/Ubuntu
apt install liburing-dev
# 或从源码编译
git clone https://github.com/axboe/liburing.git
cd liburing
./configure && make -j$(nproc) && sudo make install
3.2 初始化 io_uring 实例
io_uring 的生命周期管理:
#include <liburing.h>
struct io_uring ring;
// 基础初始化:32 个 SQE,默认标志
int ret = io_uring_queue_init(32, &ring, 0);
if (ret < 0) {
fprintf(stderr, "io_uring init failed: %s\n", strerror(-ret));
return 1;
}
// ... 使用 io_uring ...
io_uring_queue_exit(&ring);
常用初始化标志:
IORING_SETUP_IOPOLL:开启轮询模式(需要块设备支持)IORING_SETUP_SQPOLL:内核线程轮询 SQ(提交 SQE 后可以完全不调用 enter)IORING_SETUP_SQ_AFF:指定内核轮询线程绑核IORING_SETUP_CQSIZE:自定义 CQ 大小(默认 = SQ 的 2 倍)IORING_SETUP_ATTACH_WQ:绑定到已有 workqueue(多 ring 共享内核线程)
3.3 提交 I/O 请求:完整流程
// 1. 获取一个空闲 SQE
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
if (!sqe) {
// SQ 满了,需要先提交并收割部分 CQE
fprintf(stderr, "SQ full, need to submit first\n");
return;
}
// 2. 填充 SQE:异步 read
char buf[4096];
struct iovec iov = {
.iov_base = buf,
.iov_len = sizeof(buf)
};
io_uring_prep_readv(sqe, fd, &iov, 1, 0); // fd, iovec, count, offset
io_uring_sqe_set_data(sqe, (void*)0x1234); // 用户标识
// 3. 提交一个或多个 SQE 到内核(这里只提交 1 个)
ret = io_uring_submit(&ring);
if (ret < 0) {
fprintf(stderr, "submit failed\n");
return;
}
// 4. 等待并收割完成事件
struct io_uring_cqe *cqe;
ret = io_uring_wait_cqe(&ring, &cqe); // 阻塞等待至少 1 个 CQE
if (ret < 0) {
fprintf(stderr, "wait failed\n");
return;
}
// 5. 处理结果
if (cqe->res < 0) {
fprintf(stderr, "I/O error: %s\n", strerror(-cqe->res));
} else {
printf("Read %d bytes, user_data=%lld\n", cqe->res, cqe->user_data);
}
// 6. 必须标记已消费,更新 CQ head
io_uring_cqe_seen(&ring, cqe);
3.4 批量提交优化
io_uring 的威力在于批量操作。以下代码展示提交 N 个请求后仅需一次系统调用:
// 批量提交 64 个异步读请求
for (int i = 0; i < 64; i++) {
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe, fd, bufs[i], BUF_SIZE, i * BUF_SIZE);
io_uring_sqe_set_data(sqe, (void*)(uintptr_t)i);
}
// 64 个 SQE 一次性提交——仅需 1 次 io_uring_enter 系统调用
io_uring_submit(&ring); // 仅 1 次 syscall!
// 收割 64 个完成事件
for (int i = 0; i < 64; i++) {
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
// 处理结果
io_uring_cqe_seen(&ring, cqe);
}
四、io_uring 高级特性详解
4.1 Fixed Buffers(预注册缓冲区)
默认模式下,内核在处理 SQE 时需要将用户态缓冲区 pin 住(get_user_pages),这在每次 I/O 时都是额外开销。通过预注册缓冲区,可以一次性 pin 住所有缓冲区,后续 I/O 直接使用缓冲区索引:
// 注册缓冲区
struct iovec iov[BUFFERS_COUNT];
for (int i = 0; i < BUFFERS_COUNT; i++) {
iov[i].iov_base = bufs[i];
iov[i].iov_len = BUF_SIZE;
}
io_uring_register_buffers(&ring, iov, BUFFERS_COUNT);
// 使用 IORING_OP_READ_FIXED / WRITE_FIXED
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read_fixed(sqe, fd, bufs[buf_index], BUF_SIZE, offset, buf_index);
sqe->flags |= IOSQE_FIXED_BUFFER;
io_uring_submit(&ring);
性能收益:在大量小 I/O 场景下(如数据库缓冲池),Fixed Buffers 可减少约 10-15% 的 CPU 时间开销。
4.2 Fixed Files(预注册文件描述符)
与 Fixed Buffers 类似,预先注册 FD 数组,后续操作使用数组索引而非真实 FD,避免每次 I/O 时 fdtable 的查找开销:
int fds[] = { fd1, fd2, fd3 };
io_uring_register_files(&ring, fds, 3);
// SQE 中 fd 填写索引 0,并设置 IOSQE_FIXED_FILE 标志
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe, 0, buf, len, offset); // fd=0 表示 fds[0]
sqe->flags |= IOSQE_FIXED_FILE;
4.3 Linked SQE(操作链)
在某些场景下,需要强制执行操作顺序(如先读取 header 再读取 body)。通过 IOSQE_IO_LINK 标志可以实现操作链:链中前一个操作完成后才会执行下一个:
// 步骤1:读取文件头
struct io_uring_sqe *sqe1 = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe1, fd, &header, HEADER_SIZE, 0);
io_uring_sqe_set_data(sqe1, (void*)OP_READ_HEADER);
sqe1->flags |= IOSQE_IO_LINK; // 链接到下一个 SQE
// 步骤2:读取数据体(仅在 header 读取成功后执行)
struct io_uring_sqe *sqe2 = io_uring_get_sqe(&ring);
io_uring_prep_read(sqe2, fd, &body, BODY_SIZE, HEADER_SIZE);
io_uring_sqe_set_data(sqe2, (void*)OP_READ_BODY);
io_uring_submit(&ring);
4.4 Multi-shot Accept(多发接受)
在传统模式下,每次 accept 操作完成一个连接就返回。IORING_RECV_MULTISHOT 允许单次 SQE 持续为新连接生成 CQE,极大减少 accept 操作的 syscall 频率:
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_multishot_accept(sqe, listen_fd, &client_addr, &addr_len, 0);
io_uring_sqe_set_data(sqe, (void*)OP_MULTISHOT_ACCEPT);
io_uring_submit(&ring);
// 一次提交,持续产生 CQE
while (1) {
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
int new_fd = cqe->res; // 新连接 fd
if (new_fd >= 0) {
// 处理新连接:加入 epoll 或直接 read
handle_new_connection(new_fd);
}
io_uring_cqe_seen(&ring, cqe);
// 如果是 IORING_CQE_F_MORE 标志,说明 multi-shot 仍在继续
if (!(cqe->flags & IORING_CQE_F_MORE)) {
break; // multi-shot 结束了,需要重新提交 accept SQE
}
}
4.5 Splice:零拷贝数据管道
io_uring 提供 IORING_OP_SPLICE,在内核空间直接将数据从一个 FD 管道转移到另一个 FD,完全绕过用户态:
int pipefd[2];
pipe(pipefd);
// 从 source_fd 读数据,通过管道写入 dest_fd
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_splice(sqe, source_fd, -1, dest_fd, -1, 4096, SPLICE_F_MOVE);
io_uring_submit(&ring);
典型应用:Web Server 的 sendfile 替代品。相比 sendfile() 的同步调用,splice 可完全异步化。
五、高性能实战:基于 io_uring 的 Echo Server
以下是一个完整的、可用的 io_uring echo server 实现,包含连接管理、异步 IO 处理和优雅关闭:
5.1 完整源码
// io_uring_echo_server.c
// gcc -o echo_server io_uring_echo_server.c -luring -O2
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <errno.h>
#include <fcntl.h>
#include <netinet/in.h>
#include <sys/socket.h>
#include <liburing.h>
#define MAX_CONNECTIONS 4096
#define BUF_SIZE 4096
#define BACKLOG 128
#define ENTRIES 128
enum {
OP_ACCEPT = 1,
OP_READ,
OP_WRITE,
};
struct conn_info {
int fd;
unsigned int type;
char buf[BUF_SIZE];
unsigned int len;
};
static int setup_listening_socket(int port) {
int sock = socket(AF_INET, SOCK_STREAM, 0);
struct sockaddr_in addr = {
.sin_family = AF_INET,
.sin_addr.s_addr = htonl(INADDR_ANY),
.sin_port = htons(port),
};
int opt = 1;
setsockopt(sock, SOL_SOCKET, SO_REUSEADDR, &opt, sizeof(opt));
bind(sock, (struct sockaddr*)&addr, sizeof(addr));
listen(sock, BACKLOG);
return sock;
}
static void queue_accept(struct io_uring *ring, int listen_fd) {
struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
struct conn_info *ci = malloc(sizeof(struct conn_info));
memset(ci, 0, sizeof(*ci));
ci->fd = listen_fd;
ci->type = OP_ACCEPT;
io_uring_prep_accept(sqe, listen_fd, NULL, NULL, SOCK_NONBLOCK);
io_uring_sqe_set_data(sqe, ci);
io_uring_submit(ring);
}
static void queue_read(struct io_uring *ring, int fd) {
struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
struct conn_info *ci = malloc(sizeof(struct conn_info));
ci->fd = fd;
ci->type = OP_READ;
io_uring_prep_recv(sqe, fd, ci->buf, BUF_SIZE, 0);
io_uring_sqe_set_data(sqe, ci);
io_uring_submit(ring);
}
static void queue_write(struct io_uring *ring, int fd, void *buf, unsigned int len) {
struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
struct conn_info *ci = malloc(sizeof(struct conn_info));
ci->fd = fd;
ci->type = OP_WRITE;
memcpy(ci->buf, buf, len);
ci->len = len;
io_uring_prep_send(sqe, fd, ci->buf, len, 0);
io_uring_sqe_set_data(sqe, ci);
io_uring_submit(ring);
}
int main() {
struct io_uring ring;
io_uring_queue_init(ENTRIES, &ring, 0);
int listen_fd = setup_listening_socket(9999);
printf("io_uring echo server listening on port 9999...\n");
// 提交初始 accept
queue_accept(&ring, listen_fd);
while (1) {
struct io_uring_cqe *cqe;
int ret = io_uring_wait_cqe(&ring, &cqe);
if (ret < 0) {
fprintf(stderr, "wait_cqe: %s\n", strerror(-ret));
break;
}
struct conn_info *ci = io_uring_cqe_get_data(cqe);
int type = ci->type;
if (type == OP_ACCEPT) {
int new_fd = cqe->res;
if (new_fd >= 0) {
queue_read(&ring, new_fd); // 新连接:开始读数据
}
// 无论如何要重新提交 accept
queue_accept(&ring, listen_fd);
}
else if (type == OP_READ) {
int bytes_read = cqe->res;
if (bytes_read <= 0) {
close(ci->fd); // 连接关闭或出错
free(ci);
} else {
// 读完数据后提交写
queue_write(&ring, ci->fd, ci->buf, bytes_read);
free(ci);
}
}
else if (type == OP_WRITE) {
// 写完后重新提交读
queue_read(&ring, ci->fd);
free(ci);
}
io_uring_cqe_seen(&ring, cqe);
}
io_uring_queue_exit(&ring);
close(listen_fd);
return 0;
}
5.2 编译与测试
gcc -o echo_server io_uring_echo_server.c -luring -O2
# 后台运行
./echo_server &
# 测试
echo "Hello io_uring" | nc localhost 9999
# 性能对比(使用 wrk)
wrk -t2 -c1000 -d30s http://localhost:9999
核心设计要点
- 状态机模型:每个连接在 READ → WRITE → READ 之间循环流转,完全异步,不阻塞 event loop
- 零阻塞:accept/read/write 均为 io_uring 异步操作,单线程即可处理数千并发
- 内存管理:每个 I/O 操作使用 conn_info 结构传递 fd 和状态,完成后统一释放
- 可扩展性:可轻松扩展为 HTTP 服务器(在 read 回调中解析 HTTP 请求,write 回调中发送响应)
六、io_uring vs epoll vs IOCP vs kqueue
| 特性 | io_uring | epoll | IOCP (Windows) | kqueue (BSD/macOS) |
|---|---|---|---|---|
| 适用场景 | 全能型:文件+网络I/O | 网络事件通知 | 全能型(文件+网络) | 网络+文件系统事件 |
| 提交方式 | 共享内存批量(零 syscall) | 每次 epoll_ctl 1 syscall | 共享 OVERLAPPED | 内核事件队列 |
| 完成通知 | CQ 环形缓冲区(无需 syscall) | epoll_wait | 完成端口 | kevent |
| 文件 I/O 支持 | 真正的异步+零拷贝 | 不支持(仍需同步 read) | 真正的异步 | EVFILT_READ 需同步 |
| 轮询模式 | IOPOLL/SQPOLL 支持 | 不支持 | 支持 | 不支持 |
| 非阻塞回退 | 自动回退(NOOP_LINK 等) | 不适用 | 自动回退 | 不适用 |
| 编程复杂度 | 中等(liburing 简化) | 高(需自己管理状态机) | 中等 | 中等 |
七、生产环境部署最佳实践
7.1 SQPOLL 模式的正确使用
IORING_SETUP_SQPOLL 让内核线程持续轮询 SQ,用户态无需调用 io_uring_submit。但需要注意:
- CPU 占用:SQPOLL 线程会持续占用指定的 CPU 核心。建议使用
IORING_SETUP_SQ_AFF绑定到专用核 - 超时配置:通过
sq_thread_idle(ms)控制内核线程空闲超时,超时后线程进入睡眠,下次有 SQE 时唤醒 - 权限要求:SQPOLL 需要 CAP_SYS_ADMIN 权限(可在 /proc/sys/kernel/io_uring_disabled 中调整限制级别)
7.2 缓冲区注册策略
- 所有连接使用统一的 Buffer Pool(如 4096 个 4KB 缓冲区),比每个连接独立分配更高效
- 注册后不要 munmap,否则引发内核 page fault 崩溃
- 使用
io_uring_register_buffers_update支持动态更新
7.3 错误处理与降级
在某些内核版本中,特定操作可能返回 -EOPNOTSUPP。建议:
- 启动时探测
io_uring_probe确认可用 opcodes - 关键路径上预留同步 I/O 降级路径
- 监控 /proc/sys/kernel/io_uring_max_entries(若存在)调整队列深度
7.4 容器化部署
- Docker 默认禁用 io_uring(安全沙箱限制),需在
docker run --security-opt seccomp=unconfined或自定义 seccomp profile 中放行 - Kubernetes Pod 的 securityContext.seccompProfile.type=Unconfined 可解决
- 注意 io_uring 可能被用于内核漏洞利用(CVE-2023-2598 等),生产环境保持内核更新
八、主流 io_uring 生态项目
| 项目 | 用途 | io_uring 使用方式 |
|---|---|---|
| SPDK | 存储性能开发库 | NVMe 轮询模式(IOPOLL)+ Fixed Buffers |
| RocksDB | KV 存储引擎 | io_uring 异步 fsync/read,替代 posix AIO |
| Nginx (第三方模块) | Web 服务器 | ngx_http_io_uring_module,异步文件读写 |
| Tokio (uring-sys) | Rust 异步运行时 | io_uring 后端(替代 epoll 驱动) |
| Glommio | Rust 异步运行时 | 原生 io_uring 设计(thread-per-core) |
| PostgreSQL (实验性) | 关系型数据库 | io_uring 异步数据文件 I/O(PG 16+) |
| netty-incubator/transport-io_uring | Java NIO 框架 | io_uring transport(替代 epoll) |
九、总结
io_uring 已经从一个"新兴 Linux 特性"成长为高性能 I/O 的事实标准。其核心优势在于:
- 极低的 syscall 开销:共享内存 + 批量提交,理论上可逼近 kernel-bypass 框架(如 DPDK)的性能
- 真正的全异步 I/O:覆盖文件、网络、管道的所有 I/O 类型,不再受 POSIX AIO 的 O_DIRECT 限制
- 丰富的优化原语:Fixed Buffers、Fixed Files、Multi-shot、Linked SQE 等,覆盖了高性能场景的几乎所有需求
- 活跃的生态:从 SPDK 到 Tokio-Rust,从数据库到 Web 框架,io_uring 正逐步成为高性能后端系统的底层标配
对于任何 Linux 平台下的 I/O 密集型应用——无论是网络代理、存储引擎、数据库还是消息队列,io_uring 都值得认真评估。它不是银弹,但在 async I/O 这条赛道上,io_uring 已经是 Linux 生态中最优解。
参考资源
- 官方文档:Efficient IO with io_uring(Jens Axboe,io_uring 论文)
- liburing 源码:https://github.com/axboe/liburing
- Lord of the io_uring:深入理解 io_uring(经典入门教程)
- Kernel 文档:io_uring - Linux Kernel Documentation
- 性能基准测试:io_uring-bench

发表评论 取消回复