Linux 内核网络栈深度实战:从 epoll 到 io_uring
从阻塞 IO 到异步革命,一文读懂 Linux 高性能网络编程的演进之路。
引言:为什么我们需要高性能 IO?
在网络编程中,IO 始终是瓶颈所在。一个典型的 Web 服务器需要同时处理数千甚至数万连接,如果每个连接都占用一个阻塞线程,系统资源将迅速耗尽。Linux 内核为此提供了一系列高性能 IO 机制——select、poll、epoll,以及最新的 io_uring。
本文将从内核源码层面,深入剖析 epoll 和 io_uring 的设计哲学、数据结构与性能差异,并给出实战建议。
一、IO 模型演进:从阻塞到异步
1.1 传统阻塞 IO
最原始的 TCP 服务端模型:
int client_fd = accept(listen_fd, ...);
char buf[4096];
n = read(client_fd, buf, sizeof(buf)); // 阻塞等待数据
process(buf, n);
write(client_fd, response, len);
每个连接独占一个线程,并发一上来就崩溃。
1.2 非阻塞 IO + 轮询
设置 O_NONBLOCK,循环调用 read(),没数据就返回 EAGAIN。
问题:CPU 空转,n 个连接要轮询 n 次。
1.3 select 与 poll
// select:监视 FD 集合
fd_set rfds;
FD_ZERO(&rfds);
FD_SET(sockfd, &rfds);
select(sockfd + 1, &rfds, NULL, NULL, &timeout);
select 的局限:
- FD 集合大小固定(FD_SETSIZE = 1024)
- 每次调用需重置 fd_set
- O(n) 遍历整个集合
poll 的改进:
- 用 pollfd 数组替代位图,无 1024 限制
- 仍然是 O(n) 轮询
1.4 epoll:事件驱动的里程碑
epoll 是 Linux 2.6 引入的高性能事件通知机制,核心特性: - O(1) 事件通知(红黑树 + 就绪队列) - 支持边缘触发(ET)和水平触发(LT) - 无 FD 数量限制
二、epoll 深度解析
2.1 核心 API
// 1. 创建 epoll 实例
int epfd = epoll_create1(EPOLL_CLOEXEC);
// 2. 注册/修改/删除监听事件
struct epoll_event ev;
ev.events = EPOLLIN | EPOLLET; // 边缘触发
ev.data.ptr = my_data;
epoll_ctl(epfd, EPOLL_CTL_ADD, client_fd, &ev);
// 3. 等待事件
struct epoll_event events[MAX_EVENTS];
int nfds = epoll_wait(epfd, events, MAX_EVENTS, timeout_ms);
for (int i = 0; i < nfds; i++) {
handle_event(&events[i]);
}
2.2 内核数据结构
┌─────────────────────────────────────────────┐
│ epoll 内核架构 │
├─────────────────────────────────────────────┤
│ │
│ epoll_create1 → eventpoll 对象 │
│ ├── rbr (红黑树根):管理所有监控的 fd │
│ ├── rdllist (就绪队列):双链表,存储就绪 fd │
│ └── wq (等待队列):sleep 的进程 │
│ │
│ 当 socket 有数据到达: │
│ 1. 中断 → 协议栈收包 │
│ 2. sock_def_readable() 检查等待队列 │
│ 3. 调用 ep_poll_callback() │
│ 4. 将 epitem 加入 rdllist │
│ 5. 唤醒 epoll_wait 的进程 │
│ │
└─────────────────────────────────────────────┘
关键结构体:
// 每个被监控的 fd 对应一个 epitem
struct epitem {
struct rb_node rbn; // 红黑树节点
struct list_head rdllink;// 就绪队列链表节点
struct epoll_filefd ffd; // fd + file 指针
struct eventpoll *ep; // 所属的 epoll 实例
struct epoll_event event;// 注册的事件掩码
};
// epoll 实例本身
struct eventpoll {
struct mutex mtx; // 保护此结构
wait_queue_head_t wq; // epoll_wait 等待队列
wait_queue_head_t poll_wait;// epoll 自身被 poll 时的队列
struct list_head list_head; // 就绪队列(双链表)
struct rb_root rbr; // 红黑树根
...
};
2.3 边缘触发(ET)vs 水平触发(LT)
| 特性 | LT(默认) | ET |
|---|---|---|
| 通知条件 | fd 可读/可写就通知 | 状态变化才通知 |
| 数据处理 | 读到 EAGAIN 即可 | 必须循环 read 到 EAGAIN |
| 性能 | 适中(可能重复通知) | 高(减少系统调用次数) |
| 编程复杂度 | 低 | 高(容易遗漏事件) |
| 适用场景 | 通用 | 高并发 + 非阻塞 IO |
ET 模式下的正确写法:
void handle_et(int fd) {
while (1) {
ssize_t n = read(fd, buf, sizeof(buf));
if (n > 0) {
process(buf, n);
} else if (n == 0) {
close(fd); // 对端关闭
break;
} else { // n < 0
if (errno == EAGAIN || errno == EWOULDBLOCK) {
break; // 数据全部读完
}
handle_error(errno);
break;
}
}
}
2.4 epoll + 多线程架构
现代高性能服务器通常采用以下模式:
┌─────────────────────────────────┐
│ Main Reactor (accept) │
│ ├── epoll_wait → accept │
│ └── round-robin 分发给 Sub │
│ │
│ Sub Reactor 1 │
│ ├── epoll_wait → read/write │
│ └── 业务处理 │
│ │
│ Sub Reactor 2 │
│ └── ... │
└─────────────────────────────────┘
关键点:
- 主 Reactor 专职 accept,子 Reactor 负责 IO
- 子 Reactor 数量通常 = CPU 核心数
- 避免惊群:用 EPOLLEXCLUSIVE 或 SO_REUSEPORT
三、io_uring:异步 IO 的新纪元
3.1 epoll 的局限
epoll 解决了 IO 多路复用的问题,但它本质上是同步的:
- 数据就绪通知后,仍需调用 read()/write() 系统调用
- 每次系统调用涉及用户态/内核态切换
- 对于磁盘 IO,无法真正实现异步(Linux AIO 有诸多限制)
3.2 io_uring 的设计哲学
io_uring(Linux 5.1+)采用共享内存 + 生产者-消费者模型,消除了系统调用的开销:
┌─────────────────────────────────────────────────────┐
│ io_uring 架构 │
├─────────────────────────────────────────────────────┤
│ │
│ 用户空间 共享内存 内核空间 │
│ ┌───┐ ┌──────────────┐ ┌────────┐ │
│ │app│ ──SQEs─→│ SQ (提交队列) │ ──消费─→│Kernel │ │
│ │ │ │ │ │ │ │
│ │ │ ←─CQEs──│ CQ (完成队列) │ ←─生产──│ │ │
│ └───┘ └──────────────┘ └────────┘ │
│ │
│ 只需一次 io_uring_enter(): │
│ - 批量提交 SQEs │
│ - 批量收集 CQEs │
│ - FIXED_FILES/BUFFER:进一步减少开销 │
│ │
└─────────────────────────────────────────────────────┘
3.3 核心 API(liburing)
#include <liburing.h>
// 1. 初始化
struct io_uring ring;
io_uring_queue_init(QUEUE_DEPTH, &ring, 0);
// 2. 获取一个 SQE(提交队列条目)
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
// 3. 准备一个 read 操作
io_uring_prep_read(sqe, fd, buf, len, offset);
io_uring_sqe_set_data(sqe, my_ctx); // 关联上下文
// 4. 批量提交(可一次提交多个 SQE)
io_uring_submit(&ring);
// 5. 收割完成事件
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
handle_completion(cqe);
io_uring_cqe_seen(&ring, cqe);
3.4 Fixed Files 与 Registered Buffers
io_uring 的两大零拷贝优化:
Registered Buffers(预注册缓冲区):
```char *bufs[BUFFERS_COUNT]; // 一次性注册所有缓冲区 struct io_uring_provide_buf { __u64 addr; __u32 len; __u16 bid; }; io_uring_register_buffers(&ring, iovecs, BUFFERS_COUNT);
// 后续操作使用预注册缓冲区 io_uring_prep_read_fixed(sqe, fd, NULL, len, offset, bid);
**Fixed Files(预注册文件描述符):**
```c
// 预注册 fd 数组
int fds[] = {socket_fd, file_fd, ...};
io_uring_register_files(&ring, fds, count);
// 使用固定索引(避免每次 fget/fput 开销)
io_uring_prep_read(sqe, 0, buf, len, offset); // 使用 fds[0]
sqe->flags |= IOSQE_FIXED_FILE;
3.5 SQPOLL 模式:真正的零系统调用
struct io_uring_params params = {
.flags = IORING_SETUP_SQPOLL,
.sq_thread_idle = 2000, // 空闲 2s 后睡眠
};
io_uring_queue_init_params(QUEUE_DEPTH, &ring, ¶ms);
在 SQPOLL 模式下:
- 内核线程主动轮询 SQ
- 用户空间只需填充 SQ 数据并写 sq->tail
- 只有需要 CQ 收割时才进入内核(可设 IORING_SETUP_SQ_AFF)
3.6 io_uring vs epoll 性能对比
| 维度 | epoll | io_uring |
|---|---|---|
| 系统调用开销 | 每次 IO 2 次(epoll_wait + read/write) | 可接近 0(SQPOLL) |
| IO 类型 | 网络 IO 为主 | 网络 + 磁盘 + 一切 |
| 异步能力 | 伪异步(通知后仍需同步IO) | 真正异步 |
| 零拷贝支持 | 无 | Registered Buffers + Fixed Files |
| 适用场景 | 网络高并发 | 全能型(server、DB、存储) |
| 内核版本要求 | 2.6+ | 5.1+(5.10+ 功能完整) |
四、实战:io_uring echo server
下面是一个基于 io_uring 的最小 TCP echo server:
#define QUEUE_DEPTH 4096
#define BUF_SIZE 2048
#define BUF_GROUP 131
struct conn_info {
int fd;
unsigned type; // 0=accept, 1=read, 2=write
char buf[BUF_SIZE];
};
static struct io_uring ring;
void add_accept(int listen_fd, struct sockaddr *client_addr,
socklen_t *client_len) {
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_accept(sqe, listen_fd, client_addr, client_len, 0);
sqe->flags |= IOSQE_IO_LINK; // 链式:accept → recv
struct conn_info *info = malloc(sizeof(*info));
info->fd = listen_fd;
info->type = 0;
io_uring_sqe_set_data(sqe, info);
}
void add_recv(int fd) {
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_recv(sqe, fd, NULL, BUF_SIZE, 0);
sqe->buf_group = BUF_GROUP;
sqe->flags |= IOSQE_BUFFER_SELECTION; // 自动选缓冲区
sqe->flags |= IOSQE_IO_LINK; // 链式:recv → send
struct conn_info *info = malloc(sizeof(*info));
info->fd = fd;
info->type = 1;
io_uring_sqe_set_data(sqe, info);
}
void add_send(int fd, void *buf, size_t len, unsigned bgid) {
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
io_uring_prep_send(sqe, fd, buf, len, 0);
sqe->flags |= IOSQE_BUFFER_SELECTION | IOSQE_IO_LINK;
sqe->buf_group = bgid;
struct conn_info *info = malloc(sizeof(*info));
info->fd = fd;
info->type = 2;
io_uring_sqe_set_data(sqe, info);
}
int main() {
// 1. 初始化 io_uring
struct io_uring_params params = {0};
params.cq_entries = QUEUE_DEPTH * 4;
io_uring_queue_init_params(QUEUE_DEPTH, &ring, ¶ms);
// 2. 注册多组 buffer(multishot buffer select)
struct iovec iovecs[BUF_GROUP];
for (int i = 0; i < BUF_GROUP; i++) {
iovecs[i].iov_base = malloc(BUF_SIZE);
iovecs[i].iov_len = BUF_SIZE;
}
io_uring_register_buffers(&ring, iovecs, BUF_GROUP);
io_uring_register_buf_ring(&ring, iovecs, BUF_GROUP, BUF_GROUP);
// 3. 创建 listen socket
int listen_fd = socket(AF_INET, SOCK_STREAM, 0);
// ... bind, listen ...
// 4. 启动 accept
add_accept(listen_fd, NULL, NULL);
io_uring_submit(&ring);
// 5. 事件循环
while (1) {
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
struct conn_info *info = io_uring_cqe_get_data(cqe);
int res = cqe->res;
switch (info->type) {
case 0: // accept 完成
add_accept(listen_fd, NULL, NULL); // 继续 accept
if (res > 0) add_recv(res); // 新连接开始读
break;
case 1: // recv 完成
if (res > 0) {
add_send(info->fd, NULL, res, cqe->buf_group_id);
add_recv(info->fd); // 继续读
} else {
close(info->fd);
}
break;
case 2: // send 完成
// 缓冲区自动归还 buffer ring
break;
}
io_uring_cqe_seen(&ring, cqe);
free(info);
}
}
4.1 链式 SQE(IOSQE_IO_LINK)
io_uring 支持将多个操作串联起来: - SQE A → SQE B → SQE C - 前一个完成后才执行下一个 - 避免了事件循环中的状态机编程
4.2 Multishot 模式(Linux 5.19+)
// Multishot accept:一次注册,多次触发
io_uring_prep_multishot_accept(sqe, listen_fd, ..., 0);
// Multishot recv:数据到达时自动选 buffer 并重发
io_uring_prep_recv_multishot(sqe, fd, NULL, 0, 0);
相比传统模式减少了 50% 的 SQE 提交。
五、选型建议
5.1 什么时候用 epoll?
- 纯网络高并发(如 Redis、Nginx)
- 连接数 < 100K,每一连接 IO 频率适中
- 快速开发部署,生态成熟
- 需要兼容 Linux 2.6 ~ 各项老内核
5.2 什么时候用 io_uring?
- 磁盘 IO 重负载(数据库、对象存储)
- 要求极低延迟(HFT、游戏服务器)
- 连接数 > 100K,需要极大减少系统调用
- 已有 liburing 或封装库支持(如 Tokio 的 io_uring 后端)
5.3 混合架构
┌────────────────────────────────────────┐
│ Nginx / Envoy │
│ TLS 终止 + 负载均衡 │
│ (epoll 驱动) │
├────────────────────────────────────────┤
│ 业务逻辑层 │
│ 低延迟走 io_uring (SQPOLL) │
│ 高吞吐走 epoll (线程池) │
├────────────────────────────────────────┤
│ 存储层 │
│ 数据库引擎 (io_uring direct IO) │
│ kv 引擎 (io_uring + SPDK) │
└────────────────────────────────────────┘
六、性能调优参数
6.1 epoll 相关
# 单进程可打开的最大 fd 数
sysctl -w fs.file-max=2097152
sysctl -w fs.nr_open=2097152
ulimit -n 1048576
# TCP backlog
sysctl -w net.core.somaxconn=65535
sysctl -w net.ipv4.tcp_max_syn_backlog=65535
# TCP 缓冲区
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.wmem_max=16777216
sysctl -w net.ipv4.tcp_rmem="4096 87380 16777216"
sysctl -w net.ipv4.tcp_wmem="4096 65536 16777216"
6.2 io_uring 相关
# 最大 CQ 条目(影响完成队列大小)
sysctl -w kernel.io_uring_max_cq=2147483647
# SQPOLL 线程优先级
chrt -f 50 /path/to/your/io_uring_app
# 使用 Registered Buffers 减少 get_user_pages 开销
# 使用 Polling IO(需要块设备支持)
七、总结
从 select 的 O(n) 遍历,到 epoll 的 O(1) 事件通知,再到 io_uring 的批处理异步提交,Linux 性能优化始终是"减少系统调用"这一核心矛盾的演进。
epoll 仍是网络服务器的主流选择,但 io_uring 正逐步进入主流。PhotonLibOS、Glommio、Tokio-uring 等项目已经证明,io_uring 可以带来 30%~50% 的吞吐提升。对于新项目,建议直接评估 io_uring 方案。
参考资料: - Linux 内核源码
fs/io_uring.c、fs/eventpoll.c- Jens Axboe, "Efficient IO with io_uring" (kernel.dk) - liburing 官方文档: https://axboe.dk/liburing/ - "The rapid growth of io_uring" (LWN.net)

发表评论 取消回复