Linux 内核网络栈深度实战:从 epoll 到 io_uring

从阻塞 IO 到异步革命,一文读懂 Linux 高性能网络编程的演进之路。

引言:为什么我们需要高性能 IO?

在网络编程中,IO 始终是瓶颈所在。一个典型的 Web 服务器需要同时处理数千甚至数万连接,如果每个连接都占用一个阻塞线程,系统资源将迅速耗尽。Linux 内核为此提供了一系列高性能 IO 机制——select、poll、epoll,以及最新的 io_uring。

本文将从内核源码层面,深入剖析 epoll 和 io_uring 的设计哲学、数据结构与性能差异,并给出实战建议。

一、IO 模型演进:从阻塞到异步

1.1 传统阻塞 IO

最原始的 TCP 服务端模型:

int client_fd = accept(listen_fd, ...);
char buf[4096];
n = read(client_fd, buf, sizeof(buf));  // 阻塞等待数据
process(buf, n);
write(client_fd, response, len);

每个连接独占一个线程,并发一上来就崩溃。

1.2 非阻塞 IO + 轮询

设置 O_NONBLOCK,循环调用 read(),没数据就返回 EAGAIN。

问题:CPU 空转,n 个连接要轮询 n 次。

1.3 select 与 poll

// select:监视 FD 集合
fd_set rfds;
FD_ZERO(&rfds);
FD_SET(sockfd, &rfds);
select(sockfd + 1, &rfds, NULL, NULL, &timeout);

select 的局限: - FD 集合大小固定(FD_SETSIZE = 1024) - 每次调用需重置 fd_set - O(n) 遍历整个集合

poll 的改进: - 用 pollfd 数组替代位图,无 1024 限制 - 仍然是 O(n) 轮询

1.4 epoll:事件驱动的里程碑

epoll 是 Linux 2.6 引入的高性能事件通知机制,核心特性: - O(1) 事件通知(红黑树 + 就绪队列) - 支持边缘触发(ET)和水平触发(LT) - 无 FD 数量限制

二、epoll 深度解析

2.1 核心 API

// 1. 创建 epoll 实例
int epfd = epoll_create1(EPOLL_CLOEXEC);

// 2. 注册/修改/删除监听事件
struct epoll_event ev;
ev.events = EPOLLIN | EPOLLET;  // 边缘触发
ev.data.ptr = my_data;
epoll_ctl(epfd, EPOLL_CTL_ADD, client_fd, &ev);

// 3. 等待事件
struct epoll_event events[MAX_EVENTS];
int nfds = epoll_wait(epfd, events, MAX_EVENTS, timeout_ms);
for (int i = 0; i < nfds; i++) {
    handle_event(&events[i]);
}

2.2 内核数据结构

┌─────────────────────────────────────────────┐
│  epoll 内核架构                                │
├─────────────────────────────────────────────┤
│                                               │
│  epoll_create1 → eventpoll 对象               │
│    ├── rbr (红黑树根):管理所有监控的 fd       │
│    ├── rdllist (就绪队列):双链表,存储就绪 fd │
│    └── wq (等待队列):sleep 的进程             │
│                                               │
│  当 socket 有数据到达:                        │
│    1. 中断 → 协议栈收包                        │
│    2. sock_def_readable() 检查等待队列         │
│    3. 调用 ep_poll_callback()                  │
│    4. 将 epitem 加入 rdllist                   │
│    5. 唤醒 epoll_wait 的进程                   │
│                                               │
└─────────────────────────────────────────────┘

关键结构体:

// 每个被监控的 fd 对应一个 epitem
struct epitem {
    struct rb_node rbn;      // 红黑树节点
    struct list_head rdllink;// 就绪队列链表节点
    struct epoll_filefd ffd; // fd + file 指针
    struct eventpoll *ep;    // 所属的 epoll 实例
    struct epoll_event event;// 注册的事件掩码
};

// epoll 实例本身
struct eventpoll {
    struct mutex mtx;           // 保护此结构
    wait_queue_head_t wq;       // epoll_wait 等待队列
    wait_queue_head_t poll_wait;// epoll 自身被 poll 时的队列
    struct list_head list_head; // 就绪队列(双链表)
    struct rb_root rbr;         // 红黑树根
    ...
};

2.3 边缘触发(ET)vs 水平触发(LT)

特性 LT(默认) ET
通知条件 fd 可读/可写就通知 状态变化才通知
数据处理 读到 EAGAIN 即可 必须循环 read 到 EAGAIN
性能 适中(可能重复通知) 高(减少系统调用次数)
编程复杂度 低 高(容易遗漏事件)
适用场景 通用 高并发 + 非阻塞 IO

ET 模式下的正确写法:

void handle_et(int fd) {
    while (1) {
        ssize_t n = read(fd, buf, sizeof(buf));
        if (n > 0) {
            process(buf, n);
        } else if (n == 0) {
            close(fd);  // 对端关闭
            break;
        } else { // n < 0
            if (errno == EAGAIN || errno == EWOULDBLOCK) {
                break;  // 数据全部读完
            }
            handle_error(errno);
            break;
        }
    }
}

2.4 epoll + 多线程架构

现代高性能服务器通常采用以下模式:

┌─────────────────────────────────┐
│  Main Reactor (accept)          │
│  ├── epoll_wait → accept        │
│  └── round-robin 分发给 Sub      │
│                                   │
│  Sub Reactor 1                  │
│  ├── epoll_wait → read/write    │
│  └── 业务处理                    │
│                                   │
│  Sub Reactor 2                  │
│  └── ...                        │
└─────────────────────────────────┘

关键点: - 主 Reactor 专职 accept,子 Reactor 负责 IO - 子 Reactor 数量通常 = CPU 核心数 - 避免惊群:用 EPOLLEXCLUSIVE 或 SO_REUSEPORT

三、io_uring:异步 IO 的新纪元

3.1 epoll 的局限

epoll 解决了 IO 多路复用的问题,但它本质上是同步的: - 数据就绪通知后,仍需调用 read()/write() 系统调用 - 每次系统调用涉及用户态/内核态切换 - 对于磁盘 IO,无法真正实现异步(Linux AIO 有诸多限制)

3.2 io_uring 的设计哲学

io_uring(Linux 5.1+)采用共享内存 + 生产者-消费者模型,消除了系统调用的开销:

┌─────────────────────────────────────────────────────┐
│  io_uring 架构                                       │
├─────────────────────────────────────────────────────┤
│                                                       │
│  用户空间              共享内存              内核空间  │
│  ┌───┐         ┌──────────────┐         ┌────────┐  │
│  │app│ ──SQEs─→│  SQ (提交队列) │ ──消费─→│Kernel  │  │
│  │   │         │              │         │        │  │
│  │   │ ←─CQEs──│  CQ (完成队列) │ ←─生产──│        │  │
│  └───┘         └──────────────┘         └────────┘  │
│                                                       │
│  只需一次 io_uring_enter():                          │
│  - 批量提交 SQEs                                      │
│  - 批量收集 CQEs                                      │
│  - FIXED_FILES/BUFFER:进一步减少开销                 │
│                                                       │
└─────────────────────────────────────────────────────┘

3.3 核心 API(liburing)

#include <liburing.h>

// 1. 初始化
struct io_uring ring;
io_uring_queue_init(QUEUE_DEPTH, &ring, 0);

// 2. 获取一个 SQE(提交队列条目)
struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);

// 3. 准备一个 read 操作
io_uring_prep_read(sqe, fd, buf, len, offset);
io_uring_sqe_set_data(sqe, my_ctx);  // 关联上下文

// 4. 批量提交(可一次提交多个 SQE)
io_uring_submit(&ring);

// 5. 收割完成事件
struct io_uring_cqe *cqe;
io_uring_wait_cqe(&ring, &cqe);
handle_completion(cqe);
io_uring_cqe_seen(&ring, cqe);

3.4 Fixed Files 与 Registered Buffers

io_uring 的两大零拷贝优化:

Registered Buffers(预注册缓冲区):

```char *bufs[BUFFERS_COUNT]; // 一次性注册所有缓冲区 struct io_uring_provide_buf { __u64 addr; __u32 len; __u16 bid; }; io_uring_register_buffers(&ring, iovecs, BUFFERS_COUNT);

// 后续操作使用预注册缓冲区 io_uring_prep_read_fixed(sqe, fd, NULL, len, offset, bid);

**Fixed Files(预注册文件描述符):**

```c
// 预注册 fd 数组
int fds[] = {socket_fd, file_fd, ...};
io_uring_register_files(&ring, fds, count);

// 使用固定索引(避免每次 fget/fput 开销)
io_uring_prep_read(sqe, 0, buf, len, offset);  // 使用 fds[0]
sqe->flags |= IOSQE_FIXED_FILE;

3.5 SQPOLL 模式:真正的零系统调用

struct io_uring_params params = {
    .flags = IORING_SETUP_SQPOLL,
    .sq_thread_idle = 2000,  // 空闲 2s 后睡眠
};
io_uring_queue_init_params(QUEUE_DEPTH, &ring, &params);

在 SQPOLL 模式下: - 内核线程主动轮询 SQ - 用户空间只需填充 SQ 数据并写 sq->tail - 只有需要 CQ 收割时才进入内核(可设 IORING_SETUP_SQ_AFF)

3.6 io_uring vs epoll 性能对比

维度 epoll io_uring
系统调用开销 每次 IO 2 次(epoll_wait + read/write) 可接近 0(SQPOLL)
IO 类型 网络 IO 为主 网络 + 磁盘 + 一切
异步能力 伪异步(通知后仍需同步IO) 真正异步
零拷贝支持 无 Registered Buffers + Fixed Files
适用场景 网络高并发 全能型(server、DB、存储)
内核版本要求 2.6+ 5.1+(5.10+ 功能完整)

四、实战:io_uring echo server

下面是一个基于 io_uring 的最小 TCP echo server:

#define QUEUE_DEPTH 4096
#define BUF_SIZE 2048
#define BUF_GROUP 131

struct conn_info {
    int fd;
    unsigned type;  // 0=accept, 1=read, 2=write
    char buf[BUF_SIZE];
};

static struct io_uring ring;

void add_accept(int listen_fd, struct sockaddr *client_addr,
                socklen_t *client_len) {
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    io_uring_prep_accept(sqe, listen_fd, client_addr, client_len, 0);
    sqe->flags |= IOSQE_IO_LINK;  // 链式:accept → recv

    struct conn_info *info = malloc(sizeof(*info));
    info->fd = listen_fd;
    info->type = 0;
    io_uring_sqe_set_data(sqe, info);
}

void add_recv(int fd) {
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    io_uring_prep_recv(sqe, fd, NULL, BUF_SIZE, 0);
    sqe->buf_group = BUF_GROUP;
    sqe->flags |= IOSQE_BUFFER_SELECTION;  // 自动选缓冲区
    sqe->flags |= IOSQE_IO_LINK;  // 链式:recv → send

    struct conn_info *info = malloc(sizeof(*info));
    info->fd = fd;
    info->type = 1;
    io_uring_sqe_set_data(sqe, info);
}

void add_send(int fd, void *buf, size_t len, unsigned bgid) {
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    io_uring_prep_send(sqe, fd, buf, len, 0);
    sqe->flags |= IOSQE_BUFFER_SELECTION | IOSQE_IO_LINK;
    sqe->buf_group = bgid;

    struct conn_info *info = malloc(sizeof(*info));
    info->fd = fd;
    info->type = 2;
    io_uring_sqe_set_data(sqe, info);
}

int main() {
    // 1. 初始化 io_uring
    struct io_uring_params params = {0};
    params.cq_entries = QUEUE_DEPTH * 4;
    io_uring_queue_init_params(QUEUE_DEPTH, &ring, &params);

    // 2. 注册多组 buffer(multishot buffer select)
    struct iovec iovecs[BUF_GROUP];
    for (int i = 0; i < BUF_GROUP; i++) {
        iovecs[i].iov_base = malloc(BUF_SIZE);
        iovecs[i].iov_len = BUF_SIZE;
    }
    io_uring_register_buffers(&ring, iovecs, BUF_GROUP);
    io_uring_register_buf_ring(&ring, iovecs, BUF_GROUP, BUF_GROUP);

    // 3. 创建 listen socket
    int listen_fd = socket(AF_INET, SOCK_STREAM, 0);
    // ... bind, listen ...

    // 4. 启动 accept
    add_accept(listen_fd, NULL, NULL);
    io_uring_submit(&ring);

    // 5. 事件循环
    while (1) {
        struct io_uring_cqe *cqe;
        io_uring_wait_cqe(&ring, &cqe);

        struct conn_info *info = io_uring_cqe_get_data(cqe);
        int res = cqe->res;

        switch (info->type) {
        case 0:  // accept 完成
            add_accept(listen_fd, NULL, NULL);  // 继续 accept
            if (res > 0) add_recv(res);          // 新连接开始读
            break;
        case 1:  // recv 完成
            if (res > 0) {
                add_send(info->fd, NULL, res, cqe->buf_group_id);
                add_recv(info->fd);  // 继续读
            } else {
                close(info->fd);
            }
            break;
        case 2:  // send 完成
            // 缓冲区自动归还 buffer ring
            break;
        }

        io_uring_cqe_seen(&ring, cqe);
        free(info);
    }
}

4.1 链式 SQE(IOSQE_IO_LINK)

io_uring 支持将多个操作串联起来: - SQE A → SQE B → SQE C - 前一个完成后才执行下一个 - 避免了事件循环中的状态机编程

4.2 Multishot 模式(Linux 5.19+)

// Multishot accept:一次注册,多次触发
io_uring_prep_multishot_accept(sqe, listen_fd, ..., 0);

// Multishot recv:数据到达时自动选 buffer 并重发
io_uring_prep_recv_multishot(sqe, fd, NULL, 0, 0);

相比传统模式减少了 50% 的 SQE 提交。

五、选型建议

5.1 什么时候用 epoll?

  • 纯网络高并发(如 Redis、Nginx)
  • 连接数 < 100K,每一连接 IO 频率适中
  • 快速开发部署,生态成熟
  • 需要兼容 Linux 2.6 ~ 各项老内核

5.2 什么时候用 io_uring?

  • 磁盘 IO 重负载(数据库、对象存储)
  • 要求极低延迟(HFT、游戏服务器)
  • 连接数 > 100K,需要极大减少系统调用
  • 已有 liburing 或封装库支持(如 Tokio 的 io_uring 后端)

5.3 混合架构

┌────────────────────────────────────────┐
│  Nginx / Envoy                          │
│   TLS 终止 + 负载均衡                    │
│     (epoll 驱动)                       │
├────────────────────────────────────────┤
│  业务逻辑层                               │
│    低延迟走 io_uring (SQPOLL)            │
│    高吞吐走 epoll (线程池)               │
├────────────────────────────────────────┤
│  存储层                                  │
│    数据库引擎 (io_uring direct IO)       │
│    kv 引擎 (io_uring + SPDK)             │
└────────────────────────────────────────┘

六、性能调优参数

6.1 epoll 相关

# 单进程可打开的最大 fd 数
sysctl -w fs.file-max=2097152
sysctl -w fs.nr_open=2097152
ulimit -n 1048576

# TCP backlog
sysctl -w net.core.somaxconn=65535
sysctl -w net.ipv4.tcp_max_syn_backlog=65535

# TCP 缓冲区
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.wmem_max=16777216
sysctl -w net.ipv4.tcp_rmem="4096 87380 16777216"
sysctl -w net.ipv4.tcp_wmem="4096 65536 16777216"

6.2 io_uring 相关

# 最大 CQ 条目(影响完成队列大小)
sysctl -w kernel.io_uring_max_cq=2147483647

# SQPOLL 线程优先级
chrt -f 50 /path/to/your/io_uring_app

# 使用 Registered Buffers 减少 get_user_pages 开销
# 使用 Polling IO(需要块设备支持)

七、总结

从 select 的 O(n) 遍历,到 epoll 的 O(1) 事件通知,再到 io_uring 的批处理异步提交,Linux 性能优化始终是"减少系统调用"这一核心矛盾的演进。

epoll 仍是网络服务器的主流选择,但 io_uring 正逐步进入主流。PhotonLibOS、Glommio、Tokio-uring 等项目已经证明,io_uring 可以带来 30%~50% 的吞吐提升。对于新项目,建议直接评估 io_uring 方案。


参考资料: - Linux 内核源码 fs/io_uring.c、fs/eventpoll.c - Jens Axboe, "Efficient IO with io_uring" (kernel.dk) - liburing 官方文档: https://axboe.dk/liburing/ - "The rapid growth of io_uring" (LWN.net)

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿
网站二维码

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部
/* 跳过导航链接 (无障碍) */ position: absolute; top: -100px; left: 15px; z-index: 99999; padding: 8px 16px; background: #007bff; color: #fff; font-size: 14px; border-radius: 0 0 4px 4px; text-decoration: none; transition: top 0.2s; } top: 0; outline: 3px solid #0056b3; }