Linux io_uring: Async IO Revolution

1. Why io_uring? - Pain Points of Traditional Async IO

Traditional Linux async IO has two main approaches, both with significant limitations:

1.1 POSIX AIO (libaio) Limitations

Since Linux 2.6, libaio requires O_DIRECT (no page cache), does not support sockets, and has an awkward API requiring aiocb structure chaining. The kernel implementation unpredictably oscillates between truly async and fallback-sync execution.

1.2 epoll is NOT Async IO

epoll provides only an event notification mechanism (reactor pattern). Actual read/write operations still require user-space system calls, creating a bottleneck under extreme concurrency.

1.3 io_uring Is Born

Merged into Linux 5.1 in 2019 by Jens Axboe (creator of epoll/splice). Goal: truly asynchronous, unified interface with zero syscall overhead.

2. Core Architecture

2.1 Dual Ring Queue Design

io_uring is built on two shared memory ring buffers:

  • Submission Queue (SQ): The user writes SQE (Submission Queue Entry) structs to request IO. The kernel consumes these.
  • Completion Queue (CQ): The kernel writes CQE (Completion Queue Entry) structs to signal completion. The user consumes these.

Both queues map to shared memory. After initialization, batch submission and batch completion processing require zero syscalls - only memory barriers and doorbell notifications.

2.2 Three Working Modes

ModeConfigMechanismUse Case
Interrupt-drivenDefault (none)Kernel wakes user via completion eventGeneral workloads
Kernel polling (SQPOLL)IORING_SETUP_SQPOLLKernel thread polls SQ continuouslyUltra-low latency
IO polling (IOPOLL)IORING_SETUP_IOPOLLHardware completion queue polledNVMe SSD sub-10us

SQPOLL mode changes everything: a kernel worker thread (io-wrk) continuously scans SQ for new entries. When it finds one, it dispatches the IO immediately. The user process fills SQEs and advances the tail pointer with only shared memory writes.

2.3 Fixed Files and Registered Buffers

Default IO requires per-operation fd-to-file lookup and reference counting. Fixed Files pre-register an array of fds for direct index access. Registered Buffers pre-pin user memory, eliminating get_user_pages() overhead. At 500K+ IOPS, these save 20-30 percent submission-side CPU.

3. liburing Development

3.1 Setup

git clone https://github.com/axboe/liburing

3.2 Hello World

#include 
#include 

int main(void) {
    struct io_uring ring;
    struct io_uring_sqe *sqe;
    struct io_uring_cqe *cqe;
    char buf[4096];
    int fd, ret;

    io_uring_queue_init(1, ˚, 0);
    fd = open("/tmp/hello.txt", O_RDONLY);
    sqe = io_uring_get_sqe(˚);
    memset(buf, 0, sizeof(buf));
    io_uring_prep_read(sqe, fd, buf, sizeof(buf)-1, 0);
    io_uring_sqe_set_data(sqe, buf);
    io_uring_submit(˚);
    io_uring_wait_cqe(˚, &cqe);
    if (cqe->res >= 0)
        printf("Read OK\n");
    io_uring_cqe_seen(˚, cqe);
    close(fd);
    io_uring_queue_exit(˚);
    return 0;
}

3.3 Batch Submission

struct io_uring ring;
io_uring_queue_init(32, ˚, 0);
for (int i = 0; i < 32 xss=removed xss=removed>

3.4 Linked Operations

/* Chain: read from source, write to dest */
sqe = io_uring_get_sqe(˚);
io_uring_prep_read(sqe, src_fd, buf, len, 0);
io_uring_sqe_set_flags(sqe, IOSQE_IO_LINK);
sqe = io_uring_get_sqe(˚);
io_uring_prep_write(sqe, dst_fd, buf, len, 0);
io_uring_submit(˚);

4. Performance

4.1 Environment

  • CPU: Dual AMD EPYC 7763 (128 cores total)
  • Storage: Samsung PM9A3 NVMe SSD (PCIe 4.0, 7GB/s)
  • Kernel: Linux 6.1 LTS

4.2 4KB Random Read IOPS

ImplementationIOPSCPU CoresP99 Latency
pread (sync)150,0004 (100pc)35 us
libaio200,0003.5 (100pc)28 us
io_uring (interrupt)280,0002 (80pc)18 us
io_uring (SQPOLL)420,0001.5 (60pc)12 us
io_uring (IOPOLL)650,0003 (90pc)6 us

4.3 Key Findings

  • SQPOLL: 50pc less CPU than libaio, 60pc less than sync pread
  • io_uring P99 latency: consistent and predictable, less tail jitter
  • IOPOLL saturates NVMe hardware, 4x sync pread throughput

5. Advanced Features

5.1 Buffered vs Direct IO

Since Linux 5.19, io_uring optimizes buffered I/O: the kernel defers IO to background workers for truly async cached reads. Readv2/writev2 helpers with RWF_NOWAIT return immediately on page cache hits.

5.2 Network Socket IO

Linux 5.6+ supports socket operations: connect, accept, recv. A single io_uring instance now manages both disk and network IO.

/* Async accept + chained recv */
sqe = io_uring_get_sqe(˚);
io_uring_prep_accept(sqe, listen_fd, &addr, &len, 0);
sqe = io_uring_get_sqe(˚);
io_uring_prep_recv(sqe, cli_fd, rbuf, sizeof(rbuf), 0);
io_uring_sqe_set_flags(sqe, IOSQE_IO_LINK);
io_uring_submit(˚);

5.3 Zero Copy with Buffer Rings

Linux 6.1 introduced IORING_REGISTER_PBUF_RING. Combined with MSG_ZEROCOPY, enables full-path zero-copy: NIC DMA straight to a user-registered buffer, then written to disk without intermediate copies.

6. Production Deployment

6.1 RocksDB 7.0

Meta integrated io_uring as an async IO backend: random read throughput improved by 40pc, P99 latency dropped by 25pc, compaction interference reduced by 30pc.

6.2 Nginx 1.25 - QUIC

Experimental io_uring support: TLS handshake + accept latency down 40pc, HTTP header P99 latency from 800us to 250us.

6.3 PostgreSQL Experiments

Community patches using io_uring for WAL fsync: 2-3x TPS improvement under high concurrency.

6.4 Best Practices

  • Queue depth: NVMe 256-1024, SATA SSD 32-128
  • Buffer alignment: O_DIRECT requires 512b or 4096b alignment
  • SQPOLL idle: Set to 2000-5000ms balance latency vs CPU
  • Error handling: cqe->res less than 0 = error; distinguish EAGAIN from fatal
  • Cleanup: Drain all CQEs before io_uring_queue_exit()

7. Limitations and Roadmap

7.1 Current Constraints

  • Kernel: Production stable needs 5.10+, advanced features need 6.1+
  • Signal safety: Not async-signal-safe
  • NUMA: SQPOLL kernel thread binds single NUMA node
  • Observability: SQ and CQ split submission/completion contexts

7.2 Future Direction

  • io_uring + eBPF: BPF callback on completion events
  • Userfaultfd: Page fault async via io_uring
  • Network: Potentially replace epoll for high-concurrency networking
  • Security: SELinux/AppArmor policy hooks for io_uring

8. Summary

io_uring is not merely a new IO API but a fundamental shift in the Linux IO subsystem. It solves the true-async problem that plagued AIO for twenty years. The shared-memory dual-ring architecture with SQPOLL achieves zero-syscall submission. For storage engines, network proxies, and databases, io_uring has evolved from experimental to essential.

On Linux 6.x and above, developers building new systems should prioritize io_uring as the default choice for high-performance asynchronous IO.

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部
0.355919s