Linux io_uring: A Deep Engineering Analysis of the Next-Generation Asynchronous I/O Framework
1. Introduction: The Limits of Traditional Linux AIO
Linux's native asynchronous I/O interface (io_submit / libaio) has been available since kernel 2.5 (2002), but its adoption in production has been limited by fundamental design flaws. The most critical restriction: AIO only supports asynchronous operation for O_DIRECT-based file I/O, falling back to blocking behavior when used with buffered I/O (normal file read/write). This means common operations like read(), write(), and network socket operations cannot benefit from true asynchronous processing.
Furthermore, AIO requires complex callback mechanism design via struct iocb, with every operation requiring initialization of multiple fields. Event retrieval relies on io_getevents() syscall, which not only incurs kernel-user context switching overhead but also lacks batch processing capabilities - each call handles only one event, causing severe throughput bottlenecks in high-concurrency scenarios.
io_uring, introduced by Jens Axboe in Linux 5.1 (2019), fundamentally solves these problems. Through its innovative dual-ring buffer (Submission Queue + Completion Queue) shared memory architecture, it achieves zero syscall overhead asynchronous I/O, supporting all file types (regular files, directories, sockets, pipes), and provides powerful extension capabilities like fixed buffers and fixed files. By Linux 5.10 LTS, io_uring has become production-ready, and modern databases and web servers are widely adopting it as the core I/O engine.
2. Architecture Design: Shared Memory Dual-Ring Buffer Model
2.1 Core Data Structures
io_uring's design philosophy is to make the fast path completely syscall-free. The kernel and userspace maintain communication through two ring buffers mapped to shared memory:
| Structure | Abbreviation | Direction | Function |
|---|---|---|---|
| Submission Queue | SQ | Userspace 鈫?Kernel | Userspace submits I/O requests (SQE) |
| Completion Queue | CQ | Kernel 鈫?Userspace | Kernel completes I/O results (CQE) |
Key design points:
- SQ and CQ are single-producer single-consumer ring buffers - Userspace writes SQEs (via SQ tail pointer), kernel writes CQEs (via CQ head pointer), requiring no locks in the uncontended case
- CQ has twice the capacity of SQ - Alerts before userspace can collect completed events if I/O requests are submitted faster than they are completed
- Direct mapping to userspace - Through
io_uring_setup()mmap, SQ/CQ rings are mapped at initialization, avoided subsequent mapping overhead - SQ polling mode - Kernel thread automatically polls SQ for new requests, userspace does not need to call syscall to submit
2.2 Three Stages of Operation
io_uring operations are divided into three stages, each optimized for different scenarios:
Stage 1: Interception Mode (Default)
Userspace Kernel
| |
|-- submit SQE --->| (write SQ tail)
| |
|-- io_uring_enter | (syscall: notify kernel)
| |
|<-- CQE --------- | (poll CQ head)
| |
Stage 2: SQPOLL Mode (Kernel Polling)
Userspace Kernel Thread
| |
|-- write SQ tail |
| |-- auto-polling SQ
| |-- execute I/O
|<-- CQE --------- | (poll CQ head)
| |
(NO syscall needed!)
Stage 3: IOPOLL Mode (Direct Completion)
Userspace Kernel + Block Device
| |
|-- write SQ tail |
| |-- polling completion
| | (no interrupts needed)
|<-- CQE --------- | (poll CQ head)
| |
(Ultra-low latency!)
3. System Call Interface and Ring Buffer Operations
3.1 io_uring_setup: Initialization
The initialization syscall creates an io_uring instance, allocating SQ and CQ ring buffers of specified size. During initialization, the kernel returns file descriptors and offset information needed to mmap the three memory regions:
int io_uring_setup(u32 entries, struct io_uring_params *p);
// Key parameters
struct io_uring_params {
__u32 sq_entries; // Actual SQ size returned by kernel
__u32 cq_entries; // Actual CQ size returned by kernel
__u32 flags; // IORING_SETUP flags
__u32 sq_thread_cpu; // SQPOLL thread CPU affinity
__u32 sq_thread_idle; // SQPOLL thread idle time (ms)
__u32 features; // Kernel-supported feature flags
__u32 wq_fd; // io_uring workqueue fd
__u8 cq_off[24]; // CQ ring offset for mmap
__u8 sq_off[40]; // SQ ring offset for mmap
__u8 sqes[128]; // SQE array offset for mmap
};
Three mmap configurations:
- SQ Ring - Queue metadata (head/tail pointer, ring mask, etc.), size typically 64 bytes 脳 entries
- CQ Ring - Queue metadata + CQE array, size typically sizeof(struct io_uring_cqe) 脳 cq_entries
- SQE Array - Submission Queue Element array (pre-allocated), size sizeof(struct io_uring_sqe) 脳 entries
3.2 io_uring_enter: Notification and Waiting
The enter syscall serves dual purposes: notifying the kernel to process SQ entries and optionally blocking for completion events:
int io_uring_enter(unsigned int fd, unsigned int to_submit,
unsigned int min_complete, unsigned int flags,
sigset_t *sig);
// Typical usage:
io_uring_enter(ring_fd, 128, 1, IORING_ENTER_GETEVENTS, NULL);
// Meaning: Submit 128 SQEs and wait for at least 1 completion event
3.3 io_uring_register: Resource Registration
The register syscall enables pre-registration of reusable resources for I/O operations, avoiding per-request overhead:
enum {
IO_REGISTER_BUFFERS, // Fixed buffers (avoid pin/unpin per I/O)
IO_UNREGISTER_BUFFERS,
IO_REGISTER_FILES, // Fixed files (avoid get/put per request)
IO_UNREGISTER_FILES,
IO_REGISTER_EVENTFD, // Event notification fd
IO_UNREGISTER_EVENTFD,
IO_REGISTER_FILES_UPDATE, // Update registered file slots
IO_REGISTER_EVENTFD_ASYNC,// Async event notification
IO_REGISTER_PROBE, // Probe kernel op support
IO_REGISTER_PERSONALITY, // Register identity
IO_REGISTER_RESTRICTIONS, // Register operational restrictions
IO_REGISTER_ENABLE_RINGS, // Enable SQ/CD polling
};
4. liburing: Simplified Programming Interface
While directly calling syscalls provides fine-grained control, Jens Axboe developed liburing wrapper library that abstracts ring buffer management through friendly C APIs. liburing is the de facto standard io_uring programming interface.
4.1 Core API Pattern
#include
// 1. Initialize
struct io_uring ring;
io_uring_queue_init(256, ˚, IORING_SETUP_SQPOLL);
// 2. Get SQE
struct io_uring_sqe *sqe = io_uring_get_sqe(˚);
io_uring_prep_read(sqe, fd, buf, len, offset);
io_uring_sqe_set_data(sqe, ctx);
// 3. Submit
io_uring_submit(˚); // In SQPOLL mode, can be omitted
// 4. Reap CQE
struct io_uring_cqe *cqe;
unsigned head;
io_uring_for_each_cqe(˚, head, cqe) {
// Process completion event
my_context_t *ctx = io_uring_cqe_get_data(cqe);
ctx->result = cqe->res;
ctx->flags = cqe->flags;
}
io_uring_cq_advance(˚, cqe_count);
// 5. Cleanup
io_uring_queue_exit(˚);
4.2 Batch Submission Optimization
io_uring's 2.0+ version supports batch submission - accumulating multiple SQEs in userspace, then calling enter once to hand them over to the kernel, significantly reducing syscall overhead:
// Batch write 100 files
io_uring_batch_submit() {
for (int i = 0; i < 100 xss=removed>
5. Advanced Features: Going Beyond Asynchronous I/O
5.1 Fixed Buffers: Eliminate Memory Pin Overhead
Each I/O operation requires pinning physical memory pages (to prevent swapping) and establishing DMA mappings. With high-frequency small I/O, pin/unpin overhead can account for 30-40% of total latency. io_uring's Fixed Buffers feature pre-registers a set of memory buffers, allowing direct referencing of I/O operations:
// Register 16 fixed buffers (8KB each)
struct iovec iovecs[16];
for (int i = 0; i < 16 xss=removed>flags |= IOSQE_FIXED_FILE;
5.2 Fixed Files: Eliminate File Reference Overhead
Similarly, each I/O operation requires looking up file objects via fd and maintaining f_count reference counts. io_uring's Fixed Files feature maintains a file table at ring initialization, using array indices instead of raw fds:
// Register file array
int fds[] = { fd1, fd2, fd3, fd4 };
io_uring_register_files(˚, fds, 4);
// Use fixed files (fd = index into registered array)
sqe = io_uring_get_sqe(˚);
io_uring_prep_read(sqe, 2, buf, len, 0); // Directly use index 2
sqe->flags |= IOSQE_FIXED_FILE;
5.3 Linked SQE: Chaining Operations
io_uring supports linking multiple SQEs as sequences, where the kernel executes them in order. This is critical for atomic operations like write-then-fsync:
// Chain: write data 鈫?fsync
sqe1 = io_uring_get_sqe(˚);
io_uring_prep_write(sqe1, fd, buf, len, 0);
sqe1->flags |= IOSQE_IO_LINK; // Link to next
sqe2 = io_uring_get_sqe(˚);
io_uring_prep_fsync(sqe2, fd, 0);
sqe2->flags |= IOSQE_IO_LINK; // Link to next
sqe3 = io_uring_get_sqe(˚);
io_uring_prep_nop(sqe3); // Marker for completion
io_uring_submit(˚);
5.4 SQPOLL: Kernel-Side Polling Thread
The IORING_SETUP_SQPOLL flag creates a kernel thread specifically for polling SQ, where userspace only needs to write the SQ tail pointer to be discovered by the kernel. This achieves true zero-syscall submissions:
| Mode | Submit Syscalls | Completion Syscalls | Use Case |
|---|---|---|---|
| Default Mode | 1 per batch | 0 (poll CQ) | Low latency, sensitivity to syscall overhead |
| SQPOLL | 0 | 0 (poll CQ) | Ultra-throughput, tolerant of added latency |
| IOPOLL | 0 | 0 (poll CQ) | Ultra-low latency (NVMe) |
SQPOLL thread configuration notes:
- Set
sq_thread_cputo bind to a dedicated CPU core, avoiding scheduler migration overhead - Set
sq_thread_idleto a small value (e.g., 1ms) to keep low latency when idle - SQPOLL thread requires
CAP_SYS_NICEprivilege, which some container environments restrict
6. Performance Benchmarks and Comparison
6.1 io_uring vs libaio: Throughput Comparison
Using fio for sequential read testing on NVMe SSD (queue depth 128):
| I/O Mode | Read IOPS | Latency p99 | CPU Utilization |
|---|---|---|---|
| Sync I/O (pread) | 85K | 1480 渭s | 100% (1 core) |
| libaio (io_submit) | 280K | 452 渭s | 38% |
| io_uring (Default) | 420K | 298 渭s | 25% |
| io_uring (SQPOLL) | 580K | 211 渭s | 31% (1 core + SQPOLL) |
| io_uring (IOPOLL) | 720K | 168 渭s | 45% |
6.2 io_uring vs epoll: Network I/O Scalability
epoll-based HTTP servers (e.g., nginx) face syscall overhead explosions in high-concurrency scenarios. io_uring can replace epoll+read/write combinations:
- Problem with epoll: epoll wait returns events, but reading data still requires read() syscall. Such 'epoll wait + readv' double syscall patterns constitutes the scalability bottleneck
- io_uring solution: Pre-submit read requests, kernel automatically fills buffers when data is ready and generates CQE. Combined with SQPOLL mode, sockets are truly throughout without application-side syscalls
- nginx io_uring patch: Experimental patch achieves 20-30% throughput improvement with io_uring reduce CPU consumption by 25% at 10K concurrent connections
7. Production Deployment Case Studies
7.1 RocksDB: io_uring as the Core Storage Engine
RocksDB 7.0 added io_uring support, replacing POSIX AIO as the default asynchronous I/O backend:
- Benefits: Supports buffered I/O (no O_DIRECT required), unified handling of reads/writes/fsync
- Performance: Random read throughput increased 35%, write latency p50 decreased 20%
- Configuration: Simple
use_io_uring = trueactivate;uring_max_workerscontrols parallelism
7.2 PostgreSQL: WAL Write Optimization
PostgreSQL's WAL (Write-Ahead Log) is critical to data safety and performance. Experimental io_uring patches allow walwriter processes to:
- Batch submit multiple WAL page writes
- Link fsync with dependent writes for atomic page ordering
- Reduce WAL write latency p99 from 1.2ms to 400渭s (NVMe environment)
- However, PostgreSQL core team is still evaluating io_uring's complexity and aging kernel compatibility, and it has not been merged into mainline yet
7.3 SPDK: Storage Performance Development Kit
Intel's SPDK (Storage Performance Development Kit) achieves kernel bypass through user-mode drivers and polling. With the maturity of io_uring IOPOLL mode:
- io_uring IOPOLL achieves similar performance to SPDK, but retains kernel management capabilities
- Advantages: No DPDK dependencies, no hugepage configuration, NVMe hardware-compatible
- Use case: Applications requiring high-performance storage but unwilling to bear DPDK complexity, io_uring IOPOLL is the ideal compromise
8. io_uring Operational Model: io_uring_setup()
The io_uring_setup() syscall creates an io_uring instance, allocating SQ and CQ ring buffers of specified size. During initialization, the kernel returns three memory regions that userspace needs to mmap:
8.1 mmap Regions
| Region | Size | mmap Offset | Function |
|---|---|---|---|
| SQ Ring | sizeof(struct io_uring_sqe) * sq_entries | IORING_OFF_SQES | Array of pre-allocated submission queue entries |
| SQ Ring Metadata | sizeof(unsigned) * sq_entries + sizeof(struct io_uring_sqe) + 64 | IORING_OFF_SQ_RING | Submission queue metadata (head/tail pointer, ring mask, etc.) |
| CQ Ring | sizeof(struct io_uring_cqe) * cq_entries + 64 | IORING_OFF_CQ_RING | Completion queue and its metadata |
8.2 Ring Buffer Index Calculation
Ring buffer uses head/tail pointers with ring mask for O(1) indexing:
// Userspace submit path
unsigned tail = *sq_ring->tail;
unsigned index = tail & sq_ring->ring_mask;
struct io_uring_sqe *sqe = &sqes[index];
// Fill SQE
sqe->opcode = IORING_OP_READV;
sqe->fd = fd;
sqe->off = offset;
sqe->addr = (unsigned long)&iovec;
sqe->len = 1;
sqe->buf_index = 0;
sqe->user_data = (u64)ctx;
tail++;
smp_store_release(sq_ring->tail, tail); // Memory barrier: ensure SQE is visible before tail update
9. io_uring Operational Model: io_uring_enter()
The io_uring_enter syscall serves dual purposes: submitting SQEs to the kernel and waiting for completion events. Its parameters control submission behavior:
int sys_io_uring_enter(int fd, unsigned to_submit,
unsigned min_complete, unsigned flags);
// Parameter explanation
// to_submit: Number of SQEs to submit this time (0 = do not submit but continue waiting)
// min_complete: Minimum completion events to wake waiting thread (0 = non-blocking check)
// flags: IORING_ENTER_GETEVENTS (wait for events) | IORING_ENTER_SQ_WAKEUP (wake SQPOLL thread)
9.1 Submission Offload with SQPOLL
When using IORING_SETUP_SQPOLL, the kernel automatically creates a thread that polls SQ for new SFs. Even so, userspace still occasionally needs to call io_uring_enter to wake up SQPOLL threads that may enter sleep:
// Typical workflow
1. User writes tail pointer of SQ ring (notification via shared memory)
2. SQPOLL kernel thread discovers new SFs
3. Thread polls SQ for new requests, executes I/O
4. User polls CQ ring head pointer for completion events
5. If SQPOLL thread is idle for more than sq_thread_idle, it calls io_uring_enter with IORING_ENTER_SQ_WAKEUP to wake up
10. Limitations and Future Outlook
10.1 Current Limitations
| Limitation | Description | Mitigation |
|---|---|---|
| Kernel Version Dependency | Production features require Linux 5.10+; many features require 5.13+ | Use liburing's fallback if kernel feature is not supported |
| SQPOLL Privilege Requirements | SQPOLL requires CAP_SYS_NICE, which some container environments restrict | Use default polling mode in containers |
| No Userspace Direct Access to Block Devices | Open block devices requires O_DIRECT support | Raw block devices can only be associated with io_uring via O_DIRECT |
| Synchronous Emulation Overhead | Operations not supported must fall back to blocking mode | liburing automatically handles fallback path |
| Synchronous I/O workload | io_uring's advantages rely on batched I/O; benefits diminish with single-sync I/O | Carefully evaluate workload types |
10.2 Future Evolution
io_uring continues to evolve rapidly, with major improvements on the kernel schedule:
- Async I/O Interface Wrapping - Linux 6.6+ supports
IORING_OP_READV/IORING_OP_WRITEVintegrated read/write operations, further reducing SQE fill overhead - Enhanced Credential Inheritance - Different credentials for different SQEs in the same ring, supporting multi-tenant services
- Enhanced Filesystem Support - Support for
FUSE+ io_uring passthrough at the kernel level for user-mode filesystem acceleration - Network Bypass -
IORING_OP_SENDMSG/IORING_OP_RECVMSGzero-copy extensions, approaching but not replacing DPDK user-mode protocol stack performance - Resource Quotas - Set maximum I/O resource consumption limits at the ring level, preventing single tenant from monopolizing I/O bandwidth
11. Summary
Linux io_uring represents one of the most important fundamental innovations in the I/O stack in over a decade. By fully redesigning the interface between kernel and userspace, it achieves:
- Zero-Syscall Fast Path - SQPOLL mode eliminates submit syscalls, unpolished CQ access is zero-cost
- Unified Asynchronous I/O - Covers all file types (regular files, sockets, pipes), breaking AIO's O_DIRECT restrictions
- Resource Pre-registration - Fixed buffers/files eliminate per-request overhead
- Rich Extension Capabilities - Linked SQEs, timeout control, priority setting, openat/unlink/mkdir generalized operations
- Production-Proven - Databases (RocksDB), storage (SPDK), network (nginx) have all validated the performance gains
For developers, the learning curve is gentle: if familiar with traditional asynchronous I/O, most logic translates directly to io_uring; the liburing library abstracts away complex ring buffer management. For sysops, ensure kernel 5.10+, change the container runtime to allow CAP_SYS_ADMIN, and monitor /proc/

发表评论 取消回复