XDP Deep Dive: eXpress Data Path for High-Performance Packet Processing at the Driver Layer
Modern cloud infrastructure demands packet processing at rates of 100Gbps and beyond — a scale where traditional kernel networking stacks struggle to keep up. XDP (eXpress Data Path), built on top of eBPF technology, provides a programmable, high-performance packet processing framework that operates at the lowest point in the software stack: the network interface card (NIC) driver layer. This article explores the XDP architecture from the ground up, examining its execution model, action types, map integrations, and real-world production deployments that achieve millions of packets per second per core.
1. The Need for a Faster Data Path
The Linux kernel networking stack, while robust and feature-rich, introduces significant latency and overhead for each packet:
- Memory allocation: Every packet requires a
sk_buffallocation (including the ~230-byte metadata structure) - Context switches: Hardware interrupt → softirq → network stack → protocol layers → userspace
- Layer traversal: Each layer (Ethernet → IP → TCP/UDP → socket buffer) adds processing overhead
- Copy operations: Data must traverse kernel-userspace boundaries for applications
At 10Gbps line rate, a system processing minimum-size (64-byte) packets must handle ~14.88 million packets per second (Mpps). Each sk_buff allocation alone can consume hundreds of CPU cycles, making it nearly impossible to reach line rate with standard stack processing on commodity hardware.
2. XDP Architecture: Processing Before the Stack
XDP inserts an eBPF hook at the earliest possible point in the RX path — directly inside the NIC driver's receive function, before the kernel allocates an sk_buff. This enables zero-copy, zero-allocation packet processing capable of handling 25+ Mpps per core.
// --- XDP execution point in the RX path ---
// NIC Driver → [XDP Hook (eBPF)] → sk_buff allocation → netif_receive_skb() → Protocol Stack
//
// XDP sees raw packet data in a BPF context:
struct xdp_md {
__u32 data; // Start of packet data
__u32 data_end; // End of packet data
__u32 data_meta; // Metadata pointer
__u32 rx_queue_index; // Incoming RX queue index
__u32 ingress_ifindex; // Interface index
};
2.1 Three Attachment Modes
| Mode | Description | Performance |
|---|---|---|
| Native XDP | Driver-level hook (ndo_xdp_xmit). Requires driver support (i40, mlx5, ixgbe, etc.) | Highest — runs before sk_buff allocation |
| Generic XDP | Fallback in core network stack. Allocates sk_buff first, then runs XDP. | Moderate — useful for unsupported NICs |
| Offloaded XDP | Program loaded directly into NIC firmware (Netronome/Corstone SmartNIC). | Absolute maximum — BPF runs on NIC hardware |
3. XDP Actions: Directing Packet Fate
Every XDP program must return one of five action codes that determines the kernel's handling of the packet:
3.1 XDP_ABORTED (0)
Indicates a program error. The packet is dropped, and a trace event is recorded for debugging. This should never happen in production code — it indicates verifier-detected issues or runtime exceptions.
3.2 XDP_PASS
Pass the packet to the normal kernel networking stack for standard processing. Use this when XDP can't handle the packet (e.g., non-IPv4 traffic, packets that need full stack inspection) or as a learning mode action.
3.3 XDP_DROP
Drop the packet immediately in the driver. This is the fastest possible "dispose" action — ideal for DDoS mitigation where unwanted packets must be discarded as early as possible.
3.4 XDP_TX
Transmit the (possibly modified) packet back out through the same NIC it arrived on. Useful for:
- Layer 2/L3 rewriting (MAC address translation)
- Simple load balancers that rewrite destination and bounce back
- Reflection attacks (in testing scenarios)
3.5 XDP_REDIRECT
The most powerful action — redirect the packet to:
- Another network interface (cross-NIC forwarding)
- A different CPU core for further processing (via
XSKMAPto AF_XDP sockets) - A remote CPU via
cpumapfor userspace consumption
4. BPF Maps for XDP: State and Communication
XDP programs leverage BPF maps to maintain state and communicate with userspace. The most critical map types for XDP are:
4.1 DEVMAP (BPF_MAP_TYPE_DEVMAP)
ALongside bpf_redirect_map(), DEVMAP enables XDP programs to forward packets to other NIC interfaces. The map keys are interface indices, and values are bpf_devmap_val structures containing the target interface and CPU information.
// --- DEVMAP definition and redirect usage ---
struct {
__uint(type, BPF_MAP_TYPE_DEVMAP);
__uint(max_entries, 64);
__type(key, __u32);
__type(value, struct bpf_devmap_val);
} devmap SEC(".maps");
// In the XDP program:
SEC("xdp")
int xdp_l2fwd(struct xdp_md *ctx) {
__u32 key = DST_IFINDEX;
return bpf_redirect_map(&devmap, key, XDP_PASS);
}
4.2 CPUMAP (BPF_MAP_TYPE_CPUMAP)
For cases where packets need to be processed by specific CPU cores (e.g., for RSS queue steering), CPUMAP redirects packets to per-CPU socket queues. This is the foundation for XDP-based userspace packet processing pipelines.
4.3 XSKMAP (BPF_MAP_TYPE_XSKMAP)
AF_XDP sockets combined with XSKMAP enable zero-copy packet delivery from XDP directly to userspace applications, completely bypassing the kernel networking stack. This is how high-performance DPDK-like processing is achieved while retaining the safety and programmability of the kernel.
4.4 Array and Hash Maps
Standard BPF array and hash maps are ubiquitous for:
- Configuration data (allowlists, blocklists, rate limits)
- Statistics (packet counters per flow)
- Routing tables (longest-prefix-match lookups)
5. Building a Production DDoS Mitigation System
XDP's primary production success story is in high-performance DDoS mitigation. A typical architecture:
// --- Simplified DDoS XDP filter ---
#define MAX_RULES 4096
struct {
__uint(type, BPF_MAP_TYPE_HASH);
__uint(max_entries, MAX_RULES);
__type(key, __u32); // Source IP
__type(value, __u64); // Bytes per second counter
} blocklist SEC(".maps");
struct {
__uint(type, BPF_MAP_TYPE_ARRAY);
__uint(max_entries, 1);
__type(key, __u32);
__type(value, __u64); // Current second timestamp
} current_sec SEC(".maps");
SEC("xdp")
int xdp_ddos_filter(struct xdp_md *ctx) {
void *data_end = (void *)(long)ctx->data_end;
void *data = (void *)(long)ctx->data;
struct ethhdr *eth = data;
struct iphdr *ip;
// Bounds check (required by verifier)
if ((void *)(eth + 1) > data_end)
return XDP_PASS;
if (eth->h_proto != bpf_htons(ETH_P_IP))
return XDP_PASS; // Let stack handle non-IPv4
ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end)
return XDP_PASS;
// Check blocklist
__u32 src_ip = bpf_ntohl(ip->saddr);
__u64 *blocked = bpf_map_lookup_elem(&blocklist, &src_ip);
if (blocked && *blocked > THRESHOLD)
return XDP_DROP; // Drop in driver layer
// Rate limiting logic...
return XDP_PASS;
}
Production results: Cloudflare reported processing 100Gbps of DDoS attack traffic using XDP on a single server, dropping 6M+ malicious packets per second per core with just 10% CPU utilization.
6. AF_XDP: Zero-Copy Userspace Access
AF_XDP is the mechanism that allows userspace applications to receive packets directly from XDP with zero memory copies. The UMEM (User Memory Manager) model works as follows:
- Allocation: Userspace allocates a contiguous memory region (UMEM), divided into equal-sized frames.
- Ring buffers: Four rings connect kernel and userspace:
- Fill ring: Userspace gives empty frames to kernel
- Completion ring: Kernel returns used frames after TX
- RX ring: Kernel delivers received packets
- TX ring: Userspace submits frames for transmission
- Zero-copy: Packet data is written directly into UMEM by the NIC (via DMA), with both kernel and userspace accessing the same memory.
This architecture achieves remarkable performance: Facebook's Katran L4 load balancer processes 100M+ packets per second using XDP + AF_XDP.
7. XDP vs DPDK vs AF_PACKET
| Aspect | XDP + eBPF | DPDK | AF_PACKET |
|---|---|---|---|
| Flexibility | High — reprogrammable at runtime, verifier-safe | Moderate — requires application restart for changes | Low — raw socket capture |
| Performance | 10-25 Mpps/core | 25-80 Mpps/core | 1-3 Mpps/core |
| Driver Support | Requires DPDK-compatible NIC | Works with any NIC | |
| Safety | Kernel verifies program safety — no crashes | Unsafe — bug = kernel panic or data corruption | Safe but slow |
| Memory Model | Shared kernel/userspace (AF_XDP) | Hugepages + userspace drivers | sk_buff copy to userspace |
| Deployment Complexity | Low — works with existing kernel | High — requires DPDK setup, hugepages, NUMA tuning | Minimal |
| Update in Production | Hot-reload eBPF program without downtime | Requires service restart + reinitialization | N/A |
8. Advanced Techniques
8.1 Packet Header Rewrite
XDP can modify packet headers in-place using safe helper functions. The verifier ensures only safe modifications are allowed:
// --- NAT-like rewrite using bpf_xdp_adjust_head ---
SEC("xdp")
int xdp_nat(struct xdp_md *ctx) {
void *data_end = (void *)(long)ctx->data_end;
void *data = (void *)(long)ctx->data;
struct ethhdr *eth = data;
struct iphdr *ip = (void *)(eth + 1);
// Shrink packet headroom for encapsulation
if (bpf_xdp_adjust_head(ctx, -(int)sizeof(struct gre_hdr)) < 0)
return XDP_DROP;
// Rewrite GRE header, then update IPs
// ... (bounds-checked operations)
return XDP_TX; // Send modified packet back out
}
8.2 Tail Calls (bpf_tail_call)
Large XDP programs can be decomposed into smaller chained programs using tail calls. Each tail call replaces the current program's context atomically, with no function call overhead:
// --- Tail call chain ---
struct {
__uint(type, BPF_MAP_TYPE_PROG_ARRAY);
__uint(max_entries, 4);
__type(key, __u32);
__type(value, __u32);
} xdp_progs SEC(".maps");
SEC("xdp")
int xdp_router_main(struct xdp_md *ctx) {
// Parse L2, dispatch by next header
if (ip->protocol == IPPROTO_TCP)
bpf_tail_call(ctx, &xdp_progs, PROG_TCP);
else if (ip->protocol == IPPROTO_UDP)
bpf_tail_call(ctx, &xdp_progs, PROG_UDP);
return XDP_PASS;
}
8.3 Batch Processing (libbpf XDP APIs)
The latest XDP APIs support batch processing, reducing per-packet overhead by processing multiple packets in a single system call. This is critical for reaching maximum throughput on modern NICs.
9. Practical Deployment Checklist
Before deploying XDP in production, ensure:
- Driver support: Verify NIC driver supports native XDP (ethtool -k | grep xdp, or check
/sys/class/net/<iface>/xdp_*) - Ring buffer sizing: Set net.core.netdev_budget=60000 (default 300 may limit XDP performance)
- NUMA awareness: XDP program runs on the CPU handling the NIC's RX queue. Pin IRQs to local NUMA nodes.
- Hugepages (AF_XDP): AF_XDP requires hugepages for UMEM allocation. Configure
/proc/sys/vm/nr_hugepages. - Verifier compliance: All XDP programs must pass the BPF verifier. Strict bounds checks on every pointer dereference are mandatory.
- Fallback strategy: Always have a generic XDP or iptables fallback for production safety.
10. Production Case Studies
Cloudflare ddos mitigation: StackArmor system uses XDP to drop 6M+ packets/sec per core during volumetric attacks. The eBPF program maintains per-source-IP counters in hash maps, with automatic threshold detection and blocking.
Facebook Katran: Layer 4 load balancer handles 100M+ pps using XDP + AF_XDP. Uses consistent hashing in eCPF maps to distribute connections to backend servers without connection state.
LinkedIn Traffic Shaping: XDP-based traffic policers use token bucket algorithms implemented in BPF maps for per-flow rate limiting at 100Gbps.
11. Conclusion
XDP represents a paradigm shift in how we approach high-performance networking in the Linux kernel. By extending the kernel's programmable capabilities to the NIC driver layer, it achieves DPDK-class performance with kernel-class safety, flexibility, and operational simplicity. Whether you're building DDoS mitigation systems, layer 4 load balancers, network observability platforms, or programmable firewalls, XDP provides a robust, production-proven foundation that continues to evolve with each kernel release.
The synergy of eBPF programs, BPF maps, AF_XDP zero-copy delivery, and kernel-bypass techniques ensures that XDP will remain at the center of cloud-native networking innovation for years to come.

发表评论 取消回复