Table of Contents
- Introduction to eBPF Networking
- Core eBPF Concepts for Networking
- XDP: Express Data Path
- TC: Traffic Control Subsystem
- Socket-level Tracing
- TCP Protocol Stack Tracing
- DNS Query Monitoring
- Network Performance Analysis
- Production Deployment Guide
- Conclusion
Introduction to eBPF Networking
Extended Berkeley Packet Filter (eBPF) has revolutionized Linux networking by providing a safe, high-performance framework for running sandboxed programs in the kernel. Unlike kernel modules, eBPF programs pass through a verifier that ensures safety, making them ideal for production environments where stability is paramount.
eBPF networking programs operate at multiple layers: from raw packet processing at the driver level (XDP), through traffic control (TC), to socket and protocol stack tracing. This article provides a comprehensive deep-dive into eBPF networking internals, combining theory with practical production implementations.
Core eBPF Concepts for Networking
eBPF Program Types for Networking
The Linux kernel provides several specialized eBPF program types for networking:
- XDP (eXpress Data Path): Executes at the lowest point in the kernel networking stack, directly after packet reception from the NIC driver. Ideal for DDoS mitigation, load balancing, and packet filtering with minimal overhead.
- TC (Traffic Control): Attaches to the kernel's traffic control subsystem, providing access to sk_buff structures and enabling complex packet manipulation.
- kprobes/kretprobes: Dynamic tracing hooks into kernel functions, allowing observation of internal networking functions without modifying kernel source.
- tracepoints: Statically defined hooks in the kernel that provide stable ABI across versions, ideal for long-term monitoring.
- socket filters: Attach directly to sockets for application-level packet capture and filtering.
BPF Maps: Data Exchange with Userspace
BPF maps are the primary mechanism for sharing data between eBPF programs and userspace, or between different eBPF programs. Key map types for networking include:
- BPF_MAP_TYPE_HASH: Key-value storage for connection tracking, statistics counters, and configuration lookup.
- BPF_MAP_TYPE_PERCPU_HASH: Per-CPU variant that eliminates lock contention for high-frequency counters.
- BPF_MAP_TYPE_RINGBUF: High-throughput, lockless data streaming to userspace, ideal for packet capture and event logging.
- BPF_MAP_TYPE_LPM_TRIE: Longest Prefix Match trie for IP routing tables and CIDR-based filtering.
- BPF_MAP_TYPE_QUEUE/STACK: FIFO/LIFO data structures for packet queuing and ordered event delivery.
XDP: Express Data Path
Architecture and Execution Flow
XDP provides the fastest possible packet processing by hooking directly into the network driver's receive path. When the NIC receives a packet, the driver calls the XDP program before allocating an sk_buff, resulting in significantly lower overhead than traditional approaches.
XDP Actions
XDP programs return one of several action codes:
- XDP_PASS: Continue normal kernel networking stack processing.
- XDP_DROP: Silently discard the packet (ideal for DDoS mitigation).
- XDP_TX: Transmit the packet back through the same NIC.
- XDP_REDIRECT: Forward to another NIC or CPU via XDP.
Practical XDP Program Example
The following XDP program demonstrates IP-based rate limiting:
// xdp_rate_limit.bpf.c
#include <linux/bpf.h>
#include <linux/if_ether.h>
#include <linux/ip.h>
#include "bpf_helpers.h"
struct bpf_map_def SEC("maps") ip_count_map = {
.type = BPF_MAP_TYPE_PERCPU_HASH,
.key_size = sizeof(__u32),
.value_size = sizeof(struct counter),
.max_entries = 10000,
};
SEC("xdp")
int xdp_rate_limit_handler(struct xdp_md *ctx) {
void *data_end = (void *)(long)ctx->data_end;
void *data = (void *)(long)ctx->data;
struct ethhdr *eth = data;
if ((void *)(eth + 1) > data_end)
return XDP_PASS;
if (eth->h_proto != __constant_htons(ETH_P_IP))
return XDP_PASS;
struct iphdr *ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end)
return XDP_PASS;
__u32 src_ip = ip->saddr;
__u64 *counter = bpf_map_lookup_elem(&ip_count_map, &src_ip);
if (counter) {
if (*counter > RATE_LIMIT_THRESHOLD)
return XDP_DROP;
__sync_fetch_and_add(counter, 1);
}
return XDP_PASS;
}
TC: Traffic Control Subsystem
TC CLSACT Qdisc and eBPF Attachment
The TC subsystem provides more sophisticated packet handling than XDP, operating after sk_buff allocation and enabling packet modification, classification, and queuing discipline attachment. The clsact qdisc is specifically designed for eBPF classifier attachment with minimal overhead:
// Add clsact qdisc
tc qdisc add dev eth0 clsact
// Attach TC BPF program
tc filter add dev eth0 ingress bpf da obj tc_monitor.o sec ingress
sk_buff Structure Access
Unlike XDP's xdp_md context, TC programs receive access to the full sk_buff structure, providing visibility into protocol headers, metadata, packet timestamp, and socket information. This enables deeper packet inspection and modification:
struct sk_buff {
unsigned int len; // Total packet length
__u16 protocol; // Ethernet protocol type
__u32 priority; // QoS priority
void *data; // Current header pointer
struct sock *sk; // Owning socket (if any)
ktime_t tstamp; // Packet arrival timestamp
// ... additional fields
};
Socket-level Tracing
Tracing Socket Operations with kprobes
kprobes allow dynamic instrumentation of kernel socket functions. Common socket functions for tracing include:
- tcp_connect(): Track outbound TCP connection initiation.
- tcp_close(): Monitor connection termination and closure reasons.
- tcp_sendmsg(): Observe outbound data flow at the socket level.
- tcp_recvmsg(): Capture inbound data delivery to userspace applications.
- inet_accept(): Track incoming connection acceptance.
Socket Tracing Implementation
// socket_trace.bpf.c
#include "vmlinux.h"
#include "bpf_helpers.h"
#include <bpf/bpf_tracing.h>
struct socket_event {
__u64 timestamp;
__u32 pid;
__u32 uid;
__u32 saddr;
__u32 daddr;
__u16 sport;
__u16 dport;
__u8 protocol;
__u8 direction;
char comm[16];
};
struct bpf_map_def SEC("maps") events = {
.type = BPF_MAP_TYPE_RINGBUF,
.max_entries = 256 * 1024,
};
SEC("kprobe/tcp_connect")
int trace_tcp_connect(struct pt_regs *ctx, struct sock *sk) {
struct socket_event *event;
event = bpf_ringbuf_reserve(&events, sizeof(*event), 0);
if (!event)
return 0;
event->timestamp = bpf_ktime_get_ns();
event->pid = bpf_get_current_pid_tgid() >> 32;
event->uid = bpf_get_current_uid_gid();
bpf_get_current_comm(&event->comm, sizeof(event->comm));
event->saddr = BPF_CORE_READ(sk, __sk_common.skc_rcv_saddr);
event->daddr = BPF_CORE_READ(sk, __sk_common.skc_daddr);
event->dport = bpf_ntohs(BPF_CORE_READ(sk, __sk_common.skc_dport));
event->protocol = IPPROTO_TCP;
event->direction = 0; // outbound
bpf_ringbuf_submit(event, 0);
return 0;
}
TCP Protocol Stack Tracing
TCP State Machine Tracing
TCP connection states provide critical visibility into network health and application behavior. Tracing tcp_set_state() allows monitoring of all TCP state transitions:
SEC("tracepoint/sock/inet_sock_set_state")
int trace_tcp_state_change(struct trace_event_raw_inet_sock_set_state *ctx) {
struct tcp_state_event *event;
__u16 family = ctx->family;
if (family != AF_INET && family != AF_INET6)
return 0;
event = bpf_ringbuf_reserve(&tcp_events, sizeof(*event), 0);
if (!event)
return 0;
event->oldstate = ctx->oldstate;
event->newstate = ctx->newstate;
event->sport = ctx->sport;
event->dport = ctx->dport;
if (family == AF_INET) {
bpf_probe_read(event->saddr, sizeof(struct in_addr), ctx->saddr);
bpf_probe_read(event->daddr, sizeof(struct in_addr), ctx->daddr);
}
bpf_ringbuf_submit(event, 0);
return 0;
}
TCP Retransmission Analysis
TCP retransmissions are a primary indicator of network congestion, packet loss, or routing issues. Tracing retransmission events helps identify network problems at their source:
// Track TCP retransmission statistics
struct bpf_map_def SEC("maps") retransmit_stats = {
.type = BPF_MAP_TYPE_PERCPU_HASH,
.key_size = sizeof(struct flow_key),
.value_size = sizeof(struct flow_stats),
.max_entries = 100000,
};
SEC("kprobe/tcp_retransmit_skb")
int trace_tcp_retransmit(struct pt_regs *ctx) {
struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
struct flow_stats *stats, zero = {};
struct flow_key key = {};
key.saddr = BPF_CORE_READ(sk, __sk_common.skc_rcv_saddr);
key.daddr = BPF_CORE_READ(sk, __sk_common.skc_daddr);
key.sport = BPF_CORE_READ(sk, __sk_common.skc_num);
key.dport = BPF_CORE_READ(sk, __sk_common.skc_dport);
stats = bpf_map_lookup_or_try_init(&retransmit_stats, &key, &zero);
if (stats)
__sync_fetch_and_add(&stats->retransmit_count, 1);
return 0;
}
DNS Query Monitoring
Monitoring DNS at the Socket Level
DNS queries can be traced by monitoring UDP packets on port 53. This enables DNS analytics, threat detection, and troubleshooting:
// DNS monitoring via trace_udp_sendmsg
SEC("kprobe/udp_sendmsg")
int trace_dns_query(struct pt_regs *ctx) {
struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
struct msghdr *msg = (struct msghdr *)PT_REGS_PARM2(ctx);
__u16 dport = BPF_CORE_READ(sk, __sk_common.skc_dport);
if (dport != __constant_htons(53))
return 0;
struct iov_iter *iter = &msg->msg_iter;
__u32 data_len = iter->count;
if (data_len > 0 && data_len <= 512) {
__u8 buf[512];
bpf_probe_read(buf, data_len, iter->iov->iov_base);
// Parse DNS header to extract domain name
parse_dns_query(buf, data_len);
}
return 0;
}
CBPF-based DNS Capture with Socket Filters
For per-socket DNS monitoring, BPF socket filters provide a lightweight alternative to kprobes, with direct access to payload data and minimal overhead.
Network Performance Analysis
Latency Distribution Analysis
eBPF enables precise network latency measurement at multiple layers. By capturing timestamps at strategic points, we can decompose network latency into meaningful components:
struct bpf_map_def SEC("maps") latency_hist = {
.type = BPF_MAP_TYPE_HISTOGRAM,
.key_size = sizeof(u32),
};
// Measure TCP RTT at connection level
SEC("kprobe/tcp_rtt_estimator")
int trace_rtt(struct pt_regs *ctx) {
struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
struct tcp_sock *tp = (struct tcp_sock *)sk;
__u64 srtt = BPF_CORE_READ(tp, srtt_us) >> 3;
u32 slot = bpf_log2l(srtt);
hist_inc(latency_hist, slot);
return 0;
}
Throughput Measurement
Accurate network throughput measurement requires accounting for packet payload sizes and protocol headers. The following approach provides both L3 throughput and L7 application-level statistics:
// Track per-connection throughput
struct bpf_map_def SEC("maps") throughput_map = {
.type = BPF_MAP_TYPE_PERCPU_HASH,
.key_size = sizeof(struct connection_key),
.value_size = sizeof(struct throughput_stats),
.max_entries = 50000,
};
SEC("kprobe/tcp_sendmsg")
int trace_send(struct pt_regs *ctx, struct sock *sk, struct msghdr *msg, size_t size) {
struct throughput_stats *stats = lookup_throughput(sk);
if (stats) {
__sync_fetch_and_add(&stats->bytes_sent, size);
__sync_fetch_and_add(&stats->packets_sent, 1);
}
return 0;
}
Production Deployment Guide
Kernel Version Requirements
For production eBPF networking, use Linux 5.10+ with the following configuration:
CONFIG_BPF=y
CONFIG_BPF_SYSCALL=y
CONFIG_BPF_JIT=y
CONFIG_BPF_EVENTS=y
CONFIG_KPROBE_EVENTS=y
CONFIG_TRACEPOINTS=y
CONFIG_NET_CLS_BPF=m
CONFIG_NET_ACT_BPF=m
CONFIG_XDP_SOCKETS=y
CONFIG_BPF_STREAM_PARSER=y
Containerized Environments
When running in containers, ensure CAP_BPF and CAP_NET_ADMIN are granted. For Kubernetes, Pod security policies must allow privileged capabilities or use a CNI plugin with eBPF support (Cilium, Calico).
Performance Optimization
Key optimization strategies for production eBPF networking:
- Use BPF_MAP_TYPE_PERCPU_* variants for high-frequency counters to eliminate lock contention.
- Leverage BPF_MAP_TYPE_RINGBUF over perf_event_array for 2-4x throughput improvement.
- Implement batch processing to reduce userspace transition overhead.
- Use BTF (BPF Type Format) for CO-RE (Compile Once, Run Everywhere) compatibility across kernel versions.
- Enable BPF JIT compilation for 10-20x performance improvement over interpreter.
Verifying Program Safety
The eBPF verifier ensures program safety by analyzing control flow, memory access, and bounded execution. Common verifier errors and resolutions:
- Uninitialized register access: Initialize all variables before use.
- Out-of-bounds stack access: Ensure bounds checks before array access.
- Infinite loop detection: Add #pragma unroll for bounded loops.
- Function call limitations: Convert recursive functions to iterative form.
Conclusion
eBPF is transforming network observability and performance optimization in Linux. With its safe execution environment and minimal overhead, eBPF enables real-time network monitoring, security enforcement, and performance tracing that was previously impossible without kernel modifications or expensive hardware solutions.
As the Linux kernel continues to evolve, eBPF networking capabilities expand with support for socket lookup, TLS encryption observability, and high-speed packet processing at 100Gbps+. Organizations deploying eBPF-based networking solutions gain unprecedented visibility and control over their network infrastructure.
The production-ready implementations discussed in this article provide a foundation for building robust, high-performance network monitoring and optimization systems that scale from single-node deployments to large-scale distributed architectures.
References
- Kernel documentation:
Documentation/networking/filter.rst - eBPF libraries: libbpf, cilium/ebpf, libpam
- Tools: bpftrace, bpftool, tcpdump (libpcap)
- Projects: Cilium, Katran, Falco

发表评论 取消回复