Table of Contents

Introduction to eBPF Networking

Extended Berkeley Packet Filter (eBPF) has revolutionized Linux networking by providing a safe, high-performance framework for running sandboxed programs in the kernel. Unlike kernel modules, eBPF programs pass through a verifier that ensures safety, making them ideal for production environments where stability is paramount.

eBPF networking programs operate at multiple layers: from raw packet processing at the driver level (XDP), through traffic control (TC), to socket and protocol stack tracing. This article provides a comprehensive deep-dive into eBPF networking internals, combining theory with practical production implementations.

Core eBPF Concepts for Networking

eBPF Program Types for Networking

The Linux kernel provides several specialized eBPF program types for networking:

  • XDP (eXpress Data Path): Executes at the lowest point in the kernel networking stack, directly after packet reception from the NIC driver. Ideal for DDoS mitigation, load balancing, and packet filtering with minimal overhead.
  • TC (Traffic Control): Attaches to the kernel's traffic control subsystem, providing access to sk_buff structures and enabling complex packet manipulation.
  • kprobes/kretprobes: Dynamic tracing hooks into kernel functions, allowing observation of internal networking functions without modifying kernel source.
  • tracepoints: Statically defined hooks in the kernel that provide stable ABI across versions, ideal for long-term monitoring.
  • socket filters: Attach directly to sockets for application-level packet capture and filtering.

BPF Maps: Data Exchange with Userspace

BPF maps are the primary mechanism for sharing data between eBPF programs and userspace, or between different eBPF programs. Key map types for networking include:

  • BPF_MAP_TYPE_HASH: Key-value storage for connection tracking, statistics counters, and configuration lookup.
  • BPF_MAP_TYPE_PERCPU_HASH: Per-CPU variant that eliminates lock contention for high-frequency counters.
  • BPF_MAP_TYPE_RINGBUF: High-throughput, lockless data streaming to userspace, ideal for packet capture and event logging.
  • BPF_MAP_TYPE_LPM_TRIE: Longest Prefix Match trie for IP routing tables and CIDR-based filtering.
  • BPF_MAP_TYPE_QUEUE/STACK: FIFO/LIFO data structures for packet queuing and ordered event delivery.

XDP: Express Data Path

Architecture and Execution Flow

XDP provides the fastest possible packet processing by hooking directly into the network driver's receive path. When the NIC receives a packet, the driver calls the XDP program before allocating an sk_buff, resulting in significantly lower overhead than traditional approaches.

XDP Actions

XDP programs return one of several action codes:

  • XDP_PASS: Continue normal kernel networking stack processing.
  • XDP_DROP: Silently discard the packet (ideal for DDoS mitigation).
  • XDP_TX: Transmit the packet back through the same NIC.
  • XDP_REDIRECT: Forward to another NIC or CPU via XDP.

Practical XDP Program Example

The following XDP program demonstrates IP-based rate limiting:

// xdp_rate_limit.bpf.c
#include <linux/bpf.h>
#include <linux/if_ether.h>
#include <linux/ip.h>
#include "bpf_helpers.h"

struct bpf_map_def SEC("maps") ip_count_map = {
    .type = BPF_MAP_TYPE_PERCPU_HASH,
    .key_size = sizeof(__u32),
    .value_size = sizeof(struct counter),
    .max_entries = 10000,
};

SEC("xdp")
int xdp_rate_limit_handler(struct xdp_md *ctx) {
    void *data_end = (void *)(long)ctx->data_end;
    void *data = (void *)(long)ctx->data;
    
    struct ethhdr *eth = data;
    if ((void *)(eth + 1) > data_end)
        return XDP_PASS;
    
    if (eth->h_proto != __constant_htons(ETH_P_IP))
        return XDP_PASS;
    
    struct iphdr *ip = (void *)(eth + 1);
    if ((void *)(ip + 1) > data_end)
        return XDP_PASS;
    
    __u32 src_ip = ip->saddr;
    __u64 *counter = bpf_map_lookup_elem(&ip_count_map, &src_ip);
    
    if (counter) {
        if (*counter > RATE_LIMIT_THRESHOLD)
            return XDP_DROP;
        __sync_fetch_and_add(counter, 1);
    }
    
    return XDP_PASS;
}

TC: Traffic Control Subsystem

TC CLSACT Qdisc and eBPF Attachment

The TC subsystem provides more sophisticated packet handling than XDP, operating after sk_buff allocation and enabling packet modification, classification, and queuing discipline attachment. The clsact qdisc is specifically designed for eBPF classifier attachment with minimal overhead:

// Add clsact qdisc
tc qdisc add dev eth0 clsact

// Attach TC BPF program
tc filter add dev eth0 ingress bpf da obj tc_monitor.o sec ingress

sk_buff Structure Access

Unlike XDP's xdp_md context, TC programs receive access to the full sk_buff structure, providing visibility into protocol headers, metadata, packet timestamp, and socket information. This enables deeper packet inspection and modification:

struct sk_buff {
    unsigned int len;           // Total packet length
    __u16 protocol;             // Ethernet protocol type
    __u32 priority;             // QoS priority
    void *data;                 // Current header pointer
    struct sock *sk;            // Owning socket (if any)
    ktime_t tstamp;             // Packet arrival timestamp
    // ... additional fields
};

Socket-level Tracing

Tracing Socket Operations with kprobes

kprobes allow dynamic instrumentation of kernel socket functions. Common socket functions for tracing include:

  • tcp_connect(): Track outbound TCP connection initiation.
  • tcp_close(): Monitor connection termination and closure reasons.
  • tcp_sendmsg(): Observe outbound data flow at the socket level.
  • tcp_recvmsg(): Capture inbound data delivery to userspace applications.
  • inet_accept(): Track incoming connection acceptance.

Socket Tracing Implementation

// socket_trace.bpf.c
#include "vmlinux.h"
#include "bpf_helpers.h"
#include <bpf/bpf_tracing.h>

struct socket_event {
    __u64 timestamp;
    __u32 pid;
    __u32 uid;
    __u32 saddr;
    __u32 daddr;
    __u16 sport;
    __u16 dport;
    __u8 protocol;
    __u8 direction;
    char comm[16];
};

struct bpf_map_def SEC("maps") events = {
    .type = BPF_MAP_TYPE_RINGBUF,
    .max_entries = 256 * 1024,
};

SEC("kprobe/tcp_connect")
int trace_tcp_connect(struct pt_regs *ctx, struct sock *sk) {
    struct socket_event *event;
    
    event = bpf_ringbuf_reserve(&events, sizeof(*event), 0);
    if (!event)
        return 0;
    
    event->timestamp = bpf_ktime_get_ns();
    event->pid = bpf_get_current_pid_tgid() >> 32;
    event->uid = bpf_get_current_uid_gid();
    bpf_get_current_comm(&event->comm, sizeof(event->comm));
    
    event->saddr = BPF_CORE_READ(sk, __sk_common.skc_rcv_saddr);
    event->daddr = BPF_CORE_READ(sk, __sk_common.skc_daddr);
    event->dport = bpf_ntohs(BPF_CORE_READ(sk, __sk_common.skc_dport));
    event->protocol = IPPROTO_TCP;
    event->direction = 0; // outbound
    
    bpf_ringbuf_submit(event, 0);
    return 0;
}

TCP Protocol Stack Tracing

TCP State Machine Tracing

TCP connection states provide critical visibility into network health and application behavior. Tracing tcp_set_state() allows monitoring of all TCP state transitions:

SEC("tracepoint/sock/inet_sock_set_state")
int trace_tcp_state_change(struct trace_event_raw_inet_sock_set_state *ctx) {
    struct tcp_state_event *event;
    __u16 family = ctx->family;
    
    if (family != AF_INET && family != AF_INET6)
        return 0;
    
    event = bpf_ringbuf_reserve(&tcp_events, sizeof(*event), 0);
    if (!event)
        return 0;
    
    event->oldstate = ctx->oldstate;
    event->newstate = ctx->newstate;
    event->sport = ctx->sport;
    event->dport = ctx->dport;
    
    if (family == AF_INET) {
        bpf_probe_read(event->saddr, sizeof(struct in_addr), ctx->saddr);
        bpf_probe_read(event->daddr, sizeof(struct in_addr), ctx->daddr);
    }
    
    bpf_ringbuf_submit(event, 0);
    return 0;
}

TCP Retransmission Analysis

TCP retransmissions are a primary indicator of network congestion, packet loss, or routing issues. Tracing retransmission events helps identify network problems at their source:

// Track TCP retransmission statistics
struct bpf_map_def SEC("maps") retransmit_stats = {
    .type = BPF_MAP_TYPE_PERCPU_HASH,
    .key_size = sizeof(struct flow_key),
    .value_size = sizeof(struct flow_stats),
    .max_entries = 100000,
};

SEC("kprobe/tcp_retransmit_skb")
int trace_tcp_retransmit(struct pt_regs *ctx) {
    struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
    struct flow_stats *stats, zero = {};
    struct flow_key key = {};
    
    key.saddr = BPF_CORE_READ(sk, __sk_common.skc_rcv_saddr);
    key.daddr = BPF_CORE_READ(sk, __sk_common.skc_daddr);
    key.sport = BPF_CORE_READ(sk, __sk_common.skc_num);
    key.dport = BPF_CORE_READ(sk, __sk_common.skc_dport);
    
    stats = bpf_map_lookup_or_try_init(&retransmit_stats, &key, &zero);
    if (stats)
        __sync_fetch_and_add(&stats->retransmit_count, 1);
    
    return 0;
}

DNS Query Monitoring

Monitoring DNS at the Socket Level

DNS queries can be traced by monitoring UDP packets on port 53. This enables DNS analytics, threat detection, and troubleshooting:

// DNS monitoring via trace_udp_sendmsg
SEC("kprobe/udp_sendmsg")
int trace_dns_query(struct pt_regs *ctx) {
    struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
    struct msghdr *msg = (struct msghdr *)PT_REGS_PARM2(ctx);
    __u16 dport = BPF_CORE_READ(sk, __sk_common.skc_dport);
    
    if (dport != __constant_htons(53))
        return 0;
    
    struct iov_iter *iter = &msg->msg_iter;
    __u32 data_len = iter->count;
    
    if (data_len > 0 && data_len <= 512) {
        __u8 buf[512];
        bpf_probe_read(buf, data_len, iter->iov->iov_base);
        // Parse DNS header to extract domain name
        parse_dns_query(buf, data_len);
    }
    
    return 0;
}

CBPF-based DNS Capture with Socket Filters

For per-socket DNS monitoring, BPF socket filters provide a lightweight alternative to kprobes, with direct access to payload data and minimal overhead.

Network Performance Analysis

Latency Distribution Analysis

eBPF enables precise network latency measurement at multiple layers. By capturing timestamps at strategic points, we can decompose network latency into meaningful components:

struct bpf_map_def SEC("maps") latency_hist = {
    .type = BPF_MAP_TYPE_HISTOGRAM,
    .key_size = sizeof(u32),
};

// Measure TCP RTT at connection level
SEC("kprobe/tcp_rtt_estimator")
int trace_rtt(struct pt_regs *ctx) {
    struct sock *sk = (struct sock *)PT_REGS_PARM1(ctx);
    struct tcp_sock *tp = (struct tcp_sock *)sk;
    
    __u64 srtt = BPF_CORE_READ(tp, srtt_us) >> 3;
    u32 slot = bpf_log2l(srtt);
    hist_inc(latency_hist, slot);
    
    return 0;
}

Throughput Measurement

Accurate network throughput measurement requires accounting for packet payload sizes and protocol headers. The following approach provides both L3 throughput and L7 application-level statistics:

// Track per-connection throughput
struct bpf_map_def SEC("maps") throughput_map = {
    .type = BPF_MAP_TYPE_PERCPU_HASH,
    .key_size = sizeof(struct connection_key),
    .value_size = sizeof(struct throughput_stats),
    .max_entries = 50000,
};

SEC("kprobe/tcp_sendmsg")
int trace_send(struct pt_regs *ctx, struct sock *sk, struct msghdr *msg, size_t size) {
    struct throughput_stats *stats = lookup_throughput(sk);
    if (stats) {
        __sync_fetch_and_add(&stats->bytes_sent, size);
        __sync_fetch_and_add(&stats->packets_sent, 1);
    }
    return 0;
}

Production Deployment Guide

Kernel Version Requirements

For production eBPF networking, use Linux 5.10+ with the following configuration:

CONFIG_BPF=y
CONFIG_BPF_SYSCALL=y
CONFIG_BPF_JIT=y
CONFIG_BPF_EVENTS=y
CONFIG_KPROBE_EVENTS=y
CONFIG_TRACEPOINTS=y
CONFIG_NET_CLS_BPF=m
CONFIG_NET_ACT_BPF=m
CONFIG_XDP_SOCKETS=y
CONFIG_BPF_STREAM_PARSER=y

Containerized Environments

When running in containers, ensure CAP_BPF and CAP_NET_ADMIN are granted. For Kubernetes, Pod security policies must allow privileged capabilities or use a CNI plugin with eBPF support (Cilium, Calico).

Performance Optimization

Key optimization strategies for production eBPF networking:

  • Use BPF_MAP_TYPE_PERCPU_* variants for high-frequency counters to eliminate lock contention.
  • Leverage BPF_MAP_TYPE_RINGBUF over perf_event_array for 2-4x throughput improvement.
  • Implement batch processing to reduce userspace transition overhead.
  • Use BTF (BPF Type Format) for CO-RE (Compile Once, Run Everywhere) compatibility across kernel versions.
  • Enable BPF JIT compilation for 10-20x performance improvement over interpreter.

Verifying Program Safety

The eBPF verifier ensures program safety by analyzing control flow, memory access, and bounded execution. Common verifier errors and resolutions:

  • Uninitialized register access: Initialize all variables before use.
  • Out-of-bounds stack access: Ensure bounds checks before array access.
  • Infinite loop detection: Add #pragma unroll for bounded loops.
  • Function call limitations: Convert recursive functions to iterative form.

Conclusion

eBPF is transforming network observability and performance optimization in Linux. With its safe execution environment and minimal overhead, eBPF enables real-time network monitoring, security enforcement, and performance tracing that was previously impossible without kernel modifications or expensive hardware solutions.

As the Linux kernel continues to evolve, eBPF networking capabilities expand with support for socket lookup, TLS encryption observability, and high-speed packet processing at 100Gbps+. Organizations deploying eBPF-based networking solutions gain unprecedented visibility and control over their network infrastructure.

The production-ready implementations discussed in this article provide a foundation for building robust, high-performance network monitoring and optimization systems that scale from single-node deployments to large-scale distributed architectures.

References

  • Kernel documentation: Documentation/networking/filter.rst
  • eBPF libraries: libbpf, cilium/ebpf, libpam
  • Tools: bpftrace, bpftool, tcpdump (libpcap)
  • Projects: Cilium, Katran, Falco
点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部