1. eBPF Core Architecture Overview
eBPF (Extended Berkeley Packet Filter) is a revolutionary technology that allows running sandboxed programs inside the Linux kernel without modifying kernel source code or loading kernel modules. Originally designed for network packet filtering, eBPF has evolved into a versatile kernel subsystem supporting tracing, networking, security, and observability.
1.1 eBPF Program Lifecycle
The eBPF program lifecycle consists of five distinct phases:
- Compilation: eBPF programs are written in C (or Rust/Python), compiled to eBPF bytecode via LLVM/Clang with target bpf
- Verification: The kernel verifier performs static analysis to ensure program safety — no unbounded loops, no out-of-bounds access, no unreachable instructions
- JIT Compilation: Verified bytecode is translated to native machine code for optimal performance
- Attachment: Programs are hooked to kernel functions, tracepoints, kprobes, uprobes, or XDP hooks
- Execution: Triggered by kernel events, with results exported via maps or perf events
1.2 eBPF Virtual Machine
The eBPF VM operates with 11 64-bit registers (R0-R10), where R0 stores return values and R1-R5 hold function arguments, a 512-byte stack, and communicates with userspace via maps storing arbitrary data. Maps can be accessed from both kernel and user space by eBPF programs.
2. eBPF Maps: Kernel-Userspace Communication
eBPF maps are key-value stores enabling data exchange between eBPF programs and userspace, or between different eBPF programs:
| Map Type | Use Case | Key Size | Value Size |
|---|---|---|---|
| Hash | Key-value lookup, counters | Variable | Variable |
| Array | Fixed-size indexed storage | 4 bytes | Variable |
| Ring Buffer | High-throughput streaming | N/A | Variable |
| Perf Event Array | Per-CPU event output | 4 bytes | 4 bytes (fd) |
| LRU Hash | Eviction-based caching | Variable | Variable |
| LRU Per-CPU Hash | Per-CPU caching | Variable | Variable |
| CPU Map | CPU-specific data access | 4 bytes | Variable |
| Stack | LIFO trace storage | 4 bytes | Variable |
| Bloom Filter | Probabilistic membership | N/A | Variable |
2.1 Ring Buffer vs Perf Buffer
Perf Buffer (perfbuf) is the traditional mechanism requiring pre-allocated per-CPU buffers, which may lose events when full. Ring Buffer (ringbuf) is the modern alternative with automatic eviction, ordered delivery, and better memory efficiency — the recommended choice for new programs.
3. Helper Functions and Verifier Mechanisms
3.1 BTF and Cross-Kernel Compatibility
BTF (BPF Type Format) provides metadata about kernel data structures, enabling CO-RE (Compile Once, Run Everywhere). With BTF, eBPF programs can adapt to different kernel versions at runtime by reading BTF info, eliminating the need to recompile for each kernel version.
3.2 Verifier Safety Guarantees
The verifier ensures eBPF programs are safe by checking: bounded loops (max iterations typically 1 million), bounded stack usage (512 bytes), and only verified memory accesses (bpf_probe_read()). The verifier uses abstract interpretation to track register states across all possible execution paths.
4. Tracing Mechanisms
4.1 Kprobes and Kretprobes
Kprobes provide dynamic kernel instrumentation by inserting breakpoints at arbitrary kernel function entry points. Kretprobes intercept function return values. Both incur overhead due to breakpoints but offer maximum flexibility.
4.2 Tracepoints
Tracepoints are static instrumentation hooks embedded in kernel source with zero overhead when not active. They provide stable signatures across kernel versions, making them preferable for production tracing.
4.3 XDP: eXpress Data Path
XDP enables packet processing at the lowest point in the software stack, immediately after packet reception by the NIC driver. XDP programs are executed before the kernel allocates an sk_buff, achieving single-digit nanosecond per-packet latency. Return codes include XDP_PASS, XDP_DROP, XDP_TX, and XDP_REDIRECT.
5. eBPF-based Observability Tools in Production
5.1 bpftrace
bpftrace provides a high-level scripting language for eBPF, ideal for ad-hoc tracing:
# Trace execve() syscalls with timestamps
bpftrace -e 'tracepoint:syscalls:sys_enter_execve { printf("%s %d %s\n", comm, pid, str(args->filename)); }'
# Count read() syscalls by process
bpftrace -e 'tracepoint:syscalls:sys_enter_read { @[comm] = count(); }'
# Trace file opens with latency histogram
bpftrace -e 'kprobe:do_sys_openat2 { @start[tid] = nsecs; } kretprobe:do_sys_openat2 /@start[tid]/ { @latency_us = hist((nsecs - @start[tid]) / 1000); delete(@start[tid]); }'
5.2 BCC (BPF Compiler Collection)
BCC provides Python-based frontends for eBPF with embedded C code. It's ideal for complex tracing and performance analysis. For example: tracing VFS read sizes with histogram output, or monitoring TCP retransmissions by hooking tcp_retransmit_skb.
5.3 libbpf CO-RE
libbpf with BTF-based CO-RE enables portable eBPF programs that adapt to different kernel versions. Build with clang -g -O2 -target bpf -D__TARGET_ARCH_x86, and the loader applies relocations at load time based on kernel BTF.
6. Container and Network Observability
6.1 Container Tracing with cgroup
eBPF programs can be attached to cgroup hooks to monitor container I/O, network, and CPU usage. The cgroup_skb program type enables per-cgroup network filtering, while cgroup_device controls device access.
6.2 Cilium: eBPF-based CNI
Cilium leverages eBPF for networking, security, and observability in Kubernetes:
- Networking: Direct kube-proxy replacement via socket-level load balancing (BPF_SOCK_OPS, SK_LOOKUP)
- Security: Network policies enforced at XDP, TC, and socket levels
- Observability: Hubble provides flow-level visibility without sidecars
6.3 Istio Ambient Mesh
Istio's ambient mode uses eBPF for zero-sidecar mTLS and L7 policy enforcement, reducing resource overhead compared to sidecar proxies.
7. Enhanced Observability with eBPF
7.1 Network Latency Analysis
eBPF can measure TCP connection latency by hooking tcp_connect and tcp_rcv_state_process, tracking the duration from SYN to SYN-ACK at nanosecond precision. This enables P99 tail latency identification without modifying applications.
7.2 Syscall Tracing and Filtering
By attaching eBPF programs to tracepoint/syscalls/sys_enter_* and sys_exit_*, one can trace all system calls with arguments, return values, and latencies. Combined with PID/cgroup filtering, this enables per-container syscall auditing.
7.3 Off-CPU Analysis
eBPF can trace scheduler context switches via sched_switch and sched_wakeup tracepoints, building off-CPU flame graphs to identify blocking I/O, lock contention, and scheduler delays.
8. Security: LSM eBPF
LSM (Linux Security Module) BPF allows attaching programs to LSM hooks for fine-grained security policy enforcement. Compared to traditional LSM modules, eBPF LSM provides dynamic policy updates without kernel rebuilds, enabling runtime security controls like file access restrictions, network access controls, and privilege escalation prevention.
9. Performance Benchmarks
| Operation | Native Kernel Module | eBPF | Overhead |
|---|---|---|---|
| Packet Filter (XDP) | 8.2M pps | 7.9M pps | ~4% |
| Syscall Trace (open) | N/A | 1.2M events/s | ~3% CPU |
| TCP Retransmit Hook | N/A | 0.8M events/s | ~2% CPU |
| cgroup Network Filter | N/A | 5.6M pps | ~6% |
10. Future Directions
- eBPF for Scheduling: Custom schedulers via sched_ext framework, enabling user-defined scheduling policies
- eBPF for Storage: I/O schedulers and block-layer observability
- DPU/SmartNIC offload: XDP programs executed on NIC hardware for ultra-low latency
- Kernel-level AI inference: Using eBPF for lightweight inference within kernel context
- Portable eBPF: Cross-platform eBPF via user-mode execution (eBPF for Windows, etc.)

发表评论 取消回复