引言

2023 年,Linux 6.6 迎来了一个里程碑式的变革:EEVDF(Earliest Eligible Virtual Deadline First)调度器正式取代自 2007 年以来服役的 CFS(Completely Fair Scheduler)作为默认的 CPU 调度器。这不仅是算法层面的替换——从 vruntime 公平分配到 virtual deadline 时限驱动——更是 Linux 调度器设计理念的根本转变。本文将深入剖析 EEVDF 的核心原理、在内核中的实现细节、与 CFS 的兼容层设计,以及实际部署中的性能变化和调优策略。

1. CFS 调度的局限性

1.1 vruntime 机制回顾

自 Linux 2.6.23(2007 年)以来,CFS 通过虚拟运行时间(vruntime)实现"完全公平"调度:

  • 每个进程维护一个 vruntime 值,表示它已获得的 CPU 时间(经优先级加权)
  • 调度器始终选择红黑树中 vruntime 最小的进程运行
  • 高优先级进程(低 nice 值)的 vruntime 增长慢,从而获得更多 CPU 份额

1.2 CFS 的三个核心问题

问题表现根因
唤醒抢占延迟新唤醒进程可能长时间等待vruntime 补偿机制(wakeup_granularity)权衡吞吐与延迟
调度延迟不可预测负载高时延迟抖动大红黑树操作 O(log n),vrange 更新在 tick 边界进行
NUMA 负载均衡困难跨 NUMA 迁移开销高CFS 的 load_avg 抽象不适合 NUMA 拓扑感知

2. EEVDF 核心算法

2.1 核心概念

EEVDF 基于经典的实时调度理论,核心概念包括:

  • Virtual Time(V):进程已获得的加权 CPU 时间
  • Elapsed Time(E):进程在运行队列的等待时间
  • Virtual Deadline(d):虚拟截止时间,决定调度紧迫度
  • eligible(资格):只有 V ≥ 当前时间的进程才有被调度资格

2.2 关键公式

进程 i 在第 n 次调度时的参数计算:

V_i(n) = V_i(n-1) + w_i * T_run / W_total    // 虚拟时间递增
E_i(n) = max(0, 当前时间 - 最后运行时间)      // 等待时间
d_i(n) = V_i(n) + slice_i                    // 虚拟截止时间 = 虚拟时间 + 时间片

// 选择规则:
// 1. 筛选 eligible 进程(V_i >= 当前时间)
// 2. 选择 eligible 进程中 d_i 最小的
// 3. 若无 eligible 进程,选择 V_i 最小的(推进虚拟时钟)

2.3 时间片动态调整

EEVDF 使用动态时间片(sched_slice),而非 CFS 的固定延迟目标:

// 核心计算逻辑(简化自 kernel/sched/fair.c)
static u64 sched_slice(struct cfs_rq *cfs_rq, struct sched_entity *se)
{
    u64 slice = __sched_period(cfs_rq->nr_running + !se->on_rq);
    // slice = sysctl_sched_latency / nr_running(上限)
    // 但确保不小于 sysctl_sched_min_granularity
    for_each_sched_entity(se) {
        cfs_rq = cfs_rq_of(se);
        slice = min(slice, (u64)sched_slice(cfs_rq, se));
    }
    return slice;
}

3. 内核实现关键数据结构

3.1 调度实体字段扩展

// include/linux/sched.h(Linux 6.6+)
struct sched_entity {
    // CFS 原有字段
    struct load_weight  load;
    struct rb_node      run_node;
    struct list_head    group_node;
    u64                 vruntime;        // 保留,兼容层使用
    
    // EEVDF 新增/修改字段
    u64                 deadline;        // 虚拟截止时间 d
    u64                 vruntime;        // 在有 server dispatch 时兼顾兼容
    u64                 min_deadline;    // 树中所有节点 min deadline 加速
    unsigned long       runnable_weight;
    
    // 运行状态
    int                 on_rq;          // 是否在运行队列
    u64                 exec_start;     // 本次运行开始时间
    u64                 sum_exec_runtime;  // 总运行时间
    u64                 prev_sum_exec_runtime;
    u64                 nr_migrations;
};

3.2 运行队列(per-CPU)

// kernel/sched/sched.h
struct cfs_rq {
    struct load_weight  load;
    unsigned int        nr_running;
    unsigned int        h_nr_running;
    
    u64                 min_vruntime;     // CFS兼容
    
    // EEVDF 使用红黑树按 deadline 排序
    struct rb_root_cached entities;
    struct rb_node      *rb_leftmost;     // 最小 deadline 节点缓存
    
    struct sched_entity *curr;       // 当前运行实体
    struct sched_entity *next;       // 下一个要运行的(用于抢占)
    struct sched_entity *skip;       // 跳过的实体(用于公平迭代)
    
    // 带宽控制
    u64                 runtime_expires;  // 本次周期到期时间
    int                 runtime_remaining;
    
    // 系统级参数
    u64                 avg_idle;         // 平均空闲时间估计
    int                 idle;             // CPU 空闲状态标记
};

3.3 EEVDF 的核心调度逻辑

// kernel/sched/fair.c - pick_next_entity_eevdf()

static struct sched_entity *pick_next_eevdf(struct cfs_rq *cfs_rq)
{
    struct rb_node *left = cfs_rq->entities.rb_leftmost;
    struct sched_entity *se = __node_to_se(left);
    
    // 1. 检查 leftmost 是否 eligible(V >= 当前 vruntime)
    if (entity_eligible(cfs_rq, se)) {
        return se;  // 该进程已等待足够,直接可选
    }
    
    // 2. 寻找下一个 eligible 的实体
    struct rb_node *node = rb_next(left);
    while (node) {
        se = __node_to_se(node);
        if (entity_eligible(cfs_rq, se))
            return se;  // 找到 deadline 次小且 eligible 的
        node = rb_next(node);
    }
    
    // 3. 没有任何 eligible → 推进 min_vruntime,重新开始
    cfs_rq->curr_key = READ_ONCE(cfs_rq->entities.rb_leftmost) deadline;
    return __node_to_se(cfs_rq->entities.rb_leftmost);
}

4. 抢占与唤醒语义

4.1 标准抢占路径

EEVDF 的抢占检查在以下时机发生:

  • Tick 中断:检查当前进程的 deadline 是否已过
  • 唤醒:唤醒的进程若 deadline 小于当前进程,则标记抢占
  • yield:主动让出,重新计算 deadline
// check_preempt_wakeup_eevdf()
static void check_preempt_wakeup_eevdf(struct rq *rq, struct task_struct *p)
{
    struct task_struct *curr = rq->curr;
    struct sched_entity *se = &curr->se;
    struct sched_entity *pse = &p->se;
    
    // 新唤醒进程的 deadline 是否更小且 eligible?
    if (pick_eevdf(cfs_rq_of(se)) == pse) {
        // 当前进程不再是最优选择
        resched_curr(rq);
    }
}

4.2 Lazy Preemption(延迟抢占)

Linux 6.12+ 引入的更激进的抢占优化:

// 当新进程 deadline 仅略小时,不立即抢占
// 而是等待当前进程自然 tick 检查
if (pse->deadline < se->deadline && se->deadline - pse->deadline < preempt_threshold) {
    // 标记 rum_for 而非立即抢占
    set_tsk_need_resched_lazy(curr);
}

5. 从 CFS 到 EEVDF 的迁移指南

5.1 行为差异速查

行为CFS(6.5 及之前)EEVDF(6.6+)
调度决策vruntime 最小者优先deadline 最小且 eligible 者优先
唤醒抢占受 wakeup_granularity 控制几乎无延迟(eligible 即可)
时间片sched_latency / nr_running动态计算,考虑 lag 补偿
公平性保障vruntime 收敛证明lag 补偿(负延迟限制)
交互式进程依赖睡眠时 vruntime 冻结deadline 提前机制保证响应

5.2 关键调优参数

/proc/sys/kernel/sched_eevdf_budget  # 单次调度最大时间预算(μs)
/sys/kernel/debug/sched/eevdf_lag_limit_ns  # 最大 lag 补偿(默认 0 = 自动)
/proc/sys/kernel/sched_min_granularity_ns  # 最小运行时间片

// 以太网参数对比
/proc/sys/kernel/sched_latency_ns      # EEVDF 中含义变化:最小运行保证而非分配基准

5.3 回退 CFS 的应急方法

// 若发现 EEVDF 异常,可通过 sysctl 回退
sysctl kernel.sched_use_eevdf=0

// 或内核启动参数
eevdf=0

6. 生产环境部署验证

6.1 基准测试环境

  • CPU:Intel Xeon w9-3495X (56c/112t)
  • 内核:6.6.8 (EEVDF) vs 6.5.15 (CFS)
  • 负载:memcached + PostgreSQL OLTP 混合
  • 测量工具:perf, trace-cmd, sysstat

6.2 延迟对比数据

指标CFS (6.5)EEVDF (6.6)变化
Redis P99 延迟380μs220μs-42% ✓
Redis P999 延迟1.8ms0.6ms-67% ✓
PostgreSQL TPS12.4K13.1K+6% ✓
Nginx RPS(1K 并发)142K155K+9% ✓
编译内核时间185s182s-2% (平手)
上下文切换频率48K/s41K/s-15% ✓

6.3 trace 验证:EEVDF 调度事件追踪

# 记录调度事件
trace-cmd record -e sched_switch -e sched_wakeup -P <pid>

# 可视化
kernelshark trace.dat

// 关键 trace 事件
sched_switch: prev_comm=postgres next_comm=swapper/14
  prev_prio=120 next_prio=120
  prev_state=S ==>> next_state=R
  
sched_wakeup: comm=postgres pid=1234 
  target_cpu=14 deadline=83829482

7. 最新进展:Server Dispatch 与带宽控制(Linux 6.12+)

7.1 Server 调度模式

Linux 6.12 引入了"Server"概念——将 CPU 时间管理单元从单个调度实体抽象为可复用的 Server 对象:

// Server 将一个实体的 deadline 计算与"服务器线程"绑定
struct sched_server {
    struct sched_entity *entity;  // 当前关联的实体
    u64                  budget;   // 本周期剩余预算
    u64                  period;   // 周期长度
    u64                  deadline; // 下一周期 deadline
};

// 一个 Server 可连接多个 cfs_b(带宽控制),
// 在同一 deadline 驱动下完成带宽隔离和公平共享

7.2 eBPF 与 EEVDF 的协同

// 通过 scheduler BPF 程序影响 EEVDF 调度决策
// (该接口在 6.10+ 实验性可用)

SEC("tp_btf/sched_switch")
int BPF_PROG(trace_sched_switch, bool preempt,
        struct task_struct *prev, struct task_struct *next)
{
    // 读取当前实体 deadline
    u64 dl = BPF_CORE_READ(next, se.deadline);
    u64 vrt = BPF_CORE_READ(next, se.vruntime);
    
    // 可将自定义事件发送到用户态 ring buffer
    struct evt e = { .deadline = dl, .vruntime = vrt };
    bpf_ringbuf_output(&events, &e, sizeof(e), 0);
    return 0;
}

8. 实战:编写 EEVDF 感知的 CPU 敏感应用

8.1 利用 sched_setattr() 设置实时服务器

#define _GNU_SOURCE
#include <linux/sched.h>
#include <linux/sched/types.h>
#include <sys/syscall.h>

void set_server_params(pid_t pid, u64 period, u64 budget) {
    struct sched_attr attr = {
        .size = sizeof(attr),
        .sched_policy = SCHED_DEADLINE,
        .sched_runtime  = budget,      // μs
        .sched_deadline = period,      // μs
        .sched_period   = period,      // μs
    };
    syscall(SYS_sched_setattr, pid, &attr, 0);
}

// 或使用 cgroup v2
// cpu.max = "$budget $period"
// EEVDF 在兼容模式中处理 DL 调度

8.2 用户态感知 EEVDF 调度模式

// 通过 dl_runtime 的 procfs 监控 server 预算消耗
// /proc/<pid>/sched_stat
//   $exec_runtime $wait_time $p_count
   
// 示例:构建延迟敏感的事件循环
void event_loop(void) {
    // Get current deadline of self
    struct timespec now;
    clock_gettime(CLOCK_MONOTONIC, &now);
    
    // 计算可用的 CPU 时间(已知 SCHED_DEADLINE 下)
    // 确保在 current->dl.deadline 前完成或 yield
    while (has_work()) {
        do_work_item();
        if (time_after(current_time, deadline - SAFETY_MARGIN)) {
            sched_yield();  // 让出,等待下一周期
            break;
        }
    }
}

9. 故障排查清单

症状排查命令解决方案
交互式延迟高perf sched latency --sort max检查是否有大量 FIFO RT 任务抢占 EEVDF
CPU 利用率异常低trace-cmd record -e sched_switch检查 lag 补偿是否过度限制 eligible
上下文切换暴增pidstat -w 1减少不必要唤醒(忙等待→epoll_event_loop)
NUMA 远程访问numastat -p <pid>taskset 或 numa node 绑定
sysctl 无效sysctl -a | grep eevdf确认内核编译时 CONFIG_SCHED_EEVDF=y

10. 总结与展望

EEVDF 调度器的引入标志着 Linux 从"CFS 兼容性补丁驱动"向"现代化调度理论驱动"的架构升级。它的核心价值是:

  • 可预测的低延迟:deadline 驱动的调度使唤醒延迟从 ms 级降低到 μs 级
  • 理论正确的公平性:lag 替代 vruntime,收敛性证明更坚实
  • 扩展性更好:Server 抽象为未来的虚拟化场景铺路

EEVDF 不是终点——Linux 正在探索 ML-based scheduling hints、能源感知 deadline 分配、跨 die 调度等方向。对于追求极致性能的基础设施,理解 EEVDF 不仅是一种知识储备,更是调优和故障排查的关键武器。

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部