Linux 内核调度器革命:sched_ext — 用 BPF 重写 CPU 调度策略

在 Linux 6.12 之前,修改 CPU 调度策略意味着修改内核源码、重新编译内核、重启机器。sched_ext 的出现彻底改变了这一现状——它允许在运行时加载 BPF 程序来自定义调度行为,无需重启、无需内核模块、无需风险操作。本文将深入剖析 sched_ext 的架构设计、实现原理,并构建一个完整的实战案例。

一、为什么需要 sched_ext?

1.1 调度器困境:通用 vs 专用

Linux CFS(完全公平调度器)设计目标是"在各种场景下都还不错"。这种折中带来了不可避免的性能损失:

  • AI 推理服务:对尾延迟极其敏感,需要 microsecond 级别的调度精度,CFS 的毫秒级时间片太粗糙
  • 实时音视频:需要严格的优先级继承和 deadline 保证,CFS 无法满足
  • 批处理作业:需要最大化吞吐量,CFS 的公平分配导致缓存利用率低
  • 混合部署:同一台机器上同时运行延迟敏感和批处理任务,CFS 无法有效隔离

PREEMPT_RT 解决了实时性问题,EEVDF 改进了延迟公平性,但都无法解决一个根本问题:每种工作负载的最优调度策略都不同,而内核不可能内置所有策略。

1.2 传统方案的局限

方案 缺点
修改 CFS 源码 升级困难、维护成本高、风险大
编写内核模块 安全性差、稳定性风险、API 不稳定
用户态调度器 上下文切换开销大、无法直接访问内核状态
cgroup 调优 灵活性有限、无法实现自定义调度逻辑

sched_ext 的解决方案是:在内核内部运行 BPF 程序,直接访问调度器数据结构,享受 BPF 验证器的安全保证。

1.3 sched_ext 的设计哲学

sched_ext 的核心思想是将调度策略与调度机制分离:

  • 机制(Mechanism):内核提供 CPU 分配、上下文切换、负载均衡、运行队列管理等基础设施
  • 策略(Policy):用户用 BPF 编写自定义的调度决策逻辑

类似 eBPF 在网络栈的成功——XDP 程序可以自定义数据包处理,sched_ext 让调度策略可编程化。

二、架构深度剖析

2.1 整体架构

┌─────────────────────────────────────────────────────────────────┐
│                        用户态                                    │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────────────┐  │
│  │ scx_rustland  │  │ scx_lavd     │  │ 你的自定义调度器      │  │
│  │ (Rust 简单)   │  │ (游戏/桌面)  │  │ (BPF 程序)           │  │
│  └──────┬───────┘  └──────┬───────┘  └──────────┬───────────┘  │
│         │                  │                      │             │
│         └──────────────────┼──────────────────────┘             │
│                            │ BPF 加载                          │
├────────────────────────────┼────────────────────────────────────┤
│                        内核态                                    │
│                            ▼                                     │
│  ┌──────────────────────────────────────────────────────────┐  │
│  │              sched_ext 核心框架                           │  │
│  │  ┌─────────┐ ┌──────────┐ ┌───────────┐ ┌────────────┐ │  │
│  │  │ BPF 验证 │ │ 调度队列  │ │  CPU 分配  │ │ 负载均衡   │ │  │
│  │  └─────────┘ └──────────┘ └───────────┘ └────────────┘ │  │
│  └──────────────────────────────────────────────────────────┘  │
│                            │                                     │
│                            ▼                                     │
│  ┌──────────────────────────────────────────────────────────┐  │
│  │              硬件抽象层 (CPU topology)                     │  │
│  └──────────────────────────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────────┘

2.2 核心数据结构

sched_ext 的关键数据结构围绕"调度实体"(sched_ext_entity, 简称 se)构建:

struct sched_ext_entity {
    /* 调度权重 */
    u64         weight;
    
    /* 调度还有多少时间才需被抢占 */
    u64         slice;           /* 时间片长度(纳秒) */
    u64         dsq_vtime;       /* 虚拟运行时间 */
    
    /* 入队/出队时间戳 */
    u64         enqueue_time;    /* 入队时间 */
    u64         start_task_time; /* 开始运行时间 */
    
    /* 调度约束 */
    u64         cpumask[CPUS_U64_LONGS];  /* 允许运行的 CPU */
    u64         cpus_ptr;        /* 指向 cpumask 的指针 */
    
    /* 调度标志 */
    u64         flags;           /* SCX_TASK_* 标志 */
    
    /* CPU ID */
    s32         cpu;             /* 分配的 CPU */
    
    /* 负载跟踪 */
    u64         load_weight;     /* 估计的运行负载 */
    
    /* 任务组层次 */
    struct cgroup *cg;           /* 所属 cgroup */
    u64         h_weight;        /* 分层权重 */
};

2.3 BPF 钩子函数

调度器通过 BPF 程序实现一组核心钩子:

/* 必选钩子 */

/* 选取下一个要运行的任务 */
void BPF_STRUCT_OPS(select_cpu, struct task_struct *p, s32 prev_cpu, u64 wake_flags);

/* 将任务入队到调度队列 */
void BPF_STRUCT_OPS(enqueue, struct task_struct *p, u64 enq_flags);

/* 从调度队列选取任务并运行 */
void BPF_STRUCT_OPS(dispatch, s32 cpu, struct task_struct *p);

/* 任务被唤醒时调用 */
void BPF_STRUCT_OPS(runnable, struct task_struct *p, u64 enq_flags);

/* 可选钩子 */

/* 任务停止运行时 */
void BPF_STRUCT_OPS(stopping, struct task_struct *p, bool runnable);

/* 任务退出时 */
void BPF_STRUCT_OPS(release, struct task_struct *p);

/* CPU 变为空闲 */
void BPF_STRUCT_OPS(cpu_release, s32 cpu);

/* CPU 变为可用 */
void BPF_STRUCT_OPS(cpu_online, s32 cpu);

/* CPU 变为离线 */
void BPF_STRUCT_OPS(cpu_offline, s32 cpu);

/* 调度器启用 */
s32 BPF_STRUCT_OPS(init)(void);

/* 调度器退出 */
void BPF_STRUCT_OPS(exit)(struct scx_exit_info *ei);

/* tick 事件 */
void BPF_STRUCT_OPS(tick, struct task_struct *p);

/* 设置任务的 CPU 亲和性 */
bool BPF_STRUCT_OPS(cpuset_cgroup_move)(struct task_struct *p, struct cgroup *from, struct cgroup *to);

/* cgroup 配置改变的回调 */
void BPF_STRUCT_OPS(cgroup_init)(struct cgroup *cgrp, struct cgroup_subsys_state *css);
void BPF_STRUCT_OPS(cgroup_exit)(struct cgroup *cgrp);

2.4 调度队列模型

sched_ext 支持两种调度队列模式:

全局队列(DSQ, Dispatchable Scheduling Queue):

  • 所有任务放入全局 FIFO 或 vtime 排序的队列
  • 实现简单,适合大多数场景
  • 示例:SCX_DSQ_GLOBAL

每 CPU 队列(Per-CPU DSQ):

  • 每个 CPU 有独立的运行队列
  • 需要负载均衡来避免任务堆积
  • 大多数高性能调度器使用此模式
全局 DSQ 模式:
┌─────────┐
│  CPU 0  │ ←─┐
├─────────┤   │
│  CPU 1  │ ←─┼─── [ Global DSQ: T1→T2→T3→T4 ]
├─────────┤   │
│  CPU 2  │ ←─┘
└─────────┘

Per-CPU DSQ 模式:
┌─────────┐   ┌────────┐
│  CPU 0  │ ← │ DSQ #0 │ → T1, T2
├─────────┤   └────────┘
│  CPU 1  │   ┌────────┐
├─────────┤ ← │ DSQ #1 │ → T3, T4
│  CPU 2  │   └────────┘
└─────────┘   (需负载均衡)

2.5 内置辅助函数

sched_ext 提供丰富的 BPF 辅助函数:

/* 调度队列操作 */
void scx_bpf_dispatch(struct task_struct *p, u64 dsq, u64 slice, u64 enq_flags);
void scx_bpf_dispatch_vtime(struct task_struct *p, u64 dsq, u64 slice, u64 vtime, u64 enq_flags);
bool scx_bpf_dispatch_nr_queued(u64 dsq);
void scx_bpf_consume(u64 dsq);

/* CPU 相关 */
s32 scx_bpf_pick_idle_cpu(const struct cpumask *cpus_allowed, u64 flags);
s32 scx_bpf_pick_any_cpu(const struct cpumask *cpus_allowed, u64 flags);
bool scx_bpf_test_and_clear_cpu_idle(s32 cpu);
u32 scx_bpf_nr_cpu_ids(void);
const struct cpumask *scx_bpf_get_possible_cpumask(void);
const struct cpumask *scx_bpf_get_online_cpumask(void);

/* 任务信息 */
u64 scx_bpf_now(void);                    /* 当前时间(ns) */
u64 scx_bpf_task_vtime(const struct task_struct *p);
u32 scx_bpf_task_cpu(const struct task_struct *p);
u64 scx_bpf_task_cgroup_id(const struct task_struct *p);
struct cgroup *scx_bpf_task_cgroup(struct task_struct *p);

/* 负载均衡 */
void scx_bpf_kick_cpu(s32 cpu, u64 flags);
void scx_bpf_error(const char *fmt, ...);

/* 迭代器 */
struct bpf_iter_scx_dsq *scx_bpf_dsq_iter_start(struct bpf_iter_scx_dsq *it);

三、实战:构建一个延迟敏感型调度器

下面我们构建一个简化但完整的调度器,专门针对 AI 推理服务场景优化:优先调度延迟敏感任务、隔离批处理任务、最小化跨 NUMA 调度。

3.1 调度策略设计

我们的调度器需要实现以下策略:

  1. 任务分类:根据 cgroup 路径将任务分为 "latency-critical" 和 "batch"
  2. 优先级队列:延迟任务使用高优先级独立队列
  3. CPU 隔离:延迟任务只在专属 CPU 批处理用剩余 CPU
  4. 时间片控制:延迟任务使用时间片更小但优先级更高
  5. 抢占机制:延迟任务到达时立即抢占批处理任务
  6. 3.2 完整 BPF 调度器实现

    // latsched.bpf.c
    #include "vmlinux.h"
    #include <bpf/bpf_helpers.h>
    #include <bpf/bpf_tracing.h>
    #include <bpf/bpf_core_helpers.h>
    
    #define __SCX_BPF__
    #include <scx/common.bpf.h>
    
    char _license[] SEC("license") = "GPL";
    
    /* 配置参数 */
    #define LATENCY_SLICE_NS    1000000     // 1ms for latency-critical tasks
    #define BATCH_SLICE_NS      4000000     // 4ms for batch tasks
    #define PREEMPT_WEIGHT      100          // Preemption threshold
    #define LATENCY_CG_SZ       64
    
    /* 自定义 DSQ ID */
    enum {
        DSQ_LATENCY = 0,    /* 延迟敏感任务队列 */
        DSQ_BATCH,          /* 批处理任务队列 */
        DSQ_GLOBAL,         /* 全局兜底队列 */
        NUM_DSQ,
    };
    
    /* 任务统计 */
    struct {
        __uint(type, BPF_MAP_TYPE_PERCPU_ARRAY);
        __uint(max_entries, 1);
        __type(key, u32);
        __type(value, struct task_perf);
    } stats SEC(".maps");
    
    struct task_stats {
        u64     latency_dispatched;
        u64     batch_dispatched;
        u64     latency_preempted;
        u64     batch_preempted;
        u64     total_runs;
    };
    
    /* 
     * 判断任务是否为延迟敏感任务
     * 通过 cgroup 路径判断:/latency/ 子树下的为延迟敏感任务
     */
    static __always_inline bool is_latency_critical(struct task_struct *p)
    {
        struct cgroup *cg;
        const char *path;
        
        cg = scx_bpf_task_cgroup(p);
        if (!cg)
            return false;
        
        // 使用 bpf_cgroup_ancestor_id 性能更好
        // 这里简化实现:通过层次判断
        if (cg->level <= 1)
            return false;
        
        // 顶层 cgroup 为 /latency 则表示延迟敏感
        // 实际可用 BPF map 缓存此判断
        return true;
    }
    
    /* ======== 必选钩子实现 ======== */
    
    /*
     * 选取任务应该运行在哪个 CPU
     * 策略:延迟任务优先选空闲的专属 CPU,批处理任务随便
     */
    void BPF_STRUCT_OPS(select_cpu, struct task_struct *p, s32 prev_cpu, u64 wake_flags)
    {
        s32 cpu;
        bool latency = is_latency_critical(p);
        const struct cpumask *idle_mask = scx_bpf_get_idle_cpumask();
        
        if (!idle_mask)
            return;
        
        if (latency) {
            // 延迟任务:在专属 CPU (0-3) 中选闲置的
            cpumask_t latency_cpus;
            cpumask_clear(&latency_cpus);
            // 假设 CPU 0-3 为专属延迟 CPU
            for (int i = 0; i < 4; i++) {
                if (bpf_cpumask_test_cpu(i, idle_mask))
                    bpf_cpumask_set_cpu(i, &latency_cpus);
            }
            
            cpu = bpf_cpumask_any_and_distribute(p->cpus_ptr, &latency_cpus);
            if (cpu >= 0) {
                scx_bpf_dsq_insert(p, SCX_DSQ_LOCAL, 0, 0);
                return;
            }
            
            // 延迟专属 CPU 无闲置,尝试任何闲置 CPU
            cpu = scx_bpf_pick_idle_cpu(p->cpus_ptr, 0);
            if (cpu >= 0) {
                scx_bpf_dsq_insert(p, SCX_DSQ_LOCAL, 0, 0);
            }
        } else {
            // 批处理任务:选任何闲置 CPU(避免延迟专属)
            cpu = scx_bpf_pick_idle_cpu(p->cpus_ptr, 0);
            if (cpu >= 0) {
                scx_bpf_dsq_insert(p, SCX_DSQ_LOCAL, 0, 0);
            }
        }
        
        scx_bpf_put_idle_cpumask(idle_mask);
    }
    
    /*
     * 任务入队
     * 根据任务类型放入不同 DSQ,并设置时间片
     */
    void BPF_STRUCT_OPS(enqueue, struct task_struct *p, u64 enq_flags)
    {
        bool latency = is_latency_critical(p);
        u64 slice;
        
        if (latency) {
            slice = LATENCY_SLICE_NS;
            // 高优先级队列
            scx_bpf_dispatch(p, DSQ_LATENCY, slice, enq_flags);
            
            // 尝试抢占当前运行的批处理任务
            struct task_struct *curr = scx_bpf_curr_task(p->cpu);
            if (curr && !is_latency_critical(curr)) {
                // 缩短当前任务的时间片以触发抢占
                scx_bpf_curr_task_slice(curr, 0);
                scx_bpf_kick_cpu(scx_bpf_task_cpu(curr), SCX_KICK_PREEMPT);
            }
        } else {
            slice = BATCH_SLICE_NS;
            // 批处理队列
            scx_bpf_dispatch(p, DSQ_BATCH, slice, enq_flags);
        }
    }
    
    /*
     * 调度器分发:从队列中取出任务并分配到 CPU
     * 延迟队列优先处理
     */
    void BPF_STRUCT_OPS(dispatch, s32 cpu, struct task_struct *prev)
    {
        /* 优先从延迟敏感任务队列取任务 */
        if (scx_bpf_consume(DSQ_LATENCY))
            return;
        
        /* 延迟队列无任务,检查是否有 prev 任务被抢占 */
        if (prev && is_latency_critical(prev) && p->scx->slice > 0) {
            scx_bpf_dispatch(prev, DSQ_LATENCY, p->scx->slice, 0);
        }
        
        /* 尝试从批处理队列取 */
        scx_bpf_consume(DSQ_GLOBAL);
        scx_bpf_consume(DSQ_BATCH);
    }
    
    /*
     * 定时器 tick:更新时间片,触发抢占
     */
    void BPF_STRUCT_OPS(tick, struct task_struct *p)
    {
        // tick 期间可以做更多精细调整
        // 例如:动态调整时间片、更新任务频率统计等
    }
    
    /* ======== 调度器生命周期 ======== */
    
    s32 BPF_STRUCT_OPS(init)(void)
    {
        struct task_stats *s;
        u32 key = 0;
        
        s = bpf_map_lookup_elem(&stats, &key);
        if (s)
            __builtin_memset(s, 0, sizeof(*s));
        
        bpf_printk("latsched: BPF 调度器已加载\n");
        return 0;
    }
    
    void BPF_STRUCT_OPS(exit)(struct scx_exit_info *ei)
    {
        bpf_printk("latsched: BPF 调度器已退出: %s\n", ei->msg);
    }
    
    /* ======== 调度器定义 ======== */
    
    SEC(".struct_ops.link")
    struct sched_ext_ops latsched_ops = {
        .select_cpu     = (void *)select_cpu,
        .enqueue        = (void *)enqueue,
        .dispatch       = (void *)dispatch,
        .tick           = (void *)tick,
        .init           = (void *)init,
        .exit           = (void *)exit,
        .name           = "latsched",
    };

    3.3 用户态加载器

    // latsched_user.c
    #include <stdio.h>
    #include <unistd.h>
    #include <signal.h>
    #include <bpf/libbpf.h>
    #include "latsched.skel.h"
    
    static volatile bool running = true;
    
    static void sig_handler(int sig)
    {
        running = false;
    }
    
    int main(int argc, char **argv)
    {
        struct latsched_bpf *skel;
        int err;
        
        signal(SIGINT, sig_handler);
        signal(SIGTERM, sig_handler);
        
        libbpf_set_strict_mode(LIBBPF_STRICT_ALL);
        
        /* 加载 BPF 程序 */
        skel = latsched_bpf__open_and_load();
        if (!skel) {
            fprintf(stderr, "Failed to open/load BPF skeleton\n");
            return 1;
        }
        
        /* 附加调度器 */
        err = latsched_bpf__attach(skel);
        if (err) {
            fprintf(stderr, "Failed to attach BPF scheduler: %d\n", err);
            goto cleanup;
        }
        
        printf("sched_ext 调度器 'latsched' 已加载\n");
        printf("使用 cgroup /sys/fs/cgroup/latency/ 标记延迟敏感任务\n");
        printf("按 Ctrl+C 退出...\n");
        
        while (running) {
            sleep(1);
            
            /* 可以在这里读取 BPF map 获取统计信息 */
        }
        
    cleanup:
        latsched_bpf__destroy(skel);
        return 0;
    }

    3.4 Makefile

    # Makefile
    BPF_SRC     := latsched.bpf.c
    USER_SRC    := latsched_user.c
    SKELETON    := latsched.skel.h
    
    CC          := clang
    CFLAGS      := -g -O2 -Wall
    
    BPF_CFLAGS  := -target bpf -D__TARGET_ARCH_x86
    
    .PHONY: all clean
    
    all: latsched
    
    $(SKELETON): $(BPF_SRC)
    	$(CC) $(BPF_CFLAGS) -c $(BPF_SRC) -o latsched.bpf.o
    	bpftool gen skeleton latsched.bpf.o > $@
    
    latsched: $(USER_SRC) $(SKELETON)
    	$(CC) $(CFLAGS) -o $@ $(USER_SRC) -lbpf -lelf -lz
    
    clean:
    	rm -f latsched latsched.bpf.o $(SKELETON)

    四、高级特性与最佳实践

    4.1 分层调度域

    现代服务器的 NUMA 拓扑对调度性能影响巨大。sched_ext 可直接访问 NUMA 拓扑信息:

    void BPF_STRUCT_OPS(select_cpu, struct task_struct *p, s32 prev_cpu, u64 wake_flags)
    {
        const struct cpumask *idle_mask = scx_bpf_get_idle_cpumask();
        
        // 尝试在同一 NUMA 节点的闲置 CPU 上运行
        s32 cpu = scx_bpf_pick_idle_cpu_node(prev_cpu_to_node(prev_cpu), p->cpus_ptr);
        
        // 同 NUMA 无闲置,选择任何闲置 CPU
        if (cpu < 0)
            cpu = scx_bpf_pick_idle_cpu(p->cpus_ptr, 0);
        
        scx_bpf_put_idle_cpumask(idle_mask);
    }

    4.2 CPU RRD(Round-Robin Distribution)

    当特定 CPU 上的流水线状态时,可以实现 CPU 轮询来提高缓存命中率:

    /* 维护 per-CPU 上次分配的 CPU ID */
    struct {
        __uint(type, BPF_MAP_TYPE_PERCPU_ARRAY);
        __uint(max_entries, 1);
        __type(key, u32);
        __type(value, s32);
    } last_cpu SEC(".maps");
    
    s32 cpu_rrd_select(struct task_struct *p, s32 prev_cpu)
    {
        u32 key = 0;
        s32 *last = bpf_map_lookup_elem(&last_cpu, &key);
        s32 candidate;
        
        if (!last)
            return prev_cpu;
        
        // 从上次 CPU + 1 开始环形搜索
        candidate = (*last + 1) % scx_bpf_nr_cpu_ids();
        *last = candidate;
        return candidate;
    }

    4.3 自适应时间片

    根据任务运行历史动态调整时间片:

    #define SLICE_MIN_NS    500000      // 0.5ms
    #define SLICE_MAX_NS    8000000     // 8ms
    
    u64 adaptive_slice(struct task_struct *p)
    {
        struct perf *perf = get_task_perf(p);
        u64 slice = SLICE_MIN_NS;
        
        // 高频短任务:大时间片(充分利用 CPU)
        if (perf->run_freq > 1000 && perf->avg_duration < 100000)
            slice = SLICE_MAX_NS;
        // 低频长任务:小时间片(减少对其他任务阻塞)
        else if (perf->avg_duration > 10000000)
            slice = SLICE_MIN_NS;
        else
            slice = (perf->avg_duration >> 1);
        
        return clamp(slice, SLICE_MIN_NS, SLICE_MAX_NS);
    }

    4.4 用户态通信

    通过 BPF perf buffer 或 ring buffer 向用户态发送事件:

    struct {
        __uint(type, BPF_MAP_TYPE_RINGBUF);
        __uint(max_entries, 256 * 1024);
    } events SEC(".maps");
    
    struct event {
        u32     type;       // PREEMPT, WAKEUP, etc.
        u32     pid;
        s32     from_cpu;
        s32     to_cpu;
        u64     vtime;
    };
    
    void send_event(struct event *ev)
    {
        struct event *e;
        
        e = bpf_ringbuf_reserve(&events, sizeof(*ev), 0);
        if (!e)
            return;
        
        __builtin_memcpy(e, ev, sizeof(*ev));
        bpf_ringbuf_submit(e, 0);
    }

    4.5 性能优化要点

    1. 避免在热路径做内存分配:所有数据结构使用 per-cpu map
    2. 极简钩子实现:dispatch 和 enqueue 是热路径,越简单越好
    3. 预计算与缓存:cgroup 路径判断等操作在 init 中预计算缓存到 map
    4. BPF 循环限制:verifier 默认循环上界较小,注意循环次数
    5. 避免 bpf_printk 热路径:调试时可用,生产环境移除
    6. 五、实战部署与评估

      5.1 部署环境准备

      # 确认内核支持 sched_ext
      grep CONFIG_SCHED_EXT /boot/config-$(uname -r)
      # 返回 y 表示支持
      
      # 安装必要依赖
      sudo apt install -y clang llvm libbpf-dev linux-tools-$(uname -r) \
                          bpftool libelf-dev zlib1g-dev
      
      # 编译调度器
      make
      
      # 停止 CFS 对目标 CPU 的管理(假设使用 CPU 0-3)
      # 先启用调度器
      sudo ./latsched &
      SCX_PID=$!
      
      # 将任务分类放入对应 cgroup
      sudo mkdir -p /sys/fs/cgroup/latency-critical
      sudo mkdir -p /sys/fs/cgroup/batch
      
      # 启动延迟敏感服务
      echo $$ | sudo tee /sys/fs/cgroup/latency-critical/cgroup.procs
      # 然后在该 shell 中启动推理服务
      
      # 启动批处理任务
      echo $$ | sudo tee /sys/fs/cgroup/batch/cgroup.procs

      5.2 性能评估指标

      部署后需要关注以下指标:

      # 查看调度器状态
      cat /sys/kernel/debug/sched/ext
      
      # 监控上下文切换速率
      perf stat -e 'sched:sched_switch' -a sleep 10
      
      # 查看尾延迟(对推理服务至关重要)
      # 使用 BPF 工具测量从任务唤醒到实际执行的时间
      bpftrace -e '
      tracepoint:sched:sched_wakeup,
      traceway:sched:sched_waking {
          @start[args->pid] = nsecs;
      }
      tracepoint:sched:sched_switch /@start[args->prev_pid]/ {
          $lat = nsecs - @start[args->prev_pid];
          @us = hist($lat / 1000);
          delete(@start[args->prev_pid]);
      }'
      
      # 查看 CPU 利用率分布
      mpstat -P ALL 1

      5.3 生产环境注意事项

      资源规划:

      • 为延迟敏感任务预留足够专属 CPU,建议 >= 25%
      • 监控 NUMA 拓扑,避免跨 NUMA 访问成为新瓶颈
      • 设置 cgroup cpu.max 防止批处理任务耗尽 CPU

      故障恢复:

      # 优雅卸载调度器
      kill -SIGINT $SCX_PID
      
      # 或强制卸载(返回 CFS)
      sudo sysctl kernel.sched_ext=0

      调试:

      # 查看 BPF 验证器日志
      cat /sys/kernel/debug/tracing/trace_pipe
      
      # 查看 bpf_printk 输出
      sudo bpftool prog trig pipe

      六、sched_ext 与 eBPF 调度器的未来

      6.1 主要项目当前状态

      项目 语言 特点 适用场景
      scx_rustland Rust 简单、用户态辅助 学习、实验
      scx_lavd C Latency-Aware Value 桌面、游戏、通用
      scx_rlfifo C Round-Robin FIFO 简单场景
      scx_simple C 最简框架 开发模板
      scx_userd C 用户态决策辅助 复杂策略

      6.2 演进方向

      短期(Linux 6.x):

      • 更多内置调度器上线
      • NUMA 感知增强
      • 能耗感知调度(EAS 集成)

      中期(Linux 7.x):

      • 负载均衡 BPF 化(当前还是 C 实现)
      • 设备模型调度(GPU/DPU)
      • 异构核心(P-core/E-core)感知

      长期愿景:

      • 完全可编程的调度器框架
      • 机器学习驱动的自适应调度策略
      • 与 WASM 运行时集成(用户态 WASM 沙箱支持 BPF 调度)

      6.3 对内核开发生态的影响

      sched_ext 代表了 Linux 内核设计思路的重要转变:

      1. 安全扩展而非修改核心:内核社区更倾向提供安全钩子而非直接合并的策略
      2. 快速迭代:策略更新无需内核升级,缩短反馈循环
      3. 生态融合:Rust-for-Linux + BPF 可能形成新的内核编程范式
      4. 标准化接口:类似 POSIX 的标准 API,确保 BPF 调度器跨版本兼容
      5. /* 
         * 未来理想状态:内核提供框架,各场景最优策略生态化
         * 
         * ┌────────────────────────────────────────┐
         * │           应用层 (AI推理/实时音频/数据库) │
         * └────────────────┬───────────────────────┘
         *                  │ libscx API
         * ┌────────────────▼───────────────────────┐
         * │       sched_ext BPF 策略层              │
         * │  ┌─────────┐ ┌──────────┐ ┌─────────┐ │
         * │  │ AI 推理  │ │ 实时音频  │ │ 数据库   │ │
         * │  └─────────┘ └──────────┘ └─────────┘ │
         * └────────────────┬───────────────────────┘
         *                  │
         * ┌────────────────▼───────────────────────┐
         * │      BPF 运行时 + SCX 框架              │
         * └────────────────┬───────────────────────┘
         *                  │
         * ┌────────────────▼───────────────────────┐
         * │         Linux 内核调度基础设施           │
         * └────────────────────────────────────────┘
         */

        七、总结

        sched_ext 是 Linux 内核 30 年来调度器架构最大的一次变革。它将 eBPF 的安全可编程性引入 CPU 调度,解决了困扰运维工程师多年的"一刀切"调度问题。

        关键要点回顾:

        • sched_ext 通过 BPF 程序在运行时自定义调度策略,无需修改内核源码
        • 核心架构围绕DSQ调度队列和 BPF 钩子函数展开
        • 实战价值在于为特定工作负载(AI推理、实时应用)提供针对性的调度优化
        • 安全性通过 BPF 验证器保证,崩溃不会影响系统稳定性
        • 未来将与 Rust-for-Linux、WASM 等技术融合,形成新的内核编程范式

        对于从事底层性能优化的工程师而言,sched_ext 是必须掌握的新工具;对于平台架构理解而言,它代表了内核设计哲学从"一刀切"走向"可编程"的重要里程碑。


        参考资源: - sched-ext/scx GitHub — 调度器集合 - Linux Kernel Documentation: sched-ext - BPF & eBPF for Beginners - [RFC 00/15] sched: EXTensible Scheduler Class — 原始 RFC 系列
点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部