一、PSI 核心原理与架构

Pressure Stall Information (PSI) 是 Linux 4.20 引入的内核特性,通过 /proc/pressure/ 和 cgroup v2 接口精确量化系统资源的争用程度,是三阶资源监控体系的关键组成部分:

  • avg10 / avg60 / avg300:资源被部分/全部占用的指数移动平均占比(百分比)
  • total:累计等待时间(微秒),用于观察绝对阻塞量
  • some:至少一个任务被阻塞的比例
  • full:所有非空闲任务同时被阻塞的比例
some 与 full 的区别
"some=10%" 表示某时刻至少有一个任务因资源不足而停滞。"full=5%" 表示 CPU/IO/MEMORY 所有可用任务同时停滞。两者结合可区分轻度竞争与严重死锁。

PSI 在内层使用定时器触发 + 等待队列扫描机制,在每个时间窗口内统计任务在就绪队列中的停留时间。核心数据结构包括:

struct psi_group {
    struct psi_group_cpu __percpu *pgc;   // per-CPU 统计
    unsigned long avg[PSI_AVGS];          // avg10/60/300
    u64 total[NR_PSI_STATES];             // 累计等待时间
    struct delayed_work clock_work;       // 周期性统计
    struct psi_group_stats stats;         // /proc 暴露的统计
};

enum psi_states {
    PSI_IO_SOME = 0,
    PSI_IO_FULL,
    PSI_MEM_SOME,
    PSI_MEM_FULL,
    PSI_CPU_SOME,    // full 在 CPU 上语义与 some 相同
    NR_PSI_STATES
};

二、PSI 内核实现深度剖析

2.1 时间窗口与 EMA (指数移动平均) 算法

// 核心更新逻辑 (kernel/sched/psi.c)
static u64 calc_avgs(u64 total, unsigned long missed_periods)
{
    // missed_periods: 自上次更新以来的时钟周期数
    // 新 avg = (1 - decay) * old_avg + decay * (total_this_period / period)
    u64 decay = decay_calc(avg_period, ...);
    u64 val = total * 100;
    do_div(val, avg_period);
    return (old_avg * decay + val * (100 - decay)) / 100;
}

/* 三个时间窗口的 expires 值 */
#define PSI_FREQ_AVG10   (2 * HZ)       // 2秒周期
#define PSI_FREQ_AVG60   (5 * HZ)       // 5秒周期
#define PSI_FREQ_AVG300  (10 * HZ)      // 10秒周期

PSI 使用指数加权移动平均 (EWMA),确保 avg10 反映短期突发压力,avg300 捕获长期趋势。当 period=0(无等待),平均值按指数衰减至零。

2.2 内核源码核心路径

// 1. 任务进入等待状态 → psi_task_change()
void psi_task_change(struct task_struct *task, int clear, int set)
{
    // 当 task 从 TASK_RUNNING → TASK_INTERRUPTIBLE 时
    // groupc->tasks[ groupc->tasks_in_state[task]-- ]
    // 统计等待任务数,供 full/some 判断
}

// 2. 统计周期触发 → psi_avgs_work()
static void psi_avgs_work(struct work_struct *work)
{
    struct psi_group *group;
    // a. 遍历 per-CPU 状态
    // b. 计算本轮 total 增量
    // c. 应用 EMA 更新三个窗口
    // d. 更新 total 累计值
    // e. 触发阈值唤醒(polling)
}

// 3. 轮询机制 → psi_poll_worker()
static void psi_poll_worker(struct work_struct *work){
    // 当 avgX 跨越用户配置阈值时
    // 唤醒 epoll/select 等待的进程(如 systemd、cgroup 监控)
}

2.3 三种资源的 PSI 监控方式

资源类型full 含义监测等待点
cpuall tasks waiting(同 some)运行队列 rr_cpu_time_slice 耗尽
memory所有任务同时缺页直接回收 / 交换 / CMA 失败
io所有任务同时等待 I/O请求队列满 / 写回阻塞

三、PSI 与 cgroup v2 的深度集成

3.1 通过 cgroup v2 memory.pressure 接口暴露

# /sys/fs/cgroup/memory.pressure 内容 (cgroup v2)
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0

# 实时监控 (使用 poll/select/epoll)
# systemd 通过 sd_event_add_inotify / sd_event_source_set_memory_pressure()
echo "some 50000 200000" > /sys/fs/cgroup/mycgroup/memory.pressure
# 含义:如果在任意2秒窗口内存在内存压力超过50%,则唤醒等待进程

systemd 248+ 提供 LogMemoryPressure= 和 MemoryPressureThresholdSec= 配置,基于 PSI 触发 OOM 预警与服务降级。

3.2 配置 PSI 触发阈值

# systemd 服务配置 - 当内存 avg10>80% 时执行降级
[Service]
MemoryPressureThresholdSec=2s
ExecStartPost=/opt/scripts/memory_pressure_handler.sh

# Kubernetes 通过Resource QoS与admission webhook结合
# cgroup v2 路径: /sys/fs/cgroup/kubepods.slice/kubepods-besteffort.slice/.../memory.pressure

四、生产级运维实战

4.1 PSI 与 Page Cache 回收的预测模型

#!/bin/bash
# psi_oom_predictor.sh - 基于 PSI 的 OOM 预警脚本

MEM_PRESSURE_FILE=/sys/fs/cgroup/memory.pressure
THRESHOLD=70  # avg10 超过70% 触发预警

while true; do
    PRESSURE=$(grep "some" "$MEM_PRESSURE_FILE" | awk '{print $2}' | cut -d= -f2 | cut -d. -f1)
    if [ "$PRESSURE" -gt "$THRESHOLD" ]; then
        # 发送告警 / 触发缓存释放 / 扩容决策
        logger -t psi_alert "MEMORY PRESSURE avg10=${PRESSURE}% - risk of OOM"
        echo 3 > /proc/sys/vm/drop_caches  # 紧急释放可选
    fi
    sleep 2
done

4.2 PSI 与 CPU 调度的联动

# 检测 CPU 压力
cat /proc/pressure/cpu
some avg10=45.32 avg60=32.15 avg300=18.67 total=4567890123

# avg10=45.32表示:
# 在过去10秒内,稳定状态下约 45.32% 的时间存在至少一个任务等待 CPU 时间片

# Kubernetes VPA 结合 PSI 扩容
# 当 cpu.avg10 > 60 持续30秒,触发 Pod 垂直扩容
if [ "$(cat /proc/pressure/cpu | head -1 | awk '{print $2}' | cut -d. -f1)" -gt 60 ]; then
    kubectl scale deployment --replicas=$(( $(kubectl get deploy -o=jsonpath='{.spec.replicas}') + 1 )) myapp
fi

4.3 PSI 与 I/O 负载的深度诊断

# full avg300 > 30% 持续 → 磁盘队列深度长期饱和
# 典型场景:RocksDB compaction 引发 IO 风暴
# 排查链:
# 1. iostat -x 1 → 检查 %util / await / avgqu-sz
# 2. iotop → 识别瓶颈进程
# 3. cat /proc/pressure/io → 验证 PSI 趋势
# 4. cgroup v2 io.max → 限制写入带宽防止 IO full

五、PSI 监控系统设计

5.1 数据流架构

┌─────────────┐     ┌──────────────────┐     ┌─────────────────────────┐
│ /proc/pressure│────→│  Monitoring Agent │────→│  Time Series DB         │
│ /sys/fs/cgroup │     │  (Node Exporter) │     │  (Prometheus/InfluxDB)  │
└─────────────┘     └──────────────────┘     └─────────────────────────┘
                               │
                               ▼
                  ┌────────────────────────┐
                  │  Alert Manager / Grafana│
                  │  Alert Rule:            │
                  │  memory_full_avg10>40   │
                  │  for 30s → page         │
                  │  reclaim imminent       │
                  └────────────────────────┘

5.2 Prometheus 抓取与告警规则

# node-exporter 文本收集
# metrics:
#   node_pressure_cpu_waiting_seconds_total
#   node_pressure_io_waiting_seconds_total
#   node_pressure_memory_waiting_seconds_total

groups:
- name: psi_alerts
  rules:
  - alert: MemoryPressureImminent
    expr: rate(node_pressure_memory_waiting_seconds_total[1m]) > 0.4
    for: 30s
    labels:
      severity: critical
    annotations:
      summary: "内存压力濒临 OOM (instance {{ $labels.instance }})"
      description: "PSI memory full avg10={{ $value | humanizePercentage }}"

  - alert: IOPressureCritical
    expr: rate(node_pressure_io_stalled_seconds_total[1m]) > 0.3
    for: 1m
    labels:
      severity: warning
    annotations:
      summary: "I/O 压力严重 — 队列饱和度 {{ $value | humanizePercentage }}"

5.3 BPF 级 PSI 增强

// 使用 BPF 跟踪 PSI 事件的发生源
SEC("tracepoint/sched/sched_switch")
int trace_sched_contention(struct trace_event_raw_sched_switch *ctx)
{
    u32 pid = ctx->next_pid;
    u64 ts = bpf_ktime_get_ns();

    // 记录每个任务的实际执行延迟
    struct task_info *task = bpf_map_lookup_elem(&task_map, &pid);
    if (task) {
        u64 wait_time = ts - task->last_run_ts;
        bpf_map_update_elem(&latency_map, &pid, &wait_time, BPF_ANY);
    }
    return 0;
}

六、PSI 压力指标解读速查表

avg10 值含义推荐响应
0-10%健康,资源充足无需操作
10-30%轻度争用观察趋势,检查是否周期性任务触发
30-60%中度压力检查 top 10 进程资源占用,评估扩容
60-80%严重争用启动降级预案,限流保护核心服务
80%+临界点,濒临耗尽触发紧急扩容 / OOM 防护策略

七、PSI 的内部限制与最佳实践

  • 采样精度 vs 开销:PSI 在每次时钟中断时检查任务状态,单核开销约 ~1ms/秒,多核情况下约 ~0.02% CPU
  • 轮询模式 (poll) 注意:使用 epoll 等待阈值触发而非 busy-poll,推荐配合 systemd sd_event
  • NUMA 感知:cgroup v2 中 PSI 按 NUMA node 独立统计,(memory压力) 可能仅在某些 node
  • 容器逃逸风险:PSI /proc/pressure 在 Host 全量监控下可能泄露容器外进程延迟信息,建议隔离 namespace

八、PSI 的未来演进

Linux 6.x 后续版本持续增强 PSI:

  • PSI 任务级追踪 (task-level):通过 task_state_arr 精确统计每个 cgroup 内进程的等待类型
  • 与 eBPF 深度融合:BPF_MAP_TYPE_QUEUE + PSI 自定义阈值回调,实现智能弹性伸缩
  • mempcg OOM 优先级调整:内核 EAS 调度器根据 PSI 动态调整 low watermark
  • AI 驱动的预测模型:利用 PSI avg300 趋势+时间序列算法预测 30 秒后内存耗尽时间点 (OOM ETA)

最后更新:2026年10月6日 | 适用内核版本:Linux 4.20+ | 完整源码参考:kernel/sched/psi.c

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部