Linux内核Thermal子系统与散热管理:从设备树到Governor算法的深度工程实践

引言

在现代数据中心和AI训练集群中,散热管理已经从"能跑就行"变为核心性能瓶颈。一颗NVIDIA H100的TDP高达700W,一个8卡H100节点轻松突破6kW功耗——这意味着散热设计稍有不当,就会导致cpu/ gpu因过热而主动降频,直接拖慢训练吞吐。

Linux内核的Thermal子系统正是为系统级散热控制而生。它不仅是嵌入式设备里调节风扇转速的小工具,在服务器领域更是与cpufreq、RAPL(Running Average Power Limit)等机制深度耦合,构成了现代处理器的功耗-热管理闭环。

本文将从内核架构出发,深入解析Thermal子系统的三大核心组件——thermal zone、thermal governor和cooling device,结合设备树描述、governor算法实现以及在AI集群中的生产级部署,呈现一幅完整的散热管理工程图景。

一、Thermal子系统架构总览

Linux Thermal子系统定义了一个清晰的抽象层次:

┌──────────────────────────────────────────────────┐
│              User Space                           │
│  (thermald / lm-sensors / custom daemons)        │
├──────────────────────────────────────────────────┤
│              Thermal Governor                     │
│  (step_wise / power_allocator / bang_bang / ...) │
├──────────────────────────────────────────────────┤
│              Thermal Zone                         │
│  (temperature sampling, trip points)             │
├──────────────────────────────────────────────────┤
│              Cooling Device                       │
│  (cpufreq / fan / clock throttling / ...)        │
├──────────────────────────────────────────────────┤
│              Hardware Sensors                     │
│  (ACPI thermal / DTS thermal zones / hwmon)      │
└──────────────────────────────────────────────────┘

核心数据结构关系:

// thermal_zone_device — 温度采样源
struct thermal_zone_device {
    int id;
    char type[THERMAL_NAME_LENGTH];
    struct thermal_zone_device_ops *ops;  // get_trip_type/get_temp/...
    struct thermal_zone_params *tzp;      // governor参数
    struct thermal_governor *governor;    // 绑定的governor
    struct list_head thermal_instances;   // cooling device绑定列表
    struct thermal_attr *trip_temp_attrs; // trip point温度值
    ...
};

// thermal_cooling_device — 散热执行器
struct thermal_cooling_device {
    int id;
    char type[THERMAL_NAME_LENGTH];
    struct thermal_cooling_device_ops *ops;  // get_max_state/get_cur_state/set_cur_state
    unsigned long state;                      // 当前散热等级
    unsigned long max_state;                  // 最大散热等级
    ...
};

thermal zone通过trip point(温度触发点)机制与cooling device关联。当温度达到或超过某个trip point时,绑定的cooling device会按governor决策提升散热等级(即进入更深度的"降温"状态)。

二、设备树中的Thermal Zone描述

在ARM服务器和嵌入式SoC中,thermal zone通常通过设备树(Device Tree)描述。以Rockchip RK3588为例:

thermal_zones: thermal-zones {
    soc-thermal {
        polling-delay-passive = <200>;   /* ms, 被动散热轮询间隔 */
        polling-delay = <1000>;          /* ms, 主动采样间隔 */
        thermal-sensors = <&tsadc 0>;    /* 绑定到温度传感器 */

        trips {
            threshold: trip-point-0 {
                temperature = <70000>;   /* 70°C, 单位毫摄氏度 */
                hysteresis = <2000>;     /* 2°C 回差,防止抖动 */
                type = "passive";        /* 被动散热触发点 */
            };

            target: trip-point-1 {
                temperature = <85000>;
                hysteresis = <2000>;
                type = "active";         /* 主动散热触发点 */
            };

            hot: trip-point-2 {
                temperature = <95000>;
                hysteresis = <2000>;
                type = "hot";           /* 严重过热,内核告警 */
            };

            crit: trip-point-crit {
                temperature = <110000>;
                hysteresis = <0>;
                type = "critical";      /* 触发紧急关机 */
            };
        };

        cooling-maps {
            map0 {
                trip = <&threshold>;
                cooling-device = <&cpu0_cooling THERMAL_NO_LIMIT THERMAL_NO_LIMIT>;
            };
            map1 {
                trip = <&target>;
                cooling-device = <&fan0 1 3>; /* 风扇等级1~3 */
            };
        };
    };
};

关键字段解读:

  • polling-delay-passive / polling-delay:前者是有被动散热需求时的轮询间隔,后者是无散热需求时的空闲采样间隔。设计时前者通常更小,以便快速响应温度上升。
  • hysteresis:回差值。当温度下降到 trip temperature - hysteresis 以下时,才解除该trip的触发状态。这防止了风扇频繁启停或cpufreq反复调节。
  • trip type:critical 类型的trip触发内核紧急关机(kernel_power_off),hot触发内核打印告警但系统仍运行,passive 和 active 则分别对应被动降频和主动散热。

三、Thermal Governor 算法深度解析

Governor是散热决策的"大脑"。内核内置了四种实现,各有适用场景。

3.1 step_wise:最简单的阶梯式策略

// drivers/thermal/gov_step_wise.c
static void thermal_zone_temperature_update(struct thermal_zone_device *tz,
                                            int temp)
{
    struct thermal_instance *instance;
    int cur_state;

    // 遍历所有绑定的 cooling device
    list_for_each_entry(instance, &tz->thermal_instances, tz_node) {
        cur_state = thermal_get_cur_state(instance);

        if (temp >= instance->trip_temp + thys) {
            // 温度超过trip point:提升散热等级
            if (cur_state < instance->upper)
                thermal_set_cur_state(instance, cur_state + 1);
        } else if (temp <= instance->trip_temp - instance->tz->tzp->thys / 2) {
            // 温度回退到 hysteresis 以下:降低散热等级
            if (cur_state > instance->lower)
                thermal_set_cur_state(instance, cur_state - 1);
        }
    }
}

step_wise 算法逻辑直接:温度超过trip就逐级加大散热级别,低于trip就逐级减小。优点是简单可预测,缺点是响应阶梯化、容易产生温度振荡。

3.2 power_allocator:基于PID的功率分配

这是目前ARM服务器和移动SoC中最常用的governor。它的核心思路是:先将温度预算转化为"可分配的功耗预算",再按比例分发给各个cooling device。

static void power_allocator_pid(struct thermal_zone_device *tz,
                                 int temp)
{
    int target_state;

    // 计算PID误差
    err = target_temp - temp;

    // 比例项
    p_term = tz->tzp->k_po * (err - tz->tzp->prev_err);
    // 积分项
    i_term = tz->tzp->k_p * err;
    i_term = clamp(i_term, MIN_POWER_ALLOCATOR_BUDGET, 
                   MAX_POWER_ALLOCATOR_BUDGET);

    // 总功率预算 (单位: mW)
    budget = p_term + i_term + tz->tzp->ctrl_temp;

    // 按权重分配给各cooling device
    list_for_each_entry(instance, &tz->thermal_instances, tz_node) {
        cd_dev = instance->cdev;
        weight = instance->weight;  // cooling-maps中定义

        // 计算该cooling device应分配的功率配额
        quota = (budget * weight) / total_weight;

        // 根据功率-状态映射表确定目标散热等级
        target_state = thermal_cdev_state2power(cd_dev, quota) 
                       ? find_target_state_by_power(cd_dev, quota)
                       : 0;

        thermal_set_cur_state(instance, target_state);
    }
}

power_allocator需要设备树提供以下参数:

thermal-zones {
    soc-thermal {
        thermal-governor = "power_allocator";
        sustainable-power = <2000>;          /* 可持续功耗,单位mW */

        thermal-sensors = <&tsadc 0>;

        trips {
            trip-point {
                temperature = <75000>;
                type = "passive";
            };
        };

        cooling-maps {
            map0 {
                trip = <&trip-point>;
                /* weight列表:按cooling device的贡献权重分配功率 */
                cooling-device = <&cpu0_cooling 1 3>,
                                 <&gpu_cooling 1 2>;
            };
        };
    };
};

适用场景:多cooling device协调散热(如CPU+GPU共享散热器),需要精细散热控制的高端SoC。

3.3 bang_bang:简单的bang-bang控制

// 温度超过trip → 开到最大散热
// 温度低于trip且减去hysteresis → 降到最低
if (temp >= trip_temp)
    thermal_set_cur_state(instance, max_state);
else if (temp < trip_temp - hysteresis)
    thermal_set_cur_state(instance, 0);

bang-bang适用于响应速度要求极高且不需要精细调节的场景,如服务器风扇控制——要么不开,一开就最大。

3.4 user_space:将决策权交给用户态

该governor不自己做决策,而是通过thermal_zone设备节点将温度信息暴露给用户态程序(如thermald),由用户态daemon全权决定散热策略。

static void user_space_bind_to_trip(struct thermal_zone_device *tz,
                                    const struct thermal_instance *instance)
{
    struct thermal_zone_device *zone = instance->tz;
    if (!tz->gov && zone->tzp && zone->tzp->governor_name)
        thermal_set_governor(zone, TherGov);
    // 不做任何内核侧散热决策
}

user_space在Intel平台上最为常见,配合thermald daemon使用。

四、Cooling Device 与 cpufreq 的深度集成

cpufreq cooling device是thermal子系统与CPU频率调节的关键桥梁。它让散热决策直接作用于CPU频率:当温度过高时,thermal governor通过降低cpufreq cooling device的"状态"来提高散热等级,反之亦然。

// drivers/thermal/cpufreq_cooling.c
static int cpufrcooling_set_cur_state(struct thermal_cooling_device *cdev,
                                      unsigned long state)
{
    struct cpufreq_cooling_device *cpufreq_cdev = cdev->devdata;
    unsigned int clip_freq;

    // 计算状态对应的频率上限
    clip_freq = cpufreq_cdev->freq_table[state];

    // 通过 cpufreq notifier 限制最大频率
    cpus_qos_update_request(&cpufreq_cdev->q_req, clip_freq);

    cdev->last_state = clip_freq;
    return 0;
}

一个典型的cpufreq cooling device频率-状态映射表如下:

状态(level) 频率上限(MHz) 说明
0 3600 无限制,最高频
1 3200 轻微降频
2 2800 中度降频
3 2200 重度降频
4 1600 紧急散热

这意味着thermal governor通过声明"最多到状态2",就能间接将CPU锁定在2.8GHz以下。

在实际硬件中,这个映射表通常由内核的OPP(Operating Performance Points)框架自动生成:

static int __init cpufreq_cooling_init(struct device_node *np)
{
    // 从设备树读取OPP表,构建 frequency table
    count = dev_pm_opp_get_opp_count(cpu_dev);
    for (i = 0; i < count; i++) {
        freq = dev_pm_opp_get_freq(opp);
        // 索引越小 = 频率越高 = 散热等级越低
        table[count - 1 - i] = freq;
    }
}

五、生产环境中的散热管理实战

5.1 Thermal Throttling对AI训练的影响

在AI推理/训练集群中,thermal throttling是隐蔽的性能杀手。当GPU或CPU因过热被thermal限制降频时,训练的token/s或FPS会断崖式下跌。

使用perf stat或直接读取MSR可以量化热节流的影响:

# 监控thermal throttling事件
perf stat -e power/energy-pkg/, power/energy-ram/ -a sleep 10

# 检查当前是否处于PROCHOT(过热降频)状态
rdmsr -p 0 0x19c  # MSR_IA32_THERM_STATUS
# bit 0 = Thermal Status, bit 1 = Thermal Log

# 查看RAPL power limit
rapl-azure /dev/cpu/*/msr  # 需msr-tools

5.2 散热策略与cgroups的协同

在容器化的AI推理集群中,我们希望关键推理任务不被thermal throttling拖累。通过cgroup v2的cpu.max和thermal协同可以实现优先级分层:

# 高优先级推理任务:分配专用的低NUMA节点和最高功率限制
mkdir /sys/fs/cgroup/inference-critical
echo "400000 100000" > /sys/fs/cgroup/inference-critical/cpu.max
echo "max 100" > /sys/fs/cgroup/inference-critical/cpu.latency

# Batch训练任务:允许被thermal限制,因为它们更关注吞吐而非延迟
mkdir /sys/fs/cgroup/training-batch
echo "200000 100000" > /sys/fs/cgroup/training-batch/cpu.max

5.3 RAPL与ThermalGovernor的闭环联动

RAPL(Running AveragePower Limit)允许在固件层面限制CPU功耗。它与thermal governor形成了一个双层防护:

   ┌─────────────────┐
   │ Thermal Governor │ ─→ 频率限制(长周期,秒级响应)
   │ (内核侧)         │
   └─────────────────┘
            ↑ 协同
   ┌─────────────────┐
   │ RAPL            │ ─→ 功耗限制(短周期,毫秒级响应)
   │ (固件/硬件侧)    │
   └─────────────────┘

在AI训练场景中,建议设置: - RAPL限制在 TDP 的90%,给突发负载留裕量 - Thermal governor 临界温度设置高于RAPL触发点,作为最后防线

# 设置RAPL PL1 (持续功耗) 到200W
echo 200000000 > /sys/class/powercap/intel-rapl:0/constraint_0_power_limit_uw
echo 200000000 > /sys/class/powercap/intel-rapl:1/constraint_0_power_limit_uw

5.4 故障排查:热力失控的四种症状

在实际运维中,thermal导致的故障通常表现为:

症状1:cpufreq无法跑满最大频

# 查看是否被thermal限制
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_max_freq

# 对比被thermal拦截的统计
cat /sys/devices/virtual/thermal/thermal_zone0/trip_0_temp
cat /sys/devices/virtual/thermal/thermal_zone0/policy

症状2:间歇性性能下降

# 监控温度波动
watch -n 1 'cat /sys/class/thermal/thermal_zone*/temp'

# 查看thermal事件日志
dmesg | grep -i thermal
# 类似输出: "thermal thermal_zone0: critical temperature (100 C), shutting down"

症状3:风扇持续满速但温度不降 这种情况通常是散热器安装问题或风扇故障。可通过sysfs接口强制降低CPU频率临时缓解:

echo 1600000 > /sys/devices/system/cpu/cpu*/cpufreq/scaling_max_freq

症状4:thermal governor震荡切换 governor参数不合适时,cooling device状态会在相邻级别反复切换(如风扇在等级2和3间跳动)。调节k_po、k_p等PID系数,或增大trip的hysteresis即可解决。

六、sysfs接口与监控

Linux thermal子系统通过sysfs暴露了完整的运行时接口:

# 查看所有thermal zone
ls /sys/class/thermal/
thermal_zone0  thermal_zone1  cooling_device0  cooling_device1

# 温度 (单位:毫摄氏度)
cat /sys/class/thermal/thermal_zone0/temp        # 例如:65000

# 查看trip点
cat /sys/class/thermal/thermal_zone0/trip_0_type  # passive/active/hot/crit
cat /sys/class/thermal/thermal_zone0/trip_0_temp   # 例如:85000 (85°C)
cat /sys/class/thermal/thermal_zone0/trip_0_hyst   # 例如:2000 (2°C)

# 查看cooling device状态
cat /sys/class/thermal/cooling_device0/type       # Processor (cpufreq)
cat /sys/class/thermal/cooling_device0/cur_state    # 当前散热等级
cat /sys/class/thermal/cooling_device0/max_state    # 最大散热等级

# 动态切换governor
echo "power_allocator" > /sys/class/thermal/thermal_zone0/policy

# 动态修改trip温度(适用于临时调试)
echo 80000 > /sys/class/thermal/thermal_zone0/trip_0_temp

使用lm-sensors和s-tui等工具可以可视化thermal状态:

# 安装sensors
apt install lm-sensors
sensors-detect

# 实时监控(支持thermal zone温度+风扇转速+CPU频率)
s-tui --thermal_zones all

七、展望:CXL与新一代散热管理范式

随着CXL(Compute Express Link)内存扩展的普及,散热管理正在面临新挑战:内存模块(尤其是CXL附加的DRAM和HBM)的热管理需求日益凸显。传统thermal子系统主要针对CPU和GPU设计,未来需要:

  1. CXL Type3设备热管理扩展:将CXL内存模块纳入cooling device框架
  2. NUMA感知的thermal调度:跨NUMA节点的thermal感知进程迁移
  3. AI驱动的预测thermal管理:利用温度时序数据预测热点,提前主动调速
  4. P-State与Thermal Governor的融合:在异构多核(如Intel P-core/E-core + 独立GPU)上实现统一的功耗-热调度

结语

Linux内核Thermal子系统虽然代码量不大(位于drivers/thermal/,约30个文件),但它连接了硬件传感器、cpufreq频率调节、风扇/液冷执行器,构成了整个平台散热管理的"神经系统"。

对于AI基础设施工程师而言,理解thermalsub系统与cpufreq、RAPL的交互原理,是优化GPU/CPU利用率、保障SLA的基础技能。在算力密度持续攀升的今天,散热管理不再是可有可无的板级配置,而是与计算效率直接相关的核心软件栈。

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部