Linux内核Thermal子系统与散热管理:从设备树到Governor算法的深度工程实践
引言
在现代数据中心和AI训练集群中,散热管理已经从"能跑就行"变为核心性能瓶颈。一颗NVIDIA H100的TDP高达700W,一个8卡H100节点轻松突破6kW功耗——这意味着散热设计稍有不当,就会导致cpu/ gpu因过热而主动降频,直接拖慢训练吞吐。
Linux内核的Thermal子系统正是为系统级散热控制而生。它不仅是嵌入式设备里调节风扇转速的小工具,在服务器领域更是与cpufreq、RAPL(Running Average Power Limit)等机制深度耦合,构成了现代处理器的功耗-热管理闭环。
本文将从内核架构出发,深入解析Thermal子系统的三大核心组件——thermal zone、thermal governor和cooling device,结合设备树描述、governor算法实现以及在AI集群中的生产级部署,呈现一幅完整的散热管理工程图景。
一、Thermal子系统架构总览
Linux Thermal子系统定义了一个清晰的抽象层次:
┌──────────────────────────────────────────────────┐
│ User Space │
│ (thermald / lm-sensors / custom daemons) │
├──────────────────────────────────────────────────┤
│ Thermal Governor │
│ (step_wise / power_allocator / bang_bang / ...) │
├──────────────────────────────────────────────────┤
│ Thermal Zone │
│ (temperature sampling, trip points) │
├──────────────────────────────────────────────────┤
│ Cooling Device │
│ (cpufreq / fan / clock throttling / ...) │
├──────────────────────────────────────────────────┤
│ Hardware Sensors │
│ (ACPI thermal / DTS thermal zones / hwmon) │
└──────────────────────────────────────────────────┘
核心数据结构关系:
// thermal_zone_device — 温度采样源
struct thermal_zone_device {
int id;
char type[THERMAL_NAME_LENGTH];
struct thermal_zone_device_ops *ops; // get_trip_type/get_temp/...
struct thermal_zone_params *tzp; // governor参数
struct thermal_governor *governor; // 绑定的governor
struct list_head thermal_instances; // cooling device绑定列表
struct thermal_attr *trip_temp_attrs; // trip point温度值
...
};
// thermal_cooling_device — 散热执行器
struct thermal_cooling_device {
int id;
char type[THERMAL_NAME_LENGTH];
struct thermal_cooling_device_ops *ops; // get_max_state/get_cur_state/set_cur_state
unsigned long state; // 当前散热等级
unsigned long max_state; // 最大散热等级
...
};
thermal zone通过trip point(温度触发点)机制与cooling device关联。当温度达到或超过某个trip point时,绑定的cooling device会按governor决策提升散热等级(即进入更深度的"降温"状态)。
二、设备树中的Thermal Zone描述
在ARM服务器和嵌入式SoC中,thermal zone通常通过设备树(Device Tree)描述。以Rockchip RK3588为例:
thermal_zones: thermal-zones {
soc-thermal {
polling-delay-passive = <200>; /* ms, 被动散热轮询间隔 */
polling-delay = <1000>; /* ms, 主动采样间隔 */
thermal-sensors = <&tsadc 0>; /* 绑定到温度传感器 */
trips {
threshold: trip-point-0 {
temperature = <70000>; /* 70°C, 单位毫摄氏度 */
hysteresis = <2000>; /* 2°C 回差,防止抖动 */
type = "passive"; /* 被动散热触发点 */
};
target: trip-point-1 {
temperature = <85000>;
hysteresis = <2000>;
type = "active"; /* 主动散热触发点 */
};
hot: trip-point-2 {
temperature = <95000>;
hysteresis = <2000>;
type = "hot"; /* 严重过热,内核告警 */
};
crit: trip-point-crit {
temperature = <110000>;
hysteresis = <0>;
type = "critical"; /* 触发紧急关机 */
};
};
cooling-maps {
map0 {
trip = <&threshold>;
cooling-device = <&cpu0_cooling THERMAL_NO_LIMIT THERMAL_NO_LIMIT>;
};
map1 {
trip = <&target>;
cooling-device = <&fan0 1 3>; /* 风扇等级1~3 */
};
};
};
};
关键字段解读:
- polling-delay-passive / polling-delay:前者是有被动散热需求时的轮询间隔,后者是无散热需求时的空闲采样间隔。设计时前者通常更小,以便快速响应温度上升。
- hysteresis:回差值。当温度下降到 trip temperature - hysteresis 以下时,才解除该trip的触发状态。这防止了风扇频繁启停或cpufreq反复调节。
- trip type:critical 类型的trip触发内核紧急关机(kernel_power_off),hot触发内核打印告警但系统仍运行,passive 和 active 则分别对应被动降频和主动散热。
三、Thermal Governor 算法深度解析
Governor是散热决策的"大脑"。内核内置了四种实现,各有适用场景。
3.1 step_wise:最简单的阶梯式策略
// drivers/thermal/gov_step_wise.c
static void thermal_zone_temperature_update(struct thermal_zone_device *tz,
int temp)
{
struct thermal_instance *instance;
int cur_state;
// 遍历所有绑定的 cooling device
list_for_each_entry(instance, &tz->thermal_instances, tz_node) {
cur_state = thermal_get_cur_state(instance);
if (temp >= instance->trip_temp + thys) {
// 温度超过trip point:提升散热等级
if (cur_state < instance->upper)
thermal_set_cur_state(instance, cur_state + 1);
} else if (temp <= instance->trip_temp - instance->tz->tzp->thys / 2) {
// 温度回退到 hysteresis 以下:降低散热等级
if (cur_state > instance->lower)
thermal_set_cur_state(instance, cur_state - 1);
}
}
}
step_wise 算法逻辑直接:温度超过trip就逐级加大散热级别,低于trip就逐级减小。优点是简单可预测,缺点是响应阶梯化、容易产生温度振荡。
3.2 power_allocator:基于PID的功率分配
这是目前ARM服务器和移动SoC中最常用的governor。它的核心思路是:先将温度预算转化为"可分配的功耗预算",再按比例分发给各个cooling device。
static void power_allocator_pid(struct thermal_zone_device *tz,
int temp)
{
int target_state;
// 计算PID误差
err = target_temp - temp;
// 比例项
p_term = tz->tzp->k_po * (err - tz->tzp->prev_err);
// 积分项
i_term = tz->tzp->k_p * err;
i_term = clamp(i_term, MIN_POWER_ALLOCATOR_BUDGET,
MAX_POWER_ALLOCATOR_BUDGET);
// 总功率预算 (单位: mW)
budget = p_term + i_term + tz->tzp->ctrl_temp;
// 按权重分配给各cooling device
list_for_each_entry(instance, &tz->thermal_instances, tz_node) {
cd_dev = instance->cdev;
weight = instance->weight; // cooling-maps中定义
// 计算该cooling device应分配的功率配额
quota = (budget * weight) / total_weight;
// 根据功率-状态映射表确定目标散热等级
target_state = thermal_cdev_state2power(cd_dev, quota)
? find_target_state_by_power(cd_dev, quota)
: 0;
thermal_set_cur_state(instance, target_state);
}
}
power_allocator需要设备树提供以下参数:
thermal-zones {
soc-thermal {
thermal-governor = "power_allocator";
sustainable-power = <2000>; /* 可持续功耗,单位mW */
thermal-sensors = <&tsadc 0>;
trips {
trip-point {
temperature = <75000>;
type = "passive";
};
};
cooling-maps {
map0 {
trip = <&trip-point>;
/* weight列表:按cooling device的贡献权重分配功率 */
cooling-device = <&cpu0_cooling 1 3>,
<&gpu_cooling 1 2>;
};
};
};
};
适用场景:多cooling device协调散热(如CPU+GPU共享散热器),需要精细散热控制的高端SoC。
3.3 bang_bang:简单的bang-bang控制
// 温度超过trip → 开到最大散热
// 温度低于trip且减去hysteresis → 降到最低
if (temp >= trip_temp)
thermal_set_cur_state(instance, max_state);
else if (temp < trip_temp - hysteresis)
thermal_set_cur_state(instance, 0);
bang-bang适用于响应速度要求极高且不需要精细调节的场景,如服务器风扇控制——要么不开,一开就最大。
3.4 user_space:将决策权交给用户态
该governor不自己做决策,而是通过thermal_zone设备节点将温度信息暴露给用户态程序(如thermald),由用户态daemon全权决定散热策略。
static void user_space_bind_to_trip(struct thermal_zone_device *tz,
const struct thermal_instance *instance)
{
struct thermal_zone_device *zone = instance->tz;
if (!tz->gov && zone->tzp && zone->tzp->governor_name)
thermal_set_governor(zone, TherGov);
// 不做任何内核侧散热决策
}
user_space在Intel平台上最为常见,配合thermald daemon使用。
四、Cooling Device 与 cpufreq 的深度集成
cpufreq cooling device是thermal子系统与CPU频率调节的关键桥梁。它让散热决策直接作用于CPU频率:当温度过高时,thermal governor通过降低cpufreq cooling device的"状态"来提高散热等级,反之亦然。
// drivers/thermal/cpufreq_cooling.c
static int cpufrcooling_set_cur_state(struct thermal_cooling_device *cdev,
unsigned long state)
{
struct cpufreq_cooling_device *cpufreq_cdev = cdev->devdata;
unsigned int clip_freq;
// 计算状态对应的频率上限
clip_freq = cpufreq_cdev->freq_table[state];
// 通过 cpufreq notifier 限制最大频率
cpus_qos_update_request(&cpufreq_cdev->q_req, clip_freq);
cdev->last_state = clip_freq;
return 0;
}
一个典型的cpufreq cooling device频率-状态映射表如下:
| 状态(level) | 频率上限(MHz) | 说明 |
|---|---|---|
| 0 | 3600 | 无限制,最高频 |
| 1 | 3200 | 轻微降频 |
| 2 | 2800 | 中度降频 |
| 3 | 2200 | 重度降频 |
| 4 | 1600 | 紧急散热 |
这意味着thermal governor通过声明"最多到状态2",就能间接将CPU锁定在2.8GHz以下。
在实际硬件中,这个映射表通常由内核的OPP(Operating Performance Points)框架自动生成:
static int __init cpufreq_cooling_init(struct device_node *np)
{
// 从设备树读取OPP表,构建 frequency table
count = dev_pm_opp_get_opp_count(cpu_dev);
for (i = 0; i < count; i++) {
freq = dev_pm_opp_get_freq(opp);
// 索引越小 = 频率越高 = 散热等级越低
table[count - 1 - i] = freq;
}
}
五、生产环境中的散热管理实战
5.1 Thermal Throttling对AI训练的影响
在AI推理/训练集群中,thermal throttling是隐蔽的性能杀手。当GPU或CPU因过热被thermal限制降频时,训练的token/s或FPS会断崖式下跌。
使用perf stat或直接读取MSR可以量化热节流的影响:
# 监控thermal throttling事件
perf stat -e power/energy-pkg/, power/energy-ram/ -a sleep 10
# 检查当前是否处于PROCHOT(过热降频)状态
rdmsr -p 0 0x19c # MSR_IA32_THERM_STATUS
# bit 0 = Thermal Status, bit 1 = Thermal Log
# 查看RAPL power limit
rapl-azure /dev/cpu/*/msr # 需msr-tools
5.2 散热策略与cgroups的协同
在容器化的AI推理集群中,我们希望关键推理任务不被thermal throttling拖累。通过cgroup v2的cpu.max和thermal协同可以实现优先级分层:
# 高优先级推理任务:分配专用的低NUMA节点和最高功率限制
mkdir /sys/fs/cgroup/inference-critical
echo "400000 100000" > /sys/fs/cgroup/inference-critical/cpu.max
echo "max 100" > /sys/fs/cgroup/inference-critical/cpu.latency
# Batch训练任务:允许被thermal限制,因为它们更关注吞吐而非延迟
mkdir /sys/fs/cgroup/training-batch
echo "200000 100000" > /sys/fs/cgroup/training-batch/cpu.max
5.3 RAPL与ThermalGovernor的闭环联动
RAPL(Running AveragePower Limit)允许在固件层面限制CPU功耗。它与thermal governor形成了一个双层防护:
┌─────────────────┐
│ Thermal Governor │ ─→ 频率限制(长周期,秒级响应)
│ (内核侧) │
└─────────────────┘
↑ 协同
┌─────────────────┐
│ RAPL │ ─→ 功耗限制(短周期,毫秒级响应)
│ (固件/硬件侧) │
└─────────────────┘
在AI训练场景中,建议设置: - RAPL限制在 TDP 的90%,给突发负载留裕量 - Thermal governor 临界温度设置高于RAPL触发点,作为最后防线
# 设置RAPL PL1 (持续功耗) 到200W
echo 200000000 > /sys/class/powercap/intel-rapl:0/constraint_0_power_limit_uw
echo 200000000 > /sys/class/powercap/intel-rapl:1/constraint_0_power_limit_uw
5.4 故障排查:热力失控的四种症状
在实际运维中,thermal导致的故障通常表现为:
症状1:cpufreq无法跑满最大频
# 查看是否被thermal限制
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_max_freq
# 对比被thermal拦截的统计
cat /sys/devices/virtual/thermal/thermal_zone0/trip_0_temp
cat /sys/devices/virtual/thermal/thermal_zone0/policy
症状2:间歇性性能下降
# 监控温度波动
watch -n 1 'cat /sys/class/thermal/thermal_zone*/temp'
# 查看thermal事件日志
dmesg | grep -i thermal
# 类似输出: "thermal thermal_zone0: critical temperature (100 C), shutting down"
症状3:风扇持续满速但温度不降 这种情况通常是散热器安装问题或风扇故障。可通过sysfs接口强制降低CPU频率临时缓解:
echo 1600000 > /sys/devices/system/cpu/cpu*/cpufreq/scaling_max_freq
症状4:thermal governor震荡切换
governor参数不合适时,cooling device状态会在相邻级别反复切换(如风扇在等级2和3间跳动)。调节k_po、k_p等PID系数,或增大trip的hysteresis即可解决。
六、sysfs接口与监控
Linux thermal子系统通过sysfs暴露了完整的运行时接口:
# 查看所有thermal zone
ls /sys/class/thermal/
thermal_zone0 thermal_zone1 cooling_device0 cooling_device1
# 温度 (单位:毫摄氏度)
cat /sys/class/thermal/thermal_zone0/temp # 例如:65000
# 查看trip点
cat /sys/class/thermal/thermal_zone0/trip_0_type # passive/active/hot/crit
cat /sys/class/thermal/thermal_zone0/trip_0_temp # 例如:85000 (85°C)
cat /sys/class/thermal/thermal_zone0/trip_0_hyst # 例如:2000 (2°C)
# 查看cooling device状态
cat /sys/class/thermal/cooling_device0/type # Processor (cpufreq)
cat /sys/class/thermal/cooling_device0/cur_state # 当前散热等级
cat /sys/class/thermal/cooling_device0/max_state # 最大散热等级
# 动态切换governor
echo "power_allocator" > /sys/class/thermal/thermal_zone0/policy
# 动态修改trip温度(适用于临时调试)
echo 80000 > /sys/class/thermal/thermal_zone0/trip_0_temp
使用lm-sensors和s-tui等工具可以可视化thermal状态:
# 安装sensors
apt install lm-sensors
sensors-detect
# 实时监控(支持thermal zone温度+风扇转速+CPU频率)
s-tui --thermal_zones all
七、展望:CXL与新一代散热管理范式
随着CXL(Compute Express Link)内存扩展的普及,散热管理正在面临新挑战:内存模块(尤其是CXL附加的DRAM和HBM)的热管理需求日益凸显。传统thermal子系统主要针对CPU和GPU设计,未来需要:
- CXL Type3设备热管理扩展:将CXL内存模块纳入cooling device框架
- NUMA感知的thermal调度:跨NUMA节点的thermal感知进程迁移
- AI驱动的预测thermal管理:利用温度时序数据预测热点,提前主动调速
- P-State与Thermal Governor的融合:在异构多核(如Intel P-core/E-core + 独立GPU)上实现统一的功耗-热调度
结语
Linux内核Thermal子系统虽然代码量不大(位于drivers/thermal/,约30个文件),但它连接了硬件传感器、cpufreq频率调节、风扇/液冷执行器,构成了整个平台散热管理的"神经系统"。
对于AI基础设施工程师而言,理解thermalsub系统与cpufreq、RAPL的交互原理,是优化GPU/CPU利用率、保障SLA的基础技能。在算力密度持续攀升的今天,散热管理不再是可有可无的板级配置,而是与计算效率直接相关的核心软件栈。

发表评论 取消回复