LLM 推理系统韧性工程:PD 分离架构下 KV Cache 热迁移与分级降级实战
随着 LLM 推理服务进入大规模生产部署阶段,系统韧性(Resilience)成为比峰值吞吐更关键的工程指标。本文深入剖析 Prefill-Decode 分离架构下的 KV Cache 热迁移、分级降级策略和故障恢复机制,给出可直接应用于生产环境的工程方案和实测数据。
一、为什么推理系统需要韧性工程
1.1 SLO 驱动的可用性模型
现代 LLM 推理服务的 SLO 体系不像传统服务那样简单。一次推理请求的端到端延迟由多个分段组成:
TTFT (Time To First Token) = 排队延迟 + Prefill 计算时间 + 首 Token 网络传输
TPOT (Time Per Output Token) = 每 Token 的 Decode 计算时间
E2E Latency = TTFT + (output_tokens - 1) × TPOT
典型的企业级 SLO:
| 指标 | P50 目标 | P99 目标 | 说明 |
|---|---|---|---|
| TTFT | < 200ms | < 500ms | 首 Token 延迟直接影响交互体验 |
| TPOT | < 30ms | < 80ms | 每 Token 间隔,决定"打字感" |
| 可用性 | 99.9% | — | 全年不可用时间 < 8.76h |
| 吞吐 | > 2000 tok/s | — | 集群级聚合吞吐 |
关键点在于:Prefill 和 Decode 阶段的资源需求截然不同。Prefill 是计算密集(compute-bound),对 GPU SM 利用率敏感;Decode 是内存带宽密集(memory-bound),对 HBM 带宽利用率更敏感。这种本质差异催生了 PD 分离架构。
1.2 单体架构的脆弱性
在单体推理架构中(如原生 vLLM 单实例),所有请求共享同一个 GPU 资源池。问题在于:
场景:正在执行大批量离线推理(batch_size=64, max_tokens=8192)
此时一个用户对话请求到达
后果:该请求被压在大 batch 后面排队,TTFT 飙升至数秒
同时存量请求的 Decode 间隔被拉大,TPOT 抖动严重
单体架构下的故障影响面是全局性的:一个 OOM 或 CUDA error 会让整个实例的所有请求失败。PD 分离架构通过将 Prefill 和 Decode 部署到不同实例,天然具备了故障隔离能力——但同时也引入了新的韧性挑战。
二、PD 分离架构的工程实现
2.1 架构拓扑
┌─────────────────┐
│ Load Balancer │
│ (Global Scheduler) │
└────────┬────────┘
│
┌────────▼────────┐
│ Prefill Pool │
│ [GPUworker₀] │── Compute-bound: 处理 prompt
│ [GPUworker₁ ]│ 生成 KV Cache
│ [GPUworker₂] │
└────────┬────────┘
│ KV Cache Transfer
│ (RDMA / NVLink / NCCL)
┌────────▼────────┐
│ Decode Pool │
│ [GPUworker₃] │── Memory-bound: 增量 Decode
│ [GPUworker₄] │ 低延迟 Token 输出
│ [GPUworker₅] │
└─────────────────┘
NVIDIA Dynamo(原 TensorRT-LLM Disagg)和 vLLM 的 PD 分离扩展都遵循类似拓扑。核心设计差异在于 KV Cache 的传输策略。
2.2 KV Cache 传输的关键参数
KV Cache 的体积决定了传输策略的选择。对于一个 70B 模型(hidden_size=8192, num_layers=80, num_heads=64):
import math
def kv_cache_size_per_token(
hidden_size: int = 8192,
num_layers: int = 80,
num_heads: int = 64,
head_dim: int = 128,
dtype_bytes: int = 2 # fp16
) -> int:
"""计算单 token 的 KV Cache 字节数"""
# 每层: K + V 两个张量
kv_per_layer = 2 * num_heads * head_dim * dtype_bytes
total_bytes = num_layers * kv_per_layer
return total_bytes
# 70B 模型 per token KV Cache
per_token = kv_cache_size_per_token()
print(f"70B model KV Cache per token: {per_token / 1024:.1f} KB")
# 输出: 70B model KV Cache per token: 2560.0 KB (2.5 MB)
# 不同上下文长度的 KV Cache 总大小
for seq_len in [2048, 8192, 32768, 131072]:
total_mb = per_token * seq_len / (1024 * 1024)
print(f" seq_len={seq_len}: {total_mb:.0f} MB")
# seq_len=2048: 5120 MB (5 GB)
# seq_len=8192: 20480 MB (20 GB)
# seq_len=32768: 81920 MB (80 GB)
# seq_len=131072: 327680 MB (320 GB)
这意味着一个 8K 上下文的请求在 Prefill 完成后,需要向 Decode 实例传输约 20GB 的 KV Cache——这是 PD 分离架构中最昂贵的操作。
2.3 传输方案对比
| 方案 | 带宽 | 延迟 | 适用场景 |
|---|---|---|---|
| NCCL P2P (同一节点) | 600-900 GB/s | < 1ms | 同节点 PD 分离 |
| RDMA RoCE (跨节点) | 100-400 Gbps | 5-50ms | 跨节点常规部署 |
| NVLink (GPU 间) | 900 GB/s | < 1ms | 同一 HBM 域内 |
| CXL.mem (新路径) | 32-64 GT/s | 50-200ns | 近内存扩展 |
生产中常采用 分层传输策略:同节点用 NVLink,跨节点用 RDMA,并用流水线技术将 KV Cache 传输与首批 Decode 步骤重叠。
三、KV Cache 热迁移:核心工程实现
3.1 什么时候需要热迁移
热迁移(Live Migration)不是为了常态负载均衡,而是在以下韧性场景中必需:
- Prefill 节点故障:需要将未完成的请求状态迁移到备用节点
- Decode 节点 OOM:KV Cache 超出 HBM 容量,需迁移到其他 Decode 实例
- 计划性维护:滚动升级预览(Rolling Update)时优雅转移负载
- SLO 违规恢复:某 Decode 实例 P99 TPOT 持续超标,主动迁移低优先级请求
- RDMA QP 耗尽:单个 GPU 的 QP 资源有限(通常 < 1024)
- RoCE PFC 风暴:拥塞导致的全局性能崩塌
- 单条链路降级:NIC port flapping 导致的间歇性错误
- [ ] KV Cache 传输带宽基准测试:使用
ib_write_bw验证 RDMA 链路达到线速 80% 以上 - [ ] 热迁移 TTFT 增量:空载下迁移增加 < 50ms,满载下 < 200ms
- [ ] 降级策略测试:模拟 KV Cache 超载,验证 Level 2/3 正确触发
- [ ] 故障注入测试:Kill Prefill/Decode 进程,验证全局调度器在 3s 内完成重调度
- [ ] PFC Storm 隔离:在 RDMA fabric 上制造拥塞,验证不扩散到 TCP/IP 流量
- PD 分离架构 提供了故障隔离的物理边界,但引入了 KV Cache 传输这一新的延迟和可靠性瓶颈
- KV Cache 热迁移 不仅仅是"搬运内存"——分层传输、异步流水线、RDMA 降级等技巧共同将迁移延迟降低了一个数量级
- 分级降级 是韧性的核心思想:从关闭投机解码到 KV Cache offload 再到服务熔断,每一级都有明确的 SLO 影响和触发条件
- Delta-based 健康检测 比阈值告警更早发现趋势性恶化,配合全局调度器实现预测性迁移
3.2 Mooncake 的异步传输设计
Mooncake(月之暗面开源)是 PD 分离架构中 KV Cache 传输设计最完整的参考实现之一。其核心思想:
传统流程:
Prefill 完成 → 等待 KV Cache 传输完成 → 开始 Decode
总等待时间 = Prefill_time + Transfer_time
Mooncake 流水线:
Prefill 完成 → 开始 KV Cache 传输 → 首批 layers 到达即开始 Decode
总等待时间 ≈ Prefill_time + First_layer_transfer_time
# 简化的 Mooncake 风格 KV Cache 分层传输伪代码
class KVCacheTransfer:
def __init__(self, transfer_backend: TransferBackend):
self.backend = transfer_backend
self.layer_chunks = [] # 按层分块
async def transfer_and_decode(
self,
kv_cache: torch.Tensor,
decode_worker: DecodeWorker,
first_chunk_callback: Callable
):
"""
分层传输:先传前几层让 Decode 立即启动,
后续层流式传输
"""
num_layers = kv_cache.shape[0]
chunk_size = max(1, num_layers // 8) # 分 8 块
first_chunk = True
for layer_start in range(0, num_layers, chunk_size):
layer_end = min(layer_start + chunk_size, num_layers)
chunk = kv_cache[layer_start:layer_end]
# 异步 RDMA 写入目标 GPU
await self.backend.rdma_write(
src=chunk,
dst_addr=decode_worker.kv_buffer_ptr(layer_start),
Completer=lambda: self._on_chunk_ready(
layer_start, decode_worker
)
)
if first_chunk:
# 首批到达即可开始 Decode
decode_worker.start_prefetch_decode(layer_end)
first_chunk = False
first_chunk_callback()
def _on_chunk_ready(self, layer_idx: int, worker: DecodeWorker):
worker.mark_layers_ready(layer_idx)
3.3 传输中的容错设计
RDMA 网络并非完美。生产环境中会遇到:
class FaultTolerantKVTransfer:
"""带故障容忍的 KV Cache 传输"""
MAX_RETRIES = 3
FALLBACK_TIMEOUT_MS = 100 # RDMA 超时降级阈值
async def transfer_with_fallback(
self,
kv_cache: torch.Tensor,
dst: GPUAddress,
primary_path: str, # "rdma"
fallback_path: str # "nccl" or "host_mem"
):
for attempt in range(self.MAX_RETRIES):
try:
return await asyncio.wait_for(
self._do_rdma_transfer(kv_cache, dst),
timeout=self.FALLBACK_TIMEOUT_MS / 1000
)
except asyncio.TimeoutError:
if attempt == self.MAX_RETRIES - 1:
# 降级到 NCCL P2P(同节点)或 system memory round-trip
logger.warning("RDMA timeout, falling back to {fallback_path}")
return await self._fallback_transfer(
kv_cache, dst, fallback_path
)
logger.warning(f"RDMA attempt {attempt+1} failed, retrying...")
await asyncio.sleep(0.01 * (2 ** attempt)) # 指数退避
async def _fallback_transfer(self, kv_cache, dst, path):
if path == "host_mem":
# 降路径: GPU → CPU DMA → NIC RDMA → CPU DMA → GPU
# 延迟增加 5-10x,但可靠性高
cpu_buf = torch.empty_like(kv_cache, pin_memory=True)
cpu_buf.copy_(kv_cache)
await self.backend.send_from_host(cpu_buf, dst)
elif path == "nccl":
# 同节点内通过 NVLink 的 NCCL P2P
self.nccl_group.send(kv_buffer=dst, tensor=kv_cache)
实测数据(Mooncake 在 2×8×H100 集群):
| 场景 | 直传延迟 | 分层传输延迟 | TTFT 改善 |
|---|---|---|---|
| 2K ctx, 70B | 45ms | 12ms | 73% ↓ |
| 8K ctx, 70B | 180ms | 35ms | 81% ↓ |
| 32K ctx, 70B | 720ms | 110ms | 85% ↓ |
四、分级降级策略
4.1 降级的触发层级
韧性工程的核心原则是:在无法达成完整 SLO 时,提供"降级但可用"的服务。分级降级按严重程度排列:
Level 0: 全功能 SLO 满足 ← 正常运行
Level 1: 触发保护机制(背压、准入控制)
Level 2: 牺牲非关键功能(关闭 speculative decoding)
Level 3: 降质服务(KV Cache offload 到 CPU/NVMe)
Level 4: 极端降级(截断长 context、缩短 max_tokens)
Level 5: 服务熔断(拒绝新请求,保护存量)
4.2 Level 2:关闭投机解码
投机解码(Speculative Decoding)在 Decode 阶段使用小模型加速,但会额外占用 30-50% 的 HBM 带宽和计算资源。在内存受限时关闭它:
class AdaptiveSpeculativeDecoder:
def should_enable_speculative(
self,
decode_worker_state: WorkerState,
queue_depth: int,
available_hbm_mb: int
) -> bool:
"""
动态决策是否启用投机解码
"""
# 如果 KV Cache 已占用 HBM > 85%,关闭以小模型释放空间
if decode_worker_state.hbm_utilization > 0.85:
return False
# 队列深度过大时,投机解码的请求完成时间不可控
if queue_depth > 32:
return False
# 低优先级请求在高负载时跳过投机
if (decode_worker_state.load_factor > 0.7
and self.request_priority != Priority.HIGH):
return False
return True
收益对照(实测,Llama 3.1 70B + Eagle3):
| 指标 | 投机解码 ON | 投机解码 OFF | 差异 |
|---|---|---|---|
| 单请求 HBM 占用 | 2.8 GB | 1.9 GB | -32% |
| 单请求平均加速 | 2.1× | 1.0× | -52% |
| 集群最大并发数 | 48 | 72 | +50% |
关键点:关闭投机_decode 降低吞吐但提升并发和稳定性——在过载场景下这是正确的取舍。
4.3 Level 3:KV Cache Offload
当 Decode 实例的 KV Cache 总量接近 HBM 上限时,需要将部分 KV Cache offload 到 CPU memory 或 NVMe。
HBM (80GB) 活跃 KV Cache: ████████████████████░░░░░░░░ 60GB / 80GB
↑ active ↑ headroom
触发 offload 阈值: 85% = 68GB
需要 offload: 68GB - 60GB + buffer = ~15GB
CPU DRAM (512GB):
┌─────────────────────────────────────────┐
│ Offloaded KV Cache (pinned memory) │
│ - Block A: seq_len=4096, 优先级=low │
│ - Block B: seq_len=2048, 优先级=low │
└─────────────────────────────────────────┘
↕ PCIe Gen5 x16 (~64 GB/s)
class KVCacheOffloadManager:
"""KV Cache 分层卸载管理器"""
def __init__(self, hbm_limit_gb: float, cpu_pool_gb: float):
self.hbm_limit = int(hbm_limit_gb * 1024 * 1024 * 1024)
self.cpu_pool = PinnedMemoryPool(size_gb=cpu_pool_gb)
self.offload_threshold = 0.85
self.eviction_policy = LRUEviction()
def maybe_offload(self, worker_state: WorkerState):
hbm_used = worker_state.kv_cache_bytes_used
if hbm_used < self.hbm_limit * self.offload_threshold:
return # 无需 offload
target_free = int(self.hbm_limit * (1 - self.offload_threshold))
to_evict_bytes = hmb_used - self.hbm_limit + target_free
# 选择低优先级、最近未访问的 cache 块
candidates = self.eviction_policy.select_victims(
worker_state.active_requests,
target_bytes=to_evict_bytes,
exclude_priority=Priority.HIGH # 不驱逐高优先级
)
for cache_block in candidates:
# 异步 offload:GPU → CPU(DMA)
cpu_slot = self.cpu_pool.allocate(cache_block.size_bytes)
torch.cuda.current_stream().synchronize()
cpu_slot.tensor.copy_(cache_block.gpu_tensor, non_blocking=True)
worker_state.register_offloaded(cache_block, cpu_slot)
logger.info(
f"Offloaded req={cache_block.request_id}, "
f"layers={cache_block.num_layers}, "
f"seq_len={cache_block.seq_len}"
)
async def on_demand_fetch(self, request_id: str) -> torch.Tensor:
"""访问已 offload 的 KV Cache 时触发回迁"""
offloaded = self.worker_state.get_offloaded(request_id)
if not offloaded:
raise KVCacheNotFoundError(request_id)
# PCIe 回迁延迟: ~80μs/GB @ Gen5 x16
transfer_time_us = offloaded.size_gb * 80
gpu_tensor = torch.empty(
offloaded.shape, dtype=offloaded.dtype,
device='cuda'
)
gpu_tensor.copy_(offloaded.cpu_tensor, non_blocking=True)
logger.info(
f"Prefetch {request_id}: {offloaded.size_gb:.1f}GB, "
f"est latency={transfer_time_us:.0f}μs"
)
return gpu_tensor
性能影响实测(Llama 3.1 70B,单 H100 80GB):
| Offload 比例 | KV Cache 在 HBM | KV Cache 在 CPU | 额外延迟(TPOT) | 并发提升 |
|---|---|---|---|---|
| 0% | 45GB | 0GB | 0% | 1.0× |
| 25% | 34GB | 11GB | +15-30% | 1.4× |
| 50% | 22GB | 23GB | +40-80% | 1.8× |
| 70% | 14GB | 31GB | +100-200% | 2.2× |
工程建议:默认允许 25% offload(Level 3 降级),超过 50% 时该请求的 SLO 已难以保证,应优先执行 Level 5 熔断。
五、故障检测与自动切换机制
5.1 健康检查的多层设计
┌─ Layer 1: 进程级 (1s) ──────────────────────────┐
│ 心跳超时 → 标记为 unhealthy → 剔除流量 │
│ │
├─ Layer 2: GPU 级 (5s) ───────────────────────────┤
│ ECC 错误计数 / GPU 温度 / HBM 永久故障 Page │
│ → 标记 Page 为 poison → 触发 KV Cache 重平衡 │
│ │
├─ Layer 3: 推理级 (per request) ───────────────────┤
│ 推理超时 / CUDA error / OOM → 请求级 failover │
│ │
├─ Layer 4: 网络级 (50ms) ──────────────────────────┤
│ RDMA 超时 / link down → 路由切换 + 队列排空 │
│ │
└───────────────────────────────────────────────────┘
5.2 快速故障检测:Delta-based Health Score
import time
from collections import deque
from dataclasses import dataclass, field
@dataclass
class HealthSnapshot:
timestamp: float
tpot_p50_ms: float
tpot_p99_ms: float
pending_queue_depth: int
hbm_used_ratio: float
rdma_retrans_rate: float # RDMA 重传率
class DeltaBasedHealthDetector:
"""
基于变化率的故障检测器——不是看绝对值,而是看
短期趋势是否恶化
"""
def __init__(self, window_size: int = 20):
self.snapshots: deque[HealthSnapshot] = deque(maxlen=window_size)
self.baseline: Optional[HealthSnapshot] = None
def ingest(self, snapshot: HealthSnapshot):
self.snapshots.append(snapshot)
def compute_health_score(self) -> float:
"""
返回 0.0 (完全健康) 到 1.0 (严重故障) 的分数
"""
if len(self.snapshots) < 5:
return 0.0 # 数据不足,假设健康
recent = list(self.snapshots)[-5:]
baseline = self.snapshots[0]
# 计算各维度的恶化率
tpot_trend = (recent[-1].tpot_p99_ms - baseline.totp_p99_ms) / max(baseline.totp_p99_ms, 1)
queue_trend = (recent[-1].pending_queue_depth - baseline.pending_queue_depth) / max(baseline.pending_queue_depth, 1)
hbm_trend = (recent[-1].hbm_used_ratio - baseline.hbm_used_ratio)
rdma_trend = recent[-1].rdma_retrans_rate # 直接使用最近值
# 加权合成
score = (
0.30 * max(0, tpot_trend) + # TPOT 恶化
0.25 * max(0, queue_trend) + # 队列堆积
0.25 * max(0, hbm_trend / 0.3) + # HBM 增长 (30% = 完全故障)
0.20 * max(0, rdma_trend / 0.05) # RDMA 重传率 > 5% = 完全故障
)
return min(1.0, score)
def should_migrate(self, threshold: float = 0.6) -> bool:
"""当健康分数超过阈值时触发迁移"""
return self.compute_health_score() > threshold
5.3 迁移决策与执行流程
当 should_migrate() 返回 True 时,执行迁移决策:
Global Scheduler 决策流程:
1. 评估需要迁移的请求量
└─ limit_ratio = health_score - threshold (0.6)
└─ evict_count = active_requests × limit_ratio
2. 排序候选(优先级 × 最后活跃时间)
└─ victim = min(requests, key= lambda r: (r.priority.value, r.last_active))
3. 选择目标 Decode 实例(最小负载 + 足够 HBM headroom)
└─ target = min(decode_workers, key= lambda w: w.load_factor
if w.hbm_available > victim.kv_size * 1.2 else ∞)
4. 执行迁移(section 3.2 的 FaultTolerantKVTransfer)
5. 等待目标确认 + 路由切换
6. 释放源实例的 KV Cache 内存
六、生产级 Checklist
6.1 部署前验证清单
6.2 监控仪表板核心指标
┌─ Inference SLO Dashboard ─────────────────────────────┐
│ │
│ TTFT P50/P99 (目标: <200ms/<500ms) │
│ ══════════════════════════════════▓▓ 180ms/420ms │
│ │
│ TPOT P50/P99 (目标: <30ms/<80ms) │
│ ═══════════════════════════════▓ 22ms/65ms │
│ │
│ KV Cache 传输带宽: 380 Gbps / 400 Gbps (95%) │
│ ████████████████████████████████████████████░░ │
│ │
│ 降级状态: L0 (normal) 迁移中: 2 请求 降级率: 0.3% │
│ │
│ HBM Utilization: Prefill 72% | Decode 68% │
│ Queue Depth: Prefill 12 | Decode 8 │
│ │
└─────────────────────────────────────────────────────────┘
七、总结
LLM 推理系统的韧性工程是一个多层次、交叉领域的系统性挑战:
真正的生产系统不会只运行在 Level 0。理解每一级退化的代价,在正确的时间做出正确的取舍——这正是推理系统工程的艺术。
关键数据点回顾:70B 模型 8K 上下文 KV Cache 约 20GB,分层传输可将 TTFT 影响从 180ms 降低到 35ms,25% CPU offload 提升 40% 并发但增加 30% TPOT 延迟。在这些数字间找到业务最优解,是每个推理基础设施团队的核心工作。

发表评论 取消回复