LLM 推理系统韧性工程:PD 分离架构下 KV Cache 热迁移与分级降级实战

随着 LLM 推理服务进入大规模生产部署阶段,系统韧性(Resilience)成为比峰值吞吐更关键的工程指标。本文深入剖析 Prefill-Decode 分离架构下的 KV Cache 热迁移、分级降级策略和故障恢复机制,给出可直接应用于生产环境的工程方案和实测数据。

一、为什么推理系统需要韧性工程

1.1 SLO 驱动的可用性模型

现代 LLM 推理服务的 SLO 体系不像传统服务那样简单。一次推理请求的端到端延迟由多个分段组成:

TTFT (Time To First Token) = 排队延迟 + Prefill 计算时间 + 首 Token 网络传输
TPOT (Time Per Output Token) = 每 Token 的 Decode 计算时间
E2E Latency = TTFT + (output_tokens - 1) × TPOT

典型的企业级 SLO:

指标P50 目标P99 目标说明
TTFT< 200ms< 500ms首 Token 延迟直接影响交互体验
TPOT< 30ms< 80ms每 Token 间隔,决定"打字感"
可用性99.9%—全年不可用时间 < 8.76h
吞吐> 2000 tok/s—集群级聚合吞吐

关键点在于:Prefill 和 Decode 阶段的资源需求截然不同。Prefill 是计算密集(compute-bound),对 GPU SM 利用率敏感;Decode 是内存带宽密集(memory-bound),对 HBM 带宽利用率更敏感。这种本质差异催生了 PD 分离架构。

1.2 单体架构的脆弱性

在单体推理架构中(如原生 vLLM 单实例),所有请求共享同一个 GPU 资源池。问题在于:

场景:正在执行大批量离线推理(batch_size=64, max_tokens=8192)
     此时一个用户对话请求到达
后果:该请求被压在大 batch 后面排队,TTFT 飙升至数秒
     同时存量请求的 Decode 间隔被拉大,TPOT 抖动严重

单体架构下的故障影响面是全局性的:一个 OOM 或 CUDA error 会让整个实例的所有请求失败。PD 分离架构通过将 Prefill 和 Decode 部署到不同实例,天然具备了故障隔离能力——但同时也引入了新的韧性挑战。

二、PD 分离架构的工程实现

2.1 架构拓扑

                    ┌─────────────────┐
                    │  Load Balancer  │
                    │  (Global Scheduler) │
                    └────────┬────────┘
                             │
                    ┌────────▼────────┐
                    │  Prefill Pool   │
                    │  [GPUworker₀]   │── Compute-bound: 处理 prompt
                    │  [GPUworker₁   ]│    生成 KV Cache
                    │  [GPUworker₂]   │
                    └────────┬────────┘
                             │ KV Cache Transfer
                             │ (RDMA / NVLink / NCCL)
                    ┌────────▼────────┐
                    │   Decode Pool   │
                    │  [GPUworker₃]   │── Memory-bound: 增量 Decode
                    │  [GPUworker₄]   │    低延迟 Token 输出
                    │  [GPUworker₅]   │
                    └─────────────────┘

NVIDIA Dynamo(原 TensorRT-LLM Disagg)和 vLLM 的 PD 分离扩展都遵循类似拓扑。核心设计差异在于 KV Cache 的传输策略。

2.2 KV Cache 传输的关键参数

KV Cache 的体积决定了传输策略的选择。对于一个 70B 模型(hidden_size=8192, num_layers=80, num_heads=64):

import math

def kv_cache_size_per_token(
    hidden_size: int = 8192,
    num_layers: int = 80,
    num_heads: int = 64,
    head_dim: int = 128,
    dtype_bytes: int = 2  # fp16
) -> int:
    """计算单 token 的 KV Cache 字节数"""
    # 每层: K + V 两个张量
    kv_per_layer = 2 * num_heads * head_dim * dtype_bytes
    total_bytes = num_layers * kv_per_layer
    return total_bytes

# 70B 模型 per token KV Cache
per_token = kv_cache_size_per_token()
print(f"70B model KV Cache per token: {per_token / 1024:.1f} KB")
# 输出: 70B model KV Cache per token: 2560.0 KB (2.5 MB)

# 不同上下文长度的 KV Cache 总大小
for seq_len in [2048, 8192, 32768, 131072]:
    total_mb = per_token * seq_len / (1024 * 1024)
    print(f"  seq_len={seq_len}: {total_mb:.0f} MB")
# seq_len=2048:  5120 MB (5 GB)
# seq_len=8192: 20480 MB (20 GB)
# seq_len=32768: 81920 MB (80 GB)
# seq_len=131072: 327680 MB (320 GB)

这意味着一个 8K 上下文的请求在 Prefill 完成后,需要向 Decode 实例传输约 20GB 的 KV Cache——这是 PD 分离架构中最昂贵的操作。

2.3 传输方案对比

方案带宽延迟适用场景
NCCL P2P (同一节点)600-900 GB/s< 1ms同节点 PD 分离
RDMA RoCE (跨节点)100-400 Gbps5-50ms跨节点常规部署
NVLink (GPU 间)900 GB/s< 1ms同一 HBM 域内
CXL.mem (新路径)32-64 GT/s50-200ns近内存扩展

生产中常采用 分层传输策略:同节点用 NVLink,跨节点用 RDMA,并用流水线技术将 KV Cache 传输与首批 Decode 步骤重叠。

三、KV Cache 热迁移:核心工程实现

3.1 什么时候需要热迁移

热迁移(Live Migration)不是为了常态负载均衡,而是在以下韧性场景中必需:

  1. Prefill 节点故障:需要将未完成的请求状态迁移到备用节点
  2. Decode 节点 OOM:KV Cache 超出 HBM 容量,需迁移到其他 Decode 实例
  3. 计划性维护:滚动升级预览(Rolling Update)时优雅转移负载
  4. SLO 违规恢复:某 Decode 实例 P99 TPOT 持续超标,主动迁移低优先级请求
  5. 3.2 Mooncake 的异步传输设计

    Mooncake(月之暗面开源)是 PD 分离架构中 KV Cache 传输设计最完整的参考实现之一。其核心思想:

    传统流程:
      Prefill 完成 → 等待 KV Cache 传输完成 → 开始 Decode
      总等待时间 = Prefill_time + Transfer_time
    
    Mooncake 流水线:
      Prefill 完成 → 开始 KV Cache 传输 → 首批 layers 到达即开始 Decode
      总等待时间 ≈ Prefill_time + First_layer_transfer_time
    # 简化的 Mooncake 风格 KV Cache 分层传输伪代码
    class KVCacheTransfer:
        def __init__(self, transfer_backend: TransferBackend):
            self.backend = transfer_backend
            self.layer_chunks = []  # 按层分块
        
        async def transfer_and_decode(
            self, 
            kv_cache: torch.Tensor,
            decode_worker: DecodeWorker,
            first_chunk_callback: Callable
        ):
            """
            分层传输:先传前几层让 Decode 立即启动,
            后续层流式传输
            """
            num_layers = kv_cache.shape[0]
            chunk_size = max(1, num_layers // 8)  # 分 8 块
            
            first_chunk = True
            for layer_start in range(0, num_layers, chunk_size):
                layer_end = min(layer_start + chunk_size, num_layers)
                chunk = kv_cache[layer_start:layer_end]
                
                # 异步 RDMA 写入目标 GPU
                await self.backend.rdma_write(
                    src=chunk,
                    dst_addr=decode_worker.kv_buffer_ptr(layer_start),
                    Completer=lambda: self._on_chunk_ready(
                        layer_start, decode_worker
                    )
                )
                
                if first_chunk:
                    # 首批到达即可开始 Decode
                    decode_worker.start_prefetch_decode(layer_end)
                    first_chunk = False
                    first_chunk_callback()
        
        def _on_chunk_ready(self, layer_idx: int, worker: DecodeWorker):
            worker.mark_layers_ready(layer_idx)

    3.3 传输中的容错设计

    RDMA 网络并非完美。生产环境中会遇到:

    • RDMA QP 耗尽:单个 GPU 的 QP 资源有限(通常 < 1024)
    • RoCE PFC 风暴:拥塞导致的全局性能崩塌
    • 单条链路降级:NIC port flapping 导致的间歇性错误
    class FaultTolerantKVTransfer:
        """带故障容忍的 KV Cache 传输"""
        
        MAX_RETRIES = 3
        FALLBACK_TIMEOUT_MS = 100  # RDMA 超时降级阈值
        
        async def transfer_with_fallback(
            self,
            kv_cache: torch.Tensor,
            dst: GPUAddress,
            primary_path: str,    # "rdma"
            fallback_path: str    # "nccl" or "host_mem"
        ):
            for attempt in range(self.MAX_RETRIES):
                try:
                    return await asyncio.wait_for(
                        self._do_rdma_transfer(kv_cache, dst),
                        timeout=self.FALLBACK_TIMEOUT_MS / 1000
                    )
                except asyncio.TimeoutError:
                    if attempt == self.MAX_RETRIES - 1:
                        # 降级到 NCCL P2P(同节点)或 system memory round-trip
                        logger.warning("RDMA timeout, falling back to {fallback_path}")
                        return await self._fallback_transfer(
                            kv_cache, dst, fallback_path
                        )
                    logger.warning(f"RDMA attempt {attempt+1} failed, retrying...")
                    await asyncio.sleep(0.01 * (2 ** attempt))  # 指数退避
        
        async def _fallback_transfer(self, kv_cache, dst, path):
            if path == "host_mem":
                # 降路径: GPU → CPU DMA → NIC RDMA → CPU DMA → GPU
                # 延迟增加 5-10x,但可靠性高
                cpu_buf = torch.empty_like(kv_cache, pin_memory=True)
                cpu_buf.copy_(kv_cache)
                await self.backend.send_from_host(cpu_buf, dst)
            elif path == "nccl":
                # 同节点内通过 NVLink 的 NCCL P2P
                self.nccl_group.send(kv_buffer=dst, tensor=kv_cache)

    实测数据(Mooncake 在 2×8×H100 集群):

    场景直传延迟分层传输延迟TTFT 改善
    2K ctx, 70B45ms12ms73% ↓
    8K ctx, 70B180ms35ms81% ↓
    32K ctx, 70B720ms110ms85% ↓

    四、分级降级策略

    4.1 降级的触发层级

    韧性工程的核心原则是:在无法达成完整 SLO 时,提供"降级但可用"的服务。分级降级按严重程度排列:

    Level 0: 全功能 SLO 满足 ← 正常运行
    Level 1: 触发保护机制(背压、准入控制)
    Level 2: 牺牲非关键功能(关闭 speculative decoding)
    Level 3: 降质服务(KV Cache offload 到 CPU/NVMe)
    Level 4: 极端降级(截断长 context、缩短 max_tokens)
    Level 5: 服务熔断(拒绝新请求,保护存量)

    4.2 Level 2:关闭投机解码

    投机解码(Speculative Decoding)在 Decode 阶段使用小模型加速,但会额外占用 30-50% 的 HBM 带宽和计算资源。在内存受限时关闭它:

    class AdaptiveSpeculativeDecoder:
        def should_enable_speculative(
            self,
            decode_worker_state: WorkerState,
            queue_depth: int,
            available_hbm_mb: int
        ) -> bool:
            """
            动态决策是否启用投机解码
            """
            # 如果 KV Cache 已占用 HBM > 85%,关闭以小模型释放空间
            if decode_worker_state.hbm_utilization > 0.85:
                return False
            
            # 队列深度过大时,投机解码的请求完成时间不可控
            if queue_depth > 32:
                return False
            
            # 低优先级请求在高负载时跳过投机
            if (decode_worker_state.load_factor > 0.7 
                and self.request_priority != Priority.HIGH):
                return False
            
            return True

    收益对照(实测,Llama 3.1 70B + Eagle3):

    指标投机解码 ON投机解码 OFF差异
    单请求 HBM 占用2.8 GB1.9 GB-32%
    单请求平均加速2.1×1.0×-52%
    集群最大并发数4872+50%

    关键点:关闭投机_decode 降低吞吐但提升并发和稳定性——在过载场景下这是正确的取舍。

    4.3 Level 3:KV Cache Offload

    当 Decode 实例的 KV Cache 总量接近 HBM 上限时,需要将部分 KV Cache offload 到 CPU memory 或 NVMe。

    HBM (80GB) 活跃 KV Cache: ████████████████████░░░░░░░░  60GB / 80GB
                                    ↑ active    ↑ headroom
                                    
    触发 offload 阈值: 85% = 68GB
    需要 offload: 68GB - 60GB + buffer = ~15GB
    
    CPU DRAM (512GB):
      ┌─────────────────────────────────────────┐
      │ Offloaded KV Cache (pinned memory)       │
      │  - Block A: seq_len=4096, 优先级=low    │
      │  - Block B: seq_len=2048, 优先级=low    │
      └─────────────────────────────────────────┘
             ↕ PCIe Gen5 x16 (~64 GB/s)
    class KVCacheOffloadManager:
        """KV Cache 分层卸载管理器"""
        
        def __init__(self, hbm_limit_gb: float, cpu_pool_gb: float):
            self.hbm_limit = int(hbm_limit_gb * 1024 * 1024 * 1024)
            self.cpu_pool = PinnedMemoryPool(size_gb=cpu_pool_gb)
            self.offload_threshold = 0.85
            self.eviction_policy = LRUEviction()
        
        def maybe_offload(self, worker_state: WorkerState):
            hbm_used = worker_state.kv_cache_bytes_used
            
            if hbm_used < self.hbm_limit * self.offload_threshold:
                return  # 无需 offload
            
            target_free = int(self.hbm_limit * (1 - self.offload_threshold))
            to_evict_bytes = hmb_used - self.hbm_limit + target_free
            
            # 选择低优先级、最近未访问的 cache 块
            candidates = self.eviction_policy.select_victims(
                worker_state.active_requests,
                target_bytes=to_evict_bytes,
                exclude_priority=Priority.HIGH  # 不驱逐高优先级
            )
            
            for cache_block in candidates:
                # 异步 offload:GPU → CPU(DMA)
                cpu_slot = self.cpu_pool.allocate(cache_block.size_bytes)
                torch.cuda.current_stream().synchronize()
                cpu_slot.tensor.copy_(cache_block.gpu_tensor, non_blocking=True)
                worker_state.register_offloaded(cache_block, cpu_slot)
                
                logger.info(
                    f"Offloaded req={cache_block.request_id}, "
                    f"layers={cache_block.num_layers}, "
                    f"seq_len={cache_block.seq_len}"
                )
        
        async def on_demand_fetch(self, request_id: str) -> torch.Tensor:
            """访问已 offload 的 KV Cache 时触发回迁"""
            offloaded = self.worker_state.get_offloaded(request_id)
            if not offloaded:
                raise KVCacheNotFoundError(request_id)
            
            # PCIe 回迁延迟: ~80μs/GB @ Gen5 x16
            transfer_time_us = offloaded.size_gb * 80
            
            gpu_tensor = torch.empty(
                offloaded.shape, dtype=offloaded.dtype, 
                device='cuda'
            )
            gpu_tensor.copy_(offloaded.cpu_tensor, non_blocking=True)
            
            logger.info(
                f"Prefetch {request_id}: {offloaded.size_gb:.1f}GB, "
                f"est latency={transfer_time_us:.0f}μs"
            )
            return gpu_tensor

    性能影响实测(Llama 3.1 70B,单 H100 80GB):

    Offload 比例KV Cache 在 HBMKV Cache 在 CPU额外延迟(TPOT)并发提升
    0%45GB0GB0%1.0×
    25%34GB11GB+15-30%1.4×
    50%22GB23GB+40-80%1.8×
    70%14GB31GB+100-200%2.2×

    工程建议:默认允许 25% offload(Level 3 降级),超过 50% 时该请求的 SLO 已难以保证,应优先执行 Level 5 熔断。

    五、故障检测与自动切换机制

    5.1 健康检查的多层设计

    ┌─ Layer 1: 进程级 (1s) ──────────────────────────┐
    │  心跳超时 → 标记为 unhealthy → 剔除流量           │
    │                                                   │
    ├─ Layer 2: GPU 级 (5s) ───────────────────────────┤
    │  ECC 错误计数 / GPU 温度 / HBM 永久故障 Page       │
    │  → 标记 Page 为 poison → 触发 KV Cache 重平衡       │
    │                                                   │
    ├─ Layer 3: 推理级 (per request) ───────────────────┤
    │  推理超时 / CUDA error / OOM → 请求级 failover     │
    │                                                   │
    ├─ Layer 4: 网络级 (50ms) ──────────────────────────┤
    │  RDMA 超时 / link down → 路由切换 + 队列排空       │
    │                                                   │
    └───────────────────────────────────────────────────┘

    5.2 快速故障检测:Delta-based Health Score

    import time
    from collections import deque
    from dataclasses import dataclass, field
    
    @dataclass
    class HealthSnapshot:
        timestamp: float
        tpot_p50_ms: float
        tpot_p99_ms: float
        pending_queue_depth: int
        hbm_used_ratio: float
        rdma_retrans_rate: float  # RDMA 重传率
    
    class DeltaBasedHealthDetector:
        """
        基于变化率的故障检测器——不是看绝对值,而是看
        短期趋势是否恶化
        """
        
        def __init__(self, window_size: int = 20):
            self.snapshots: deque[HealthSnapshot] = deque(maxlen=window_size)
            self.baseline: Optional[HealthSnapshot] = None
        
        def ingest(self, snapshot: HealthSnapshot):
            self.snapshots.append(snapshot)
        
        def compute_health_score(self) -> float:
            """
            返回 0.0 (完全健康) 到 1.0 (严重故障) 的分数
            """
            if len(self.snapshots) < 5:
                return 0.0  # 数据不足,假设健康
            
            recent = list(self.snapshots)[-5:]
            baseline = self.snapshots[0]
            
            # 计算各维度的恶化率
            tpot_trend = (recent[-1].tpot_p99_ms - baseline.totp_p99_ms) / max(baseline.totp_p99_ms, 1)
            queue_trend = (recent[-1].pending_queue_depth - baseline.pending_queue_depth) / max(baseline.pending_queue_depth, 1)
            hbm_trend = (recent[-1].hbm_used_ratio - baseline.hbm_used_ratio)
            rdma_trend = recent[-1].rdma_retrans_rate  # 直接使用最近值
            
            # 加权合成
            score = (
                0.30 * max(0, tpot_trend) +       # TPOT 恶化
                0.25 * max(0, queue_trend) +       # 队列堆积
                0.25 * max(0, hbm_trend / 0.3) +   # HBM 增长 (30% = 完全故障)
                0.20 * max(0, rdma_trend / 0.05)   # RDMA 重传率 > 5% = 完全故障
            )
            
            return min(1.0, score)
        
        def should_migrate(self, threshold: float = 0.6) -> bool:
            """当健康分数超过阈值时触发迁移"""
            return self.compute_health_score() > threshold

    5.3 迁移决策与执行流程

    当 should_migrate() 返回 True 时,执行迁移决策:

    Global Scheduler 决策流程:
                        
      1. 评估需要迁移的请求量
         └─ limit_ratio = health_score - threshold (0.6)
         └─ evict_count = active_requests × limit_ratio
    
      2. 排序候选(优先级 × 最后活跃时间)
         └─ victim = min(requests, key= lambda r: (r.priority.value, r.last_active))
    
      3. 选择目标 Decode 实例(最小负载 + 足够 HBM headroom)
         └─ target = min(decode_workers, key= lambda w: w.load_factor 
                         if w.hbm_available > victim.kv_size * 1.2 else ∞)
    
      4. 执行迁移(section 3.2 的 FaultTolerantKVTransfer)
      
      5. 等待目标确认 + 路由切换
      
      6. 释放源实例的 KV Cache 内存

    六、生产级 Checklist

    6.1 部署前验证清单

    • [ ] KV Cache 传输带宽基准测试:使用 ib_write_bw 验证 RDMA 链路达到线速 80% 以上
    • [ ] 热迁移 TTFT 增量:空载下迁移增加 < 50ms,满载下 < 200ms
    • [ ] 降级策略测试:模拟 KV Cache 超载,验证 Level 2/3 正确触发
    • [ ] 故障注入测试:Kill Prefill/Decode 进程,验证全局调度器在 3s 内完成重调度
    • [ ] PFC Storm 隔离:在 RDMA fabric 上制造拥塞,验证不扩散到 TCP/IP 流量

    6.2 监控仪表板核心指标

    ┌─ Inference SLO Dashboard ─────────────────────────────┐
    │                                                         │
    │  TTFT P50/P99 (目标: <200ms/<500ms)                     │
    │  ══════════════════════════════════▓▓  180ms/420ms     │
    │                                                         │
    │  TPOT P50/P99 (目标: <30ms/<80ms)                       │
    │  ═══════════════════════════════▓  22ms/65ms           │
    │                                                         │
    │  KV Cache 传输带宽: 380 Gbps / 400 Gbps (95%)           │
    │  ████████████████████████████████████████████░░        │
    │                                                         │
    │  降级状态: L0 (normal)  迁移中: 2 请求  降级率: 0.3%     │
    │                                                         │
    │  HBM Utilization:  Prefill 72% | Decode 68%             │
    │  Queue Depth:       Prefill 12    | Decode 8             │
    │                                                         │
    └─────────────────────────────────────────────────────────┘

    七、总结

    LLM 推理系统的韧性工程是一个多层次、交叉领域的系统性挑战:

    1. PD 分离架构 提供了故障隔离的物理边界,但引入了 KV Cache 传输这一新的延迟和可靠性瓶颈
    2. KV Cache 热迁移 不仅仅是"搬运内存"——分层传输、异步流水线、RDMA 降级等技巧共同将迁移延迟降低了一个数量级
    3. 分级降级 是韧性的核心思想:从关闭投机解码到 KV Cache offload 再到服务熔断,每一级都有明确的 SLO 影响和触发条件
    4. Delta-based 健康检测 比阈值告警更早发现趋势性恶化,配合全局调度器实现预测性迁移
    5. 真正的生产系统不会只运行在 Level 0。理解每一级退化的代价,在正确的时间做出正确的取舍——这正是推理系统工程的艺术。

      关键数据点回顾:70B 模型 8K 上下文 KV Cache 约 20GB,分层传输可将 TTFT 影响从 180ms 降低到 35ms,25% CPU offload 提升 40% 并发但增加 30% TPOT 延迟。在这些数字间找到业务最优解,是每个推理基础设施团队的核心工作。
点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部