LLM 推理引擎 Continuous Batching 与 PagedAttention:从 vLLM 原理到生产级部署实战
引言:大模型推理的工程学挑战
大语言模型(LLM)推理服务与训练有本质区别:训练是批量处理可控的数据集,推理则面临请求到达的不可预测性、生成序列长度的不确定性、显存碎片化三大根本挑战。传统的 Static Batching ——将多个请求打包成统一张量——在推理场景下效率极低:当一个请求完成生成后,整个批次必须等待最慢的请求结束才能释放资源。
vLLM 的 PagedAttention 和 Continuous Batching 机制正是解决这些挑战的核心工程创新。它们将操作系统内存管理的分页思想引入 GPU KV Cache 管理,使推理吞吐量提升了 2-4 倍。本文将从原理到实现,完整拆解这两项技术。
一、推理服务的基础架构
1.1 Prefill 与 Decode 两阶段
LLM 推理分为两个截然不同的计算阶段:
Prefill(预填充):一次性处理完整输入 Prompt,复杂度为 O(n²)(n 为 prompt 长度)。这是一个计算密集型(Compute-bound)操作,GPU 利用率高。
Decode(解码):逐 token 自回归生成,每一步仅产生一个 token,复杂度为 O(n)。这是一个内存带宽受限(Memory-bandwidth-bound)操作,瓶颈在于从 HBM 加载 KV Cache。
Prefill 阶段:
Input tokens: [t1, t2, t3, ..., tn]
↓
一次性并行计算,生成完整 KV Cache
↓
产生首 token
Decode 阶段:
Step 1: [t1...tn] + new token → KV Cache [1..n+1] → next token
Step 2: [t1...tn+1] + new token → KV Cache [1..n+2] → next token
...
Step k: ... → EOS token(终止)
1.2 KV Cache 的本质
KV Cache 是 Transformer 自注意力机制的核心优化。在自回归生成中,每个新 token 的自注意力需要与之前所有 token 的 Key 和 Value 进行点积计算。如果不缓存这些向量,每一步都需要重新计算,复杂度将从 O(n) 飙升到 O(n²)。
对于 Llama-2-7B 模型(隐藏维度 4096,32 层,32 个注意力头):
- 每 token 的 KV 大小:2(K+V)× 32 层 × 32 头 × 128 维 × 2 字节(FP16)= 524,288 字节 ≈ 0.5 MB/token
- 对于 4096 token 的序列:4096 × 0.5 MB = 2 GB
这就是为什么 KV Cache 管理是推理引擎的核心问题——单个长序列就能消耗数 GB 的 GPU 显存。
1.3 传统 Static Batching 的问题
Static Batching 的核心缺陷在于「木桶效应」:
Batch 中的三个请求:
Request A: 输入 100 tokens,生成 50 tokens(短请求)
Request B: 输入 200 tokens,生成 500 tokens(长请求)
Request C: 输入 150 tokens,生成 100 tokens(中请求)
┌──────────────────────────────────────────────┐
│█████████ A done ███████│ │ A 完成但资源被占用
│████████████████████████│██████████████████ │ B 还在生成,浪费 GPU
│███████████████████ C done ████ │ C 完成但无法释放
└──────────────────────────────────────────────┘
Time: |-----prefill-----|--------decode---------|
问题清单:
- GPU 空转:最快完成的请求等待最慢请求,期间 GPU 无法处理新请求
- 显存预分配浪费:必须按最大生成长度预分配连续显存块
- 突发流量响应差:批次满了就必须等,排队延迟急剧上升
- 批量大小固定:无法动态调整适应负载变化
- 经验模型:
estimated_gen_len = min(input_len × factor, max_gen_len) - 轻量小模型预测
- 默认:256
- 高吞吐场景:512
- 低延迟场景:64
- 默认:0.90
- 长序列场景:0.95(减少被其他进程抢占的概率)
- 混合场景:0.85(保留余量给其他模型)
- 默认值:通常是 GPU 最大可并行 tokens 数(由 profiling 决定)
- 调大:提高 Prefill 吞吐
- 调小:减少单步延迟
- 长序列高并发时推荐开启
- 会让 TTFT 更平滑,但略微增加总吞吐
- System Prompt 共享场景强烈推荐
- 显存充足时建议开启
- Continuous Batching:迭代级调度,请求进出自如,GPU 利用率从 50% 提升到 90%+
- PagedAttention:虚拟分页管理 KV Cache,显存利用率从 70% 提升到 98%+
- Chunked Prefill:长 Prefill 分块执行,TTFT 平滑,不阻塞 Decode
- Prefix Caching:哈希去重共享 KV Blocks,System Prompt 重复场景节省 50%+ 显存
- 分离式架构:Prefill/Decode 分集群部署,资源成本最优
二、Continuous Batching(迭代级批处理)
2.1 核心思想
Continuous Batching(也叫 Iteration-level Batching)的核心是:在每个 decode step 的边界处进行批次调整。当一个请求完成时,立即释放其资源,并插入新请求。
Continuous Batching 时间线:
Step 1: [A(prefill), B(prefill), C(prefill)] → decode
Step 2: [A(decode), B(decode), C(decode)]
Step 3: [A(done)→D(prefill), B(decode), C(decode)] ← A完成,D立即插入
Step 4: [D(decode), B(decode), C(done)→E(prefill)] ← C完成,E立即插入
Step 5: [D(decode), B(done)→F(prefill), E(prefill)]
...
2.2 实现关键点
Prefill 与 Decode 混合批次:这是精妙之处。新请求的 Prefill(计算密集)可以与在途请求的 Decode(内存带宽密集)混合在同一批次中执行,因为它们的瓶颈不同,互不干扰 GPU 资源。
迭代级调度器:调度器的粒度不是「请求级别」而是「模型前向传播级别」。每次 forward() 调用后,调度器重新评估批次组成:
class ContinuousBatchingScheduler:
def __init__(self, max_batch_size, max_num_tokens):
self.running_queue = [] # 正在 decode 的请求
self.waiting_queue = [] # 等待 prefill 的新请求
self.max_batch_size = max_batch_size
self.max_num_tokens = max_num_tokens
def step(self):
"""每个 decode step 调用一次"""
# 1. 收集已完成的请求
finished = [r for r in self.running_queue if r.is_finished()]
for r in finished:
self._free_kv_cache(r)
self.running_queue.remove(r)
# 2. 从等待队列插入新请求
while (len(self.running_queue) < self.max_batch_size
and self.waiting_queue):
new_req = self.waiting_queue.pop(0)
if self._can_allocate(new_req):
self.running_queue.append(new_req)
else:
# 显存不足,回首继续等待
self.waiting_queue.insert(0, new_req)
break
# 3. 构建批次:混合 prefill 和 decode
batch = self._build_batch()
return batch
def _build_batch(self):
"""构建混合 prefill/decode 批次"""
prefill_reqs = [r for r in self.running_queue if r.needs_prefill]
decode_reqs = [r for r in self.running_queue if not r.needs_prefill]
# 优先插入 prefill,因为 decode 可以延后
batch = []
total_tokens = 0
for r in prefill_reqs:
tokens = len(r.prompt_tokens)
if total_tokens + tokens <= self.max_num_tokens:
batch.append(r)
total_tokens += tokens
for r in decode_reqs:
if len(batch) < self.max_batch_size:
batch.append(r)
total_tokens += 1 # decode 每步只处理1个token
return batch
2.3 TGI 与 vLLM 的 Continuous Batching 对比
| 特性 | HuggingFace TGI | vLLM |
|---|---|---|
| Cross-attention KV 共享 | 不支持 | 支持(Prefix Caching) |
| Chunked Prefill | 支持 | 支持 |
| 显存管理 | 连续预分配 | PagedAttention |
| GPU 利用率 | 中等 | 高 |
| 最大吞吐 | ~3x baseline | ~4x baseline |
三、PagedAttention:GPU 上的虚拟内存分页
3.1 操作系统分页的启示
传统操作系统解决物理内存碎片化的方案是虚拟内存 + 分页:逻辑连续的地址空间被映射到离散的物理页帧。进程无需关心物理布局,MMU 通过页表完成地址转换。
PagedAttention 将这一思想引入 KV Cache 管理:
传统方式(连续分配):
Request A: [████████████____] ← 预分配4096 slots,只用1200
Request B: [████████████████] ← 几乎用满
Request C: [██______________] ← 预分配4096 slots,只用300
显存浪费 ≈ (4096-1200) + (4096-300) × 0.5MB/token = 3.1 GB
PagedAttention(分页分配):
Request A: [███] [███] [██] ← 只分配实际需要的页
Request B: [███] [███] [███] [██] ← 按需增长
Request C: [██] ← 精确分配
显存浪费 ≈ 仅最后一页的部分空间(< 2%)
3.2 核心数据结构
class BlockManager:
"""PagedAttention 块管理器 - 类比 OS 内存管理器"""
def __init__(self, num_gpu_blocks, block_size):
self.block_size = block_size # 每块存储的 token 数(如16)
self.num_gpu_blocks = num_gpu_blocks
# 空闲块列表(类比 OS 空闲页帧链表)
self.free_blocks = list(range(num_gpu_blocks))
# 块表:request_id → 物理块号列表(类比页表)
self.block_tables = {}
def allocate(self, request_id, num_tokens):
"""为请求分配 KV Cache 块"""
num_blocks_needed = ceil(num_tokens / self.block_size)
if len(self.free_blocks) < num_blocks_needed:
return False # OOM
blocks = [self.free_blocks.pop() for _ in range(num_blocks_needed)]
self.block_tables[request_id] = blocks
return True
def append_token(self, request_id):
"""生成新 token 时追加一个 slot"""
blocks = self.block_tables[request_id]
current_len = self._get_seq_len(request_id)
last_block_idx = (current_len - 1) // self.block_size
# 如果当前块未满,直接写入
if current_len % self.block_size != 0:
self._record_token(request_id, current_len)
return True
# 需要分配新块
if not self.free_blocks:
return False # OOM
new_block = self.free_blocks.pop()
blocks.append(new_block)
self._record_token(request_id, current_len)
return True
def free(self, request_id):
"""释放请求的所有块"""
if request_id in self.block_blocks:
self.free_blocks.extend(self.block_tables[request_id])
del self.block_tables[request_id]
def get_block_table(self, request_id):
"""获取物理块地址列表(用于 GPU kernel)"""
return self.block_tables.get(request_id, [])
3.3 GPU Kernel 中的分页读取
PagedAttention 核心是一个 CUDA kernel,它根据 block table 间接寻址读取 KV Cache:
// 简化版 PagedAttention CUDA kernel
// grid: (num_heads, num_seqs)
// block: (num_threads)
__global__ void paged_attention_kernel(
const half* __restrict__ k_cache, // [num_blocks, block_size, head_dim]
const half* __restrict__ v_cache, // [num_blocks, block_size, head_dim]
const int* __restrict__ block_tables, // [max_num_blocks_per_seq]
const int* __restrict__ seq_lens, // [num_seqs]
half* __restrict__ out, // 输出
const int scale,
int num_heads,
int head_dim,
int block_size,
int max_num_blocks_per_seq
) {
int head_idx = blockIdx.x;
int seq_idx = blockIdx.y;
int thread_idx = threadIdx.x;
int seq_len = seq_lens[seq_idx];
extern __shared__ float sram[];
// 计算当前 query token 对应的所有 KV 位置
float max_val = -INFINITY;
float sum_exp = 0.0f;
for (int block_offset = 0; block_offset < seq_len; block_offset += block_size) {
// 通过 block table 间接寻址
int physical_block = block_tables[seq_idx * max_num_blocks_per_seq
+ block_offset / block_size];
int tokens_in_block = min(block_size, seq_len - block_offset);
for (int i = thread_idx; i < tokens_in_block; i += blockDim.x) {
const half* k_ptr = k_cache + physical_block * block_size * head_dim
+ i * head_dim + head_idx * head_dim;
// 计算 attention score: q · k
float score = 0.0f;
for (int d = 0; d < head_dim; d++) {
score += __half2float(q[d]) * __half2float(k_ptr[d]);
}
score *= scale;
// Online softmax
float new_max = fmaxf(max_val, score);
sum_exp = sum_exp * expf(max_val - new_max) + expf(score - new_max);
max_val = new_max;
sram[i] = score; // 保存用于后续加权求和
}
}
// 第二阶段:加权求和 Value
// ...(省略归一化和 V 加权部分)
}
3.4 Block Size 的选择
Block Size 需要在内存碎片和Kernel 效率之间权衡:
| Block Size | 每请求最大浪费 | 间接寻址开销 | 推荐场景 |
|---|---|---|---|
| 1 | 0 tokens/token | 高(每个 token 一次查表) | 极致显存效率 |
| 8 | 7 tokens | 中等 | 短序列为主 |
| 16 | 15 tokens | 低 | 通用默认 |
| 32 | 31 tokens | 很低 | 长序列场景 |
公式:浪费率 = (block_size / 2) / avg_seq_len
对于平均序列长度 2048 tokens、block_size=16 的情况:浪费率 = 8 / 2048 ≈ 0.4%。
四、高级内存管理技术
4.1 KV Cache 的 Shared Prefix(前缀共享)
当多个请求共享相同的 System Prompt 时,可以共享它们的 KV Cache:
System Prompt: "你是一个专业的Python编程助手..."
Request 1: [System + "实现快排"]
Request 2: [System + "实现归并排序"]
PagedAttention 前缀共享:
System KV ───→ [Block_0, Block_1, Block_2] ← 只存一份
├─→ Request 1: [Block_3, Block_4]
└─→ Request 2: [Block_5, Block_6]
在 vLLM 中,这通过全局哈希表实现:每个块的 token 内容被哈希,相同内容的块只存一份,引用计数管理生命周期。
4.2 Automatic Prefix Caching(APC)
Automatic Prefix Caching 不需要显式指定「哪些请求共享前缀」,而是通过 LRU 缓存最近使用的 KV Blocks:
class PrefixCachingBlockManager(BlockManager):
"""带自动前缀缓存的块管理器"""
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
# 哈希表: hash(tokens) → [physical_block_ids]
self.cached_blocks = {}
# LRU 驱逐队列
self.eviction_queue = OrderedDict()
def allocate(self, request_id, tokens):
"""先查找缓存,再分配新块"""
token_hashes = self._compute_hash(tokens)
cached_ranges = self._find_cached_blocks(token_hashes)
if cached_ranges:
# 重用缓存块
for block_ids in cached_ranges:
for bid in block_ids:
self._add_ref(bid)
self.block_tables[request_id] = cached_ranges + new_blocks
else:
# 无缓存,正常分配
super().allocate(request_id, len(tokens))
return True
4.3 Chunked Prefill
当单个请求的 Prompt 很长(如 8K tokens)时,Prefill 会阻塞 Decode 阶段其他请求的处理。Chunked Prefill 将 Prefill 拆分为多个 Chunk,插入 Decode Steps 之间:
传统(长 Prefill 阻塞 Decode):
[Request A: prefill 8000 tokens ][ ][ ][ ][ ]
[Request B: ............waiting................][d][d][d]
Chunked Prefill:
[A: chunk_1 ][B: chunk_1][decode batch][A: chunk_2][B: chunk_2][decode]...
这样 B 无需等 A 完整 prefill 就能开始,延迟(TTFT)更平滑
4.4 Swap:CPU offload
当 GPU 显存极度紧张时,可以将 KV Cache swap 到 CPU 内存:
class SwappingManager:
def __init__(self, num_gpu_blocks, num_cpu_blocks, block_size):
self.gpu_block_mgr = BlockManager(num_gpu_blocks, block_size)
self.cpu_block_mgr = BlockManager(num_cpu_blocks, block_size) # 用CPU内存
self.swap_map = {} # request_id → {gpu_blocks, cpu_blocks}
def swap_out(self, request_id):
"""将不活跃的 KV Cache swap 到 CPU"""
blocks = self.gpu_block_mgr.get_block_table(request_id)
cudaMemcpy(cpu_ptr, gpu_ptr, ...) # D2H
self.gpu_block_mgr.free(request_id)
self.cpu_block_mgr.allocate(request_id, len(blocks) * self.block_size)
def swap_in(self, request_id):
"""将即将活跃的请求 swap 回 GPU"""
# 先 swap 出其他请求腾出空间
self._evict_one_request()
cudaMemcpy(gpu_ptr, cpu_ptr, ...) # H2D
五、调度策略
5.1 First-Come-First-Serve(FCFS)
最基本的调度策略,按请求到达顺序处理。缺点:一个超长请求可能被后面的短请求饿死。
5.2 Shortest-Job-First(SJF)
预估 Token 数量,最短请求优先。预估方法:
5.3 Watermark 调度
vLLM 默认使用 Watermark 策略:
class WatermarkScheduler:
"""基于显存使用水位线的动态调度"""
def schedule(self):
results = []
total_tokens = 0
for request in self.waiting_queue:
# 检查显存使用是否低于高水位线
if self._gpu_memory_usage() > self.high_watermark:
break
needed_blocks = ceil((request.input_len + request.max_gen_len)
/ self.block_size)
# 检查是否有足够空闲块
if len(self.gpu_block_mgr.free_blocks) >= needed_blocks:
total_tokens += request.input_len
if total_tokens > self.max_num_tokens_in_batch:
break
results.append(request)
self.waiting_queue.pop(0)
return results
5.4 Preemption(抢占)
当 GPU 显存耗尽时,需要抢占低优先级请求:
def preempt(self):
"""抢占策略:从运行队列末端腾出空间"""
num_preemptions = 0
while self.gpu_block_mgr.free_blocks == 0 and self.running_queue:
# 策略1:抢占最近加入的(Recompute)
victim = self.running_queue.pop(-1)
self.waiting_queue.insert(0, victim)
self.gpu_block_mgr.free(victim.id)
num_preemptions += 1
# 策略2:Swap to CPU
# self.swap_manager.swap_out(victim.id)
return num_preemptions
六、生产级部署工程实践
6.1 vLLM 部署架构
┌─────────────────────┐
│ Load Balancer │
│ (Nginx/Traefik) │
└────────┬────────────┘
│
┌──────────────┼──────────────┐
│ │ │
┌─────┴─────┐ ┌─────┴─────┐ ┌─────┴─────┐
│ vLLM Node │ │ vLLM Node │ │ vLLM Node │
│ (GPU 0) │ │ (GPU 1) │ │ (GPU 2-3) │
└─────┬─────┘ └─────┬─────┘ └─────┬─────┘
│ │ │
┌─────┴──────────────┴──────────────┴─────┐
│ Prometheus + Grafana │
│ (Monitoring: TTFT/TPOT/Queue/GPU mem) │
└─────────────────────────────────────────┘
6.2 关键性能监控指标
| 指标 | 含义 | 目标值 |
|---|---|---|
| TTFT | Time To First Token(首 token 延迟) | < 200ms |
| TPOT | Time Per Output Token(每 token 生成时间) | < 30ms |
| E2E Latency | 端到端延迟 | 取决于任务 |
| Throughput | 总吞吐 (tokens/sec) | 越高越好 |
| Queue Size | 排队请求数 | 趋近 0 |
| GPU KV Cache Usage | KV Cache 使用率 | < 90% |
| Preemptions/s | 每秒抢占次数 | 趋近 0 |
| TTFT-P99 | 首 token P99 延迟 | < 500ms |
6.3 参数调优指南
max_num_seqs:同时处理的最大序列数。增大可以提高吞吐但增加延迟。
gpu_memory_utilization:KV Cache 占 GPU 显存比例。
max_num_batched_tokens:单次前向传播的最大 tokens 数。
enable_chunked_prefill:是否启用 Chunked Prefill。
enable_prefix_caching:是否启用自动前缀缓存。
6.4 典型部署配置示例
# 单卡 A100-80G 部署 Llama-2-7B
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-chat-hf \
--dtype float16 \
--max-model-len 4096 \
--gpu-memory-utilization 0.92 \
--max-num-seqs 256 \
--max-num-batched-tokens 8192 \
--enable-chunked-prefill \
--enable-prefix-caching \
--port 8000
# 多卡 Tensor Parallel 部署 Llama-2-70B
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-70b-chat-hf \
--tensor-parallel-size 4 \
--dtype float16 \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--enable-chunked-prefill \
--enable-prefix-caching \
--port 8000
6.5 压力测试与容量规划
# 使用 vLLM 自带 benchmark 工具
python benchmarks/benchmark_serving.py \
--model meta-llama/Llama-2-7b-chat-hf \
--dataset ShareGPT_V3_unfiltered_cleaned_split.json \
--request-rate inf \
--num-prompts 1000
# 关键输出示例:
# = Benchmark Results =
# Successful requests: 1000
# Failed requests: 0
# Benchmark duration: 96.21 s
# Total input tokens: 223,417
# Total generated tokens: 198,203
# Request throughput: 10.39 requests/s
# Output token throughput: 2060.18 tokens/s
# Total Token throughput: 4383.53 tokens/s
# ---------------Time to First Token----------------
# Mean TTFT: 98.23 ms
# Median TTFT: 76.41 ms
# P99 TTFT: 312.05 ms
# -----Time per Output Token (excl. 1st token)------
# Mean TPOT: 18.42 ms
# Median TPOT: 16.78 ms
# P99 TPOT: 45.21 ms
# ---------------Inter-token Latency----------------
# Mean ITL: 15.12 ms
# Median ITL: 12.89 ms
# P99 ITL: 67.34 ms
七、前沿演进
7.1 分离式架构(Disaggregated Serving)
传统部署将 Prefill 和 Decode 混在同一个 GPU 上。NVIDIA 和学术界提出的分离式架构将两者拆分到不同 GPU:
Disaggregative Serving 架构:
┌─────────────────────┐ KV Transfer ┌─────────────────────┐
│ Prefill Cluster │ ──────────────→ │ Decode Cluster │
│ (Compute-optimized) │ (RDMA/NVLink) │ (Memory-optimized) │
│ NVIDIA A100 │ │ NVIDIA L40S │
│ GPU HBM 充裕 │ │ 大带宽显存 │
└─────────────────────┘ └─────────────────────┘
优势:
- Prefill 集群以算力为中心,Decode 集群以带宽为中心
- 各取所需,资源成本优化
- 可以独立弹性伸缩
DistServe(OSDI'24)和 Splitwise 展示了这一架构的潜力。vLLO 也在朝这个方向演进。
7.2 Speculative Draft 加速
使用小模型(Draft Model)快速生成候选 tokens,大模型并行验证:
Draft Model (小): 快速生成 γ 个 tokens → [d1, d2, d3, d4, d5]
↓
Target Model (大): 并行验证,接受前缀 → [d1, d2, d3] accepted, d4 rejected
↓
最终输出: [d1, d2, d3, t4'] (3步生成4个有效token)
投机率:accepted / total_draft ≈ 0.7-0.9(取决于大小模型相似度)
加速比:1 / (1 - acceptance_rate) ≈ 2-5x
7.3 Mooncake:月之暗面的 KV Cache 传输引擎
Mooncake(Kimi 推理引擎)将 PagedAttention 思想进一步延伸:将 KV Cache 视为可全局调度的「一等资源」,实现了跨节点 KV Cache 的异步传输与预取,为 Thousand-GPU 级推理集群的 KV Cache 管理奠定基础。
八、总结
PagedAttention 和 Continuous Batching 代表了 AI 推理系统的底层架构革新。它们从操作系统经典理论中汲取灵感,将虚拟内存分页、迭代级调度、大小核分离等思想应用于 GPU 推理场景,解决了大模型服务中显存碎片化和 GPU 利用率低的核心问题。
随着 LLM 从 7B 走向 70B、405B,推理系统也在从单机向多机、从混部向分离式架构演进。理解这些底层机制,是构建下一代 AI 基础设施的必经之路。
核心要点回顾:

发表评论 取消回复