LLM 推理引擎 Continuous Batching 与 PagedAttention:从 vLLM 原理到生产级部署实战

引言:大模型推理的工程学挑战

大语言模型(LLM)推理服务与训练有本质区别:训练是批量处理可控的数据集,推理则面临请求到达的不可预测性、生成序列长度的不确定性、显存碎片化三大根本挑战。传统的 Static Batching ——将多个请求打包成统一张量——在推理场景下效率极低:当一个请求完成生成后,整个批次必须等待最慢的请求结束才能释放资源。

vLLM 的 PagedAttention 和 Continuous Batching 机制正是解决这些挑战的核心工程创新。它们将操作系统内存管理的分页思想引入 GPU KV Cache 管理,使推理吞吐量提升了 2-4 倍。本文将从原理到实现,完整拆解这两项技术。

一、推理服务的基础架构

1.1 Prefill 与 Decode 两阶段

LLM 推理分为两个截然不同的计算阶段:

Prefill(预填充):一次性处理完整输入 Prompt,复杂度为 O(n²)(n 为 prompt 长度)。这是一个计算密集型(Compute-bound)操作,GPU 利用率高。

Decode(解码):逐 token 自回归生成,每一步仅产生一个 token,复杂度为 O(n)。这是一个内存带宽受限(Memory-bandwidth-bound)操作,瓶颈在于从 HBM 加载 KV Cache。

Prefill 阶段:
  Input tokens: [t1, t2, t3, ..., tn]
  ↓
  一次性并行计算,生成完整 KV Cache
  ↓
  产生首 token

Decode 阶段:
  Step 1: [t1...tn] + new token → KV Cache [1..n+1] → next token
  Step 2: [t1...tn+1] + new token → KV Cache [1..n+2] → next token
  ...
  Step k: ... → EOS token(终止)

1.2 KV Cache 的本质

KV Cache 是 Transformer 自注意力机制的核心优化。在自回归生成中,每个新 token 的自注意力需要与之前所有 token 的 Key 和 Value 进行点积计算。如果不缓存这些向量,每一步都需要重新计算,复杂度将从 O(n) 飙升到 O(n²)。

对于 Llama-2-7B 模型(隐藏维度 4096,32 层,32 个注意力头):

  • 每 token 的 KV 大小:2(K+V)× 32 层 × 32 头 × 128 维 × 2 字节(FP16)= 524,288 字节 ≈ 0.5 MB/token
  • 对于 4096 token 的序列:4096 × 0.5 MB = 2 GB

这就是为什么 KV Cache 管理是推理引擎的核心问题——单个长序列就能消耗数 GB 的 GPU 显存。

1.3 传统 Static Batching 的问题

Static Batching 的核心缺陷在于「木桶效应」:

Batch 中的三个请求:
  Request A: 输入 100 tokens,生成 50 tokens(短请求)
  Request B: 输入 200 tokens,生成 500 tokens(长请求)
  Request C: 输入 150 tokens,生成 100 tokens(中请求)
  
  ┌──────────────────────────────────────────────┐
  │█████████ A done ███████│                      │  A 完成但资源被占用
  │████████████████████████│██████████████████     │  B 还在生成,浪费 GPU
  │███████████████████ C done ████                │  C 完成但无法释放
  └──────────────────────────────────────────────┘
  Time:  |-----prefill-----|--------decode---------|

问题清单:

  1. GPU 空转:最快完成的请求等待最慢请求,期间 GPU 无法处理新请求
  2. 显存预分配浪费:必须按最大生成长度预分配连续显存块
  3. 突发流量响应差:批次满了就必须等,排队延迟急剧上升
  4. 批量大小固定:无法动态调整适应负载变化
  5. 二、Continuous Batching(迭代级批处理)

    2.1 核心思想

    Continuous Batching(也叫 Iteration-level Batching)的核心是:在每个 decode step 的边界处进行批次调整。当一个请求完成时,立即释放其资源,并插入新请求。

    Continuous Batching 时间线:
      Step 1: [A(prefill), B(prefill), C(prefill)] → decode
      Step 2: [A(decode),  B(decode),  C(decode)]
      Step 3: [A(done)→D(prefill), B(decode), C(decode)]  ← A完成,D立即插入
      Step 4: [D(decode),   B(decode), C(done)→E(prefill)]  ← C完成,E立即插入
      Step 5: [D(decode),   B(done)→F(prefill), E(prefill)]
      ...

    2.2 实现关键点

    Prefill 与 Decode 混合批次:这是精妙之处。新请求的 Prefill(计算密集)可以与在途请求的 Decode(内存带宽密集)混合在同一批次中执行,因为它们的瓶颈不同,互不干扰 GPU 资源。

    迭代级调度器:调度器的粒度不是「请求级别」而是「模型前向传播级别」。每次 forward() 调用后,调度器重新评估批次组成:

    class ContinuousBatchingScheduler:
        def __init__(self, max_batch_size, max_num_tokens):
            self.running_queue = []  # 正在 decode 的请求
            self.waiting_queue = []  # 等待 prefill 的新请求
            self.max_batch_size = max_batch_size
            self.max_num_tokens = max_num_tokens
        
        def step(self):
            """每个 decode step 调用一次"""
            # 1. 收集已完成的请求
            finished = [r for r in self.running_queue if r.is_finished()]
            for r in finished:
                self._free_kv_cache(r)
                self.running_queue.remove(r)
            
            # 2. 从等待队列插入新请求
            while (len(self.running_queue) < self.max_batch_size 
                   and self.waiting_queue):
                new_req = self.waiting_queue.pop(0)
                if self._can_allocate(new_req):
                    self.running_queue.append(new_req)
                else:
                    # 显存不足,回首继续等待
                    self.waiting_queue.insert(0, new_req)
                    break
            
            # 3. 构建批次:混合 prefill 和 decode
            batch = self._build_batch()
            return batch
        
        def _build_batch(self):
            """构建混合 prefill/decode 批次"""
            prefill_reqs = [r for r in self.running_queue if r.needs_prefill]
            decode_reqs = [r for r in self.running_queue if not r.needs_prefill]
            
            # 优先插入 prefill,因为 decode 可以延后
            batch = []
            total_tokens = 0
            
            for r in prefill_reqs:
                tokens = len(r.prompt_tokens)
                if total_tokens + tokens <= self.max_num_tokens:
                    batch.append(r)
                    total_tokens += tokens
            
            for r in decode_reqs:
                if len(batch) < self.max_batch_size:
                    batch.append(r)
                    total_tokens += 1  # decode 每步只处理1个token
            
            return batch

    2.3 TGI 与 vLLM 的 Continuous Batching 对比

    特性 HuggingFace TGI vLLM
    Cross-attention KV 共享 不支持 支持(Prefix Caching)
    Chunked Prefill 支持 支持
    显存管理 连续预分配 PagedAttention
    GPU 利用率 中等 高
    最大吞吐 ~3x baseline ~4x baseline

    三、PagedAttention:GPU 上的虚拟内存分页

    3.1 操作系统分页的启示

    传统操作系统解决物理内存碎片化的方案是虚拟内存 + 分页:逻辑连续的地址空间被映射到离散的物理页帧。进程无需关心物理布局,MMU 通过页表完成地址转换。

    PagedAttention 将这一思想引入 KV Cache 管理:

    传统方式(连续分配):
      Request A: [████████████____]  ← 预分配4096 slots,只用1200
      Request B: [████████████████]  ← 几乎用满
      Request C: [██______________]  ← 预分配4096 slots,只用300
      
      显存浪费 ≈ (4096-1200) + (4096-300) × 0.5MB/token = 3.1 GB
    
    PagedAttention(分页分配):
      Request A: [███] [███] [██]        ← 只分配实际需要的页
      Request B: [███] [███] [███] [██]  ← 按需增长
      Request C: [██]                     ← 精确分配
      
      显存浪费 ≈ 仅最后一页的部分空间(< 2%)

    3.2 核心数据结构

    class BlockManager:
        """PagedAttention 块管理器 - 类比 OS 内存管理器"""
        
        def __init__(self, num_gpu_blocks, block_size):
            self.block_size = block_size  # 每块存储的 token 数(如16)
            self.num_gpu_blocks = num_gpu_blocks
            # 空闲块列表(类比 OS 空闲页帧链表)
            self.free_blocks = list(range(num_gpu_blocks))
            # 块表:request_id → 物理块号列表(类比页表)
            self.block_tables = {}
        
        def allocate(self, request_id, num_tokens):
            """为请求分配 KV Cache 块"""
            num_blocks_needed = ceil(num_tokens / self.block_size)
            if len(self.free_blocks) < num_blocks_needed:
                return False  # OOM
            blocks = [self.free_blocks.pop() for _ in range(num_blocks_needed)]
            self.block_tables[request_id] = blocks
            return True
        
        def append_token(self, request_id):
            """生成新 token 时追加一个 slot"""
            blocks = self.block_tables[request_id]
            current_len = self._get_seq_len(request_id)
            last_block_idx = (current_len - 1) // self.block_size
            
            # 如果当前块未满,直接写入
            if current_len % self.block_size != 0:
                self._record_token(request_id, current_len)
                return True
            
            # 需要分配新块
            if not self.free_blocks:
                return False  # OOM
            new_block = self.free_blocks.pop()
            blocks.append(new_block)
            self._record_token(request_id, current_len)
            return True
        
        def free(self, request_id):
            """释放请求的所有块"""
            if request_id in self.block_blocks:
                self.free_blocks.extend(self.block_tables[request_id])
                del self.block_tables[request_id]
        
        def get_block_table(self, request_id):
            """获取物理块地址列表(用于 GPU kernel)"""
            return self.block_tables.get(request_id, [])

    3.3 GPU Kernel 中的分页读取

    PagedAttention 核心是一个 CUDA kernel,它根据 block table 间接寻址读取 KV Cache:

    // 简化版 PagedAttention CUDA kernel
    // grid: (num_heads, num_seqs)
    // block: (num_threads)
    
    __global__ void paged_attention_kernel(
        const half* __restrict__ k_cache,     // [num_blocks, block_size, head_dim]
        const half* __restrict__ v_cache,     // [num_blocks, block_size, head_dim]
        const int* __restrict__ block_tables, // [max_num_blocks_per_seq]
        const int* __restrict__ seq_lens,     // [num_seqs]
        half* __restrict__ out,              // 输出
        const int scale,
        int num_heads,
        int head_dim,
        int block_size,
        int max_num_blocks_per_seq
    ) {
        int head_idx = blockIdx.x;
        int seq_idx = blockIdx.y;
        int thread_idx = threadIdx.x;
        
        int seq_len = seq_lens[seq_idx];
        extern __shared__ float sram[];
        
        // 计算当前 query token 对应的所有 KV 位置
        float max_val = -INFINITY;
        float sum_exp = 0.0f;
        
        for (int block_offset = 0; block_offset < seq_len; block_offset += block_size) {
            // 通过 block table 间接寻址
            int physical_block = block_tables[seq_idx * max_num_blocks_per_seq 
                                              + block_offset / block_size];
            
            int tokens_in_block = min(block_size, seq_len - block_offset);
            
            for (int i = thread_idx; i < tokens_in_block; i += blockDim.x) {
                const half* k_ptr = k_cache + physical_block * block_size * head_dim 
                                    + i * head_dim + head_idx * head_dim;
                
                // 计算 attention score: q · k
                float score = 0.0f;
                for (int d = 0; d < head_dim; d++) {
                    score += __half2float(q[d]) * __half2float(k_ptr[d]);
                }
                score *= scale;
                
                // Online softmax
                float new_max = fmaxf(max_val, score);
                sum_exp = sum_exp * expf(max_val - new_max) + expf(score - new_max);
                max_val = new_max;
                sram[i] = score;  // 保存用于后续加权求和
            }
        }
        
        // 第二阶段:加权求和 Value
        // ...(省略归一化和 V 加权部分)
    }

    3.4 Block Size 的选择

    Block Size 需要在内存碎片和Kernel 效率之间权衡:

    Block Size 每请求最大浪费 间接寻址开销 推荐场景
    1 0 tokens/token 高(每个 token 一次查表) 极致显存效率
    8 7 tokens 中等 短序列为主
    16 15 tokens 低 通用默认
    32 31 tokens 很低 长序列场景

    公式:浪费率 = (block_size / 2) / avg_seq_len

    对于平均序列长度 2048 tokens、block_size=16 的情况:浪费率 = 8 / 2048 ≈ 0.4%。

    四、高级内存管理技术

    4.1 KV Cache 的 Shared Prefix(前缀共享)

    当多个请求共享相同的 System Prompt 时,可以共享它们的 KV Cache:

    System Prompt: "你是一个专业的Python编程助手..."
    
    Request 1: [System + "实现快排"]
    Request 2: [System + "实现归并排序"]
    
    PagedAttention 前缀共享:
      System KV ───→ [Block_0, Block_1, Block_2] ← 只存一份
                          ├─→ Request 1: [Block_3, Block_4]
                          └─→ Request 2: [Block_5, Block_6]

    在 vLLM 中,这通过全局哈希表实现:每个块的 token 内容被哈希,相同内容的块只存一份,引用计数管理生命周期。

    4.2 Automatic Prefix Caching(APC)

    Automatic Prefix Caching 不需要显式指定「哪些请求共享前缀」,而是通过 LRU 缓存最近使用的 KV Blocks:

    class PrefixCachingBlockManager(BlockManager):
        """带自动前缀缓存的块管理器"""
        
        def __init__(self, *args, **kwargs):
            super().__init__(*args, **kwargs)
            # 哈希表: hash(tokens) → [physical_block_ids]
            self.cached_blocks = {}
            # LRU 驱逐队列
            self.eviction_queue = OrderedDict()
        
        def allocate(self, request_id, tokens):
            """先查找缓存,再分配新块"""
            token_hashes = self._compute_hash(tokens)
            cached_ranges = self._find_cached_blocks(token_hashes)
            
            if cached_ranges:
                # 重用缓存块
                for block_ids in cached_ranges:
                    for bid in block_ids:
                        self._add_ref(bid)
                self.block_tables[request_id] = cached_ranges + new_blocks
            else:
                # 无缓存,正常分配
                super().allocate(request_id, len(tokens))
            
            return True

    4.3 Chunked Prefill

    当单个请求的 Prompt 很长(如 8K tokens)时,Prefill 会阻塞 Decode 阶段其他请求的处理。Chunked Prefill 将 Prefill 拆分为多个 Chunk,插入 Decode Steps 之间:

    传统(长 Prefill 阻塞 Decode):
      [Request A: prefill 8000 tokens ][  ][  ][  ][  ]
      [Request B: ............waiting................][d][d][d]
      
    Chunked Prefill:
      [A: chunk_1 ][B: chunk_1][decode batch][A: chunk_2][B: chunk_2][decode]...
      
      这样 B 无需等 A 完整 prefill 就能开始,延迟(TTFT)更平滑

    4.4 Swap:CPU offload

    当 GPU 显存极度紧张时,可以将 KV Cache swap 到 CPU 内存:

    class SwappingManager:
        def __init__(self, num_gpu_blocks, num_cpu_blocks, block_size):
            self.gpu_block_mgr = BlockManager(num_gpu_blocks, block_size)
            self.cpu_block_mgr = BlockManager(num_cpu_blocks, block_size)  # 用CPU内存
            self.swap_map = {}  # request_id → {gpu_blocks, cpu_blocks}
        
        def swap_out(self, request_id):
            """将不活跃的 KV Cache swap 到 CPU"""
            blocks = self.gpu_block_mgr.get_block_table(request_id)
            cudaMemcpy(cpu_ptr, gpu_ptr, ...)  # D2H
            self.gpu_block_mgr.free(request_id)
            self.cpu_block_mgr.allocate(request_id, len(blocks) * self.block_size)
        
        def swap_in(self, request_id):
            """将即将活跃的请求 swap 回 GPU"""
            # 先 swap 出其他请求腾出空间
            self._evict_one_request()
            cudaMemcpy(gpu_ptr, cpu_ptr, ...)  # H2D

    五、调度策略

    5.1 First-Come-First-Serve(FCFS)

    最基本的调度策略,按请求到达顺序处理。缺点:一个超长请求可能被后面的短请求饿死。

    5.2 Shortest-Job-First(SJF)

    预估 Token 数量,最短请求优先。预估方法:

    • 经验模型:estimated_gen_len = min(input_len × factor, max_gen_len)
    • 轻量小模型预测

    5.3 Watermark 调度

    vLLM 默认使用 Watermark 策略:

    class WatermarkScheduler:
        """基于显存使用水位线的动态调度"""
        
        def schedule(self):
            results = []
            total_tokens = 0
            
            for request in self.waiting_queue:
                # 检查显存使用是否低于高水位线
                if self._gpu_memory_usage() > self.high_watermark:
                    break
                
                needed_blocks = ceil((request.input_len + request.max_gen_len) 
                                     / self.block_size)
                
                # 检查是否有足够空闲块
                if len(self.gpu_block_mgr.free_blocks) >= needed_blocks:
                    total_tokens += request.input_len
                    if total_tokens > self.max_num_tokens_in_batch:
                        break
                    results.append(request)
                    self.waiting_queue.pop(0)
            
            return results

    5.4 Preemption(抢占)

    当 GPU 显存耗尽时,需要抢占低优先级请求:

    def preempt(self):
        """抢占策略:从运行队列末端腾出空间"""
        num_preemptions = 0
        
        while self.gpu_block_mgr.free_blocks == 0 and self.running_queue:
            # 策略1:抢占最近加入的(Recompute)
            victim = self.running_queue.pop(-1)
            self.waiting_queue.insert(0, victim)
            self.gpu_block_mgr.free(victim.id)
            num_preemptions += 1
            
            # 策略2:Swap to CPU
            # self.swap_manager.swap_out(victim.id)
        
        return num_preemptions

    六、生产级部署工程实践

    6.1 vLLM 部署架构

                        ┌─────────────────────┐
                        │   Load Balancer     │
                        │  (Nginx/Traefik)    │
                        └────────┬────────────┘
                                 │
                  ┌──────────────┼──────────────┐
                  │              │              │
            ┌─────┴─────┐ ┌─────┴─────┐ ┌─────┴─────┐
            │ vLLM Node │ │ vLLM Node │ │ vLLM Node │
            │ (GPU 0)   │ │ (GPU 1)   │ │ (GPU 2-3) │
            └─────┬─────┘ └─────┬─────┘ └─────┬─────┘
                  │              │              │
            ┌─────┴──────────────┴──────────────┴─────┐
            │          Prometheus + Grafana           │
            │  (Monitoring: TTFT/TPOT/Queue/GPU mem) │
            └─────────────────────────────────────────┘

    6.2 关键性能监控指标

    指标 含义 目标值
    TTFT Time To First Token(首 token 延迟) < 200ms
    TPOT Time Per Output Token(每 token 生成时间) < 30ms
    E2E Latency 端到端延迟 取决于任务
    Throughput 总吞吐 (tokens/sec) 越高越好
    Queue Size 排队请求数 趋近 0
    GPU KV Cache Usage KV Cache 使用率 < 90%
    Preemptions/s 每秒抢占次数 趋近 0
    TTFT-P99 首 token P99 延迟 < 500ms

    6.3 参数调优指南

    max_num_seqs:同时处理的最大序列数。增大可以提高吞吐但增加延迟。

    • 默认:256
    • 高吞吐场景:512
    • 低延迟场景:64

    gpu_memory_utilization:KV Cache 占 GPU 显存比例。

    • 默认:0.90
    • 长序列场景:0.95(减少被其他进程抢占的概率)
    • 混合场景:0.85(保留余量给其他模型)

    max_num_batched_tokens:单次前向传播的最大 tokens 数。

    • 默认值:通常是 GPU 最大可并行 tokens 数(由 profiling 决定)
    • 调大:提高 Prefill 吞吐
    • 调小:减少单步延迟

    enable_chunked_prefill:是否启用 Chunked Prefill。

    • 长序列高并发时推荐开启
    • 会让 TTFT 更平滑,但略微增加总吞吐

    enable_prefix_caching:是否启用自动前缀缓存。

    • System Prompt 共享场景强烈推荐
    • 显存充足时建议开启

    6.4 典型部署配置示例

    # 单卡 A100-80G 部署 Llama-2-7B
    python -m vllm.entrypoints.openai.api_server \
        --model meta-llama/Llama-2-7b-chat-hf \
        --dtype float16 \
        --max-model-len 4096 \
        --gpu-memory-utilization 0.92 \
        --max-num-seqs 256 \
        --max-num-batched-tokens 8192 \
        --enable-chunked-prefill \
        --enable-prefix-caching \
        --port 8000
    
    # 多卡 Tensor Parallel 部署 Llama-2-70B
    python -m vllm.entrypoints.openai.api_server \
        --model meta-llama/Llama-2-70b-chat-hf \
        --tensor-parallel-size 4 \
        --dtype float16 \
        --max-model-len 8192 \
        --gpu-memory-utilization 0.90 \
        --enable-chunked-prefill \
        --enable-prefix-caching \
        --port 8000

    6.5 压力测试与容量规划

    # 使用 vLLM 自带 benchmark 工具
    python benchmarks/benchmark_serving.py \
        --model meta-llama/Llama-2-7b-chat-hf \
        --dataset ShareGPT_V3_unfiltered_cleaned_split.json \
        --request-rate inf \
        --num-prompts 1000
    
    # 关键输出示例:
    # = Benchmark Results =
    # Successful requests:                               1000
    # Failed requests:                                   0
    # Benchmark duration:                                96.21 s
    # Total input tokens:                                223,417
    # Total generated tokens:                            198,203
    # Request throughput:                                10.39 requests/s
    # Output token throughput:                           2060.18 tokens/s
    # Total Token throughput:                           4383.53 tokens/s
    # ---------------Time to First Token----------------
    # Mean TTFT:                                         98.23 ms
    # Median TTFT:                                      76.41 ms
    # P99 TTFT:                                        312.05 ms
    # -----Time per Output Token (excl. 1st token)------
    # Mean TPOT:                                        18.42 ms
    # Median TPOT:                                      16.78 ms
    # P99 TPOT:                                         45.21 ms
    # ---------------Inter-token Latency----------------
    # Mean ITL:                                         15.12 ms
    # Median ITL:                                       12.89 ms
    # P99 ITL:                                          67.34 ms

    七、前沿演进

    7.1 分离式架构(Disaggregated Serving)

    传统部署将 Prefill 和 Decode 混在同一个 GPU 上。NVIDIA 和学术界提出的分离式架构将两者拆分到不同 GPU:

    Disaggregative Serving 架构:
    ┌─────────────────────┐    KV Transfer    ┌─────────────────────┐
    │  Prefill Cluster     │ ──────────────→   │  Decode Cluster      │
    │ (Compute-optimized)  │   (RDMA/NVLink)   │  (Memory-optimized)  │
    │  NVIDIA A100        │                   │  NVIDIA L40S         │
    │  GPU HBM 充裕       │                   │  大带宽显存          │
    └─────────────────────┘                   └─────────────────────┘
    
    优势:
    - Prefill 集群以算力为中心,Decode 集群以带宽为中心
    - 各取所需,资源成本优化
    - 可以独立弹性伸缩

    DistServe(OSDI'24)和 Splitwise 展示了这一架构的潜力。vLLO 也在朝这个方向演进。

    7.2 Speculative Draft 加速

    使用小模型(Draft Model)快速生成候选 tokens,大模型并行验证:

    Draft Model (小): 快速生成 γ 个 tokens → [d1, d2, d3, d4, d5]
                                             ↓
    Target Model (大): 并行验证,接受前缀 → [d1, d2, d3] accepted, d4 rejected
                                             ↓
    最终输出: [d1, d2, d3, t4']  (3步生成4个有效token)
    
    投机率:accepted / total_draft ≈ 0.7-0.9(取决于大小模型相似度)
    加速比:1 / (1 - acceptance_rate) ≈ 2-5x

    7.3 Mooncake:月之暗面的 KV Cache 传输引擎

    Mooncake(Kimi 推理引擎)将 PagedAttention 思想进一步延伸:将 KV Cache 视为可全局调度的「一等资源」,实现了跨节点 KV Cache 的异步传输与预取,为 Thousand-GPU 级推理集群的 KV Cache 管理奠定基础。

    八、总结

    PagedAttention 和 Continuous Batching 代表了 AI 推理系统的底层架构革新。它们从操作系统经典理论中汲取灵感,将虚拟内存分页、迭代级调度、大小核分离等思想应用于 GPU 推理场景,解决了大模型服务中显存碎片化和 GPU 利用率低的核心问题。

    随着 LLM 从 7B 走向 70B、405B,推理系统也在从单机向多机、从混部向分离式架构演进。理解这些底层机制,是构建下一代 AI 基础设施的必经之路。

    核心要点回顾:

    • Continuous Batching:迭代级调度,请求进出自如,GPU 利用率从 50% 提升到 90%+
    • PagedAttention:虚拟分页管理 KV Cache,显存利用率从 70% 提升到 98%+
    • Chunked Prefill:长 Prefill 分块执行,TTFT 平滑,不阻塞 Decode
    • Prefix Caching:哈希去重共享 KV Blocks,System Prompt 重复场景节省 50%+ 显存
    • 分离式架构:Prefill/Decode 分集群部署,资源成本最优
点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿
网站二维码

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部
/* 跳过导航链接 (无障碍) */ position: absolute; top: -100px; left: 15px; z-index: 99999; padding: 8px 16px; background: #007bff; color: #fff; font-size: 14px; border-radius: 0 0 4px 4px; text-decoration: none; transition: top 0.2s; } top: 0; outline: 3px solid #0056b3; }