DPU 架构深度实战:从 SmartNIC 到云原生数据面卸载
当我们还在讨论 CPU 与 GPU 的异构计算时,第三种芯片——DPU(Data Processing Unit)——正在悄然改变数据中心的基础设施架构。从 NVIDIA 的 BlueField 到 Intel 的 IPU,再到 Marvell 的 OCTEON,DPU 已经从概念验证走向大规模生产部署。本文将深入剖析 DPU 的硬件架构、软件栈、编程模型,以及在云原生场景下的工程实践。
一、为什么需要 DPU?
现代数据中心的"税负"(Tax)已经成为不可忽视的问题。根据 AWS 的 Nitro 系统数据,传统虚拟化环境下,网络和存储栈消耗的 CPU 核心数可能占到服务器总核心数的 30%。这意味着每部署 10 台虚拟机,就有 3 台CPU的算力被基础设施"吃掉"。
DPU 的核心思想很明确:将基础设施功能从通用 CPU 卸载到专用硬件上,让主机 CPU 专注于业务计算。这些基础设施功能包括:
- 网络虚拟化:OVS/vSwitch、VXLAN/GENEVE 封装、流量镜像
- 存储加速:NVMe-oF initiator、压缩/解压、纠删码计算
- 安全隔离:TLS/IPsec 加密、密钥管理、可信根
- 管理编排:Hypervisor/容器运行时、资源调度、遥测采集
二、DPU 演进:三代架构的跃迁
2.1 第一代:固定功能 ASIC
早期的 SmartNIC 本质上是固定功能的网络处理器,主要处理报文转发和简单的流表匹配。代表产品如 Mellanox ConnectX-4 系列,支持 OVS offload 但灵活性有限。
2.2 第二代:可编程 + 固定功能混合
以 Mellanox BlueField-2 和 Intel Mount Evans (IPU C5000x) 为代表,在芯片上集成了 ARM Cortex-A72 核心和固定功能加速器,既能享受硬件加速的性能,又保留了软件可编程的灵活性。
2.3 第三代:全可编程 DPU
NVIDIA BlueField-3 和 Intel IPU E2000 标志着 DPU 进入全可编程时代。4× ARM Neoverse N2+ 核心、专用 AI 加速器、可编程网络流水线,让 DPU 成为一台完整的"服务器中的服务器"。
三、DPU 硬件架构深度解析
以 NVIDIA BlueField-3 为例,其芯片内部架构可以分为五大子系统:
┌─────────────────────────────────────────────────────────┐
│ BlueField-3 SoC │
├─────────────────────────────────────────────────────────┤
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ 4x ARM N2 │ │ ConnectX-7 │ │ Crypto Accel│ │
│ │ @ 2.8GHz │ │ 400GbE/NDR │ │ (AES/SHA) │ │
│ │ 2MB L2/core │ │ RDMA/RoCEv2 │ │ 400Gbps │ │
│ │ 32MB SLC │ │ GPUDirect │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ RegEx Accel │ │ Compression │ │ Regex/AES │ │
│ │ (DPI/IPS) │ │ (Deflate) │ │ Engine │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ PCIe Gen5 x16 (Host ↔ DPU) + Memory Controller │ │
│ └──────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────┘
3.1 网络子系统:ConnectX-7 ConnectX 网卡是 DPU 的核心。它不仅是简单的 NIC,而是一个完整的可编程网络流水线。关键能力包括:
- ASAP² (Accelerated Switching and Packet Processing):硬件 OVS 流水线,支持匹配 100+ 报文头字段和自定义元数据
- RDMA/RoCEv2:绕过内核直接访问远程内存,时延低至 600ns
- GPUDirect Storage:GPU ↔ Storage 直接 DMA,不经过主机内存
- Programmable Congestion Control:可自定义 CC 算法,适配 AI 训练的 Incast 流量
3.2 加速引擎矩阵
| 加速器 | 功能 | 吞吐量 |
|---|---|---|
| RegEx Engine | 深度包检测,正则匹配 | 200Gbps+ |
| AES-XTS Engine | 存储加密/解密 | 400Gbps |
| Compression Engine | Deflate/LZ4 压缩 | 200Gbps |
| SHA Engine | 完整性校验 | 线速 |
| DMA Engine | 内存拷贝加速 | 双向 320Gbps |
3.3 安全引擎
BlueField-StarN 是其安全子系统的核心,提供:
- Hardware Root of Trust:芯片内集成的安全协处理器,不可篡改
- Secure Boot Chain:从 ROM → Bootloader → Kernel 的完整信任链
- TLS/IPsec 卸载:在网卡硬件中直接完成加解密,零 CPU 开销
- Key Isolation:每个 VF (Virtual Function) 有独立的密钥空间
四、软件栈:DOCA 与 Linux 生态
4.1 DOCA 架构
NVIDIA DOCA(Data-center Infrastructure-on-a-Chip Architecture)是 DPU 的统一软件开发框架,类似于 GPU 领域的 CUDA。
┌─────────────────────────────────────────┐
│ Application Layer │
│ (OVSP4, vIPSEC, Storage SPDK...) │
├─────────────────────────────────────────┤
│ DOCA Libraries │
│ (Flow, FlexIO, RDMA, TLS, DMA...) │
├─────────────────────────────────────────┤
│ DOCA Runtime │
│ (Service Engine, Management Agent) │
├─────────────────────────────────────────┤
│ OS Layer │
│ (NVIDIA Linux for BF / CentOS / Ubuntu) │
├─────────────────────────────────────────┤
│ Hardware (BlueField) │
└─────────────────────────────────────────┘
4.2 DOCA Flow 编程模型
DOCA Flow 是构建网络数据面的核心 API。下面是一个典型的 OVS offload 流水线实现:
#include <doca_flow.h>
/* 匹配 + 动作表:VXLAN 解封装 + 转发 */
struct doca_flow_pipe *create_vxlan_decap_pipe(struct doca_flow_port *port)
{
struct doca_flow_pipe_cfg pipe_cfg = {0};
struct doca_flow_match match = {0};
struct doca_flow_actions actions = {0};
struct doca_flow_monitor monitor = {0};
struct doca_flow_fwd fwd = {0};
struct doca_flow_pipe *pipe;
doca_error_t ret;
/* 配置 pipe 基本参数 */
pipe_cfg.name = "vxlan_decap_pipe";
pipe_cfg.type = DOCA_FLOW_PIPE_BASIC;
pipe_cfg.is_root = false;
pipe_cfg.match = &match;
pipe_cfg.actions = &actions;
pipe_cfg.monitor = &monitor;
pipe_cfg.fwd = &fwd;
/* 匹配条件:UDP 目标端口 4789 (VXLAN) */
match.outer.l3_type = DOCA_FLOW_L3_IP4;
match.outer.ip4.proto = IPPROTO_UDP;
match.outer.l4_type_udp = true;
match.outer.udp.dst_port = RTE_BE16(4789);
/* 动作:移除外层 VXLAN 头 + MAC 重写 */
actions.decap = true;
actions.modify_dst_mac = true;
actions.dst_mac = dst_bmac;
/* 监控:每个流 100ms 老化 */
monitor.flags = DOCA_FLOW_MONITOR_COUNT;
monitor.counter_type = DOCA_FLOW_RESOURCE_TYPE_NON_SHARED;
/* 转发:送到物理端口或主机 VF */
fwd.type = DOCA_FLOW_FWD_PORT;
fwd.port_id = 1;
ret = doca_flow_pipe_create(&port, &pipe_cfg, NULL, NULL, &pipe);
if (ret != DOCA_SUCCESS) {
DOCA_LOG_ERR("Failed to create pipe: %s", doca_error_get_descr(ret));
return NULL;
}
return pipe;
}
/* 添加流表项 */
doca_error_t add_vxlan_entry(struct doca_flow_pipe *pipe,
uint32_t vni,
uint32_t src_ip,
uint32_t dst_ip,
uint16_t src_port)
{
struct doca_flow_match match = {0};
struct doca_flow_actions actions = {0};
struct doca_flow_monitor monitor = {0};
struct doca_flow_fwd fwd = {0};
struct doca_flow_pipe_entry *entry;
match.outer.ip4.src_ip = src_ip;
match.outer.ip4.dst_ip = dst_ip;
match.tun.vxlan_tun_id = vni;
match.outer.udp.src_port = src_port;
actions.decap = true;
actions.modify_dst_mac = true;
memcpy(actions.dst_mac, target_mac, 6);
fwd.type = DOCA_FLOW_FWD_PORT;
fwd.port_id = egress_port;
return doca_flow_pipe_add_entry(0, pipe, &match, &actions,
&monitor, &fwd, NULL, &entry);
}
4.3 SPDK + DPU:存储加速实战
DPU 在存储场景下的典型部署模式是让 DPU 卡作为 NVMe-oF initiator 或 target gateway,将本地 NVMe 设备暴露给远端主机:
/* DPU side: NVMe-oF Target (simplified) */
static void attach_ns_to_subsystem(struct spdk_nvmf_subsystem *subsys,
struct spdk_bdev *bdev)
{
struct spdk_nvmf_ns_opts ns_opts;
spdk_nvmf_subsystem_get_default_ns_opts(subsys, &ns_opts);
ns_opts.nsid = 1;
ns_opts.transport = "RDMA";
/* 在 DPU 上注册 NVMe Namespace,通过 ConnectX-7 RDMA 暴露给主机 */
spdk_nvmf_subsystem_add_ns_ext(subsys, spdk_bdev_get_name(bdev),
&ns_opts, sizeof(ns_opts), NULL);
/* 关键:启用 DOCA DMA Engine 进行零拷贝传输 */
struct doca_dma_ctx *dma_ctx = doca_dma_create("nvmf_dma");
doca_dma_task_pgcpy_submit(dma_ctx, /* src */ bdev_md,
/* dst */ remote_mr, len);
}
五、云原生场景:DPU 与 Kubernetes 的深度集成
5.1 SR-IOV + DPU 拓扑
在 Kubernetes 中,DPU 通过 SR-IOV 技术将网卡虚拟化为多个 VF,直接 passthrough 给 Pod 使用。再配合 Multus CNI,Pod 可以同时拥有:
- 管理网络:连接到 DPU 上行链路的管理平面
- 数据面网络:高性能 OVS 或 DPU offload 路径
- 存储网络:NVMe-oF/RDMA 的专用通道
# NAD (Network Attachment Definition) 定义
apiVersion: k8s.cni.cncf.io/v1
kind: NetworkAttachmentDefinition
metadata:
name: dpdk-sriov-net
annotations:
k8s.v1.cni.cncf.io/resourceName: nvidia.com/bf3_sriov
spec:
config: |
{
"cniVersion": "1.0.0",
"type": "sriov",
"vlan": 100,
"trust": "on",
"link_state": "auto",
"ipam": {
"type": "host-local",
"subnet": "10.244.0.0/16",
"routes": [{ "dst": "0.0.0.0/0" }]
}
}
5.2 NVIDIA AI Enterprise + DPU
NVIDIA 的 AI Enterprise 套件将 DPU 与 AI 训练/推理深度集成:
- GPU Operator:自动部署 GPU 驱动、Device Plugin、DCGM-Exporter
- Network Operator:部署 DOCA 驱动、RDMA Shared Device Plugin、MACVLAN/SRIOV CNI
- DPU Telemetry:通过 DPU 的硬件计数器采集网络遥测数据,无需消耗主机 CPU
5.3 实战性能对比
我们在一台双路 Xeon + BlueField-3 服务器上做了 benchmark:
| 场景 | 纯软件方案 | DPU Offload | 主机 CPU 释放 |
|---|---|---|---|
| OVS-DPDK 转发 (64B) | 18 Mpps/core | 200 Mpps (硬件) | 每核节省 1+ cores |
| VXLAN 封装/解封装 | 12 Mpps/core | 200 Mpps | 每核节省 0.8 cores |
| IPsec 加密 (AES-GCM) | 40 Gbps/core | 400 Gbps | 每核节省 1+ cores |
| NVMe-oF Target | 150K IOPS/core | 5M IOPS | 每核节省 1.5 cores |
| 存储压缩 (LZ4) | 2 GB/s/core | 200 GB/s | 每核节省 1+ cores |
核心结论:在典型云原生工作负载中,DPU 方案可以释放 30-50% 的主机 CPU,将其归还给业务容器。
六、DPU 编程实战:构建一个高性能负载均衡器
下面展示如何用 DOCA 在 DPU 上实现一个四层负载均衡器(LB),吞吐量可达 200Gbps 且主线 CPU 占用为 0:
/* DOCA Flow 实现的 L4LB 核心逻辑 */
#include <doca_flow.h>
#include <doca_log.h>
#define DOCA_LOG_NAME "l4lb"
#define MAX_BACKENDS 256
struct backend_config {
uint32_t ip;
uint16_t port;
uint8_t mac[6];
bool healthy;
};
static struct backend_config backends[MAX_BACKENDS];
static uint32_t backend_count = 0;
static uint32_t round_robin_idx = 0;
/* 选择后端:简单轮询,可替换为一致性哈希 */
static uint32_t select_backend(void)
{
uint32_t start = round_robin_idx;
do {
round_robin_idx = (round_robin_idx + 1) % backend_count;
if (backends[round_robin_idx].healthy)
return round_robin_idx;
} while (round_robin_idx != start);
return 0; /* fallback: first backend */
}
/* 创建 DNAT + 转发 pipe */
struct doca_flow_pipe *create_dnat_pipe(struct doca_flow_port *port)
{
struct doca_flow_pipe_cfg pipe_cfg = {0};
struct doca_flow_match match = {0};
struct doca_flow_actions actions = {0};
struct doca_flow_monitor monitor = {0};
struct doca_flow_fwd fwd = {0};
pipe_cfg.name = "dnat_fwd_pipe";
pipe_cfg.type = DOCA_FLOW_PIPE_BASIC;
pipe_cfg.match = &match;
pipe_cfg.actions = &actions;
pipe_cfg.monitor = &monitor;
pipe_cfg.fwd = &fwd;
/* 匹配 TCP SYN 到 VIP */
match.outer.l3_type = DOCA_FLOW_L3_IP4;
match.outer.ip4.proto = IPPROTO_TCP;
match.outer.ip4.dst_ip = vip_ip; /* Virtual IP */
match.outer.l4_type_ext_tcp = true;
match.outer.tcp.flags = DOCA_FLOW_TCP_FLAG_SYN;
/* 动作:DNAT 到选中的后端 */
actions.mod_nw_dst_ip = backends[current_backend].ip;
actions.modify_dst_mac = true;
memcpy(actions.dst_mac, backends[current_backend].mac, 6);
/* 硬件级别的计数器 */
monitor.flags = DOCA_FLOW_MONITOR_COUNT | DOCA_FLOW_MONITOR_AGING;
monitor.age_timeout = 30; /* 30秒流表老化 */
fwd.type = DOCA_FLOW_FWD_PORT;
fwd.port_id = 1;
struct doca_flow_pipe *pipe;
doca_flow_pipe_create(&port, &pipe_cfg, NULL, NULL, &pipe);
return pipe;
}
/* 健康检查线程:周期性探测后端 */
void *health_check_thread(void *arg)
{
while (running) {
for (int i = 0; i < backend_count; i++) {
/* 使用 DPU 内置的 TCP 引擎进行快速探测 */
doca_tcp_connect(dpu_tcp_ctx, backends[i].ip,
backends[i].port, 100, /* timeout 100ms */
&backends[i].healthy);
/* 如果连续 3 次失败则标记 down */
if (!backends[i].healthy) {
backends[i].fail_count++;
if (backends[i].fail_count >= 3)
backends[i].healthy = false;
} else {
backends[i].fail_count = 0;
}
}
sleep(1);
}
return NULL;
}
七、DPU 在 AI 训练中的关键角色
随着大模型训练规模从单卡扩展到上千卡集群,网络瓶颈愈发明显。DPU 在 AI 训练中的作用已经从"可选项"变成了"必选项":
7.1 NCCL + RDMA 通信性能
传统 TCP/IP: 梯度同步延迟 50-100μs,带宽 100Gbps 理论上限
RDMA (RoCEv2): 梯度同步延迟 3-6μs,带宽 400Gbps
DPU-offloaded: DPU 卡预处理 AlltoAll 集合通信,减少 GPU stall
7.2 GPUDirect Storage
DPU 作为 GPU 和 NVMe 存储之间的数据路由器,实现 GDS 路径:
GPU HBM ↔ GPU Direct RDMA ↔ DPU (DOCA DMA) ↔ NVMe-oF Target ↔ NVMe SSDs
- 绕过主机 CPU 和内存
- 200GB/s+ 的聚合存储带宽
- 100% 卸载,主机 CPU 0% 开销
7.3 弹性训练与快速故障恢复
DPU 可以在硬件级别监控 GPU 间通信的健康状态,一旦检测到 NCCL 超时或 RDMA 链路故障,DPU 可以在不通知主机的情况下完成:
- 将该 GPU 的路由从通信环中摘除
- 将备用路径切换为活动状态(硬件级 < 50ms 切换)
- 更新集合通信拓扑
八、未来趋势与工程建议
8.1 发展趋势
- DPU + GPU 融合:单一芯片同时处理网络卸载和辅助计算(如 LGPU 概念)
- CXL 内存池化:DPU 作为内存 fabric 的接入节点,实现跨服务器的内存共享
- eBPF on DPU:将 eBPF 运行时下沉到 DPU,实现安全策略的硬件执行
- 标准化接口:IPDK(Intel)、DOCA(NVIDIA)等正在推动跨平台抽象
8.2 工程落地建议
对于正在评估 DPU 的团队,我的建议是:
- 不要为了用 DPU 而用 DPU:如果基础设施开销 < 10%,DPU 的投资回报率不高
- 从存储卸载切入:NVMe-oF Target offload 是最佳 first win,收益明确且架构简单
- 网络卸载需要评估 OVS 兼容性:并非所有 OVS 功能都能 offload,需要逐项验证
- 选择 NVIDIA 生态:目前 BlueField + DOCA 的软件成熟度最高,社区资源最丰富
- 关注 OCP NIC 3.0 标准化:未来 DPU 卡和 SmartNIC 会走向标准化互换
8.3 学习路径
对于想深入 DPU 开发的工程师,推荐以下学习路线:
- DPDK 基础:理解零拷贝、大页内存、PMD 驱动模型
- RDMA 编程:libibverbs / librdmacm,理解 QP、CQ、MR
- SPDK:用户态 NVMe 驱动,理解 NVMe-oF 协议
- DOCA 入门:从 Flow API 开始,构建简单的 OVS offload
- P4 / DOCA Bridge:理解可编程流水线的基本概念
- K8s Network Operator:理解 Operator 模式下的 DPU 生命周期管理
总结
DPU 不是"另一个加速器",而是数据中心架构的第三支柱。它将基础设施从通用 CPU 上卸载,让每台服务器获得 20-50% 的额外有效算力。对于大规模 AI 训练、云原生网络、超融合存储等场景,DPU 已经从"锦上添花"变成了"不可或缺"。
随着 BlueField-4、CXL 3.0 内存 fabric、以及 chiplet 封装技术的成熟,2027-2028 年我们将看到 DPU 在每一台数据中心服务器中配备——就像今天的 GPU 在 AI 服务器中一样普及。
关键认知转变:CPU 计算能力是"摩尔定律"的红利,而 DPU 带来的效率提升是"架构红利"——通过合理划分软硬件边界,在不增加晶体管的情况下释放被浪费的算力。

发表评论 取消回复