一、调度器架构全景:不只是"找一个 Node 跑 Pod"

Kubernetes Scheduler(kube-scheduler)是整个集群的"调度大脑",负责将新创建的 Pod 分配到合适的 Node 上运行。看似简单的"匹配"背后,是一套精密的多阶段过滤-评分-绑定流程。

调度器的核心工作流程可以分为四个阶段:

1. 调度队列(Scheduling Queue)

所有待调度的 Pod 首先进入 schedulingQueue,队列分为三个子队列:activeQ(就绪待调度)、unschedulableQ(暂时无法调度)、backoffQ(调度失败退避中)。调度器使用优先队列(Heap)实现,Pod 的 PriorityClass 直接影响出队顺序。

2. 调度循环(Scheduling Cycle)

调度器以 Goroutine 并发执行多个调度循环,但同一时刻对同一个 Pod 只运行一个循环。核心循环通过 schedulingCycle() 调度框架的 Plugins 执行扩展点逻辑。

3. 绑定循环(Binding Cycle)

当选定目标 Node 后,绑定操作通过异步 Goroutine 执行 bindingCycle(),向 API Server 发送 Binding 对象完成绑定。此阶段也支持 PreBind、PostBind、Permit 等绑定周期插件。

4. 缓存与快照(NodeInfo Snapshot)

调度器不会每次去 API Server 拉取集群状态,而是维护一份内存快照 NodeInfo Snapshot。这份快照在每轮调度循环开始时通过 Dump() 获取,保证单次调度周期内视图一致性。

二、调度队列源码解析:activeQ / backoffQ / unschedulableQ

调度队列的实现位于 pkg/scheduler/internal/queue/scheduling_queue.go,核心结构体如下:

type SchedulingQueue struct {
    // 禁止标志位中的 Pod 调度
    stopCond   atomic.Value
    clock      clock.Clock

    // 活跃队列 - 可立即调度的 Pod
    activeQ *heap.Heap

    // 退避队列 - 调度失败的 Pod 等待重试
    podBackoffQ *heap.Heap

    // 不可调度队列 - 暂时无法满足资源需求的 Pod
    unschedulablePods *UnschedulablePodsMap
}

退避策略:Pod 首次调度失败后,backoff 时间从 1 秒开始指数增长(2s, 4s, 8s...上限 60s)。带有 spec.schedulerName 的 Pod 在 backoff 时间到期后会被移动到 activeQ 重新尝试。

不可调度 Pod 的激活:两种方式触发重新调度——(1) 集群状态变更事件(如 Node 资源释放)触发 MoveAllToActiveOrBackoffQueue();(2) 插件级别的 PodBackoffEvents 白名单机制,允许插件注册特定事件触发唤醒。

UnschedulablePodsMap 的内部实现:使用 map[string]*framework.QueuedPodInfo + 最小堆(按 Pod 的 Timestamp 排序),保证最早进入的 Pod 优先被激活。

三、Predicate 预选算法:7 大过滤器的底层原理

Predicate(预选)阶段的核心目标是"过滤掉不满足硬性条件的 Node"。每个 Predicate 插件返回 Success 或具体错误原因。以下是 Kubernetes 内置的关键 Predicate:

1. NodeResourcesFit(资源匹配)

func (f *NodeResourcesFit) Filter(ctx context.Context,
    state *framework.CycleState, pod *v1.Pod,
    nodeInfo *framework.NodeInfo) *framework.Status {

    // 1. 检查 Node 是否有 Pod 的 NodeSelector 匹配
    // 2. 检查 Node 的 Allocatable 资源是否 ≥ Pod 的 Requests 总和
    // 3. 检查 hugepage 资源是否满足
    // 4. 检查 volumes 拓扑约束(VolumeLimits)

    insufficientResources := fitsRequest(pod, nodeInfo)
    if len(insufficientResources) != 0 {
        return framework.NewStatus(
            framework.Unschedulable,
            fmt.Sprintf("Insufficient %v", insufficientResources),
        )
    }
    return nil
}

这里的计算逻辑需要注意:Allocatable ≠ Capacity。Allocatable = Capacity - KubeReserved - SystemReserved - HardEvictionThreshold,确保系统组件和 kubelet 自身有足够的预留资源。

2. NodeAffinity(节点亲和性)

实现基于 NodeSelector 的精确匹配和 PreferredDuringSchedulingIgnoredDuringExecution 的加权评分。注意 Predicate 阶段只处理 requiredDuringSchedulingIgnoredDuringExecution(硬性要求),preferred 部分放到 Score 阶段处理。

3. PodAffinity / PodAntiAffinity(Pod 亲和与反亲和)

这是最复杂的 Predicate 之一,需要遍历所有已调度 Pod,检查与当前 Pod 的标签匹配关系。Kubernetes 使用 topology domains(如 zone、host)做匹配计数。原生实现是 O(n²) 复杂度,在大规模集群中可能成为瓶颈,因此提供了 labelSelector manager 和 taints manager 做索引加速。

4. TaintToleration(污点容忍)

通过遍历 Node 的所有 Taint,检查 Pod 的 Toleration 是否匹配。支持 Equal 和 Exists 两种 Operator。带有 NoExecute 效果的 Taint 还会驱逐已运行的 Pod。

5. 其他 Predicate 插件:

  • NodeUnschedulable:过滤掉设置了 spec.unschedulable=true 的 Node
  • NodeVolumeLimitsCSI:CSI 驱动定义的单节点最大 PV 数量
  • VolumeBinding:检查 PVC 是否可绑定到对应 Node 对应的 Zone
  • VolumeZone:确保 Volume 和 Node 在同一个 Zone
  • NodePorts:检查 Node 上的 HostPort 是否已被占用
  • PodTopologySpread(预选阶段):当新 Pod 会导致拓扑分布不平衡到不可接受程度时过滤

四、Score 评分算法:如何从"可用"中找到"最优"

Score 阶段通过 0-100 的评分系统对所有通过 Predicate 的 Node 进行排序。关键评分插件:

1. NodeResourcesBalancedAllocation(均衡分配)

func balancedResourceScorer(requestedToCapacityRatio resourceToValueMap,
    resourceToWeightMap resourceToWeightMap) scoreNodeFn {
    return func(node string, requestedToCapacityRatio resourceToValueMap) int64 {
        // 核心思想:优先选择资源利用率最均衡的 Node
        // 而非资源最空闲的 Node(避免"甜点效应")
        score := 0
        for resource := range requestedToCapacityRatio {
            // 计算放置后的资源比率
            resourceToCapacityRatioAfterPodAdded := ...
            // 各资源比率方差越小,评分越高(越均衡)
            // 60 分满分
            score += ... 
        }
        return score
    }
}

这个插件与 NodeResourcesLeastAllocated(优先最空闲)形成互补:前者避免把所有 Pod 堆在少数几台大机器上导致资源碎片,后者提高资源利用率。

2. ImageLocality(镜像本地优先)

如果 Node 上已经存在 Pod 所需容器镜像,则给予最高 100 分。分数与该镜像占总大小成比例——需求镜像越多、越大,得分越高。这能显著减少 Pod 启动延迟。

3. InterPodAffinity(Pod 亲和评分)

处理 preferredDuringSchedulingIgnoredDuringExecution 的软亲和需求。对于每个匹配的 topology domain,按权值累加分数。内置实现使用 topologyPairToPodMatches 索引加速查找。

4. NodeAffinity(节点亲和评分)

与 Predicate 阶段的硬性过滤不同,此处处理 preferredDuringSchedulingIgnoredDuringExecution 的加权评分。每个 WeightedTerm 根据 Node 标签匹配情况累加 Weight 值。

5. NodePreferAvoidPods(避免系统关键 Pod 抢占)

处理 scheduler.alpha.kubernetes.io/preferAvoidPods annotation,给 kube-system 中标记了关键守护进程的 Node 打 0 分,避免用户 Pod 抢占系统资源。

6. PodTopologySpread(拓扑分布)

根据 topologySpreadConstraints 评估新 Pod 放置后各拓扑域的分布均匀程度。分布越均匀,评分越高。这是实现跨 AZ/Region 高可用的关键插件。

7. TaintToleration 评分

容忍的 Taint 越少评分越高,鼓励 Pod 尽量避开有 NodeProblem/DataDisk 等 Taint 的节点。

五、调度框架(Scheduling Framework):扩展点详解

Kubernetes 1.16+ 引入的 Scheduling Framework 将调度流程拆分为多个 Pipeline Stage(扩展点),每个阶段可以注册多个 Plugin 来执行自定义逻辑:

QueueSort → PreFilter → Filter → PostFilter → PreScore → Score → Reserve → Permit → PreBind → Bind → PostBind

各扩展点详解:

QueueSort:定义 Pod 在 activeQ 中的排序策略。默认按 Priority 降序 → CreationTimestamp 升序排序。只能有一个激活的 QueueSort 插件。

PreFilter:在 Filter 之前执行,用于预计算和缓存信息。典型用例是在此阶段计算 Pod 的 ResourceRequest 并存入 CycleState,避免重复计算。

Filter(即 Predicate):硬性条件过滤。所有 Filter 插件必须全部通过,否则 Pod 被标记为 Unschedulable。注意 Filter 之间是"短路"逻辑——一旦有一个失败就立即返回。

PostFilter:1.24+ 引入(Alpha),在所有 Filter 都失败时执行。典型用例是 Descheduler 模式——通过抢占(Preemption)来强行腾出资源。FIFOPreemption 插件在此扩展点实现抢占逻辑。

PreScore:评分前的预处理,如清理过期缓存、准备评分所需的辅助数据结构。

Score:所有激活的 Score 插件按 Weight 加权求和归一化到 0-100。不同的 Profile 可以配置不同的插件和权重组合。

Reserve:在绑定前"预留"资源,防止并发调度时的资源冲突。VolumeScheduling 插件在此阶段做 Volume 的预绑定标记。

Permit:控制 Pod 是否可以进入绑定阶段,支持三种动作:approve(批准)、deny(拒绝)、wait(等待,可设超时)。自定义调度器可用此实现"配额审批流"。

PreBind:绑定前的准备工作,如创建 VolumeAttachment 对象、初始化 CNI 网络配置。

Bind:将 Pod 绑定到目标 Node 的实际操作。默认通过创建 Binding 对象发送到 API Server。只有一个 Bind 插件可被激活。

PostBind:绑定成功后的清理工作,如推送事件通知、更新监控指标。

六、KubeSchedulerConfiguration 高级调优

生产环境的调度器配置通常需要通过 KubeSchedulerConfiguration(v1beta3/v1)进行细粒度控制:

apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
parallelism: 16                    # 并行调度 Goroutine 数(默认 16)
clientConnection:
  kubeconfig: /etc/kubernetes/scheduler.conf
  qps: 100                         # 调度器对 API Server 的 QPS(默认 50)
  burst: 200                       #  Burst 上限(默认 100)
profiles:
  - schedulerName: default-scheduler
    plugins:
      multiPoint:
        enabled:
          - name: NodeResourcesFit
            weight: 2            # Score 阶段的权重
          - name: NodeResourcesBalancedAllocation
            weight: 1
          - name: PodTopologySpread
            weight: 2
        disabled:
          - name: NodeResourcesLeastAllocated  # 关闭最空闲策略
      preScore:
        disabled:
          - name: '*'             # 禁用所有 PreScore
      score:
        enabled:
          - name: MyCustomScorer
            weight: 3             # 自定义评分器权重
    pluginConfig:
      - name: NodeResourcesFit
        args:
          scoringStrategy:
            type: RequestedToCapacityRatio   # 使用请求-容量比率策略
            resources:
              - name: cpu
                weight: 1
              - name: memory
                weight: 1
              - name: nvidia.com/gpu
                weight: 2       # GPU 资源权重翻倍
      - name: KubeSchedulerProfileName
        args:
          percentageOfNodesToScore: 50  # 只对 50% 的 Node 做评分

关键调优参数详解:

parallelism:并行调度 Pod 的 Goroutine 数量。设置过高会导致锁竞争加剧(调度器全局锁),建议不超过 32。最佳实践:大规模集群(1000+ Node)设为 16-24,小集群保持默认 16。

percentageOfNodesToScore:评分阶段评估的 Node 比例。默认值:10 个节点以下为 100%,50 节点约为 50%,1000 节点以上约为 10%。降低此值可提升调度吞吐量,但可能影响调度质量。

RequestedToCapacityRatio:将资源请求-容量比率与 BalancedAllocation 结合,实现基于资源利用率的评分。适合混合负载场景(在线 + 离线)。

七、抢占(Preemption)机制全链路

当 Pod 因资源不足无法调度时,抢占机制允许高优先级 Pod 驱逐低优先级 Pod 来获取运行机会。抢占流程详解:

抢占触发条件:Pending Pod 的 Priority ≥ 1000000000(system-cluster-critical)或 PriorityClass 的 value 大于目标 Pod 的 value

抢占流程(FIFOPreemption PostFilter 插件实现):

Step 1:找候选 Node

从所有 Node 中选择"理论可以调度"的 Node——即仅因为资源不足导致 Filter 失败的 Node(而非 NodeAffinity/PortConflict 等硬性约束不满足的 Node)。

Step 2:选 Victim Pod

对候选 Node 上的所有 Pod,按 Priority 降序排列。只有 Priority 低于待调度 Pod 的 Pod 才能成为被抢占目标。如果某 Node 上所有 Pod 的 Priority 都 ≥ 待调度 Pod,则该 Node 不可作为抢占目标。

Step 3:模拟调度(Simulate)

构造一个"假设驱逐"场景——从 NodeInfo 中减去 Victim Pod 的资源占用量,然后重新对 Pending Pod 执行 Filter/Score。只有模拟成功才真正执行驱逐。

Step 4:优雅驱逐

对 Victim Pod 发送 DELETE 请求,尊重 terminationGracePeriodSeconds。被抢占 Pod 将重新进入 SchedulingQueue,如果仍无足够资源则进入 backoffQ。

PodDisruptionBudget(PDB)保护:抢占会尊重 PDB 约束,不会导致健康的 Pod 副本数低于 minAvailable。但如果 PDB 阻止了所有可能 Victim,抢占将失败。

八、自定义调度器开发实战

当内置调度器无法满足需求时,可以使用 Go 语言编写自定义调度器。以下是完整实现框架:

package main

import (
    "context"
    "fmt"
    "time"

    v1 "k8s.io/api/core/v1"
    metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
    "k8s.io/client-go/informers"
    "k8s.io/client-go/kubernetes"
    "k8s.io/client-go/tools/clientcmd"
    "k8s.io/client-go/tools/leaderelection"
    "k8s.io/client-go/tools/leaderelection/resourcelock"
)

func main() {
    config, err := clientcmd.BuildConfigFromFlags("", "/etc/kubernetes/scheduler.conf")
    if err != nil {
        panic(err.Error())
    }

    clientset, err := kubernetes.NewForConfig(config)
    if err != nil {
        panic(err.Error())
    }

    // 1. 创建 Informer Factory 监听 Pod 和 Node 事件
    informerFactory := informers.NewSharedInformerFactory(clientset, 0)
    podInformer := informerFactory.Core().V1().Pods()
    nodeInformer := informerFactory.Core().V1().Nodes()

    // 2. 自定义调度器结构
    customScheduler := &CustomScheduler{
        clientset:    clientset,
        podInformer:  podInformer,
        nodeInformer: nodeInformer,
        nodeLister:   nodeInformer.Lister(),
        podLister:    podInformer.Lister(),
    }

    // 3. 启动 Leader Election(HA 模式)
    rl, _ := resourcelock.New(
        resourcelock.LeasesResourceLock,
        "kube-system",
        "custom-scheduler",
        clientset.CoreV1(),
        clientset.CoordinationV1(),
        resourcelock.ResourceLockConfig{
            Identity: hostname,
        },
    )

    leaderelection.RunOrDie(context.TODO(), leaderelection.LeaderElectionConfig{
        Lock:            rl,
        ReleaseOnCancel: true,
        LeaseDuration:   15 * time.Second,
        RenewDeadline:   10 * time.Second,
        RetryPeriod:     2 * time.Second,
        Callbacks: leaderelection.LeaderCallbacks{
            OnStartedLeading: func(ctx context.Context) {
                // 成为 Leader 后启动调度循环
                informerFactory.Start(ctx.Done())
                customScheduler.Run(ctx.Done())
            },
            OnStoppedLeading: func() {
                klog.Fatalf("leader election lost")
            },
        },
    })
}

type CustomScheduler struct {
    clientset    *kubernetes.Clientset
    podInformer  cache.SharedIndexInformer
    nodeInformer cache.SharedIndexInformer
    nodeLister   listers.NodeLister
    podLister    listers.PodLister
}

func (s *CustomScheduler) Run(stopCh <-chan struct{}) {
    wait.Until(func() {
        s.scheduleOne()
    }, time.Second, stopCh)
}

// scheduleOne 完成单 Pod 的调度全流程
func (s *CustomScheduler) scheduleOne() {
    // Phase 1: 从Queue获取待调度 Pod
    pods, _ := s.podLister.Pods(metav1.NamespaceAll).List(
        labels.SelectorFromSet(labels.Set{
            "spec.schedulerName": "custom-scheduler",
        }))
    
    if len(pods) == 0 {
        return
    }
    targetPod := pods[0] // 取队首

    // Phase 2: 执行 Filter(自定义规则)
    nodes, _ := s.nodeLister.List(labels.Everything())
    feasibleNodes := make([]*v1.Node, 0)
    for _, node := range nodes {
        if s.podFitsNode(targetPod, node) {
            feasibleNodes = append(feasibleNodes, node)
        }
    }

    if len(feasibleNodes) == 0 {
        return // 无可用 Node,进入 backoff
    }

    // Phase 3: 执行 Score + 选最优
    bestNode := s.bestScore(targetPod, feasibleNodes)

    // Phase 4: 执行 Bind
    s.bindPod(targetPod, bestNode)
}

// podFitsNode 自定义 Filter 逻辑
func (s *CustomScheduler) podFitsNode(pod *v1.Pod, node *v1.Node) bool {
    // Filter 1: 检查资源是否足够
    allocatable := node.Status.Allocatable
    podRequest := resource.Requests(pod)
    
    if podRequest.Cpu().MilliValue() > allocatable.Cpu().MilliValue() {
        return false
    }
    if podRequest.Memory().Value() > allocatable.Memory().Value() {
        return false
    }

    // Filter 2: 自定义 Taint 检查
    for _, taint := range node.Spec.Taints {
        if taint.Key == "node.kubernetes.io/unschedulable" {
            return false
        }
    }

    // Filter 3: 自定义 Node Label 匹配
    if requiredGPU, ok := pod.Annotations["requires-gpu"]; ok {
        if requiredGPU == "true" {
            if node.Labels["gpu-type"] == "" {
                return false
            }
        }
    }

    return true
}

// bindPod 执行绑定
func (s *CustomScheduler) bindPod(pod *v1.Pod, node *v1.Node) {
    binding := &v1.Binding{
        ObjectMeta: metav1.ObjectMeta{
            Name:      pod.Name,
            Namespace: pod.Namespace,
        },
        Target: v1.ObjectReference{
            Kind: "Node",
            Name: node.Name,
        },
    }
    err := s.clientset.CoreV1().Pods(pod.Namespace).Bind(
        context.TODO(), binding, metav1.CreateOptions{})
    if err != nil {
        klog.Errorf("Failed to bind pod %s to node %s: %v",
            pod.Name, node.Name, err)
    }
}

以上代码展示了一个最小可用自定义调度器的核心结构。实际生产级的自定义调度器还需要处理:并发调度安全、抢占、缓存预热、调度延迟监控等。

九、Scheduler Extender:无需编码的扩展方案

如果不想编写完整的自定义调度器,可以使用 HTTP Extender 方式扩展调度器功能。通过在 KubeSchedulerConfiguration 中配置 extender,让 kube-scheduler 在特定阶段回调 HTTP Webhook:

apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
extenders:
  - urlPrefix: "http://scheduler-extender.kube-system:8080"
    filterVerb: "filter"
    scoreVerb: "score"
    preemptVerb: "preempt"
    bindVerb: "bind"
    enableHTTPS: true
    nodeCacheCapable: true       # Extender 自己缓存 Node 状态
    weight: 5                    # Score 权重
    managedResources:
      - name: "example.com/sriov"
        ignoredByScheduler: true  # kube-scheduler 不检查此资源
    ignorable: true              # 不可用时跳过(非关键)

Extender 的优缺点分析:

  • 优点:任何语言实现、独立部署、无需重新编译调度器、可灰度切换
  • 缺点:HTTP 调用增加延迟(5-50ms)、序列化/序列化开销、故障域增大、不参与多调度器并发

适合场景:特殊硬件调度(FPGA/GPU/InfiniBand)、Pod 放置与拓扑映射、企业内部资源配额系统的集成。

十、调度器生产运维:监控、调试与性能调优

1. 关键监控指标

# 调度延迟分布(P50/P90/P99)
histogram_quantile(0.99,
    rate(scheduler_e2e_scheduling_duration_seconds_bucket[5m]))

# 调度失败速率(Unschedulable 原因分类)
rate(scheduler_schedule_attempts_total{result="unschedulable"}[5m])
rate(scheduler_schedule_attempts_total{result="error"}[5m])

# 绑定失败次数
rate(scheduler_binding_duration_seconds_count{status="failure"}[5m])

# 活跃队列 Pod 堆积
scheduler_scheduler_queue_incoming_pods_total -
  scheduler_schedule_attempts_total

# 抢占次数
rate(scheduler_preemption_attempts_total{result="successful"}[10m])

2. 调度延迟的组成

e2e_scheduling_duration = scheduling_cycle_time + bind_cycle_time + delay_between_scheduling_and_binding。当集群规模超过 5000 Node 时,percentageOfNodesToScore 是主要调优手段。

3. 调度器 Pod 排查命令

# 查看调度器日志 - 筛选特定 Pod 的调度过程
kubectl logs -n kube-system kube-scheduler-master \
  | grep "Attempting to schedule pod\|Failed to schedule\|Successfully bound"

# 查看 Pending Pod 的调度失败原因
kubectl describe pod <pending-pod> | grep -A 10 "Events"

# 查看调度器 Leader 状态
kubectl get endpoints kube-scheduler \
  -n kube-system -o yaml

# 快速定位调度瓶颈(哪个阶段耗时最长)
curl -k https://localhost:10259/metrics | grep "scheduling_duration"

4. 常见调度问题与解决方案

  • Pod 长期 Pending 但资源显示充足:检查 Node 是否设置了 Taint、NodeSelector 不匹配、PVC 未绑定
  • 调度延迟 P99 突增:检查 Node 数量是否暴增(percentageOfNodesToScore 未调整)、API Server 是否过载(QPS/Burst 需调高)
  • 资源碎片导致大 Pod 无法调度:启用 RequestedToCapacityRatio + MostAllocated
    策略,结合 Descheduler 做重平衡
  • Preemption 风暴:设置 PriorityClass 时确保优先级层次清晰,PDB 设置合理的 minAvailable
  • 自定义调度器绑定失败:检查 RBAC 权限(create binding, delete pods)、Leader Election 是否正常

十一、前沿趋势:Kueue、Volcano 与 Koordinator

云原生调度生态在 Kubernetes 原生调度器之上,针对批量任务、混部、异构计算三大场景催生了新一代调度系统:

Kueue:Kubernetes 原生的 Job 队列管理系统,支持 Resource Flavor(等价 Node 组)、Cohort(资源借用)、Single-queue 多租户配额。适合 ML 训练、CI 等批量工作负载。与 Cluster Autoscaler 深度集成实现"无作业缩 Node 到零"。

Volcano:华为开源的批量计算调度框架,与 Kueue 竞争定位。支持 DAG 工作流、Gang Scheduling(整存整取)、Binpack/Spread 混合策略。深度学习训练场景下与 PyTorch Operator / TFJob 集成良好。

Koordinator:阿里开源的混部调度方案,专注于离线任务与在线服务的"错峰混部"。核心能力包括:QoS 感知的资源超卖(Burst QoS)、CPU Burst 调度策略、精细化 NUMA 拓扑感知、资源画像预测降低装箱抖动。

三者的共同演进趋势是:从"单 Pod 调度"到"Job 级编排 + Qos 感知 + 经济模型驱动",调度不再仅是技术问题,而是成本优化的核心引擎。

十二、总结

Kubernetes 调度器是一个精密的多层过滤-评分-绑定系统。理解其架构和扩展点,是构建高性能、高可用集群的关键基础。

核心要点回顾:

  • 调度队列使用三级有序结构 + 指数 backoff 退避策略管理待调度 Pod
  • Predicate 7 大过滤器按"资源→亲和→污点→拓扑"逐层收敛候选 Node
  • Score 评分按插件权重加权归一,BalancedAllocation 策略避免资源碎片
  • Scheduling Framework 的 12 个扩展点覆盖了完整的调度生命周期
  • PostFilter 阶段的抢占机制保障高优先级 Pod 的服务等级
  • 自定义调度器通过 Leader Election + Informer + Binding API 实现
  • 生产环境的关键调优参数是 parallelism、percentageOfNodesToScore、QPS/Burst
  • Kueur/Volcano/Koordinator 代表了调度系统向混部/队列/经济模型演进的方向

掌握这些知识后,你将能够根据业务场景灵活设计调度策略,在资源利用率与服务质量之间找到最佳平衡点。

点赞(0) 打赏

评论列表 共有 0 条评论

暂无评论
立即
投稿

微信公众账号

微信扫一扫加关注

发表
评论
返回
顶部