Topology-Aware Workload Scheduling

Topology-Aware Workload Scheduling

FEATURE STATE: Kubernetes v1.36 [alpha](disabled by default)

Topology-Aware Scheduling (TAS) is a feature of the Workload API that optimizes the placement of pods within the cluster.

TAS ensures that all pods within a PodGroup are co-located into a specific topology domain, such as a single server rack or zone. This minimizes inter-pod communication latency and prevents workload fragmentation across the cluster infrastructure.

Topology-aware scheduling with gang scheduling policy

When applied to PodGroups with gang scheduling policy, TAS simulates the potential assignment (placement) of the full group of pods at once. It guarantees that at least the specified minCount pods can fit together into the same topology domain before committing resources. If no feasible placement is found, the entire PodGroup becomes unschedulable.

This is the recommended approach for workloads like distributed AI and ML training that strictly require proximity to minimize inter-pod communication latency.

If new pods are added to the PodGroup where some pods are already scheduled (for example, if pods are recreated), the scheduler will force all new incoming pods to land on the exact same topology domain where the existing pods currently reside. If that specific domain lacks sufficient capacity for the new pods, the pods will remain pending - even if it means that less than minCount pods are scheduled at this point.

Note:

As of v1.36 Topology-Aware Scheduling does not trigger workload or pod preemption. If no feasible placement can be found without triggering preemption, the PodGroup becomes unschedulable.

Topology-aware scheduling with basic scheduling policy

Using TAS with basic scheduling policy may exhibit inconsistent behavior. The scheduler may only observe a subset of pods when entering the PodGroup scheduling cycle - therefore placement feasibility is only evaluated for the observed pods, rather than the entire PodGroup. To partially mitigate this limitation, you can use scheduling gates to hold off PodGroup scheduling until all pods within the PodGroup are in the scheduling queue.

If no feasible placement is found for the entire PodGroup, only a subset of pods may be scheduled, and they are guaranteed to meet the scheduling constraints.

If new pods are added to the PodGroup where some pods are already scheduled, the scheduler will act the same as in case of gang policy - forcing the new pods into the same domain, unless there is insufficient capacity (in which case the new pods will remain pending).

API configuration: scheduling constraints

Every PodGroup (or PodGroupTemplate) may optionally declare the schedulingConstraints field, which is interpreted by the placement-based PodGroup scheduling algorithm. If constraints are defined in PodGroupTemplate, they will be copied to referencing PodGroups.

As of Kubernetes v1.36, the API supports topology constraints.

Note:

As of Kubernetes v1.36, you can specify only a single topology constraint in each PodGroup.

Topology constraint

To define a topology constraint for a PodGroup you need to set a key, which corresponds to a Kubernetes node label, representing the target topology domain (for example, a rack or a zone). The scheduler strictly enforces that all pods within the PodGroup are placed onto nodes that share the exact same value for this specified label.

Here is an example of a PodGroup configured with a topology constraint:

apiVersion: scheduling.k8s.io/v1beta1
kind: PodGroup
metadata:
  name: example-podgroup
spec:
  schedulingPolicy:
    gang:
      minCount: 4
  schedulingConstraints:
    topology:
      - key: topology.example.com/rack

Multi-level topology-aware scheduling

FEATURE STATE: Kubernetes v1.37 [alpha](disabled by default)

Complex workloads might require co-location of their Pods at different levels of the cluster infrastructure. For example, an entire workload may need to run within a single availability zone, while different parts of that workload may require strict co-location within specific server racks.

Such multi-level co-location requirements can be expressed using the CompositePodGroup API and by specifying topology constraints at different levels of a group hierarchy.

Using the CompositePodGroup API requires enabling the CompositePodGroup feature gate and the scheduling.k8s.io/v1alpha3 API group.

Multi-level topology constraints resolution

Every group inside a CompositePodGroup hierarchy can specify a topology constraint which guarantees that all descendant Pods of that group will be scheduled in the same topology domain, matching that group's constraint.

During hierarchical scheduling, the scheduler resolves these constraints in a top-down manner. Specifically, topology domains that are considered during scheduling of a child group are confined within a topology domain that corresponds to the placement assumed by the parent group.

Kubernetes does not impose any strict requirements on the physical hierarchy of topology labels - topology keys are arbitrary node labels. However, the order in which you specify topology constraints from parent to child determines the order in which the scheduler subdivides topology domains.

Note:

As of Kubernetes v1.37, you can specify only a single topology constraint in each CompositePodGroup.

Example

The following example configures a Workload where the parent CompositePodGroupTemplate constrains the entire workload to a single availability zone (topology.example.com/zone), while two child PodGroupTemplate entries (workers and driver) constrain their respective Pods to server racks (topology.example.com/rack) within that zone:

apiVersion: scheduling.k8s.io/v1alpha3
kind: Workload
metadata:
  name: example-workload
spec:
  compositePodGroupTemplates:
  - name: root
    schedulingPolicy:
      gang:
        minGroupCount: 2
    schedulingConstraints:
      topology:
      - key: topology.example.com/zone
    podGroupTemplates:
    - name: workers
      schedulingPolicy:
        gang:
          minCount: 8
      schedulingConstraints:
        topology:
        - key: topology.example.com/rack
    - name: driver
      schedulingPolicy:
        gang:
          minCount: 1
      schedulingConstraints:
        topology:
        - key: topology.example.com/rack

After creating the Workload object, the corresponding group objects are created as follows:

  • Root CompositePodGroup referencing the root template.
  • Two child PodGroup objects (workers and driver), each referencing the root CompositePodGroup as their parent group.
apiVersion: scheduling.k8s.io/v1alpha3
kind: CompositePodGroup
metadata:
  name: workload-root
spec:
  workloadRef:
    workloadName: example-workload
    templateName: root
  schedulingPolicy:
    gang:
      minGroupCount: 2
  schedulingConstraints:
    topology:
    - key: topology.example.com/zone
---
apiVersion: scheduling.k8s.io/v1beta1
kind: PodGroup
metadata:
  name: workload-workers
spec:
  parentCompositePodGroupName: workload-root
  workloadRef:
    workloadName: example-workload
    templateName: workers
  schedulingPolicy:
    gang:
      minCount: 8
  schedulingConstraints:
    topology:
    - key: topology.example.com/rack
---
apiVersion: scheduling.k8s.io/v1beta1
kind: PodGroup
metadata:
  name: workload-driver
spec:
  parentCompositePodGroupName: workload-root
  workloadRef:
    workloadName: example-workload
    templateName: driver
  schedulingPolicy:
    gang:
      minCount: 1
  schedulingConstraints:
    topology:
    - key: topology.example.com/rack

During scheduling, the scheduler first selects an availability zone for workload-root. It then subdivides the nodes in that zone by rack to find feasible rack placements for workload-workers and workload-driver within the selected zone.

For example, consider a cluster with five nodes labeled as follows:

Nodetopology.example.com/zonetopology.example.com/rack
node-azone-1rack-1
node-bzone-1rack-1
node-czone-1rack-2
node-dzone-2rack-1
node-ezone-2rack-3

When processing workload-root, the scheduler evaluates candidate placements across all cluster nodes based on the topology.example.com/zone topology key:

Evaluated candidate placementNodes in candidate placement
zone-1node-a, node-b, node-c
zone-2node-d, node-e

When evaluating candidate placements for workload-workers, the scheduler subdivides only the nodes within the placement assumed by workload-root based on the topology.example.com/rack topology key:

Parent placementEvaluated candidate placementNodes in candidate placement
zone-1rack-1node-a, node-b
zone-1rack-2node-c
zone-2rack-1node-d
zone-2rack-3node-e

Candidate placements generated for the sibling workload-driver PodGroup are identical to those generated for workload-workers, since both groups specify the same topology key (topology.example.com/rack).

What's next

Last modified July 24, 2026 at 3:44 PM PST: Update docs for CompositePodGroup API (5d3723fc78)