Kubernetes v1.37 [beta](disabled by default)The Workload API resource defines the scheduling requirements and structure of a multi-Pod
application. While workload controllers such as Job
manage the application's runtime state, the Workload specifies how groups of Pods
should be scheduled. The Job controller is the only built-in controller that creates
PodGroup objects from the Workload's
PodGroupTemplates at runtime.
The Workload API resource is part of the scheduling.k8s.io/v1beta1
API group
and your cluster must have that API group enabled, as well as the GenericWorkload
feature gate,
before you can use this API.
A Workload is a static, long-lived policy template. It defines what scheduling
policies should be applied to groups of Pods, but does not track runtime state itself.
Runtime scheduling state is maintained by PodGroup
objects, which controllers create from the Workload's PodGroupTemplates.
A Workload consists of two fields: a list of PodGroupTemplates and an optional controller
reference. The entire Workload spec is immutable after creation: you cannot modify
existing templates, add new templates, or remove templates from podGroupTemplates.
The spec.podGroupTemplates list defines the distinct components of your workload.
For example, a machine learning job might have a driver template and a worker template.
Each entry in podGroupTemplates must have:
name that will be used to reference the template in the PodGroup's spec.podGroupTemplateRef.basic or gang).Each entry can also have priority and disruption mode fields.
WorkloadAwarePreemption
feature gate. This gate was merged into
GenericWorkload in v1.37.The maximum number of PodGroupTemplates in a single Workload is 8.
apiVersion: scheduling.k8s.io/v1beta1
kind: Workload
metadata:
name: training-job-workload
namespace: some-ns
spec:
controllerRef:
apiGroup: batch
kind: Job
name: training-job
podGroupTemplates:
- name: workers
schedulingPolicy:
gang:
# The gang is schedulable only if 4 pods can run at once
minCount: 4
priorityClassName: high-priority
disruptionMode:
all: {}
When a workload controller creates a PodGroup from one of these templates, it copies the
schedulingPolicy into the PodGroup's own spec. Changes to the Workload only affect
newly created PodGroups, not existing ones.
The controllerRef field links the Workload back to the specific high-level object defining the application,
such as a Job or a custom CRD. This is useful for observability and tooling.
This data is not used to schedule or manage the Workload.
Kubernetes v1.37 [alpha](disabled by default)When the CompositePodGroup
feature gate and the scheduling.k8s.io/v1alpha3 API group
are enabled, you can use CompositePodGroupTemplates to define multi-level, hierarchical scheduling
requirements in a Workload. These requirements can include enforcing nested topology constraints
across different layers of cluster infrastructure (multi-level topology-aware scheduling),
all-or-nothing scheduling across child groups (multi-level gang scheduling), or group-level
disruption policies.
CompositePodGroupTemplates can be defined using the spec.compositePodGroupTemplates field in the
Workload API. At runtime, workload controllers create CompositePodGroup
and PodGroup objects from these templates to maintain the runtime
scheduling state of the hierarchy. While PodGroup objects manage groups of Pods at the leaves,
CompositePodGroup objects represent non-leaf groups that enforce scheduling policies across child groups.
Workload specification, spec.compositePodGroupTemplates and spec.podGroupTemplates
fields form a union: a Workload must define either spec.podGroupTemplates (for flat
workloads) or spec.compositePodGroupTemplates (for hierarchical workloads), but cannot
specify both.The spec.compositePodGroupTemplates field defines non-leaf templates in a group-template
hierarchy tree. Each entry represents a template for a CompositePodGroup and can contain:
CompositePodGroupTemplates (for intermediate non-leaf
groups) or PodGroupTemplates (for leaf groups containing Pods).basic: Child groups are admitted and scheduled independently.gang: Enforces multi-level all-or-nothing scheduling across child groups.
Requires minGroupCount, which specifies the minimum number of child groups
that must be schedulable simultaneously for the composite group to be feasible.priorityClassName,
disruptionMode (Single or All) and preemptionPolicy for
workload-aware preemption.To ensure cluster stability and control-plane efficiency, the group-template hierarchy enforces the following limits:
compositePodGroupTemplates and podGroupTemplates list is strictly
capped at a maximum of 8 items.CompositePodGroupTemplates. You can only change
the minCount value in the gang scheduling policy defined in the leaf PodGroupTemplates.The following example defines a hierarchical Workload with a CompositePodGroup template
that enforces gang scheduling across two child PodGroup templates (minGroupCount: 2),
each specifying its own gang scheduling policy:
apiVersion: scheduling.k8s.io/v1alpha3
kind: Workload
metadata:
name: gang-of-gangs-workload
namespace: default
spec:
compositePodGroupTemplates:
- name: root
schedulingPolicy:
gang:
# Requires both child PodGroups to be schedulable together
minGroupCount: 2
podGroupTemplates:
- name: workers-a
schedulingPolicy:
gang:
# Requires 4 Pods in this group to be schedulable
minCount: 4
- name: workers-b
schedulingPolicy:
gang:
# Requires 4 Pods in this group to be schedulable
minCount: 4
Kubernetes v1.36 [alpha](disabled by default)When the
WorkloadWithJob
feature gate is enabled, the
Job controller compiles a Job's
.spec.scheduling configuration into Workload and PodGroup objects before it
creates any Pods. You opt into gang scheduling by setting
.spec.scheduling.schedulingPolicy.gang on the Job; an omitted gang.minCount
defaults to the Job's .spec.parallelism, so all Pods must be schedulable together
before any of them are bound to nodes.
When .spec.scheduling is omitted, the Job defaults to the basic policy, which
preserves standard pod-by-pod scheduling. Either way the Job controller creates the
Workload and PodGroup for you, so you do not need to create them yourself.
Other workload controllers (such as JobSet) may manage their own Workload and
PodGroup objects independently.
For the full set of scheduling fields and examples, see Integrate with Workload APIs.