← All posts Kubernetes

Spread Kubernetes pods across failure domains without making them unschedulable

The Tryssh team ·

For Kubernetes topology spread constraints, the safest approach is a bounded operational change, not a command pasted without context. This runbook starts with effective state, shows the smallest candidate action, and finishes by repeating the real user or system path.

TL;DR: Topology spread constraints compare matching pod counts across eligible domains and can reject scheduling or permit skew depending on whenUnsatisfiable. Capture a baseline with kubectl get nodes -L topology.kubernetes.io/zone,kubernetes.io/hostname, make the reviewed change only when the evidence matches, then verify with kubectl -n example rollout status deployment/app and keep the rollback ready.

Audience: cluster operators comfortable with kubectl contexts, namespaces, workload controllers, and declarative manifests. This guide assumes familiarity with Kubernetes operations and a change window appropriate to the system.

Spread Kubernetes pods across failure domains without making them unschedulable observe, interpret, change, and verify workflow

The direct answer

Replicas distribute across eligible zones within the chosen skew while capacity remains schedulable. That is the success condition for Kubernetes topology spread constraints; command completion by itself is not enough.

The important boundary is desired API object, admission, scheduling, node execution, service routing, and controller reconciliation. Topology spread constraints compare matching pod counts across eligible domains and can reject scheduling or permit skew depending on whenUnsatisfiable. If an observation does not identify which side of that boundary failed, collect a narrower observation before changing state.

How the mechanism works

For Kubernetes topology spread constraints, use this mental model: Kubernetes stores desired state in the API and multiple controllers converge actual objects toward it; a successful write does not prove scheduling, readiness, routing, or application behavior. The model prevents a common mistake—treating configuration text, control-plane acceptance, process state, and end-user behavior as the same proof.

Follow four stages:

  1. Observe: identify the exact host, object, version, owner, and active configuration.
  2. Interpret: write the expected result, the abnormal result, and what would remain inconclusive.
  3. Change: apply one reviewed action at the narrowest layer that contradicts the baseline.
  4. Verify: repeat the original path and compare the same evidence, including adjacent safety controls.

Preflight and safety boundary

Confirm the current context and namespace, save the live object, and understand which controller owns it before applying, deleting, draining, or rolling back.

Before Kubernetes topology spread constraints, record UTC time, the current version or digest, the exact target, recent changes, and who owns the workload. The rollback for this runbook is: restore the saved deployment template, wait for the controller to converge, and verify serving replicas before retrying a softer policy.

Do not continue if the target identity is ambiguous, the current state cannot be saved, the only recovery session would be at risk, or the proposed command affects more objects than the brief names.

Capture the read-only baseline

Run these commands one at a time. Replace example names and addresses deliberately; do not paste production secrets into a transcript.

kubectl get nodes -L topology.kubernetes.io/zone,kubernetes.io/hostname
kubectl -n example get pods -l app=app -o wide
kubectl -n example get deployment app -o yaml

Interpret the baseline before moving on:

  • Expected: replicas distribute across eligible zones within the chosen skew while capacity remains schedulable.
  • Abnormal: missing labels, insufficient zones, taints, affinity, or capacity leave new pods Pending.
  • Inconclusive: missing output can also mean the wrong context, permissions, namespace, log window, binary, or target. Prove those assumptions before treating absence as health.

Save the decisive output, exit status, and timestamp. Redact credentials, customer data, private topology, tokens, and complete environment dumps.

Apply the smallest candidate change

The following is state-changing example syntax, not an instruction to run it unchanged:

kubectl -n example patch deployment app --type merge -p '{\"spec\":{\"template\":{\"spec\":{\"topologySpreadConstraints\":[{\"maxSkew\":1,\"topologyKey\":\"topology.kubernetes.io/zone\",\"whenUnsatisfiable\":\"DoNotSchedule\",\"labelSelector\":{\"matchLabels\":{\"app\":\"app\"}}}]}}}}'

For Kubernetes topology spread constraints, the proposed change is acceptable only when the read-only baseline predicts its effect and the rollback is available. The key risk is: strict spread during a zone outage can reject replacement pods and reduce availability further.

Prefer an immutable artifact, validated configuration, dry-run, transaction, candidate object, or staged target when the tool supports one. Record the exact command and UTC time so later telemetry can be correlated to the change.

Verify the result from the outside in

kubectl -n example rollout status deployment/app
kubectl -n example get pods -l app=app -o wide
kubectl -n example get events --sort-by=.lastTimestamp

Verification for Kubernetes topology spread constraints has three layers:

  1. The control plane or command reports the intended effective state.
  2. The process, resource, or data path reflects that state without a new pressure signal.
  3. The original user-visible or dependent-system path succeeds from an independent vantage point.

If kubectl -n example rollout status deployment/app succeeds but the original path still fails, stop. The change may have repaired a local symptom while DNS, policy, routing, caching, dependency, or client state remains broken.

Failure branches

The baseline does not match this runbook

When missing labels, insufficient zones, taints, affinity, or capacity leave new pods Pending, do not force the candidate command. Return to identity and scope, compare a healthy peer only through effective settings, and name a new falsifiable mechanism.

The change succeeds but behavior does not

A successful kubectl -n example patch deployment app --type merge -p '{\"spec\":{\"template\":{\"spec\":{\"topologySpreadConstraints\":[{\"maxSkew\":1,\"topologyKey\":\"topology.kubernetes.io/zone\",\"whenUnsatisfiable\":\"DoNotSchedule\",\"labelSelector\":{\"matchLabels\":{\"app\":\"app\"}}}]}}}}' proves that one interface accepted a request. It does not prove convergence, readiness, data compatibility, external routing, or client recovery. Re-run the same evidence at each downstream boundary.

The change makes the system worse

Execute the written rollback: restore the saved deployment template, wait for the controller to converge, and verify serving replicas before retrying a softer policy. Preserve the failed candidate and relevant logs long enough to explain the outcome; do not destroy the evidence with broad cleanup or repeated restarts.

Operator checklist

  • Confirm the exact target, context, identity, version, and active owner.
  • Capture the read-only baseline and one disconfirming observation.
  • Label kubectl -n example patch deployment app --type merge -p '{\"spec\":{\"template\":{\"spec\":{\"topologySpreadConstraints\":[{\"maxSkew\":1,\"topologyKey\":\"topology.kubernetes.io/zone\",\"whenUnsatisfiable\":\"DoNotSchedule\",\"labelSelector\":{\"matchLabels\":{\"app\":\"app\"}}}]}}}}' as state-changing during review.
  • Keep recovery access and rollback independent of the path being edited.
  • Change one layer, record UTC time, and wait for its real convergence boundary.
  • Verify the original path, adjacent controls, resource pressure, and persistence.
  • Update the runbook when observed behavior differs from the source-reviewed model.

Investigate it in Tryssh

Tryssh keeps the command, approval boundary, output, and verification beside the host conversation.

Tryssh can preserve this evidence loop and show a state-changing command for human approval. It does not make the operator's identity, recovery access, rollback, or platform authority decisions.

Evidence and review status

This Kubernetes topology spread constraints runbook was source-reviewed on 2026-07-29 against current first-party documentation. The commands are illustrative and use example targets. The page does not claim that the change was reproduced across every distribution, managed service, version, network, or workload.

Limitations and trade-offs

Strict spread during a zone outage can reject replacement pods and reduce availability further. Managed platforms may generate configuration, restrict privileges, replace local state, or expose a different control plane than the upstream project. Confirm the installed version and provider contract before applying a repair.

Search visibility is not proof of operational correctness. Treat this page as a decision aid, preserve independent recovery, and stop when the evidence contradicts its assumptions.

Related operator runbooks

Continue with Kubernetes priority and preemption, Kubernetes PodDisruptionBudget, the Kubernetes operations foundation guide, and the SSH hardening checklist.

Sources and further reading