Kubernetes: Control Loops, Workloads, and Failure Modes
The reconciliation model behind the Kubernetes API, the workload objects you actually use, how scheduling and probes really behave, and the failure modes that dominate real clusters.
Kubernetes is a cluster-level control system for running containerised workloads. You submit a description of what should be true — five replicas of this image, reachable at this name, with these resources — and a set of controllers works continuously to make the cluster match. Understanding Kubernetes is mostly understanding that everything in it is a control loop over declared state, and that almost every confusing behaviour is a loop doing exactly what it was told.
What Kubernetes Is
The unit of execution is the Pod: one or more containers that share a network namespace and can share volumes, scheduled together onto one node. You rarely create Pods directly. Instead you create a higher-level object, and a controller creates and replaces Pods on your behalf. That indirection is the whole design: you describe an outcome, and something else owns the steps and keeps owning them after the initial action succeeds.
The Control Plane
API server The only component that writes to the datastore. Everything —
kubectl, controllers, the kubelet — talks to it. Authentication,
authorization (RBAC), admission control all happen here.
etcd The consistent key-value store holding all cluster state.
Its backups are the cluster's backups.
scheduler Watches for Pods with no node assigned; picks a node based on
resource requests, affinity rules, taints, and topology.
controller Runs the built-in reconciliation loops: Deployment → ReplicaSet
manager → Pods, node lifecycle, endpoints, and many more.
kubelet Runs on every node. Watches for Pods assigned to its node,
starts containers, reports status, runs the probes.
kube-proxy Programs node-level routing so a Service name reaches a healthy
/ CNI backing Pod. The network plugin supplies Pod networking.
The loop is always the same shape: observe current state, compare to desired state, take one step toward desired, repeat. There is no transaction and no rollback. If you delete a Pod, a controller creates a replacement because the declared replica count still says it should exist — not because deletion failed, but because the loop is doing its job.
The Objects You Actually Use
| Object | What it gives you | Use it for |
|---|---|---|
| Deployment | A replica count plus a controlled rollout and rollback of Pod template changes | Stateless services; the default choice |
| StatefulSet | Stable ordinal names, stable per-replica storage, ordered rollout | Workloads whose identity and disk must persist across restarts |
| DaemonSet | One Pod per node, including new nodes | Log shippers, node agents, CNI components |
| Job / CronJob | Run to completion, with retries; on a schedule | Batch work, migrations, scheduled tasks |
| Service | A stable virtual address and load balancing across matching Pods | In-cluster addressing; the ClusterIP is the common case |
| Ingress / Gateway | HTTP routing, TLS termination, host and path rules | External traffic into the cluster |
| ConfigMap / Secret | Configuration and credentials injected as env vars or files | Environment-specific values, kept out of the image |
| PVC / StorageClass | A claim on durable storage, satisfied dynamically | Anything that must survive a Pod restart |
| Namespace + RBAC | A scope for names and quotas, plus who may do what | Multi-team separation and least privilege |
Two things about Secrets are worth stating plainly, because both surprise people. They are base64 encoded, not encrypted, so anyone who can read the object can read the value; protecting them means RBAC plus encryption at rest in the datastore. And a Service selects Pods by label, not by which Deployment created them — a mismatched or accidentally shared label selector routes traffic somewhere unintended and is a common cause of the bug where "some requests go to the old version".
Scheduling, Requests, and Limits
A request is what the scheduler reserves for a container when choosing a node; it is a claim on capacity. A limit is what the runtime enforces at execution time. They are different mechanisms and confusing them causes most capacity incidents.
The Asymmetry Between CPU and Memory
CPU is compressible
Exceeding a CPU limit throttles the container. It gets slower.
Symptom: latency rises, timeouts appear, nothing crashes,
and the dashboard shows plenty of "free" CPU on the node.
Memory is not compressible
Exceeding a memory limit kills the container. OOMKilled, restart,
and if the cause is a genuine working-set size, it happens again.
No requests set
The scheduler assumes the Pod needs almost nothing, packs the node,
and every workload on it competes. This is the classic noisy-neighbour
cluster, and it is caused by omission rather than by misconfiguration.
Related knobs are worth knowing before you need them: taints and tolerations keep workloads off nodes they do not belong on; affinity and topology spread constraints prevent every replica landing in one failure domain; PodDisruptionBudgets stop voluntary operations such as a node drain from taking down more replicas than you can afford; and the horizontal autoscaler adds replicas while the cluster autoscaler adds nodes, which means a scaling event needs both to work or Pods simply sit in Pending.
Probes: Liveness, Readiness, Startup
A readiness probe controls traffic: failing it removes the Pod from Service endpoints but leaves it running. A liveness probe controls life: failing it restarts the container. A startup probe suppresses the other two until a slow-starting application has finished initialising.
The most damaging probe mistake in production is a liveness probe that checks a dependency. If your liveness endpoint queries the database, then a database slowdown causes every replica to fail liveness at once, restart in unison, lose their warm state, and hammer the recovering database as they come back. Liveness should answer one question only: is this process wedged such that a restart is the correct remedy? Dependency health belongs in readiness, where the consequence is removal from load balancing rather than a restart storm.
Custom Resources and Operators
A CustomResourceDefinition adds a new object type to the API, and a controller written against it becomes an operator: software that encodes the operational knowledge for a specific workload as a reconciliation loop. This is how databases, certificate issuance, message brokers, and progressive delivery are commonly managed in-cluster.
The pattern is powerful and frequently overused. An operator is a distributed system you now maintain, with upgrade paths, version skew against the cluster, and failure modes of its own. Adopt one when the operational knowledge it encodes is genuinely complex and well tested; write one only when the loop you need does not exist and the alternative is a human running a runbook repeatedly.
Failure Modes
| Symptom | Usual cause | Where to look |
|---|---|---|
| Pod stuck Pending | No node satisfies the requests, taints, or volume topology | Scheduling events on the Pod |
| CrashLoopBackOff | The process exits on start: bad config, missing secret, failed migration | Logs of the previous container instance |
| ImagePullBackOff | Wrong tag, wrong registry, or missing pull credentials | Pod events |
| Periodic OOMKilled | Memory limit below real working set, or a leak | Container restart reason and memory trend |
| Restart storms under load | Liveness probe checking a dependency, or timeouts too tight | Probe definition |
| Traffic to the wrong version | Service selector matches more Pods than intended | Labels on Pods and Service selector |
| Mutable image tags | Two nodes run different code under the same tag | Pin by digest |
| Single replica plus rolling update | Brief outage on every deploy | Replica count and disruption budget |
| Data loss on reschedule | Local or ephemeral storage treated as durable | Volume type and StorageClass |
For the day-to-day commands behind that table, see the kubectl reference.
When Not to Use Kubernetes
Kubernetes solves problems that come with running many services across many machines: scheduling, service discovery, rollout control, and a uniform API for automation. If you have a handful of services, a small team, and no need to schedule across a fleet, it adds a control plane to operate, an upgrade treadmill, a networking model to learn, and an access-control surface to secure — in exchange for capabilities you are not using. A managed application platform, a container service without cluster management, or a few well-configured machines will be cheaper to run and far easier to debug.
Stateful data stores deserve a separate decision. Kubernetes can run them, and StatefulSets plus mature operators make it viable, but the operational burden of backups, failover, and version upgrades does not disappear — it moves to you. Unless you have that expertise on the team, a managed database is usually the better trade.
Where Kubernetes is the right answer, the surrounding practices matter as much as the cluster: infrastructure as code for the cluster itself, GitOps for what runs inside it, and platform engineering so application teams do not each have to learn all of the above.