Demystifying the Kubernetes control plane: what happens under the hood
Most developers can deploy a pod. Fewer can say exactly how Kubernetes turns one command into a running container. This walks through the machinery step by step, from a single kubectl apply to a pod on a node.
You run kubectl apply -f deployment.yaml, and a few seconds later a container is running somewhere in your cluster. The steps in between are not mysterious once you see them: a small set of control-plane components pass a piece of desired state between them, each doing one job. Following a single command through them is the clearest way to understand how the control plane works.
01A single command, end to end
Here is the whole path before we slow it down. Your kubectl apply becomes an HTTPS request to the API server, which authenticates it, validates it, and writes the desired object into etcd. A controller notices the new Deployment and creates the Pods it implies. The scheduler notices those Pods have no node and assigns each one. The kubelet on that node sees a Pod assigned to it and tells the container runtime to start the container.1 Each step is one component reacting to a change in stored state, rather than a chain of direct commands.
# The command that starts the process
kubectl apply -f deployment.yaml
# What it becomes: an authenticated HTTPS write to the API server,
# which persists the object in etcd. Everything else reacts to that.02The API server: the only way in
The kube-apiserver is the front end of the cluster, and the only one: every other component — the scheduler, the controllers, the kubelets on each node, and your kubectl — reaches etcd through it, never directly.1 That single entry point is what makes the cluster governable. Every request passes through authentication (who are you?), authorization (are you allowed?), and admission control (is this object acceptable, and should it be modified or rejected?) before anything is stored.2
Because it is the only writer to the datastore, the API server is also where consistency is enforced. It is stateless itself, which is deliberate: you can run several replicas behind a load balancer and lose one without losing the cluster, because none of them holds state. The state lives in etcd.
03etcd: the cluster's source of truth
Everything Kubernetes knows about itself — every object, its spec, and its status — lives in etcd, a distributed, strongly consistent key-value store.3 This is the part worth remembering: etcd holds the cluster's real state. The nodes, the API server replicas, and the controllers can all be rebuilt. etcd holds the one thing that cannot be rebuilt, which is the record of what should exist.
etcd stays consistent across its replicas using the Raft consensus algorithm, which needs a majority (a quorum) of members to agree before a write is committed.8 That is why etcd clusters use odd numbers, usually three or five: a three-member cluster keeps working if one member fails, because two of three is still a majority, but loses the ability to write if two fail. Kubernetes' documentation is clear that you need a backup plan, because a cluster whose etcd data is lost and unrecoverable is effectively gone.4 Regular etcd snapshots are not optional housekeeping; they are the difference between a bad afternoon and rebuilding from scratch.
# The backup that matters most
ETCDCTL_API=3 etcdctl snapshot save snapshot.db
# Lose your nodes and you reschedule. Lose etcd with no snapshot
# and you have lost the cluster's record of what should exist.04The scheduler: assigning pods to nodes
When a Pod is created it starts out unscheduled: it exists as desired state but has no node. The kube-scheduler watches for these unassigned Pods and, for each one, runs a two-stage decision: filtering (which nodes could run this Pod, given its resource requests, node selectors, taints and tolerations, and affinity rules?) and then scoring (of the feasible nodes, which is best?).5 It picks a node and writes that choice back. It does not start anything itself; it only decides where.
The pattern is worth noting: the scheduler's whole job is to fill in one field on the Pod, the node assignment. Something else acts on that decision.
05Controllers: the reconciliation loop
The kube-controller-manager runs a set of controllers, and controllers are the core of how Kubernetes works. Each one runs the same loop continuously: observe the current state of some part of the system, compare it to the desired state, and take one step to close the gap, then repeat.6 The Deployment controller notices you asked for three replicas but only two exist, and creates one. When a node fails and its Pods disappear, the relevant controllers see the shortfall and create replacements, which the scheduler then places.
This reconciliation loop is why Kubernetes is self-healing rather than only automated. You do not tell it how to recover from a failure; you tell it what should be true, and controllers work to make the running state match. It is also why kubectl apply only declares intent rather than issuing step-by-step commands.6
Kubernetes does not run your deployment as a script. It records what you want, and a set of control loops work continuously to make the cluster match that description, including after a machine fails.
06Why the control plane needs to be highly available
If the whole control plane is down, already-running Pods keep serving traffic for a while, because the kubelets on each node continue with what they were last told. But the cluster can no longer react: nothing new can be scheduled, no failure can be healed, and no change can be applied, because the only entry point and the only stored state are both unreachable. A control-plane outage is not a minor inconvenience; it is the loss of the cluster's ability to respond.
That is why production clusters run the control plane highly available: multiple API server replicas and a multi-member etcd cluster, spread across failure domains, so losing one machine costs nothing.7 For etcd, high availability and regular snapshots together are the survival plan, and the documented disaster-recovery path depends on having those snapshots.4 The control plane is the part to protect most carefully, because the nodes are replaceable but the control plane holds the record of what the cluster is supposed to be.