Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Explain the differences between Deployment and StatefulSet. Provide a scenario (e.g., a database cluster) that requires StatefulSet semantics and explain which StatefulSet features (stable network IDs, ordinal indexes, persistent volumes) are necessary.
Sample Answer
Deployment and StatefulSet both manage a set of pods from one template, but a Deployment treats every pod as interchangeable, while a StatefulSet gives each pod a stable, ordinal identity and its own dedicated storage. Reach for StatefulSet only when the application itself needs to know which replica it is, not just how many replicas exist.
Direct comparison
| Deployment | StatefulSet | |
|---|---|---|
| Pod naming | random suffix, no ordering | ordinal suffix (pod-0, pod-1, ...), stable across restarts |
| Network identity | pod IP changes on recreation; normally fronted by a load-balancing Service | stable hostname per ordinal, usually via a headless Service (clusterIP: None), so pod-0.svc.namespace.svc.cluster.local always resolves to the same replica |
| Storage | usually shared, external, or none | one PersistentVolumeClaim (PVC, a request for storage) bound to a PersistentVolume (PV, the actual storage resource) per ordinal, via volumeClaimTemplates; deleting the StatefulSet does not delete these PVCs |
| Startup and scale order | no ordering guarantee | default OrderedReady policy starts and stops one ordinal at a time, lowest-to-highest on scale-up and highest-to-lowest on scale-down; a Parallel policy is available when the app does not need this ordering |
Worked example: a PostgreSQL primary-replica cluster
A database cluster with one primary and several read replicas needs exactly the guarantees a Deployment cannot give:
- Stable hostnames: replication configuration and client connection strings can target
postgres-0.postgres-svcby name, and that name always means the same replica, even across restarts. - Ordinal indexes: an init script can branch on the pod's ordinal (ordinal 0 initializes as primary, ordinals 1 and above join as replicas) and rely on ordered startup so replicas do not attempt to join before the primary exists.
- Per-ordinal persistent volumes:
volumeClaimTemplatesties each replica's data directory to its own ordinal, sopostgres-1always reattaches topostgres-1's data after a restart, never to another replica's.
Trade-offs and pitfalls
- Not everything with 'stateful' in its name needs a StatefulSet. Kafka brokers do: each broker owns local partition log segments on disk and needs a stable identity for replica assignment. Kafka Connect workers do not: a Connect worker's task state and offsets live in Kafka topics, not on local disk, so the worker pool is effectively stateless and runs fine as a plain Deployment. Confirm where the durable state actually lives before defaulting to StatefulSet.
- StatefulSet's ordered rollout is slower and more conservative than a Deployment's by design; if a stateful workload's replicas are genuinely independent (sharded, no leader election), that ordering guarantee adds rollout latency with no real benefit, and a
Parallelpod-management policy, or even a Deployment with per-shard volumes mounted individually, may fit better. - Deleting a StatefulSet leaves its PVCs behind on purpose, to protect data from an accidental delete. Teams that do not know this get surprised in both directions: unexpected data loss if they assumed cleanup was automatic, or orphaned storage cost if they never clean it up.
Describe what a Pod is in Kubernetes and why it is considered the smallest deployable unit. Explain when you would run multiple containers in a single pod, how containers inside a pod share network and volumes, and trade-offs of co-locating containers such as sidecar patterns versus separate pods.
Sample Answer
A pod is the smallest unit Kubernetes schedules and manages: one or more containers that always run together on the same node, sharing a network namespace and, optionally, storage volumes. It is the unit, rather than the individual container, because the kubelet, the scheduler, and every controller reason about placement, restarts, and networking at that granularity, not below it.
Why the pod, not the container, is the atomic unit
- Co-scheduling guarantee: every container in a pod is placed on the same node and started, stopped, and restarted as a unit by the kubelet.
- One IP per pod: regardless of how many containers it holds, a pod gets exactly one IP address; Services and DNS resolve to that pod IP, not to an individual container.
- Shared network namespace: containers in a pod talk to each other over
localhost, and must not claim conflicting ports, because they share the same network namespace. - Shared volumes: volumes are declared once at the pod spec level and can be mounted into more than one container, which is how a sidecar can read or write files the main container produces without a network hop.
When to run more than one container in a pod
- Sidecar pattern: a helper process that shares the fate and resources of the main container, most commonly a log shipper, a proxy, or a metrics exporter.
- Init containers: containers that run to completion, in order, before the main containers start, typically for one-time setup like a schema migration or config templating.
- Native sidecar containers: since the
SidecarContainersfeature became enabled by default in Kubernetes 1.29 (stable as of 1.33), you can declare a sidecar underinitContainerswithrestartPolicy: Always. Kubernetes then starts it before the main container, keeps it running for the pod's whole life, and stops it after the main container on shutdown, which fixes the older ordering problem where a plain extra container might not be ready before the app started, or might outlive a completed Job's main container instead of shutting down with it.
A small example, an app container and a log-forwarding sidecar sharing a volume instead of a network call:
containers:
- name: app
image: myapp:1.4
volumeMounts:
- name: logs
mountPath: /var/log/app
- name: log-shipper
image: fluent-bit:latest
volumeMounts:
- name: logs
mountPath: /var/log/app
readOnly: true
volumes:
- name: logs
emptyDir: {}
The app writes to /var/log/app; the sidecar tails the same path through the shared emptyDir volume, no network hop involved.
Worked example: what a failing sidecar looks like
If the log-shipper container above enters a crash loop while app keeps running fine, the two containers' fates are independent even though they share a pod: app keeps serving traffic, but the log pipeline drops. This is visible directly in the pod list, where the READY column reports containers-ready over containers-total, not a single number:
NAME READY STATUS RESTARTS AGE
myapp-6d947f8db8-x2z1p 1/2 Running 4 (30s ago) 6m
Here 1/2 means only one of the pod's two containers is passing its readiness state, and the restart count of 4 belongs to the sidecar, not to app.
Trade-offs and pitfalls
- Co-location couples lifecycle: a sidecar cannot be scaled independently of the app the way a separate Deployment could be. Native sidecar containers loosen this slightly by giving the sidecar its own restart behavior, but it is still tied to the pod's schedule and node.
- Every container in the pod counts toward the pod's total resource footprint; forgetting the sidecar's own requests and limits when sizing the node is a common under-provisioning mistake.
- Splitting a component into a separate pod (plus a Service) regains independent scaling and blast-radius isolation, at the cost of a network hop and losing the localhost and shared-volume convenience. Reach for separate pods when the components genuinely differ in scaling or failure characteristics, for example a stateless API versus a shared cache, rather than defaulting to a sidecar for anything that happens to run alongside the main app.
Describe the Kubernetes pod lifecycle and common pod states (Pending, ContainerCreating, Running, Succeeded, Failed, Unknown, CrashLoopBackOff). For each state explain what it implies and list kubectl commands and API resources you would inspect to diagnose a pod that is not in Running state.
Sample Answer
A pod moves through Pending (accepted but not yet running), ContainerCreating (scheduled, kubelet is pulling images and mounting volumes), Running (at least one container is up), and finally either Succeeded or Failed for a pod that's meant to terminate, or a restart loop such as CrashLoopBackOff for one that keeps failing and coming back. Unknown is different in kind from the rest: it means the API server has simply lost contact with the node, not that anything about the pod itself is known to be wrong.
States, what each implies, and where to look
| State | What it implies | First place to look |
|---|---|---|
Pending | Not yet scheduled (capacity, affinity, or taint mismatch), or scheduled but the image can't be pulled yet | kubectl describe pod Events; look for FailedScheduling or ImagePullBackOff |
ContainerCreating | Scheduled; kubelet is pulling the image, attaching volumes, or setting up the pod's network namespace via the CNI (container network interface, the plugin that wires a pod into the cluster network) | kubectl describe pod; a stall here usually points to a slow registry pull or a PersistentVolumeClaim (PVC) that hasn't bound |
Running | At least one container is up (does not by itself mean the app is healthy or serving traffic; that's what readiness probes are for) | kubectl get pod -o yaml for status.conditions; kubectl logs |
Succeeded | All containers exited 0 (normal for a Job, unusual for a long-running Deployment pod) | kubectl logs <pod>; the owning Job's status for completion count |
Failed | A container exited non-zero and the pod's restartPolicy did not restart it | kubectl describe pod; kubectl logs --previous for the last container's output before it died |
CrashLoopBackOff | The container keeps exiting and Kubernetes is backing off between restart attempts with an increasing delay | kubectl logs --previous; kubectl describe pod for the exit code and reason under lastState |
Unknown | The API server can't get a status update from the node's kubelet | kubectl get nodes for NotReady; node/kubelet logs, not the pod itself |
A sample of what the Events section actually looks like for a crash-looping pod:
$ kubectl describe pod worker-6b9f-x2z4p -n prod
...
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Ready: False
Restart Count: 6
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning BackOff 38s (x5 over 90s) kubelet Back-off restarting failed container
OOMKilled here means the kernel's out-of-memory killer terminated the container because it exceeded its memory limit, not that the application itself crashed; that distinction changes the fix (raise the memory limit or reduce the footprint, versus debug application logic).
Graceful termination, and how it relates to Failed vs a clean stop
When a pod is deleted (including during a rolling update), Kubernetes first removes it from Service endpoints so it stops receiving new traffic, then sends SIGTERM to the container, runs any preStop hook first if one is defined, and waits up to terminationGracePeriodSeconds (30 seconds by default) before sending SIGKILL. A container that ignores SIGTERM and needs the full grace period before being force-killed can look, from the outside, like a slow or stuck termination rather than a clean stop; tuning the grace period and handling SIGTERM in the application are what make a shutdown graceful instead of abrupt.
Trade-offs and pitfalls
Runningis frequently mistaken for "healthy." A container can beRunningwhile its readiness probe fails continuously, meaning it's alive but not receiving traffic, which shows up instatus.conditions(Ready: False), not in the pod phase.FailedversusCrashLoopBackOffis really aboutrestartPolicy: withAlways(the Deployment default) a failing container becomesCrashLoopBackOffbecause Kubernetes keeps retrying with backoff; withNeverorOnFailurein the wrong combination a similar failure surfaces asFailedinstead, which changes which command shows you the useful evidence (--previouslogs only apply once a restart has actually happened).
Explain liveness, readiness, and startup probes in Kubernetes. For each type describe when it is evaluated, what consequences a failing probe has on pod lifecycle and traffic routing, and list best practices for implementing probes for a typical HTTP-based web service.
Sample Answer
Liveness, readiness, and startup probes all ask whether a container is okay, but each answer drives a different Kubernetes action: a failing liveness probe gets the container restarted, a failing readiness probe gets the pod pulled out of Service traffic without touching the container at all, and a startup probe simply delays the other two until the app has had time to boot.
What each probe gates
| Probe | Evaluated | Consequence on failure | Effect on traffic |
|---|---|---|---|
| Liveness | continuously, after the container starts | kubelet kills the container; it is recreated per the pod's restartPolicy | indirect only, through the restart |
| Readiness | continuously, independent of liveness | pod is marked NotReady and removed from the Service's Endpoints and EndpointSlices, the objects that track which pod IPs actually receive traffic | direct: no new requests are routed to it until it passes again |
| Startup | only until it first succeeds | container is killed and restarted if it fails before ever succeeding; liveness and readiness are not evaluated at all until it does | none directly, but it prevents liveness from killing a still-booting container |
Worked example: sizing a startup budget by workload archetype
What 'booting' means differs a lot by workload, and the startup probe has to be sized for the actual archetype, not guessed at: a machine learning (ML) inference service loading model weights into memory might need several minutes; a batch worker doing asynchronous Java Virtual Machine (JVM) warmup, classloading, and connection-pool initialization for an extract-transform-load (ETL) job might need under a minute; a stateless HTTP handler might be ready in under a second. Whichever number applies, it has to be encoded as periodSeconds times failureThreshold. Budgeting 5 minutes of startup headroom with a 10-second check interval for the ML case:
10×30=300s=5 min
means periodSeconds: 10 and failureThreshold: 30. Too tight in this calculation and the startup probe itself kills a healthy-but-slow container before it ever gets a chance to serve; too loose, and a genuinely stuck container burns minutes before anything reacts.
Trade-offs and pitfalls
- Swapping liveness and readiness is the classic mistake: pointing liveness at a deep dependency check (database reachability) means a transient database blip restarts every application pod at once instead of simply pulling them from rotation, turning a recoverable dependency issue into a self-inflicted outage.
- Using liveness as a substitute for a startup probe on a slow-booting app causes a restart loop before the app ever finishes initializing, since the container never survives long enough to pass a liveness check tuned for steady-state behavior.
- A readiness probe that is too permissive, common with the JVM-async pattern where the process starts accepting connections before its dependency pools are actually warm, reports the pod as ready while real requests still fail; that failure mode never shows up as a probe failure at all, only as user-visible errors.
Explain the differences between a Pod, ReplicaSet, Deployment, StatefulSet, and DaemonSet in Kubernetes. For each resource, describe its primary use cases, how it manages lifecycle and scaling, how updates/rollbacks are handled, and provide concrete examples of when you'd choose each resource (stateless web service, per-node agent, stateful database, etc.).
Sample Answer
Pods are the runtime unit; everything else here is a controller deciding how many pods to run and how to replace them. A ReplicaSet just keeps N identical pods alive. A Deployment wraps a ReplicaSet with declarative rolling updates and rollback history. A StatefulSet keeps stable per-replica identity and storage for workloads that care which replica they are. A DaemonSet guarantees exactly one pod per matching node. Job and CronJob run pods to completion, once or on a schedule, rather than keeping them running indefinitely.
Comparison
| Object | Identity model | Storage | Update mechanics | Reach for it when |
|---|---|---|---|---|
| ReplicaSet | anonymous, interchangeable pods | none built in | no native rolling update | almost never directly; it is what a Deployment manages underneath |
| Deployment | anonymous | typically shared, external, or none | rolling update via surge and unavailable settings, with full revision history and rollback | stateless services: web tiers, API servers |
| StatefulSet | stable ordinal identity (pod-0, pod-1, ...) and hostname | one PersistentVolumeClaim per ordinal via volumeClaimTemplates | ordered rolling update by default, highest ordinal first | stateful systems that need to know which replica they are: databases, brokers |
| DaemonSet | one pod per matching node | node-local, if any | rolling update per node, bounded by maxUnavailable | node-level agents: log collectors, CNI (Container Network Interface) plugins, monitoring exporters |
| Job | runs to completion, tracked by completions, parallelism, backoffLimit | ephemeral | retried on failure up to backoffLimit, not a rolling concept | one-off or batch work |
| CronJob | creates a Job on a schedule | ephemeral | governed by concurrencyPolicy (Allow, Forbid, or Replace) | scheduled batch work: nightly reports, cleanup jobs |
Worked example: rolling-update arithmetic
For a Deployment with replicas: 10, maxSurge: 25%, and maxUnavailable: 25%, Kubernetes rounds surge up and unavailable down:
surge=⌈10×0.25⌉=3
unavailable=⌊10×0.25⌋=2
So at the busiest point of the rollout the Deployment can briefly run up to 13 pods (10 plus 3 surge) while guaranteeing at least 8 stay available (10 minus 2 unavailable). This asymmetric rounding is why setting surge and unavailable to the same percentage still lets the total pod count grow slightly during a rollout instead of staying flat.
DaemonSets interact with node taints in a way that is easy to miss: they commonly carry an explicit toleration for the control-plane taint (conventionally node-role.kubernetes.io/control-plane:NoSchedule) so node-level agents like a CNI plugin or kube-proxy itself still run on control-plane nodes that ordinary application pods are excluded from. That toleration is a manifest choice the DaemonSet author makes, not automatic behavior.
Trade-offs and pitfalls
- Choosing Deployment for a workload that actually needs stable identity (a small quorum service, a primary-replica database) forces the team to rebuild that identity logic in application code; StatefulSet gives it for free, at the cost of slower, ordered rollouts.
- A stuck rollout is usually a readiness-probe failure hiding behind the surge and unavailable math: if new pods never pass readiness, the Deployment can stall indefinitely partway through, with old and new pods coexisting. Check pod events and readiness status before assuming a controller bug.
- A PodDisruptionBudget set too strictly (for example,
minAvailableequal to the full replica count) can make a Deployment's own rolling update unable to proceed, because the update's own temporary unavailability trips the budget meant to protect against unrelated disruptions. - DaemonSets bypass normal replica-count thinking: adding nodes silently adds pods and cost. Forgetting a DaemonSet exists across a large node pool is a common source of untracked resource consumption.
Unlock Full Question Bank
Get access to all 13 Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.