Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Describe what a Pod is in Kubernetes and why it is considered the smallest deployable unit. Explain when you would run multiple containers in a single pod, how containers inside a pod share network and volumes, and trade-offs of co-locating containers such as sidecar patterns versus separate pods.
Sample Answer
A pod is the smallest unit Kubernetes schedules and manages: one or more containers that always run together on the same node, sharing a network namespace and, optionally, storage volumes. It is the unit, rather than the individual container, because the kubelet, the scheduler, and every controller reason about placement, restarts, and networking at that granularity, not below it.
Why the pod, not the container, is the atomic unit
- Co-scheduling guarantee: every container in a pod is placed on the same node and started, stopped, and restarted as a unit by the kubelet.
- One IP per pod: regardless of how many containers it holds, a pod gets exactly one IP address; Services and DNS resolve to that pod IP, not to an individual container.
- Shared network namespace: containers in a pod talk to each other over
localhost, and must not claim conflicting ports, because they share the same network namespace. - Shared volumes: volumes are declared once at the pod spec level and can be mounted into more than one container, which is how a sidecar can read or write files the main container produces without a network hop.
When to run more than one container in a pod
- Sidecar pattern: a helper process that shares the fate and resources of the main container, most commonly a log shipper, a proxy, or a metrics exporter.
- Init containers: containers that run to completion, in order, before the main containers start, typically for one-time setup like a schema migration or config templating.
- Native sidecar containers: since the
SidecarContainersfeature became enabled by default in Kubernetes 1.29 (stable as of 1.33), you can declare a sidecar underinitContainerswithrestartPolicy: Always. Kubernetes then starts it before the main container, keeps it running for the pod's whole life, and stops it after the main container on shutdown, which fixes the older ordering problem where a plain extra container might not be ready before the app started, or might outlive a completed Job's main container instead of shutting down with it.
A small example, an app container and a log-forwarding sidecar sharing a volume instead of a network call:
containers:
- name: app
image: myapp:1.4
volumeMounts:
- name: logs
mountPath: /var/log/app
- name: log-shipper
image: fluent-bit:latest
volumeMounts:
- name: logs
mountPath: /var/log/app
readOnly: true
volumes:
- name: logs
emptyDir: {}
The app writes to /var/log/app; the sidecar tails the same path through the shared emptyDir volume, no network hop involved.
Worked example: what a failing sidecar looks like
If the log-shipper container above enters a crash loop while app keeps running fine, the two containers' fates are independent even though they share a pod: app keeps serving traffic, but the log pipeline drops. This is visible directly in the pod list, where the READY column reports containers-ready over containers-total, not a single number:
NAME READY STATUS RESTARTS AGE
myapp-6d947f8db8-x2z1p 1/2 Running 4 (30s ago) 6m
Here 1/2 means only one of the pod's two containers is passing its readiness state, and the restart count of 4 belongs to the sidecar, not to app.
Trade-offs and pitfalls
- Co-location couples lifecycle: a sidecar cannot be scaled independently of the app the way a separate Deployment could be. Native sidecar containers loosen this slightly by giving the sidecar its own restart behavior, but it is still tied to the pod's schedule and node.
- Every container in the pod counts toward the pod's total resource footprint; forgetting the sidecar's own requests and limits when sizing the node is a common under-provisioning mistake.
- Splitting a component into a separate pod (plus a Service) regains independent scaling and blast-radius isolation, at the cost of a network hop and losing the localhost and shared-volume convenience. Reach for separate pods when the components genuinely differ in scaling or failure characteristics, for example a stateless API versus a shared cache, rather than defaulting to a sidecar for anything that happens to run alongside the main app.
Describe the Kubernetes pod lifecycle and common pod states (Pending, ContainerCreating, Running, Succeeded, Failed, Unknown, CrashLoopBackOff). For each state explain what it implies and list kubectl commands and API resources you would inspect to diagnose a pod that is not in Running state.
Sample Answer
A pod moves through Pending (accepted but not yet running), ContainerCreating (scheduled, kubelet is pulling images and mounting volumes), Running (at least one container is up), and finally either Succeeded or Failed for a pod that's meant to terminate, or a restart loop such as CrashLoopBackOff for one that keeps failing and coming back. Unknown is different in kind from the rest: it means the API server has simply lost contact with the node, not that anything about the pod itself is known to be wrong.
States, what each implies, and where to look
| State | What it implies | First place to look |
|---|---|---|
Pending | Not yet scheduled (capacity, affinity, or taint mismatch), or scheduled but the image can't be pulled yet | kubectl describe pod Events; look for FailedScheduling or ImagePullBackOff |
ContainerCreating | Scheduled; kubelet is pulling the image, attaching volumes, or setting up the pod's network namespace via the CNI (container network interface, the plugin that wires a pod into the cluster network) | kubectl describe pod; a stall here usually points to a slow registry pull or a PersistentVolumeClaim (PVC) that hasn't bound |
Running | At least one container is up (does not by itself mean the app is healthy or serving traffic; that's what readiness probes are for) | kubectl get pod -o yaml for status.conditions; kubectl logs |
Succeeded | All containers exited 0 (normal for a Job, unusual for a long-running Deployment pod) | kubectl logs <pod>; the owning Job's status for completion count |
Failed | A container exited non-zero and the pod's restartPolicy did not restart it | kubectl describe pod; kubectl logs --previous for the last container's output before it died |
CrashLoopBackOff | The container keeps exiting and Kubernetes is backing off between restart attempts with an increasing delay | kubectl logs --previous; kubectl describe pod for the exit code and reason under lastState |
Unknown | The API server can't get a status update from the node's kubelet | kubectl get nodes for NotReady; node/kubelet logs, not the pod itself |
A sample of what the Events section actually looks like for a crash-looping pod:
$ kubectl describe pod worker-6b9f-x2z4p -n prod
...
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Ready: False
Restart Count: 6
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning BackOff 38s (x5 over 90s) kubelet Back-off restarting failed container
OOMKilled here means the kernel's out-of-memory killer terminated the container because it exceeded its memory limit, not that the application itself crashed; that distinction changes the fix (raise the memory limit or reduce the footprint, versus debug application logic).
Graceful termination, and how it relates to Failed vs a clean stop
When a pod is deleted (including during a rolling update), Kubernetes first removes it from Service endpoints so it stops receiving new traffic, then sends SIGTERM to the container, runs any preStop hook first if one is defined, and waits up to terminationGracePeriodSeconds (30 seconds by default) before sending SIGKILL. A container that ignores SIGTERM and needs the full grace period before being force-killed can look, from the outside, like a slow or stuck termination rather than a clean stop; tuning the grace period and handling SIGTERM in the application are what make a shutdown graceful instead of abrupt.
Trade-offs and pitfalls
Runningis frequently mistaken for "healthy." A container can beRunningwhile its readiness probe fails continuously, meaning it's alive but not receiving traffic, which shows up instatus.conditions(Ready: False), not in the pod phase.FailedversusCrashLoopBackOffis really aboutrestartPolicy: withAlways(the Deployment default) a failing container becomesCrashLoopBackOffbecause Kubernetes keeps retrying with backoff; withNeverorOnFailurein the wrong combination a similar failure surfaces asFailedinstead, which changes which command shows you the useful evidence (--previouslogs only apply once a restart has actually happened).
Explain the Kubernetes networking model in detail for a DevOps team unfamiliar with it. Describe the IP-per-pod concept, the flat cluster network assumption, how pod-to-pod communication works across nodes, the role of the Container Network Interface (CNI) and kube-proxy, and any common limitations or implicit assumptions operators should be aware of when designing cluster networking.
Sample Answer
Kubernetes gives every Pod its own IP address on one flat, cluster-wide network where any Pod can reach any other Pod's IP directly, without port-mapping or NAT (Network Address Translation) in the middle. That single assumption is the whole model; everything else (Services, kube-proxy, the CNI plugin) exists to make that flat network real and to add load-balancing and service discovery on top of it.
The core assumption: IP-per-Pod, flat and routable
- Every Pod gets its own IP, allocated from the cluster's Pod network, not from the node's own network.
- Containers inside the same Pod share that IP and can reach each other over
localhost. - Any Pod's IP is expected to be reachable from any node in the cluster, as if the whole cluster were one big L3 (Layer 3, meaning IP-address-level) network, even though physically it spans many separate machines.
This is a deliberate simplification: application code never has to deal with port-mapping or think about which node it's running on to reach another Pod. The cost of that simplification is pushed down into the networking layer that has to make the "flat network" illusion actually true.
How Pod-to-pod traffic actually crosses nodes
Two Pods on the same node talk through a local bridge or virtual interface, no different in spirit from two processes on one machine. Two Pods on different nodes need their packets to physically cross the network between those nodes, and that's the job of the CNI (Container Network Interface) plugin, using one of a few common approaches:
- Overlay/encapsulation (for example VXLAN, used by Flannel by default): the CNI wraps the Pod packet inside another packet addressed node-to-node, then unwraps it on arrival. Simple to run, but every packet pays an encapsulation cost.
- Native routing (for example Calico's BGP mode): nodes advertise routes to each other's Pod IP ranges directly, so packets travel as ordinary routed IP traffic with no wrapping.
- In-kernel eBPF forwarding (for example Cilium): programs attached to the kernel's networking hooks forward Pod traffic with less per-packet overhead than either of the above.
Whichever mechanism is used, kube-proxy is not involved in raw Pod-to-Pod traffic; that's entirely the CNI's job. kube-proxy only comes into play for Service traffic, described next.
Where kube-proxy fits in
A Pod's IP is not stable: Pods get rescheduled, restarted, and replaced with new IPs constantly. A Service gives client code one stable address (a ClusterIP) that represents a group of Pods, and kube-proxy is the component that makes connections to that stable address actually land on one of the current, healthy backend Pods. It does this by programming the node's kernel with forwarding rules (commonly iptables or IPVS, an in-kernel load-balancing feature) built from the Service's list of ready backend Pods. Put simply: the CNI plugin makes any Pod reachable by its own IP; kube-proxy makes a stable Service IP resolve to whichever real Pod IP should currently handle the traffic.
Common limitations and implicit assumptions to watch for
- The underlying network must actually support it. Cloud VPC (Virtual Private Cloud) routing tables, security groups, or on-prem firewalls have to allow the CNI's chosen traffic pattern (raw routed IP, VXLAN-encapsulated UDP, or eBPF-forwarded packets) between every pair of nodes, or the "flat network" assumption breaks silently.
- IP address planning matters. The Pod CIDR (the address range Pods are allocated from) needs to be sized for cluster growth and must not overlap with the VPC's own address space, or routing becomes ambiguous.
- NetworkPolicy is opt-in. Without it, any Pod can reach any other Pod by default (east-west traffic is wide open); the flat-network model is about reachability, not isolation, and isolation has to be added deliberately.
- Encapsulation costs latency and throughput. An overlay CNI adds a real, if small, per-packet tax versus native routing or eBPF; this matters more as traffic volume grows.
- kube-proxy itself has a scaling ceiling. iptables-based rule chains get slower to evaluate as the number of Services and endpoints grows; IPVS or an eBPF-based CNI's own Service handling scales better at high counts.
Worked example
A Pod on Node A (IP 10.244.1.7) calling a Service backed by a Pod on Node B (IP 10.244.2.4):
- The calling Pod connects to the Service's ClusterIP, say
10.96.10.20:80. - kube-proxy's node-local rules (built from the Service's endpoint list) DNAT that connection to
10.244.2.4:8080, the real backend Pod. - The CNI plugin now has to deliver a packet addressed to
10.244.2.4, a Pod IP on a different node, across the physical network: an overlay CNI wraps it in a VXLAN packet destined for Node B's real IP; a routed CNI simply routes it there directly using advertised Pod-subnet routes. - Node B unwraps (if needed) and delivers the packet to the backend Pod over its local bridge.
Trade-offs and pitfalls for a team new to this model
- Don't assume Pod IPs are stable enough to hardcode anywhere; always address other workloads through a Service name, never a Pod IP directly.
- Introducing a service mesh or NetworkPolicy changes this baseline model by adding sidecar proxies or enforcement points into the path described above; treat this explanation as the foundation those layers build on, not the final picture.
- Multi-cluster or hybrid-cloud setups often break the single-flat-network assumption outright (two clusters don't share one Pod CIDR space by default), which is exactly why solutions in that space exist to stitch separate flat networks back together.
Explain liveness, readiness, and startup probes in Kubernetes. For each type describe when it is evaluated, what consequences a failing probe has on pod lifecycle and traffic routing, and list best practices for implementing probes for a typical HTTP-based web service.
Sample Answer
Liveness, readiness, and startup probes all ask whether a container is okay, but each answer drives a different Kubernetes action: a failing liveness probe gets the container restarted, a failing readiness probe gets the pod pulled out of Service traffic without touching the container at all, and a startup probe simply delays the other two until the app has had time to boot.
What each probe gates
| Probe | Evaluated | Consequence on failure | Effect on traffic |
|---|---|---|---|
| Liveness | continuously, after the container starts | kubelet kills the container; it is recreated per the pod's restartPolicy | indirect only, through the restart |
| Readiness | continuously, independent of liveness | pod is marked NotReady and removed from the Service's Endpoints and EndpointSlices, the objects that track which pod IPs actually receive traffic | direct: no new requests are routed to it until it passes again |
| Startup | only until it first succeeds | container is killed and restarted if it fails before ever succeeding; liveness and readiness are not evaluated at all until it does | none directly, but it prevents liveness from killing a still-booting container |
Worked example: sizing a startup budget by workload archetype
What 'booting' means differs a lot by workload, and the startup probe has to be sized for the actual archetype, not guessed at: a machine learning (ML) inference service loading model weights into memory might need several minutes; a batch worker doing asynchronous Java Virtual Machine (JVM) warmup, classloading, and connection-pool initialization for an extract-transform-load (ETL) job might need under a minute; a stateless HTTP handler might be ready in under a second. Whichever number applies, it has to be encoded as periodSeconds times failureThreshold. Budgeting 5 minutes of startup headroom with a 10-second check interval for the ML case:
10×30=300s=5 min
means periodSeconds: 10 and failureThreshold: 30. Too tight in this calculation and the startup probe itself kills a healthy-but-slow container before it ever gets a chance to serve; too loose, and a genuinely stuck container burns minutes before anything reacts.
Trade-offs and pitfalls
- Swapping liveness and readiness is the classic mistake: pointing liveness at a deep dependency check (database reachability) means a transient database blip restarts every application pod at once instead of simply pulling them from rotation, turning a recoverable dependency issue into a self-inflicted outage.
- Using liveness as a substitute for a startup probe on a slow-booting app causes a restart loop before the app ever finishes initializing, since the container never survives long enough to pass a liveness check tuned for steady-state behavior.
- A readiness probe that is too permissive, common with the JVM-async pattern where the process starts accepting connections before its dependency pools are actually warm, reports the pod as ready while real requests still fail; that failure mode never shows up as a probe failure at all, only as user-visible errors.
A pod remains in Pending and the scheduler does not bind it. Describe commands and checks to determine why the pod is unscheduled: inspect resource requests/limits, node allocatable capacity, taints/tolerations, affinity rules, and namespace quotas. Mention concrete kubectl commands you would run.
Sample Answer
Pending means the API server has accepted the pod object but the scheduler has not bound it to a node. Almost always the fastest path to the answer is reading the scheduler's own FailedScheduling event, which names the exact reason in plain text; everything below is really just "go read that message, then verify it against the right kubectl output."
Reading the signal
kubectl describe pod <pod> -n <ns>
kubectl get events -n <ns> --sort-by=.metadata.creationTimestamp
Typical event text and what it maps to:
| Event message pattern | What it means | Where to look next |
|---|---|---|
0/3 nodes are available: 3 Insufficient cpu | Every node lacks enough allocatable CPU for the pod's request | kubectl describe node, compare Allocated resources against the pod's resources.requests |
0/3 nodes are available: 3 node(s) had untolerated taint {key: value} | Nodes are tainted and the pod has no matching toleration | kubectl describe node (Taints section) vs pod.spec.tolerations |
0/3 nodes are available: 3 node(s) didn't match Pod's node affinity/selector | nodeSelector or nodeAffinity excludes every node | pod.spec.affinity / pod.spec.nodeSelector vs node labels |
pod has unbound immediate PersistentVolumeClaims | A PersistentVolumeClaim (PVC) the pod needs isn't bound to a PersistentVolume (PV) yet | kubectl get pvc -n <ns>; check the StorageClass and provisioner (a full storage debugging pass is its own topic, this is just the scheduling-side symptom) |
Checking each dimension directly
Resources.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.containers[*].resources}'
kubectl describe node <node>
kubectl top nodes # if metrics-server is installed
Compare the pod's requests (what the scheduler actually reasons about, not limits) against each node's allocatable capacity minus what's already committed.
Taints and tolerations.
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.taints}{"\n"}{end}'
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.tolerations}'
Affinity.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.affinity}'
A requiredDuringSchedulingIgnoredDuringExecution rule is a hard constraint; if no node satisfies it, the pod stays Pending indefinitely. A preferredDuringSchedulingIgnoredDuringExecution rule is a soft hint and will not by itself cause Pending.
Namespace-level limits: two different mechanisms, easy to conflate. A ResourceQuota is enforced at admission time, when the pod object is created, not at scheduling time. If a create request would push the namespace over its quota, the API server rejects it outright: kubectl apply returns an error immediately, and no Pending pod is ever created. So if you're looking at a Pending pod, ResourceQuota is not why. What can still bite you at scheduling time is a LimitRange: if the namespace has a default request injected by a LimitRange, a pod that specified no request at all can end up with a larger effective request than you intended, which then genuinely fails to fit any node and produces the Insufficient cpu event above.
kubectl get resourcequota -n <ns>
kubectl get limitrange -n <ns> -o yaml
Trade-offs and pitfalls
- The scheduler reasons about
requests, neverlimits; a pod with a tiny request and a huge limit schedules easily and can still get evicted or throttled later, which is a different problem from Pending. - A
nodeSelectortypo (a label key or value that doesn't exist on any node) produces the exact same symptom as a genuine capacity shortage; always check the label actually exists before assuming you need more nodes. - Don't chase ResourceQuota when you see a Pending pod; check LimitRange defaults and the scheduler's own event message instead, since quota violations block creation, not scheduling.
Unlock Full Question Bank
Get access to all 13 Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.