Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
The Kubernetes API server is experiencing increased request latency. What metrics, logs, and traces would you collect to diagnose whether the bottleneck is etcd, admission controllers, or API server CPU/memory? Provide a prioritized triage checklist and remedial actions for each root cause.
Sample Answer
Direct answer
Increased kube-apiserver latency almost always traces to one of three places: etcd itself (disk or network bound), a slow admission webhook sitting in the request path, or the apiserver process running short of CPU or memory (including its own request-concurrency limiting kicking in). The fastest way to tell them apart is to look at where time is spent inside a single slow request, etcd round trip versus admission call versus everything else, rather than guessing from symptoms alone.
Structured elaboration
Metrics to pull first
| Signal | What it tells you |
|---|---|
apiserver_request_duration_seconds (histogram, by verb/resource) | Overall request latency, and whether it is one resource type or global |
apiserver_current_inflight_requests | How close the server is to its concurrency ceiling |
apiserver_flowcontrol_rejected_requests_total, apiserver_flowcontrol_current_inqueue_requests, apiserver_flowcontrol_current_executing_requests | Whether API Priority and Fairness (APF, stable since Kubernetes 1.29, the mechanism that classifies and queues requests by priority) is queuing, executing, or rejecting requests for a given priority level |
apiserver_admission_webhook_admission_duration_seconds | Per-webhook admission latency, split by mutating/validating |
etcd_disk_wal_fsync_duration_seconds, etcd_disk_backend_commit_duration_seconds | Disk-bound etcd latency; sustained spikes point at disk contention |
etcd_server_has_leader, etcd_server_leader_changes_seen_total | Whether etcd has a stable leader or is re-electing |
Triage order
- Scope it: slice
apiserver_request_duration_secondsby resource and verb. If only one resource type or client is slow, an etcd-wide or apiserver-wide problem is unlikely; look at that resource's admission webhooks first. - Check APF rejection reasons:
apiserver_flowcontrol_rejected_requests_totallabeledreason="queue-full"orreason="concurrency-limit"means the server is intentionally shedding load under its configured priority levels. The 429 responses clients see in that case are the mechanism doing its job, not a hidden bug; the real bottleneck is upstream (etcd or CPU), not APF itself. - If the slowdown is global, every resource, every client, check etcd's disk metrics and leader stability before touching the apiserver. A slow etcd backend shows up as apiserver latency because every write and every quorum read waits on it.
- If etcd looks healthy but apiserver CPU or memory is pegged, it is apiserver resource pressure.
Remedial actions per root cause
- etcd bound: faster disks (etcd is fsync-latency sensitive, not throughput sensitive), a disk dedicated to etcd separate from other I/O, defragmentation, and checking
etcd_mvcc_db_total_size_in_bytes(a bloated database slows every commit). Reduce write volume from noisy controllers before adding etcd members: more members raise the replication cost of every write, they do not spread load the way a read replica would. - Admission webhook bound: tune webhook
timeoutSecondsandfailurePolicycarefully, scale the webhook backend, and reconsider whether the check belongs in a webhook at all versus static OpenAPI schema validation. A webhook earns its cost when the rule needs data outside the request object, cross-field logic, or an external lookup; anything expressible as a plain schema constraint should live in the CRD's (Custom Resource Definition's) OpenAPI validation instead, since that costs nothing at admission time. - apiserver resource bound: unlike etcd, kube-apiserver is stateless, so horizontally scaling it (more apiserver replicas behind the control-plane load balancer) is a legitimate, common fix, not a workaround. Also check audit log verbosity and watch cardinality; a controller opening many broad watches is a frequent, overlooked CPU driver.
Worked example (concept, not a fabricated benchmark)
Suppose apiserver_request_duration_seconds p99 for PATCH pods is elevated but p99 for every other verb and resource is flat. That shape alone rules out an etcd-wide or apiserver-wide bottleneck, because both would show up across every resource type, and points at something specific to pod patches, almost always a mutating webhook registered on Pods (for example a sidecar injector). Confirming it takes one more step: check apiserver_admission_webhook_admission_duration_seconds filtered to that webhook's name. A rising p99 there, correlated with the PATCH pods latency, closes the loop without needing to touch etcd or CPU metrics at all.
Trade-offs and pitfalls
- Do not disable webhooks blind as a first move.
failurePolicy: Ignoreon a security-relevant mutating webhook (one that injects a sidecar or a required label, for example) can silently change what gets admitted, not just how fast. - Adding etcd members to "spread the load" is a common but wrong instinct: every additional voting member adds replication overhead to every write. It improves fault tolerance, not throughput.
- 429 responses from APF are a symptom of an upstream bottleneck, not a target to eliminate by raising limits. Raising a priority level's concurrency share without fixing the underlying etcd or CPU constraint just moves where the queue backs up.
Describe the core components of the Kubernetes control plane (API server, etcd, scheduler, controller-manager, cloud-controller-manager). For each component explain its primary responsibility, how it persists or interacts with cluster state, typical failure modes, and what operational metrics you would monitor to detect trouble.
Sample Answer
Kubernetes splits cluster management into a control plane, which decides and records desired state, and worker nodes, which run it. The control plane's core pieces are the kube-apiserver (the front door that validates and serves every request), etcd (the single source of truth for cluster state), the scheduler (decides which node a new pod lands on), the controller-manager (a bundle of reconciliation loops that push actual state toward desired state), and, on cloud-hosted clusters, the cloud-controller-manager (the seam that keeps cloud-specific logic like load balancer provisioning out of core Kubernetes). Every one of these follows the same pattern: watch the API server for objects it cares about, and reconcile until observed state matches spec.
Control plane components
| Component | Primary responsibility | How it touches state | Common failure signature |
|---|---|---|---|
| kube-apiserver | validates, authenticates, and serves the cluster API | reads/writes every object through etcd; the only component that talks to etcd directly | rising p99 on apiserver_request_duration_seconds, climbing 4xx/5xx rates, certificate expiry |
| etcd | strongly-consistent key-value store for all cluster objects | is the persistence layer itself | quorum loss, disk I/O saturation, rapid leader churn |
| kube-scheduler | assigns unscheduled pods to a node | watches the API server for unbound pods, writes the binding back through it | growing count of Pending pods, rising scheduling latency |
| kube-controller-manager | runs the reconciliation loops (ReplicaSet, node lifecycle, endpoints, and more) | watches and updates objects through the API server | stuck reconciliation, leader-election flapping in an HA control plane |
| cloud-controller-manager | integrates cloud-specific logic (load balancers, routes, node lifecycle) | talks to both the API server and the cloud provider's API | a Service stuck without an external address, provisioning errors surfacing as cloud API failures |
flowchart TD
Client[kubectl / clients] --> API[kube-apiserver]
API --> ETCD[(etcd)]
Sched[kube-scheduler] -->|watch unscheduled pods, write bindings| API
CM[controller-manager] -->|watch + reconcile| API
CCM[cloud-controller-manager] --> API
CCM -->|provision LB, routes| Cloud[Cloud provider API]
API --> Kubelet[kubelet, per node]
Kubelet --> CRI[container runtime, via CRI]
The API server's gatekeeping
Every request passes through three stages before it touches etcd: authentication (who are you: client certificate, bearer token, or an external identity provider via OIDC), authorization (are you allowed: almost always Role-Based Access Control, RBAC, checking your identity against Roles and RoleBindings), and admission (should this specific object be allowed or modified: built-in admission controllers plus optional mutating and validating admission webhooks, which is also the mechanism behind things like automatic sidecar injection). A gap in any one of the three shows up as a very different symptom: authentication failures look like connection refusals, authorization failures return a 403, and admission failures reject an otherwise well-formed object with a specific rejection reason from the webhook or controller.
Worker-node components: kubelet, kube-proxy, and the container runtime
The control plane decides; the node executes. The kubelet is the agent on every node that watches the API server for pods assigned to that node and drives the container lifecycle through the CRI (Container Runtime Interface), a plugin boundary that lets Kubernetes talk to any compliant runtime (containerd and CRI-O are the common choices today; Docker itself was removed as a supported CRI implementation in Kubernetes 1.24). kube-proxy implements the networking side of a Service on each node; the mechanics of that (iptables, IPVS, or the newer nftables backend) belong to Service and networking questions rather than control-plane architecture, but it is worth knowing kube-proxy is a node component, not a control-plane one.
Stepping back, the reason Kubernetes is built this way rather than as a single monolithic scheduler is the reconciliation model itself: every component only has to compare desired state to observed state and take one corrective step, repeatedly, which is what makes the system self-healing and declarative rather than a one-shot deployment tool.
Worked example: reasoning about an etcd quorum failure
etcd tolerates the loss of a minority of its members because it uses a majority-vote (Raft) protocol; for a cluster of n members it needs:
quorum=⌊n/2⌋+1
For the common 5-member etcd cluster:
⌊5/2⌋+1=2+1=3
so it tolerates 2 simultaneous member failures while still accepting writes. This is also why an operator should never round a fault-tolerance target up to an even member count: a 4-member cluster still only tolerates 1 failure (quorum is 3), the same as a 3-member cluster, but pays for a fourth voter with no extra fault tolerance. Two metrics tell you this is happening before it becomes an outage: a sustained rise in etcd_server_leader_changes_seen_total (frequent leader changes usually mean the disk cannot keep up with etcd's Raft heartbeat interval) and a simultaneous rise in apiserver_request_duration_seconds p99, since every write now waits on a less stable etcd leader.
Trade-offs and pitfalls
- Stacked etcd (co-located with control-plane nodes) is simpler to run but ties etcd's failure domain to the same nodes serving the API; an external etcd cluster isolates that blast radius at the cost of more infrastructure to operate.
- On managed clusters (Amazon Elastic Kubernetes Service, Google Kubernetes Engine, Azure Kubernetes Service) the provider hides and operates the control plane entirely; you cannot inspect etcd directly, so day-to-day monitoring shifts to the provider's exposed control-plane metrics and SLA rather than self-run dashboards.
- cloud-controller-manager problems are easy to misdiagnose as networking bugs: a Service stuck in
<pending>for its external address is very often a cloud-controller-manager or cloud-API quota issue, not a kube-proxy or CNI (Container Network Interface) problem, so check its logs before chasing the wrong component.
Explain Pod Disruption Budgets (PDBs). How do PDBs interact with rolling updates, cluster autoscaler evictions, and maintenance operations? Provide scenarios where an incorrect PDB could block upgrades or autoscaling and how to fix those issues.
Sample Answer
A PodDisruptionBudget (PDB) tells the cluster how much of a replicated workload is allowed to be taken down at once by a voluntary disruption, things a human or controller chooses to do, such as a node drain, a cluster upgrade, or the cluster autoscaler removing a node. It has no effect on involuntary disruptions: a node crashing, a container getting OOMKilled (killed by the kernel's out-of-memory killer for exceeding its memory limit), or hardware failure all bypass it entirely, because there's no eviction request for the PDB to block in those cases.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web-pdb
spec:
selector:
matchLabels: { app: web }
minAvailable: 2
(policy/v1 is the stable, current API group and version for this object; older manifests using policy/v1beta1 are targeting a removed API and will fail to apply on a current cluster.)
How PDBs interact with each operation
- Rolling updates: the Deployment controller's own
maxUnavailable/maxSurgesettings already bound disruption during a rollout, and a PDB adds an independent floor on top that applies across all voluntary disruptions hitting that workload at once, not just the rollout in isolation. If a rollout and a node drain happen to overlap, the PDB is what prevents their combined effect from dropping available replicas too far. - Cluster Autoscaler (CA) scale-down: before removing a node, CA evicts the pods on it through the eviction API, and that eviction is refused if it would violate a pod's PDB. CA then either finds another node to remove or skips that node until eviction becomes possible; it does not force through a PDB violation.
- Manual maintenance (
kubectl drain, akubeadmupgrade, a node reboot): these also go through the eviction API and are blocked the same way; an operator hitting a blocked drain either waits, adjusts the PDB, or in genuine emergencies deletes pods directly (bypassing the eviction API, which also bypasses the PDB, and should be a deliberate, audited exception rather than a routine workaround).
Problem scenarios and fixes
| Scenario | Why it blocks | Fix |
|---|---|---|
minAvailable: 2 on a 2-replica Deployment | Zero disruption tolerance; nothing can ever be evicted | Add a third replica, or switch to maxUnavailable: 1 if a 2-replica service can genuinely tolerate a brief single-replica window |
| A StatefulSet's PDB blocks Cluster Autoscaler from ever draining its last node | CA can't evict any of the StatefulSet's pods without breaching the budget, so it leaves that node running indefinitely | Loosen the PDB to allow at least one voluntary disruption, or, if the workload genuinely cannot tolerate any drop, keep it on dedicated (non-scaled-down) nodes instead of fighting the autoscaler |
| Many services' PDBs collectively stall a cluster-wide upgrade | Each PDB is individually reasonable, but the combined effect blocks progress node by node | Stagger the upgrade, use percentage-based minAvailable/maxUnavailable so budgets scale with replica count, and give the upgrade a maintenance window with clear escalation if it stalls past expected duration |
Trade-offs and pitfalls
- Prefer percentage-based values (
maxUnavailable: 25%) over absolute counts for workloads whose replica count changes with load; an absoluteminAvailable: 3on a Deployment that autoscales down to 2 replicas becomes an unsatisfiable budget that blocks every voluntary disruption until it scales back up. - A PDB with
minAvailable: 100%(or an absolute count equal to current replicas) reads as "maximally safe" but actually means zero voluntary disruptions are ever permitted, which silently blocks every future drain and upgrade until someone notices and loosens it, often under time pressure during an incident. - PDBs protect against voluntary disruption stacking up in ways a single controller's own settings can't see; they are not a substitute for redundancy. A workload with only one replica and a PDB requiring
minAvailable: 1gains nothing from the PDB, since that one pod being unavailable for any reason, voluntary or not, is already an outage.
List common cloud and network-backed storage options used with Kubernetes (examples: AWS EBS, AWS EFS, GCE PD, Azure Disk, NFS) and briefly describe trade-offs in terms of performance, durability, multi-node attach, and typical use-cases.
Sample Answer
Cloud and network-backed storage for Kubernetes splits into two families: block storage (AWS EBS, GCE PD, Azure Disk), which is fast and durable but normally attachable to only one node at a time, and network filesystems (AWS EFS, Azure Files, self-managed NFS), which are shareable across many nodes at once but pay a latency and throughput cost for that flexibility. Picking between them is really picking whether the workload needs raw single-writer performance or multi-node shared access.
Comparison
| Option | Performance | Durability | Multi-node attach | Typical use case |
|---|---|---|---|---|
| AWS EBS (Elastic Block Store) | High IOPS (input/output operations per second) and throughput on provisioned tiers; low latency | Replicated within the Availability Zone by AWS | Single-writer (ReadWriteOnce) for ordinary use; a Multi-Attach mode exists for specific volume types but requires a cluster-aware filesystem and is the exception, not the default | Databases, single-node stateful workloads |
| AWS EFS (Elastic File System) | Network filesystem; throughput scales with configured mode but per-operation latency is higher and more variable than block storage | Replicated across multiple Availability Zones by AWS | ReadWriteMany: many Pods across many nodes can mount concurrently | Shared config/assets, CI caches, content shared across replicas |
| GCE PD (Persistent Disk) | Strong block performance; low latency within a zone | Zonal by default; a regional PD variant replicates synchronously across two zones for higher availability | Single-writer for normal use; a multi-writer mode exists on specific disk types but is restricted and still expects the application to coordinate writes itself, since it is not a cluster filesystem | Databases, single-node stateful apps |
| Azure Disk | High IOPS/throughput on Premium/Ultra tiers | Replicated within the region/zone by Azure | Single-writer (ReadWriteOnce) | Block storage for VMs/Pods needing high, predictable performance |
| Azure Files | SMB/NFS network filesystem semantics | Managed, replicated by Azure | ReadWriteMany | Shared config, home directories, app assets |
| NFS (self-managed) | Depends entirely on the server and network path; can become a shared bottleneck | Depends on how the operator makes the NFS server itself highly available; no built-in durability beyond what you build | ReadWriteMany | Simple shared storage, legacy applications expecting a shared filesystem |
How to choose
- Single-writer, latency-sensitive, durable (a relational database's primary, a message queue's log): block storage (EBS, GCE PD, Azure Disk). Access mode ReadWriteOnce, sized and provisioned for the IOPS the workload actually needs.
- Shared, multi-reader-or-writer, latency-tolerant (shared configuration, static assets, a CI build cache used by many concurrent jobs): a managed network filesystem (EFS, Azure Files) if available on your cloud, or self-managed NFS if not, understanding that NFS's durability and availability are now your responsibility to engineer.
- Regional or multi-zone resilience for a block-storage workload: look at the provider's own cross-zone replication option (GCE's regional Persistent Disk is the clearest example) rather than assuming ordinary zonal block storage survives a zone failure; ordinary zonal EBS/PD/Azure Disk does not.
Worked example
A team needs (a) a primary Postgres (Postgres) volume and (b) a shared directory of report templates read by 20 replica Pods across multiple nodes.
- (a) is single-writer and latency-sensitive: provision an EBS/GCE PD/Azure Disk volume through a StorageClass with
ReadWriteOnce, sized for the database's IOPS profile. - (b) needs concurrent multi-node reads: provision an EFS/Azure Files/NFS volume through a StorageClass supporting
ReadWriteMany, since a block-storage volume cannot satisfy that access pattern at all, regardless of performance tier.
Trade-offs and pitfalls
- Don't reach for a network filesystem by default "to be safe" for multi-node access; if the workload is genuinely single-writer, block storage's lower latency is the better fit and the shared-filesystem's variability is pure downside.
- Zonal block storage (the common case for EBS/GCE PD/Azure Disk) does not survive the loss of its Availability Zone; if that's a real requirement, either use the provider's cross-zone replicated variant where one exists, or handle replication at the application layer (e.g., a database's own streaming replication to a replica in another zone) rather than assuming the storage layer covers it.
- A "multi-writer" flag on a block-storage product is not the same guarantee as a real shared filesystem: it typically still requires a cluster-aware filesystem and application-level write coordination, so verify exactly what's supported for your disk type before relying on it, rather than assuming ReadWriteMany-equivalent behavior.
A pod remains in Pending and the scheduler does not bind it. Describe commands and checks to determine why the pod is unscheduled: inspect resource requests/limits, node allocatable capacity, taints/tolerations, affinity rules, and namespace quotas. Mention concrete kubectl commands you would run.
Sample Answer
Pending means the API server has accepted the pod object but the scheduler has not bound it to a node. Almost always the fastest path to the answer is reading the scheduler's own FailedScheduling event, which names the exact reason in plain text; everything below is really just "go read that message, then verify it against the right kubectl output."
Reading the signal
kubectl describe pod <pod> -n <ns>
kubectl get events -n <ns> --sort-by=.metadata.creationTimestamp
Typical event text and what it maps to:
| Event message pattern | What it means | Where to look next |
|---|---|---|
0/3 nodes are available: 3 Insufficient cpu | Every node lacks enough allocatable CPU for the pod's request | kubectl describe node, compare Allocated resources against the pod's resources.requests |
0/3 nodes are available: 3 node(s) had untolerated taint {key: value} | Nodes are tainted and the pod has no matching toleration | kubectl describe node (Taints section) vs pod.spec.tolerations |
0/3 nodes are available: 3 node(s) didn't match Pod's node affinity/selector | nodeSelector or nodeAffinity excludes every node | pod.spec.affinity / pod.spec.nodeSelector vs node labels |
pod has unbound immediate PersistentVolumeClaims | A PersistentVolumeClaim (PVC) the pod needs isn't bound to a PersistentVolume (PV) yet | kubectl get pvc -n <ns>; check the StorageClass and provisioner (a full storage debugging pass is its own topic, this is just the scheduling-side symptom) |
Checking each dimension directly
Resources.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.containers[*].resources}'
kubectl describe node <node>
kubectl top nodes # if metrics-server is installed
Compare the pod's requests (what the scheduler actually reasons about, not limits) against each node's allocatable capacity minus what's already committed.
Taints and tolerations.
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.taints}{"\n"}{end}'
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.tolerations}'
Affinity.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.affinity}'
A requiredDuringSchedulingIgnoredDuringExecution rule is a hard constraint; if no node satisfies it, the pod stays Pending indefinitely. A preferredDuringSchedulingIgnoredDuringExecution rule is a soft hint and will not by itself cause Pending.
Namespace-level limits: two different mechanisms, easy to conflate. A ResourceQuota is enforced at admission time, when the pod object is created, not at scheduling time. If a create request would push the namespace over its quota, the API server rejects it outright: kubectl apply returns an error immediately, and no Pending pod is ever created. So if you're looking at a Pending pod, ResourceQuota is not why. What can still bite you at scheduling time is a LimitRange: if the namespace has a default request injected by a LimitRange, a pod that specified no request at all can end up with a larger effective request than you intended, which then genuinely fails to fit any node and produces the Insufficient cpu event above.
kubectl get resourcequota -n <ns>
kubectl get limitrange -n <ns> -o yaml
Trade-offs and pitfalls
- The scheduler reasons about
requests, neverlimits; a pod with a tiny request and a huge limit schedules easily and can still get evicted or throttled later, which is a different problem from Pending. - A
nodeSelectortypo (a label key or value that doesn't exist on any node) produces the exact same symptom as a genuine capacity shortage; always check the label actually exists before assuming you need more nodes. - Don't chase ResourceQuota when you see a Pending pod; check LimitRange defaults and the scheduler's own event message instead, since quota violations block creation, not scheduling.
Unlock Full Question Bank
Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.