Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Design a multi-tenant strategy for a Kubernetes cluster that will host several internal teams. Discuss the use of namespaces, RBAC roles, resource quotas, network policies, and cost allocation. Provide pros and cons of single-cluster multi-tenant vs multiple clusters per team, and when you'd recommend each approach.
Sample Answer
A single well-governed multi-tenant cluster is usually the right starting point for internal teams: namespaces provide the boundary, RBAC (Role-Based Access Control) scopes who can act inside it, ResourceQuota and LimitRange keep one team's usage from starving another, and NetworkPolicy restricts which pods can talk to which. Move a team to its own dedicated cluster only when a specific pressure, compliance, blast radius, or a competing upgrade cadence, makes shared governance the wrong trade, not as a default posture.
What a namespace actually scopes
This decides where every other control applies. A namespace is a logical partition inside one cluster used to scope names and access, not a security boundary by itself: pods in different namespaces still share the same kernel, kubelet, and node pool unless further isolation is added.
| Scoped to a namespace | Cluster-scoped (shared across every namespace) |
|---|---|
| Pod, Deployment, Service, ConfigMap, Secret | Node |
| ResourceQuota, LimitRange | PersistentVolume (the claim, PersistentVolumeClaim, is namespaced; the underlying volume is not) |
| Role, RoleBinding | ClusterRole, ClusterRoleBinding |
| NetworkPolicy | StorageClass |
| CustomResourceDefinition (CRD, the schema that defines a custom object type; the custom objects it defines can themselves be namespaced or cluster-scoped) | |
| Namespace itself |
Getting this table wrong is a common onboarding mistake: teams write a NetworkPolicy assuming it also restricts node-level traffic, or expect a namespaced RBAC Role to grant node access, when Node is cluster-scoped and untouched by namespace-level RBAC.
Building the tenant boundary
- Namespaces: one namespace per team plus separate namespaces for shared platform services (ingress controller, logging agents, CI/CD runners). Avoid namespace-per-project-per-team sprawl unless its lifecycle is also automated; an unmanaged namespace is an onboarding shortcut that becomes an offboarding liability.
- RBAC: define a small set of reusable Role templates (namespace-admin, developer, read-only) and bind them via RoleBinding to groups from an identity provider (an external system, typically integrated via OpenID Connect, OIDC, that authenticates users and asserts group membership) rather than to individual users, so team-membership changes never require touching Kubernetes RBAC directly. Reserve ClusterRole grants for the platform team.
- ResourceQuota and LimitRange: cap each namespace's aggregate CPU/memory (ResourceQuota) and set sane per-pod defaults and maximums (LimitRange) so one team cannot silently consume the whole node pool. The exact admission-time interaction between the two is its own deep mechanism; the summary a platform designer needs here is that LimitRange fills in and bounds individual pods, while ResourceQuota bounds the namespace's total.
- NetworkPolicy: default-deny at the namespace boundary, then explicitly allow the flows a team actually needs (its own pods, plus named shared services like an internal registry or logging endpoint). This requires a CNI (Container Network Interface, the plugin layer responsible for pod networking) that enforces NetworkPolicy; Calico and Cilium both do, but not every CNI plugin does.
- Cost allocation: label every workload with its owning team, scrape actual usage with kube-state-metrics into Prometheus, and attribute cost with a tool built for it (Kubecost or OpenCost are common choices) rather than hand-rolling chargeback from raw node-hour billing.
Worked example: proportional chargeback
Suppose a cluster's monthly infrastructure bill is $12,000, and three teams' namespaces show the following measured CPU-hour consumption over the month (measured usage, not their static quota, since actual usage is what should drive cost):
| Team | Measured core-hours | Share of total |
|---|---|---|
| Payments | 6,000 | 50% |
| Search | 3,000 | 25% |
| Internal tools | 3,000 | 25% |
| Total | 12,000 | 100% |
Payments’ bill=$12,000×0.50=$6,000
Search’s bill=$12,000×0.25=$3,000
Internal tools’ bill=$12,000×0.25=$3,000
This proportional-usage model scales the same way whether the cluster has 3 tenants or 1,000: the mechanism, label-based usage scraped continuously and aggregated, does not change; only the reporting layer has to move from a handful of dashboards a human reads to an automated nightly rollup a finance system consumes, since nobody reviews a thousand individual namespace dashboards by hand.
Single-cluster vs. per-team clusters
| Single cluster, multi-tenant | Cluster per team | |
|---|---|---|
| Utilization | High: teams share unused capacity | Lower: each cluster needs its own headroom |
| Operational overhead | One control plane, one upgrade path, one set of platform tooling | N control planes, N upgrade schedules, duplicated platform tooling |
| Blast radius | A bad cluster-wide change (CNI upgrade, admission webhook bug) affects every team at once | Contained to one team |
| Isolation strength | Depends entirely on RBAC/NetworkPolicy/quota discipline; the shared kubelet and node kernel remain a real, if narrow, attack surface | Strongest: separate control plane and, if desired, separate node pools |
| Fits best when | Teams are internally trusted, cost efficiency matters, the platform team can enforce policy centrally | A team needs a different Kubernetes version, has a hard compliance boundary, or its noisy-neighbor risk is unacceptable to others |
The same axis reappears one level down inside a single cluster: many teams that keep dev, staging, and production as one shared multi-tenant cluster for environments still break production out into its own cluster regardless of how they handle team tenancy, because a shared production control-plane incident is categorically worse than a shared-namespace tenant issue in a lower environment.
Trade-offs and pitfalls
- Quotas without LimitRange defaults mean a team that forgets to set requests/limits on a Deployment can be rejected outright at pod creation rather than merely scheduled sub-optimally; this surprises teams the first time it happens.
- NetworkPolicy default-deny with no shared-services allowance breaks logging and monitoring silently: pods that used to reach a cluster-wide logging endpoint stop being able to, and the failure shows up as missing logs rather than an error in the app itself.
- "Start single-cluster, split later" only works if the platform team retains an actual migration path (namespace export, workload relabeling); treat that migration tooling as part of the initial design, not a someday problem.
HorizontalPodAutoscaler (HPA) is not scaling a deployment even though CPU usage is above the configured target. Provide a troubleshooting plan including which metrics endpoints and API resources to check, how to validate metrics-server or custom metrics adapters, and common misconfigurations that prevent scaling.
Sample Answer
When CPU usage is genuinely above target but the Horizontal Pod Autoscaler (HPA) will not scale, the fault is almost always in the metrics plumbing between the container and the HPA controller (missing resource requests, a metrics-server or custom-metrics adapter that is not actually returning data, or a scaleTargetRef/selector mismatch) rather than in the scaling decision logic itself, which is a small, deterministic formula once it has correct inputs.
How the HPA actually decides (not scaling strategy, the mechanics)
The controller (autoscaling/v2 API, the current stable version; the earlier autoscaling/v2beta2 was removed in Kubernetes 1.26, so a cluster still authoring v2beta2 manifests is itself a currency problem worth checking for) recomputes replica count on a fixed sync period (default 15s) using:
with a tolerance band (default 10%, --horizontal-pod-autoscaler-tolerance) below which it intentionally does nothing, to avoid thrashing on noise. This tolerance and sync period explain most "it's above target but nothing happens" reports that turn out to be working as designed: a ratio of 1.05 (5% over target) is inside the default tolerance band and will not trigger a scale event.
Checklist, in order of how often each one is the actual cause
- HPA state and events, the fastest signal:
kubectl describe hpa myapp-hpa -n myapp
Look for the Conditions block and events such as:
Warning FailedGetResourceMetric hpa-controller
failed to get cpu utilization: unable to get metrics for resource cpu:
missing request for cpu
That specific event means the fix is a resource request, not anything about metrics-server health.
2. CPU utilization requires requests, always. HPA's Utilization metric type is computed as usage divided by the container's CPU request, not its limit. If a container has no resources.requests.cpu, the HPA controller cannot compute a percentage at all and reports the event above, regardless of how busy the pod actually is.
3. Metrics API is actually serving data:
kubectl top pods -n myapp
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/namespaces/myapp/pods" | jq
If kubectl top returns nothing, metrics-server itself is the suspect, not the HPA:
kubectl get pods -n kube-system -l k8s-app=metrics-server
kubectl logs -n kube-system <metrics-server-pod>
- Custom or external metrics, if the HPA targets something other than CPU/memory:
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | jq
An empty or erroring response here means the metrics adapter (commonly the Prometheus Adapter) is not registered correctly as an APIService, or the metric name in the HPA spec does not match the name the adapter actually exposes (a very common typo-class bug: the HPA references http_requests_per_second while the adapter exposes requests_per_second).
5. scaleTargetRef correctness: confirm .spec.scaleTargetRef.kind/.name in the HPA actually matches the Deployment being scaled; a stale reference (left over from a renamed Deployment) means the HPA is quietly watching nothing.
6. Pod readiness: the HPA only counts Ready pods toward currentReplicas and current metric averaging; a pod stuck failing its readiness probe is invisible to the calculation even though it is consuming CPU.
7. minReplicas/maxReplicas bounds: confirm maxReplicas is actually above the current replica count; a forgotten low ceiling silently caps scaling with no error at all.
Worked example
Suppose myapp-hpa targets 50% average CPU utilization, is currently running 4 replicas, and kubectl top pods shows the average pod is using CPU equal to 82% of its request. Plugging into the formula:
If the HPA is healthy, kubectl describe hpa should show current / target: 82% / 50% and a scaling event moving the Deployment toward 7 replicas. If instead the describe output shows no current value at all (a blank or <unknown>), the problem is upstream of this formula entirely: metrics are not reaching the controller, so there is nothing for the formula to compute against.
Trade-offs and pitfalls
- This checklist deliberately does not touch which metric to scale on or whether reactive CPU-based scaling is the right strategy for the workload; that is a scaling-strategy design decision, separate from why a correctly-configured HPA fails to act mechanically.
- A common wrong turn is restarting metrics-server or the HPA controller as a first move; that only helps if the actual event says the adapter is unresponsive, and it resets nothing if the real cause is a missing CPU request or a
scaleTargetReftypo. - Combining HPA with a Vertical Pod Autoscaler (VPA) on the same CPU metric can fight itself: VPA changing requests up or down changes the denominator HPA's utilization percentage is computed against, producing scaling behavior that looks erratic but is actually two controllers correctly reacting to each other.
- The tolerance band means "barely above target" is not a bug; check the actual ratio before assuming the controller is broken.
The Kubernetes API server is experiencing increased request latency. What metrics, logs, and traces would you collect to diagnose whether the bottleneck is etcd, admission controllers, or API server CPU/memory? Provide a prioritized triage checklist and remedial actions for each root cause.
Sample Answer
Direct answer
Increased kube-apiserver latency almost always traces to one of three places: etcd itself (disk or network bound), a slow admission webhook sitting in the request path, or the apiserver process running short of CPU or memory (including its own request-concurrency limiting kicking in). The fastest way to tell them apart is to look at where time is spent inside a single slow request, etcd round trip versus admission call versus everything else, rather than guessing from symptoms alone.
Structured elaboration
Metrics to pull first
| Signal | What it tells you |
|---|---|
apiserver_request_duration_seconds (histogram, by verb/resource) | Overall request latency, and whether it is one resource type or global |
apiserver_current_inflight_requests | How close the server is to its concurrency ceiling |
apiserver_flowcontrol_rejected_requests_total, apiserver_flowcontrol_current_inqueue_requests, apiserver_flowcontrol_current_executing_requests | Whether API Priority and Fairness (APF, stable since Kubernetes 1.29, the mechanism that classifies and queues requests by priority) is queuing, executing, or rejecting requests for a given priority level |
apiserver_admission_webhook_admission_duration_seconds | Per-webhook admission latency, split by mutating/validating |
etcd_disk_wal_fsync_duration_seconds, etcd_disk_backend_commit_duration_seconds | Disk-bound etcd latency; sustained spikes point at disk contention |
etcd_server_has_leader, etcd_server_leader_changes_seen_total | Whether etcd has a stable leader or is re-electing |
Triage order
- Scope it: slice
apiserver_request_duration_secondsby resource and verb. If only one resource type or client is slow, an etcd-wide or apiserver-wide problem is unlikely; look at that resource's admission webhooks first. - Check APF rejection reasons:
apiserver_flowcontrol_rejected_requests_totallabeledreason="queue-full"orreason="concurrency-limit"means the server is intentionally shedding load under its configured priority levels. The 429 responses clients see in that case are the mechanism doing its job, not a hidden bug; the real bottleneck is upstream (etcd or CPU), not APF itself. - If the slowdown is global, every resource, every client, check etcd's disk metrics and leader stability before touching the apiserver. A slow etcd backend shows up as apiserver latency because every write and every quorum read waits on it.
- If etcd looks healthy but apiserver CPU or memory is pegged, it is apiserver resource pressure.
Remedial actions per root cause
- etcd bound: faster disks (etcd is fsync-latency sensitive, not throughput sensitive), a disk dedicated to etcd separate from other I/O, defragmentation, and checking
etcd_mvcc_db_total_size_in_bytes(a bloated database slows every commit). Reduce write volume from noisy controllers before adding etcd members: more members raise the replication cost of every write, they do not spread load the way a read replica would. - Admission webhook bound: tune webhook
timeoutSecondsandfailurePolicycarefully, scale the webhook backend, and reconsider whether the check belongs in a webhook at all versus static OpenAPI schema validation. A webhook earns its cost when the rule needs data outside the request object, cross-field logic, or an external lookup; anything expressible as a plain schema constraint should live in the CRD's (Custom Resource Definition's) OpenAPI validation instead, since that costs nothing at admission time. - apiserver resource bound: unlike etcd, kube-apiserver is stateless, so horizontally scaling it (more apiserver replicas behind the control-plane load balancer) is a legitimate, common fix, not a workaround. Also check audit log verbosity and watch cardinality; a controller opening many broad watches is a frequent, overlooked CPU driver.
Worked example (concept, not a fabricated benchmark)
Suppose apiserver_request_duration_seconds p99 for PATCH pods is elevated but p99 for every other verb and resource is flat. That shape alone rules out an etcd-wide or apiserver-wide bottleneck, because both would show up across every resource type, and points at something specific to pod patches, almost always a mutating webhook registered on Pods (for example a sidecar injector). Confirming it takes one more step: check apiserver_admission_webhook_admission_duration_seconds filtered to that webhook's name. A rising p99 there, correlated with the PATCH pods latency, closes the loop without needing to touch etcd or CPU metrics at all.
Trade-offs and pitfalls
- Do not disable webhooks blind as a first move.
failurePolicy: Ignoreon a security-relevant mutating webhook (one that injects a sidecar or a required label, for example) can silently change what gets admitted, not just how fast. - Adding etcd members to "spread the load" is a common but wrong instinct: every additional voting member adds replication overhead to every write. It improves fault tolerance, not throughput.
- 429 responses from APF are a symptom of an upstream bottleneck, not a target to eliminate by raising limits. Raising a priority level's concurrency share without fixing the underlying etcd or CPU constraint just moves where the queue backs up.
A pod remains in Pending and the scheduler does not bind it. Describe commands and checks to determine why the pod is unscheduled: inspect resource requests/limits, node allocatable capacity, taints/tolerations, affinity rules, and namespace quotas. Mention concrete kubectl commands you would run.
Sample Answer
Pending means the API server has accepted the pod object but the scheduler has not bound it to a node. Almost always the fastest path to the answer is reading the scheduler's own FailedScheduling event, which names the exact reason in plain text; everything below is really just "go read that message, then verify it against the right kubectl output."
Reading the signal
kubectl describe pod <pod> -n <ns>
kubectl get events -n <ns> --sort-by=.metadata.creationTimestamp
Typical event text and what it maps to:
| Event message pattern | What it means | Where to look next |
|---|---|---|
0/3 nodes are available: 3 Insufficient cpu | Every node lacks enough allocatable CPU for the pod's request | kubectl describe node, compare Allocated resources against the pod's resources.requests |
0/3 nodes are available: 3 node(s) had untolerated taint {key: value} | Nodes are tainted and the pod has no matching toleration | kubectl describe node (Taints section) vs pod.spec.tolerations |
0/3 nodes are available: 3 node(s) didn't match Pod's node affinity/selector | nodeSelector or nodeAffinity excludes every node | pod.spec.affinity / pod.spec.nodeSelector vs node labels |
pod has unbound immediate PersistentVolumeClaims | A PersistentVolumeClaim (PVC) the pod needs isn't bound to a PersistentVolume (PV) yet | kubectl get pvc -n <ns>; check the StorageClass and provisioner (a full storage debugging pass is its own topic, this is just the scheduling-side symptom) |
Checking each dimension directly
Resources.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.containers[*].resources}'
kubectl describe node <node>
kubectl top nodes # if metrics-server is installed
Compare the pod's requests (what the scheduler actually reasons about, not limits) against each node's allocatable capacity minus what's already committed.
Taints and tolerations.
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.taints}{"\n"}{end}'
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.tolerations}'
Affinity.
kubectl get pod <pod> -n <ns> -o jsonpath='{.spec.affinity}'
A requiredDuringSchedulingIgnoredDuringExecution rule is a hard constraint; if no node satisfies it, the pod stays Pending indefinitely. A preferredDuringSchedulingIgnoredDuringExecution rule is a soft hint and will not by itself cause Pending.
Namespace-level limits: two different mechanisms, easy to conflate. A ResourceQuota is enforced at admission time, when the pod object is created, not at scheduling time. If a create request would push the namespace over its quota, the API server rejects it outright: kubectl apply returns an error immediately, and no Pending pod is ever created. So if you're looking at a Pending pod, ResourceQuota is not why. What can still bite you at scheduling time is a LimitRange: if the namespace has a default request injected by a LimitRange, a pod that specified no request at all can end up with a larger effective request than you intended, which then genuinely fails to fit any node and produces the Insufficient cpu event above.
kubectl get resourcequota -n <ns>
kubectl get limitrange -n <ns> -o yaml
Trade-offs and pitfalls
- The scheduler reasons about
requests, neverlimits; a pod with a tiny request and a huge limit schedules easily and can still get evicted or throttled later, which is a different problem from Pending. - A
nodeSelectortypo (a label key or value that doesn't exist on any node) produces the exact same symptom as a genuine capacity shortage; always check the label actually exists before assuming you need more nodes. - Don't chase ResourceQuota when you see a Pending pod; check LimitRange defaults and the scheduler's own event message instead, since quota violations block creation, not scheduling.
You have mixed hardware in the cluster: GPU nodes for machine learning, on-demand nodes for critical services, and spot instances for low-priority batch jobs. Explain how you'd use node labels, taints, tolerations, node selectors or affinity, and PodTopologySpread to ensure correct scheduling and protect critical workloads from being placed on spot instances.
Sample Answer
Direct answer
Node labels, taints/tolerations, affinity, and topology spread each answer a different scheduling question, and this scenario needs all of them combined rather than one chosen over the others. Taints on the volatile pool (spot) provide the exclusion guarantee that keeps critical workloads off it by default. Affinity provides the positive pull that sends the right workload to the right pool, and for GPU (graphics processing unit) nodes that pull has to be paired with the device-plugin's extended resource accounting, not just a label. Topology spread (or pod anti-affinity) spreads critical replicas so a single node or zone loss never removes more than one replica.
Structured elaboration
| Mechanism | Question it answers | Example in this scenario |
|---|---|---|
| Node labels | What kind of node is this | hardware=gpu / ondemand / spot |
| Taints + tolerations | Which pods are allowed here at all (an exclusion gate) | spot=true:NoSchedule, only batch pods tolerate it |
| Node affinity / nodeSelector | Where does this specific pod want to go (a positive pull) | requiredDuringSchedulingIgnoredDuringExecution on hardware=gpu |
| Extended resources (device plugin) | How many units of a scarce, non-CPU resource this pod needs | resources.limits: {nvidia.com/gpu: 1} |
| topologySpreadConstraints | How replicas of one workload are spread | maxSkew: 1 across topology.kubernetes.io/zone |
| PriorityClass + preemption | Who survives when capacity is scarce | Critical pods get a higher priorityClassName than batch |
GPU nodes need more than a label
Labeling a node hardware=gpu is necessary but not sufficient. The scheduler only knows a node has usable GPU capacity once the NVIDIA device plugin (or a vendor equivalent) DaemonSet advertises it as an extended resource in the node's allocatable, for example nvidia.com/gpu: 4. A pod requests it like any other resource:
resources:
limits:
nvidia.com/gpu: 1
Without the resource request, a pod with only nodeSelector: {hardware: gpu} can land on a GPU node without ever actually reserving a device, or worse, several pods could expect the same already-claimed device. Pair the label or affinity rule (which steers the pod to the right node family) with the extended-resource request (which makes the scheduler actually reserve a device).
Excluding spot from critical workloads
kubectl taint nodes -l node-lifecycle=spot spot=true:NoSchedule
Only pods that explicitly tolerate it can land there:
tolerations:
- key: "spot"
operator: "Equal"
value: "true"
effect: "NoSchedule"
Critical services carry no such toleration, so the scheduler treats spot nodes as invisible to them by default, with no risk of a nodeSelector typo accidentally placing one there.
Spreading critical replicas
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: critical-service
whenUnsatisfiable: DoNotSchedule makes this a hard constraint; ScheduleAnyway makes it best-effort. For a service that must survive a zone loss, DoNotSchedule is the right choice even though it can leave a pod Pending if the under-represented zone runs out of capacity, which is a signal to add capacity, not to loosen the constraint.
Priority for the pool that can vanish without warning
Spot capacity can disappear with only seconds of notice. Give critical workloads a higher priorityClassName so, if critical and batch pods ever land in the same resource pool during a capacity burst, the scheduler preempts lower-priority batch pods rather than the reverse.
Worked example
A cluster has 2 GPU nodes (4 GPUs each, 8 total), 3 on-demand nodes for critical services, and 4 spot nodes for batch. A training job requests 2 GPUs; with the device plugin installed, allocatable nvidia.com/gpu across the 2 GPU nodes totals 8, and the scheduler places the pod on whichever GPU node currently has 2 or more free, refusing to schedule (Pending, correctly) if all 8 are already claimed elsewhere. A critical service asks for 3 replicas with topologySpreadConstraints maxSkew: 1 across the 3 on-demand zones: with exactly 3 zones and 3 replicas the constraint is satisfiable exactly, one per zone, so losing any single zone loses exactly 1 of 3 replicas, never more.
Trade-offs and pitfalls
- A taint on the spot pool only stops pods without the toleration; it does not, by itself, stop a batch pod from landing on an on-demand node. If batch must also be kept off on-demand capacity, that pool needs its own taint plus a matching toleration on batch pods, not just a one-directional rule.
- Required affinity and required topology spread can leave pods unschedulable when capacity is tight. Monitor scheduling events (
PendingplusFailedScheduling) rather than reflexively loosening a hard rule to preferred, which quietly reintroduces the correlated-failure risk the rule existed to prevent. - GPU device counts are not overcommittable the way CPU or memory can be with limits set above requests. Plan capacity on whole-device granularity.
Unlock Full Question Bank
Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.