Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Describe the kubectl commands and rollout strategies you would use to perform a safe rolling restart of a Deployment, view rollout history, and rollback to a previous revision. Include examples using kubectl and explain how you would avoid causing cascading failures during a restart of a consumer‑facing service.
Sample Answer
A safe restart uses kubectl rollout restart, which recreates pods through the normal RollingUpdate strategy rather than deleting them directly, so the same availability guarantees that protect a routine deployment protect the restart too.
Commands
Trigger and watch a rolling restart:
kubectl rollout restart deployment my-app -n prod
kubectl rollout status deployment my-app -n prod --watch
View rollout history and inspect a specific revision:
kubectl rollout history deployment my-app -n prod
kubectl rollout history deployment my-app -n prod --revision=3
Roll back:
kubectl rollout undo deployment my-app -n prod --to-revision=3
kubectl rollout status deployment my-app -n prod
Ship an image change with a recorded reason (the --record flag some older references use for this is deprecated; annotate explicitly instead):
kubectl set image deployment/my-app my-app=registry/app:1.2.3 -n prod
kubectl annotate deployment my-app kubernetes.io/change-cause="bump to 1.2.3, ticket OPS-441" --overwrite -n prod
What keeps a restart from becoming a cascading failure
- RollingUpdate parameters:
maxUnavailableandmaxSurge(both default to 25% of desired replicas) bound how many old pods can be down and how many extra new pods can exist at once. For a consumer-facing service, a conservative setting (for examplemaxUnavailable: 0, maxSurge: 1) never drops capacity below the current replica count during the restart, at the cost of briefly running more pods than the steady-state count. - Readiness probes: a Service only sends traffic to pods that pass their readiness probe, so a newly restarted pod that's still initializing doesn't receive requests it can't yet handle. This is the single biggest lever against a restart-induced error spike; without a readiness probe, the rollout has no signal that a "new" pod is actually ready and can start routing traffic to it immediately.
- PodDisruptionBudget (PDB): guarantees a minimum number (or percentage) of replicas stay available throughout the restart, independent of the Deployment's own
maxUnavailablesetting, which matters when other voluntary disruptions (a node drain, a cluster upgrade) happen to overlap with the restart window. - Graceful shutdown: a
preStophook plus aterminationGracePeriodSecondslong enough for in-flight requests to finish, combined with the Service removing the pod's endpoint before the container actually stops, avoids dropping requests that were already in progress when the restart began. - Staged rollout for risk-sensitive services: restarting (or deploying) to a small subset first, watching error rate and latency, then proceeding, catches a bad new revision before it reaches full traffic; this is a general staged-rollout practice, not a specific traffic-splitting mechanism (traffic-splitting techniques like weighted canary routing are a load-balancing/ingress-layer concern, not something the Deployment object itself provides).
Trade-offs and pitfalls
kubectl rollout restartonly recreates pods; it does not change the Deployment's spec, sorollout historyrecords it as a new revision with the same template, which is easy to forget when later trying toundoyour way back past a restart that changed nothing.- Setting
maxUnavailable: 0guarantees no capacity loss but requires enough spare cluster capacity formaxSurgeextra pods to schedule; on a tightly packed cluster this can leave the rollout stuck Pending on the surge pods instead of proceeding. - A rollback only restores the pod template (image, env, resource requests, and so on). If the Deployment reads a ConfigMap or Secret by a fixed name and that ConfigMap was edited in place rather than replaced with a new name or hash-suffixed name, rolling the Deployment back does not restore the old configuration content, only the old pod template pointing at the same (already-mutated) ConfigMap. This is the most common way a rollback fails to actually roll back.
- The same gap applies to a PersistentVolumeClaim (PVC): a Deployment's rollback restores the pod template's volume mount references, not the data on the volume itself. If the new version wrote a schema migration or otherwise mutated data in place on that volume, rolling the Deployment back gives you the old code pointing at already-changed data, not the old data. Anything stateful needs its own restore path (a volume snapshot or application-level backup) alongside the Deployment rollback, not instead of thinking about it separately.
Design a secure multi-tenant Kubernetes platform. Discuss the pros and cons of cluster-per-tenant versus namespace-based multi-tenancy, and detail how you'd implement network isolation, RBAC boundaries, resource quotas, Pod Security (Seccomp/AppArmor), image scanning, runtime detection (e.g., Falco), and audit/logging to meet strong isolation and compliance requirements.
Sample Answer
For strong isolation and compliance requirements, choose cluster-per-tenant over namespace-based multi-tenancy whenever a tenant needs a genuinely separate blast radius (a compromised or noisy tenant must not be able to reach another tenant's control plane, kubelet, or node kernel) or needs to be independently certified against a specific compliance regime; use hardened namespace-based multi-tenancy for internally trusted tenants where utilization and onboarding speed matter more than a hard boundary. In practice, the strongest designs are hybrid: a small number of dedicated, hardened clusters for high-compliance tenants, plus a shared, heavily guarded cluster for everyone else.
Cluster-per-tenant vs. namespace-based
| Cluster-per-tenant | Namespace-based | |
|---|---|---|
| Isolation boundary | Separate control plane, etcd, kubelet, and usually node pool per tenant | Shared control plane and kubelet; boundary enforced entirely by RBAC (Role-Based Access Control), NetworkPolicy, and admission policy |
| Compliance story | Easier to certify: an auditor can point at one cluster and one tenant | Harder to certify: must demonstrate every layer of enforced isolation holds under adversarial conditions |
| Onboarding speed | Slower: a new cluster to provision, register, and integrate with platform tooling | Fast: a new namespace plus policy templates |
| Cost | Higher: per-cluster control-plane and headroom overhead multiplies with tenant count | Lower: tenants share unused capacity |
| Best for | Regulated tenants, adversarial or untrusted tenants, tenants needing a different Kubernetes version | Trusted internal tenants, cost-sensitive scale, fast-moving product teams |
Implementing each control
Network isolation. Use a CNI (Container Network Interface, the plugin layer responsible for pod networking) that enforces NetworkPolicy, such as Calico or Cilium. Default-deny ingress and egress per tenant namespace, then explicitly allow the flows a tenant actually needs. Cilium's eBPF (extended Berkeley Packet Filter, a Linux kernel technology for programmable packet processing) data path additionally supports layer-7-aware policy, useful for restricting a tenant to specific HTTP paths or gRPC methods on a shared internal API, not just IP-and-port pairs.
RBAC boundaries. RBAC (Role-Based Access Control) is Kubernetes' native authorization model. Grant only namespaced Role/RoleBinding pairs to tenant users; block ClusterRole and cluster-admin bindings for tenant service accounts entirely, enforced by an admission policy (OPA Gatekeeper or Kyverno) rather than by convention, since convention alone does not survive a misconfigured pipeline.
Resource quotas. ResourceQuota per tenant namespace plus a LimitRange for per-pod defaults and maximums, as in the general multi-tenancy case. In a shared cluster, also set a PriorityClass per tenant tier so that, under real node pressure, preemption and the scheduler favor higher-tier tenants' pods over lower-tier ones instead of resolving contention arbitrarily. This is the fairness mechanism that matters even when every tenant is well inside its own quota: quotas cap each tenant individually, but they do nothing to arbitrate contention for genuinely scarce cluster-wide capacity during a spike, which is what PriorityClass-driven preemption is for.
Pod security. Enforce the Pod Security Admission controller (the built-in mechanism that replaced PodSecurityPolicy, which was deprecated in Kubernetes 1.21 and removed in 1.25) at the restricted level for tenant namespaces: this denies privileged containers, host namespaces, and most capabilities by default. Layer a curated seccomp (secure computing mode, a Linux kernel feature that filters which system calls a process may make) profile and an AppArmor profile (a Linux kernel security module that restricts, per program, which files, network access, and capabilities it may use via a loaded policy) on top for defense in depth beyond what Pod Security Admission alone checks.
Image scanning and supply chain. Require every image to come from a private registry, scanned in CI before it can be deployed, and signed. cosign (part of the sigstore project) is a common tool for signing; Notation is the current CLI under the CNCF Notary Project for the same purpose (the older "notary" v1 client is the legacy predecessor). An admission-time check, the built-in ImagePolicyWebhook controller, or a Gatekeeper/Kyverno policy, verifies the signature and blocks unsigned or unscanned images from ever being scheduled.
Runtime detection. Falco watches kernel syscalls, via eBPF or a kernel module, for suspicious behavior, such as a shell spawned inside a container that never spawns shells, or a write to a path that should be read-only, and can alert or block. Cilium Hubble provides complementary network-flow visibility if Cilium is the CNI.
Audit and logging. Enable the Kubernetes API server's audit log at a verbosity that captures at least every write and every RBAC-relevant read, ship it to a tamper-evident store (object storage with retention locking, or a SIEM, a Security Information and Event Management system) separate from the cluster itself, and tag every entry with tenant identity so a compliance review can reconstruct one tenant's activity without touching another's data.
Worked example: an image-scan policy gate
Suppose the CI pipeline's scanner returns two findings for a candidate image before it is allowed into the tenant cluster:
| Finding | CVSS (Common Vulnerability Scoring System) score | Component |
|---|---|---|
| Critical: remote code execution in a base-image library | 9.8 | base OS package |
| Medium: outdated version of a dev-only dependency | 4.3 | build-time only, not shipped |
With a policy of "block on CVSS 7 or above," the first finding fails the build (9.8≥7) and the image is never pushed to the registry the admission controller trusts; the second finding (4.3<7) does not block. This is the mechanism, not the specific numbers: the policy threshold, and what counts as "shipped" (a dev-only dependency that never reaches the runtime image should not gate a build the same way a runtime dependency does), are the actual design decisions a platform team has to make, and they should be written down as policy-as-code so the same rule applies whether one engineer or a thousand submit an image.
Doing this at scale, roughly 1,000 tenants
The mechanisms above do not change in kind as tenant count grows into the hundreds or low thousands; what changes is that every manual step becomes untenable. Chargeback reporting has to run as an automated nightly job against labeled usage rather than a person building a thousand dashboards; onboarding has to be a GitOps-templated namespace-plus-policy bundle rather than a runbook a human executes by hand; and audit-log volume at that scale needs its own retention and cost budget, since a SIEM ingesting per-tenant audit trails for 1,000 tenants is a meaningfully different cost line than for ten.
Trade-offs and pitfalls
- Namespace-based isolation can be hardened close to cluster-per-tenant strength, but "close to" is doing real work in that sentence: a kernel-level container escape still reaches every tenant on that node, a risk that simply does not exist in cluster-per-tenant.
- Runtime detection (Falco) and admission-time policy (Gatekeeper/Kyverno) are complementary, not substitutes: admission policy stops known-bad configurations before they run; runtime detection catches behavior that only manifests once a workload is executing, such as a legitimate-looking image that turns malicious after a dependency is compromised post-deployment.
- Treating "namespace vs. cluster per tenant" as a single cluster-wide decision misses that different tenants can warrant different answers; the hybrid model, a few hardened dedicated clusters plus one well-governed shared cluster, is usually a better fit than picking one model for every tenant.
flowchart LR
Build[CI build] --> Scan[Image scan: CVSS gate]
Scan -->|pass| Sign[Sign image: cosign / Notation]
Scan -->|fail| Block[Build blocked]
Sign --> Registry[Private registry]
Registry --> Admission[Admission check: signature + policy]
Admission -->|pass| Run[Pod scheduled to tenant namespace]
Admission -->|fail| Reject[Pod creation rejected]
Run --> Falco[Falco: runtime syscall monitoring]
Falco --> Audit[Audit log + SIEM]
Describe what a Pod is in Kubernetes and why it is considered the smallest deployable unit. Explain when you would run multiple containers in a single pod, how containers inside a pod share network and volumes, and trade-offs of co-locating containers such as sidecar patterns versus separate pods.
Sample Answer
A pod is the smallest unit Kubernetes schedules and manages: one or more containers that always run together on the same node, sharing a network namespace and, optionally, storage volumes. It is the unit, rather than the individual container, because the kubelet, the scheduler, and every controller reason about placement, restarts, and networking at that granularity, not below it.
Why the pod, not the container, is the atomic unit
- Co-scheduling guarantee: every container in a pod is placed on the same node and started, stopped, and restarted as a unit by the kubelet.
- One IP per pod: regardless of how many containers it holds, a pod gets exactly one IP address; Services and DNS resolve to that pod IP, not to an individual container.
- Shared network namespace: containers in a pod talk to each other over
localhost, and must not claim conflicting ports, because they share the same network namespace. - Shared volumes: volumes are declared once at the pod spec level and can be mounted into more than one container, which is how a sidecar can read or write files the main container produces without a network hop.
When to run more than one container in a pod
- Sidecar pattern: a helper process that shares the fate and resources of the main container, most commonly a log shipper, a proxy, or a metrics exporter.
- Init containers: containers that run to completion, in order, before the main containers start, typically for one-time setup like a schema migration or config templating.
- Native sidecar containers: since the
SidecarContainersfeature became enabled by default in Kubernetes 1.29 (stable as of 1.33), you can declare a sidecar underinitContainerswithrestartPolicy: Always. Kubernetes then starts it before the main container, keeps it running for the pod's whole life, and stops it after the main container on shutdown, which fixes the older ordering problem where a plain extra container might not be ready before the app started, or might outlive a completed Job's main container instead of shutting down with it.
A small example, an app container and a log-forwarding sidecar sharing a volume instead of a network call:
containers:
- name: app
image: myapp:1.4
volumeMounts:
- name: logs
mountPath: /var/log/app
- name: log-shipper
image: fluent-bit:latest
volumeMounts:
- name: logs
mountPath: /var/log/app
readOnly: true
volumes:
- name: logs
emptyDir: {}
The app writes to /var/log/app; the sidecar tails the same path through the shared emptyDir volume, no network hop involved.
Worked example: what a failing sidecar looks like
If the log-shipper container above enters a crash loop while app keeps running fine, the two containers' fates are independent even though they share a pod: app keeps serving traffic, but the log pipeline drops. This is visible directly in the pod list, where the READY column reports containers-ready over containers-total, not a single number:
NAME READY STATUS RESTARTS AGE
myapp-6d947f8db8-x2z1p 1/2 Running 4 (30s ago) 6m
Here 1/2 means only one of the pod's two containers is passing its readiness state, and the restart count of 4 belongs to the sidecar, not to app.
Trade-offs and pitfalls
- Co-location couples lifecycle: a sidecar cannot be scaled independently of the app the way a separate Deployment could be. Native sidecar containers loosen this slightly by giving the sidecar its own restart behavior, but it is still tied to the pod's schedule and node.
- Every container in the pod counts toward the pod's total resource footprint; forgetting the sidecar's own requests and limits when sizing the node is a common under-provisioning mistake.
- Splitting a component into a separate pod (plus a Service) regains independent scaling and blast-radius isolation, at the cost of a network hop and losing the localhost and shared-volume convenience. Reach for separate pods when the components genuinely differ in scaling or failure characteristics, for example a stateless API versus a shared cache, rather than defaulting to a sidecar for anything that happens to run alongside the main app.
HorizontalPodAutoscaler (HPA) is not scaling a deployment even though CPU usage is above the configured target. Provide a troubleshooting plan including which metrics endpoints and API resources to check, how to validate metrics-server or custom metrics adapters, and common misconfigurations that prevent scaling.
Sample Answer
When CPU usage is genuinely above target but the Horizontal Pod Autoscaler (HPA) will not scale, the fault is almost always in the metrics plumbing between the container and the HPA controller (missing resource requests, a metrics-server or custom-metrics adapter that is not actually returning data, or a scaleTargetRef/selector mismatch) rather than in the scaling decision logic itself, which is a small, deterministic formula once it has correct inputs.
How the HPA actually decides (not scaling strategy, the mechanics)
The controller (autoscaling/v2 API, the current stable version; the earlier autoscaling/v2beta2 was removed in Kubernetes 1.26, so a cluster still authoring v2beta2 manifests is itself a currency problem worth checking for) recomputes replica count on a fixed sync period (default 15s) using:
with a tolerance band (default 10%, --horizontal-pod-autoscaler-tolerance) below which it intentionally does nothing, to avoid thrashing on noise. This tolerance and sync period explain most "it's above target but nothing happens" reports that turn out to be working as designed: a ratio of 1.05 (5% over target) is inside the default tolerance band and will not trigger a scale event.
Checklist, in order of how often each one is the actual cause
- HPA state and events, the fastest signal:
kubectl describe hpa myapp-hpa -n myapp
Look for the Conditions block and events such as:
Warning FailedGetResourceMetric hpa-controller
failed to get cpu utilization: unable to get metrics for resource cpu:
missing request for cpu
That specific event means the fix is a resource request, not anything about metrics-server health.
2. CPU utilization requires requests, always. HPA's Utilization metric type is computed as usage divided by the container's CPU request, not its limit. If a container has no resources.requests.cpu, the HPA controller cannot compute a percentage at all and reports the event above, regardless of how busy the pod actually is.
3. Metrics API is actually serving data:
kubectl top pods -n myapp
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/namespaces/myapp/pods" | jq
If kubectl top returns nothing, metrics-server itself is the suspect, not the HPA:
kubectl get pods -n kube-system -l k8s-app=metrics-server
kubectl logs -n kube-system <metrics-server-pod>
- Custom or external metrics, if the HPA targets something other than CPU/memory:
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | jq
An empty or erroring response here means the metrics adapter (commonly the Prometheus Adapter) is not registered correctly as an APIService, or the metric name in the HPA spec does not match the name the adapter actually exposes (a very common typo-class bug: the HPA references http_requests_per_second while the adapter exposes requests_per_second).
5. scaleTargetRef correctness: confirm .spec.scaleTargetRef.kind/.name in the HPA actually matches the Deployment being scaled; a stale reference (left over from a renamed Deployment) means the HPA is quietly watching nothing.
6. Pod readiness: the HPA only counts Ready pods toward currentReplicas and current metric averaging; a pod stuck failing its readiness probe is invisible to the calculation even though it is consuming CPU.
7. minReplicas/maxReplicas bounds: confirm maxReplicas is actually above the current replica count; a forgotten low ceiling silently caps scaling with no error at all.
Worked example
Suppose myapp-hpa targets 50% average CPU utilization, is currently running 4 replicas, and kubectl top pods shows the average pod is using CPU equal to 82% of its request. Plugging into the formula:
If the HPA is healthy, kubectl describe hpa should show current / target: 82% / 50% and a scaling event moving the Deployment toward 7 replicas. If instead the describe output shows no current value at all (a blank or <unknown>), the problem is upstream of this formula entirely: metrics are not reaching the controller, so there is nothing for the formula to compute against.
Trade-offs and pitfalls
- This checklist deliberately does not touch which metric to scale on or whether reactive CPU-based scaling is the right strategy for the workload; that is a scaling-strategy design decision, separate from why a correctly-configured HPA fails to act mechanically.
- A common wrong turn is restarting metrics-server or the HPA controller as a first move; that only helps if the actual event says the adapter is unresponsive, and it resets nothing if the real cause is a missing CPU request or a
scaleTargetReftypo. - Combining HPA with a Vertical Pod Autoscaler (VPA) on the same CPU metric can fight itself: VPA changing requests up or down changes the denominator HPA's utilization percentage is computed against, producing scaling behavior that looks erratic but is actually two controllers correctly reacting to each other.
- The tolerance band means "barely above target" is not a bug; check the actual ratio before assuming the controller is broken.
Explain Pod Disruption Budgets (PDBs). How do PDBs interact with rolling updates, cluster autoscaler evictions, and maintenance operations? Provide scenarios where an incorrect PDB could block upgrades or autoscaling and how to fix those issues.
Sample Answer
A PodDisruptionBudget (PDB) tells the cluster how much of a replicated workload is allowed to be taken down at once by a voluntary disruption, things a human or controller chooses to do, such as a node drain, a cluster upgrade, or the cluster autoscaler removing a node. It has no effect on involuntary disruptions: a node crashing, a container getting OOMKilled (killed by the kernel's out-of-memory killer for exceeding its memory limit), or hardware failure all bypass it entirely, because there's no eviction request for the PDB to block in those cases.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web-pdb
spec:
selector:
matchLabels: { app: web }
minAvailable: 2
(policy/v1 is the stable, current API group and version for this object; older manifests using policy/v1beta1 are targeting a removed API and will fail to apply on a current cluster.)
How PDBs interact with each operation
- Rolling updates: the Deployment controller's own
maxUnavailable/maxSurgesettings already bound disruption during a rollout, and a PDB adds an independent floor on top that applies across all voluntary disruptions hitting that workload at once, not just the rollout in isolation. If a rollout and a node drain happen to overlap, the PDB is what prevents their combined effect from dropping available replicas too far. - Cluster Autoscaler (CA) scale-down: before removing a node, CA evicts the pods on it through the eviction API, and that eviction is refused if it would violate a pod's PDB. CA then either finds another node to remove or skips that node until eviction becomes possible; it does not force through a PDB violation.
- Manual maintenance (
kubectl drain, akubeadmupgrade, a node reboot): these also go through the eviction API and are blocked the same way; an operator hitting a blocked drain either waits, adjusts the PDB, or in genuine emergencies deletes pods directly (bypassing the eviction API, which also bypasses the PDB, and should be a deliberate, audited exception rather than a routine workaround).
Problem scenarios and fixes
| Scenario | Why it blocks | Fix |
|---|---|---|
minAvailable: 2 on a 2-replica Deployment | Zero disruption tolerance; nothing can ever be evicted | Add a third replica, or switch to maxUnavailable: 1 if a 2-replica service can genuinely tolerate a brief single-replica window |
| A StatefulSet's PDB blocks Cluster Autoscaler from ever draining its last node | CA can't evict any of the StatefulSet's pods without breaching the budget, so it leaves that node running indefinitely | Loosen the PDB to allow at least one voluntary disruption, or, if the workload genuinely cannot tolerate any drop, keep it on dedicated (non-scaled-down) nodes instead of fighting the autoscaler |
| Many services' PDBs collectively stall a cluster-wide upgrade | Each PDB is individually reasonable, but the combined effect blocks progress node by node | Stagger the upgrade, use percentage-based minAvailable/maxUnavailable so budgets scale with replica count, and give the upgrade a maintenance window with clear escalation if it stalls past expected duration |
Trade-offs and pitfalls
- Prefer percentage-based values (
maxUnavailable: 25%) over absolute counts for workloads whose replica count changes with load; an absoluteminAvailable: 3on a Deployment that autoscales down to 2 replicas becomes an unsatisfiable budget that blocks every voluntary disruption until it scales back up. - A PDB with
minAvailable: 100%(or an absolute count equal to current replicas) reads as "maximally safe" but actually means zero voluntary disruptions are ever permitted, which silently blocks every future drain and upgrade until someone notices and loosens it, often under time pressure during an incident. - PDBs protect against voluntary disruption stacking up in ways a single controller's own settings can't see; they are not a substitute for redundancy. A workload with only one replica and a PDB requiring
minAvailable: 1gains nothing from the PDB, since that one pod being unavailable for any reason, voluntary or not, is already an outage.
Unlock Full Question Bank
Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.