Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Describe the role of an Ingress resource versus an Ingress Controller in Kubernetes. What does the Ingress object itself declare, what does the controller actually do with that declaration, and why does Kubernetes split the responsibility this way instead of having one object do both?
Sample Answer
An Ingress object declares what HTTP(S) routing should happen (hostnames, path rules, which Service backs each path, which TLS secret to use); an Ingress Controller is the running component that watches Ingress objects and actually implements that routing on real infrastructure. Kubernetes splits these into two objects for the same reason it splits a Deployment's spec from the controller that reconciles it: the API object stays a stable, portable declaration, while the implementation (which varies enormously between environments) is free to be swapped out without changing what you wrote.
What the Ingress object itself declares
An Ingress is pure declaration, no execution: a list of host/path rules, each pointing at a backend Service and port, plus optional TLS configuration referencing a Secret containing a certificate and key. It has no opinion about how that routing gets enforced; by itself, an Ingress object sitting in the API server does nothing.
What the Ingress Controller does with it
The controller (for example nginx-ingress, Traefik, or a cloud provider's own controller such as GKE's or AGIC for Azure) watches for Ingress objects and translates the declared rules into a real, running configuration: it opens listener ports, terminates TLS using the referenced Secret, and applies the host/path routing logic, typically by configuring an underlying proxy (like nginx) or provisioning a cloud load balancer.
Why split the two, contrasted with a Service of type LoadBalancer
A Service of type LoadBalancer works at L4 (Layer 4, the TCP/UDP transport layer): it exposes one Service on one IP and port, with no visibility into HTTP hostnames or paths. Ingress works at L7 (Layer 7, the application layer): it can route many hostnames and paths to many different backend Services through a single entry point, but only because something understands HTTP well enough to look inside the request to make that decision, which is exactly the controller's job.
| Ingress object | Ingress Controller | |
|---|---|---|
| What it is | A declarative API object | A running workload/process |
| What it knows | Routing intent: hosts, paths, TLS references | How to actually enforce that intent on real infrastructure |
| Portability | The same Ingress YAML can, in principle, work with any controller | Controller-specific behavior (annotations, feature support) varies significantly |
| Analogy | A restaurant's order ticket | The kitchen that actually cooks the order |
Why split the responsibility this way
- No single "correct" implementation exists. Different environments need genuinely different routing implementations (self-hosted nginx, a cloud provider's managed load balancer, a service mesh's gateway); a single hardcoded Ingress-to-infrastructure mapping baked into the API server couldn't serve all of them.
- The API object stays stable across environments. The same Ingress manifest can move between a local cluster running nginx-ingress and a cloud cluster running a managed controller, changing only which controller is installed, not the application team's YAML.
- It matches Kubernetes' general controller pattern. Objects declare desired state; a controller reconciles that state against reality. Ingress is just this pattern applied to L7 routing, the same separation Kubernetes already uses everywhere else (a Deployment declares desired replica state, the Deployment controller makes it real).
Trade-offs and pitfalls
- Because the controller does the real work, its feature set and its vendor-specific annotations determine what's actually possible; two clusters running different controllers can behave differently from the exact same Ingress YAML, which undermines the portability the split is meant to provide unless you stick to well-supported, portable fields.
- No Ingress Controller installed means Ingress objects sit inert; a common early mistake is writing correct Ingress rules and then wondering why nothing routes, when the real gap is a missing or misconfigured controller.
- Cross-cutting concerns that need to apply before traffic even reaches a specific controller instance (global rate limiting, a shared web application firewall) are often better handled at an edge layer in front of the cluster rather than pushed entirely into per-Ingress annotations.
You are troubleshooting PersistentVolumeClaim (PVC) not binding to any PersistentVolume (PV). Describe the steps to debug PV/PVC binding: inspect storage class, volumeMode, accessModes, reclaimPolicy, CSI provisioner logs, and cloud provider storage quotas. Provide kubectl commands and sample outputs you would expect in common failure cases.
Sample Answer
A PersistentVolumeClaim (PVC, a request for storage made by a workload) that never binds means either no matching PersistentVolume (PV, the actual piece of storage) exists yet, the StorageClass's provisioner is failing to create one, or a cloud-side constraint (quota, permissions) is blocking provisioning; the fix is to work through those three in order rather than guess.
1. Inspect the PVC itself
kubectl get pvc -n myns my-pvc -o yaml
Read status.phase, spec.storageClassName, spec.volumeMode (Filesystem or Block), spec.accessModes (ReadWriteOnce, ReadWriteMany, etc.), and spec.resources.requests.storage. A very common, entirely benign-looking state:
status:
phase: Pending
Events:
Type Reason Message
---- ------ -------
Normal WaitForFirstConsumer waiting for first consumer to be created before binding
That event is not a failure by itself: with a StorageClass whose volumeBindingMode is WaitForFirstConsumer (the current default recommendation, since it lets the scheduler pick a zone/node before the volume is provisioned there), the PVC is expected to sit Pending until a pod that uses it is scheduled. If the pod itself is also Pending (check kubectl describe pod), the PVC will never move until the pod's own scheduling problem (node selector, taint, resource fit) is resolved first.
2. Check the StorageClass
kubectl get sc <name> -o yaml
Confirm the provisioner field names a real, currently-running Container Storage Interface (CSI, the standard plugin interface storage vendors implement) driver, and note volumeBindingMode and reclaimPolicy. A StorageClass pointing at a provisioner name with no matching CSI driver deployed in the cluster will never provision anything, and the PVC's own events are usually the first place that shows up (a "no volume plugin matched" or "provisioner not found" message).
3. List existing PVs for a manual-binding mismatch
kubectl get pv -o custom-columns=NAME:.metadata.name,SC:.spec.storageClassName,CAP:.spec.capacity.storage,MODE:.spec.volumeMode,ACCESS:.spec.accessModes,CLAIM:.spec.claimRef.name
If you're using statically pre-created PVs rather than dynamic provisioning, the common mismatches are: no PV with the same storageClassName, a PV with less capacity than requested, a volumeMode mismatch (PV is Block, PVC wants Filesystem), or a PV whose claimRef already points at a different claim.
4. Read the CSI provisioner's own logs
kubectl -n kube-system logs deploy/<csi-driver-controller-deployment> -c <provisioner-container>
(the exact deployment and container name depend on which CSI driver is installed, e.g. an EBS, Persistent Disk, or Azure Disk CSI controller; find it with kubectl get pods -n kube-system | grep csi). This is where the real, specific failure usually surfaces even when the PVC's own events are vague: a resource-exhausted error from the cloud API, an IAM/permission denial on the provisioner's own service account, or a zone mismatch between the requested volume and where the pod was scheduled.
5. Cloud-side quota and permissions
If the provisioner log shows a quota or resource-limit error, check the cloud provider's own quota view for the specific volume type in the specific region, and separately confirm the CSI controller's service account or workload identity actually has permission to create volumes (a permissions error and a quota error can produce very similar-looking provisioner log lines, so read the exact error string rather than assuming which one it is).
Trade-offs and pitfalls
WaitForFirstConsumerbinding being mistaken for a stuck PVC is the single most common false alarm in this whole flow; always check whether a pod is even trying to consume the PVC before treating Pending as broken.- Static PV/PVC binding (hand-created PVs) trades away dynamic provisioning's convenience for exact control over which underlying disk a claim binds to, at the cost of every capacity or accessMode mismatch becoming a manual matching problem instead of the provisioner's job.
- A ReadWriteMany PVC (multiple pods, multiple nodes, one volume) is not supported by most cloud block-storage CSI drivers at all; that class of "PVC never binds" is a genuine design mismatch (need a shared-filesystem storage class like a managed NFS/Filestore-backed one) rather than a transient failure to retry.
Explain the differences between ConfigMap and Secret objects. Show two ways to make a Secret available to a pod (environment variables and mounted files). Discuss basic security considerations for storing secrets and recommended best practices for CI/CD pipelines.
Sample Answer
A ConfigMap and a Secret are both key-value objects for feeding configuration into a pod without baking it into the image, but a Secret is meant for sensitive values and Kubernetes handles it slightly differently. Neither is encrypted by default: Secret values are only base64-encoded in etcd (the cluster's data store), which is trivially reversible, not encryption. Treat 'it's a Secret' as an access-control and audit boundary, not as cryptographic protection, unless encryption at rest has been explicitly turned on.
ConfigMap vs Secret
| ConfigMap | Secret | |
|---|---|---|
| Intended content | non-sensitive config: feature flags, config files, environment settings | sensitive values: passwords, tokens, keys |
| Storage in etcd | plain text | base64-encoded; not encrypted unless encryption at rest is configured |
| Immutable option | immutable: true field | immutable: true field |
Two ways to expose a Secret to a pod
Environment variable:
env:
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: my-secret
key: db-password
Mounted file:
volumes:
- name: secret-vol
secret:
secretName: my-secret
containers:
- name: app
volumeMounts:
- name: secret-vol
mountPath: /etc/creds
readOnly: true
Worked example: what actually happens when a Secret's value changes
Trace what happens after my-secret's db-password key is updated, depending on how it was exposed:
- Env var: the running pod keeps using the old value until it is restarted or replaced. Environment variables are read once at container start and never update in place.
- Mounted file, no subPath: the file at
/etc/creds/db-passworddoes eventually update, once the kubelet's periodic sync catches up (on the order of a minute, not instantly), but the running process only sees the change if it re-reads the file itself, since nothing forces that. - Mounted file with subPath: it never updates at all, because a
subPathmount is bound to a specific file version at mount time.
Neither the env-var path nor the volume path forces a pod restart on a config change. The common pattern to actually guarantee a fresh pod is to hash the ConfigMap or Secret's content into a pod template annotation (Helm'schecksum/configannotation is the usual form); a content change then produces a different pod template hash, which forces a real rollout instead of relying on an in-place file update the application may not even notice.
Trade-offs and pitfalls
- Base64 is not encryption; anyone with read access to Secret objects, or to etcd's data files directly, can decode it in one command. Enabling encryption at rest (an
EncryptionConfigurationbacked by a Key Management Service, KMS, provider, the current recommended approach since the older static-key KMS v1 API was deprecated as of Kubernetes 1.28) protects the etcd-at-rest copy; Role-Based Access Control (RBAC) is what actually protects who can read the Secret object in the first place, and the two are not substitutes for each other. - Environment variables are easy to leak: they show up when describing a running pod, in crash dumps, and in some logging frameworks that log the process environment; prefer mounted files for anything sensitive when the application can read from a file path instead.
- For continuous integration and continuous delivery (CI/CD) pipelines, never let plaintext secrets sit in pipeline configuration or version control; use the pipeline platform's own secret store, scope credentials as narrowly and as short-lived as possible, and prefer pulling secrets at deploy time from an external manager (HashiCorp Vault, a cloud provider's secrets manager, or a Secrets Store CSI driver) over baking them into a committed manifest, even one covered by a
.gitignoreentry.
Describe the kubectl commands and rollout strategies you would use to perform a safe rolling restart of a Deployment, view rollout history, and rollback to a previous revision. Include examples using kubectl and explain how you would avoid causing cascading failures during a restart of a consumer‑facing service.
Sample Answer
A safe restart uses kubectl rollout restart, which recreates pods through the normal RollingUpdate strategy rather than deleting them directly, so the same availability guarantees that protect a routine deployment protect the restart too.
Commands
Trigger and watch a rolling restart:
kubectl rollout restart deployment my-app -n prod
kubectl rollout status deployment my-app -n prod --watch
View rollout history and inspect a specific revision:
kubectl rollout history deployment my-app -n prod
kubectl rollout history deployment my-app -n prod --revision=3
Roll back:
kubectl rollout undo deployment my-app -n prod --to-revision=3
kubectl rollout status deployment my-app -n prod
Ship an image change with a recorded reason (the --record flag some older references use for this is deprecated; annotate explicitly instead):
kubectl set image deployment/my-app my-app=registry/app:1.2.3 -n prod
kubectl annotate deployment my-app kubernetes.io/change-cause="bump to 1.2.3, ticket OPS-441" --overwrite -n prod
What keeps a restart from becoming a cascading failure
- RollingUpdate parameters:
maxUnavailableandmaxSurge(both default to 25% of desired replicas) bound how many old pods can be down and how many extra new pods can exist at once. For a consumer-facing service, a conservative setting (for examplemaxUnavailable: 0, maxSurge: 1) never drops capacity below the current replica count during the restart, at the cost of briefly running more pods than the steady-state count. - Readiness probes: a Service only sends traffic to pods that pass their readiness probe, so a newly restarted pod that's still initializing doesn't receive requests it can't yet handle. This is the single biggest lever against a restart-induced error spike; without a readiness probe, the rollout has no signal that a "new" pod is actually ready and can start routing traffic to it immediately.
- PodDisruptionBudget (PDB): guarantees a minimum number (or percentage) of replicas stay available throughout the restart, independent of the Deployment's own
maxUnavailablesetting, which matters when other voluntary disruptions (a node drain, a cluster upgrade) happen to overlap with the restart window. - Graceful shutdown: a
preStophook plus aterminationGracePeriodSecondslong enough for in-flight requests to finish, combined with the Service removing the pod's endpoint before the container actually stops, avoids dropping requests that were already in progress when the restart began. - Staged rollout for risk-sensitive services: restarting (or deploying) to a small subset first, watching error rate and latency, then proceeding, catches a bad new revision before it reaches full traffic; this is a general staged-rollout practice, not a specific traffic-splitting mechanism (traffic-splitting techniques like weighted canary routing are a load-balancing/ingress-layer concern, not something the Deployment object itself provides).
Trade-offs and pitfalls
kubectl rollout restartonly recreates pods; it does not change the Deployment's spec, sorollout historyrecords it as a new revision with the same template, which is easy to forget when later trying toundoyour way back past a restart that changed nothing.- Setting
maxUnavailable: 0guarantees no capacity loss but requires enough spare cluster capacity formaxSurgeextra pods to schedule; on a tightly packed cluster this can leave the rollout stuck Pending on the surge pods instead of proceeding. - A rollback only restores the pod template (image, env, resource requests, and so on). If the Deployment reads a ConfigMap or Secret by a fixed name and that ConfigMap was edited in place rather than replaced with a new name or hash-suffixed name, rolling the Deployment back does not restore the old configuration content, only the old pod template pointing at the same (already-mutated) ConfigMap. This is the most common way a rollback fails to actually roll back.
- The same gap applies to a PersistentVolumeClaim (PVC): a Deployment's rollback restores the pod template's volume mount references, not the data on the volume itself. If the new version wrote a schema migration or otherwise mutated data in place on that volume, rolling the Deployment back gives you the old code pointing at already-changed data, not the old data. Anything stateful needs its own restore path (a volume snapshot or application-level backup) alongside the Deployment rollback, not instead of thinking about it separately.
Describe how Horizontal Pod Autoscaler (HPA) can scale based on a custom metric such as queue length. Which components are required (metrics adapter, exporter), how do you expose the metric to the cluster, and what operational pitfalls should you watch for when autoscaling on custom metrics?
Sample Answer
The Horizontal Pod Autoscaler (HPA) never talks to Prometheus, your app, or a cloud queue directly: it only ever queries one of three Kubernetes metrics APIs (metrics.k8s.io, custom.metrics.k8s.io, external.metrics.k8s.io), so scaling on queue length requires something that exposes the queue's depth through one of those APIs. The standard pipeline is: the application or a sidecar exposes the metric, a metrics system collects it, and a metrics adapter (most commonly prometheus-adapter, or a project like KEDA, Kubernetes Event-Driven Autoscaling, which ships its own adapter and is now the more common current choice specifically for external event sources like queues) is registered with the API server as an aggregated API and translates queries into that metric API's shape.
The three metrics APIs
| API group | What it serves | Tied to a Kubernetes object? |
|---|---|---|
metrics.k8s.io | CPU and memory only, from metrics-server | Yes (per pod/node) |
custom.metrics.k8s.io | Any metric associated with a specific Kubernetes object (a Deployment, a Service) | Yes |
external.metrics.k8s.io | Any metric not tied to a Kubernetes object at all | No |
A message queue's depth (say, a managed queue service or a Kafka consumer-group lag) is not a property of any Kubernetes object, so it belongs under external.metrics.k8s.io and the HPA metric type: External, not Pods or Object.
Required components and how the metric gets exposed
- Instrumentation: the application (preferred) or a sidecar exporter publishes a metric, e.g. a Prometheus gauge
myapp_queue_length{queue="orders"}on/metrics. - Prometheus scrapes it on a scrape job.
prometheus-adapteris deployed with rules mapping a PromQL query to an external metric name Kubernetes will expose, and it registers anAPIService(apiregistration.k8s.io) sokubectl get --raw /apis/external.metrics.k8s.io/v1beta1returns real data. This registration needs its own RBAC: aClusterRolegranting the HPA controller's service account (system:kube-controller-managerreaching through the aggregation layer) permission to read the external metrics API, plus the adapter's own service account needing permission to read Prometheus.- The HPA references the metric by name:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
minReplicas: 2
maxReplicas: 20
metrics:
- type: External
external:
metric: {name: queue_length_orders}
target: {type: AverageValue, averageValue: "100"}
(autoscaling/v2 has been the stable API since Kubernetes 1.23; the older v2beta2 was removed in 1.26, so any manifest or tooling still referencing it is out of date.)
Operational pitfalls
- A silently wrong zero is worse than a visible failure. If the adapter genuinely can't reach its data source, the HPA does not quietly treat that as "no load": it sets the
ScalingActivecondition toFalsewith reasonFailedGetExternalMetricand shows the metric's current value as<unknown>inkubectl describe hpa, holding replica count steady. The real danger is the opposite case: an adapter that, on a query error, returns a literal0instead of erroring. That looks healthy to the HPA and drives a real scale-down during an actual outage in the metrics path, which is why adapter error-handling is worth testing explicitly rather than assumed. - Cardinality. High-cardinality labels on the underlying metric (one series per customer ID, for instance) can make Prometheus memory blow up long before the HPA ever sees a problem; keep the label set the adapter maps from small and stable.
- Flapping. A noisy queue-length signal causes replica oscillation; use
behavior.scaleDown.stabilizationWindowSeconds(part of theautoscaling/v2HPA behavior fields) or pre-aggregate with a Prometheus recording rule rather than reacting to raw noise. - Latency in the chain. Scrape interval, adapter caching, and the HPA's own sync period all stack up between a real queue-depth change and a scaling action; a 15s scrape interval plus a slow adapter cache can easily add tens of seconds of lag, which matters for a bursty queue.
- Cost. Recomputing an expensive PromQL query on every HPA sync (default every 15 seconds) across many HPAs can meaningfully load a Prometheus instance; pre-aggregate with recording rules for anything non-trivial.
Validate end to end in staging with synthetic queue load before trusting this in production, watching the full chain (producer to queue to exporter to Prometheus to adapter to HPA) rather than any single hop in isolation.
Trade-off note
Hand-rolling prometheus-adapter rules gives full control over the PromQL mapping but is fiddly YAML to maintain; KEDA trades some of that flexibility for purpose-built scalers for dozens of common event sources (queues, streams, schedules) and is usually less operational overhead for exactly this "scale on queue depth" scenario, at the cost of being one more component to run alongside (or instead of) prometheus-adapter.
Unlock Full Question Bank
Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.