Kubernetes Architecture, Operations, and Troubleshooting Questions

How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.

EasyTechnical
50 practiced

Explain the Kubernetes networking model in detail for a DevOps team unfamiliar with it. Describe the IP-per-pod concept, the flat cluster network assumption, how pod-to-pod communication works across nodes, the role of the Container Network Interface (CNI) and kube-proxy, and any common limitations or implicit assumptions operators should be aware of when designing cluster networking.

MediumTechnical
54 practiced

Describe the end-to-end service discovery and request flow when a Pod resolves a Service DNS name (e.g., 'my-service.default.svc.cluster.local'). Include CoreDNS lookup, how CoreDNS obtains service/endpoints data, kube-proxy behavior, and how traffic is routed to backend pods.

MediumSystem Design
49 practiced

Design a simple CI/CD workflow that builds a container image, runs tests, pushes to a registry, and deploys to Kubernetes. Compare an imperative pipeline that calls kubectl apply versus a GitOps approach that updates a Git repo and lets a controller (e.g., ArgoCD/Flux) reconcile the cluster.

HardTechnical
44 practiced

A control plane upgrade introduced API incompatibility with a CRD-backed controller and caused mass pod failures. Explain how you would roll back the control plane safely, mitigate the failing controller to stop further damage, validate cluster integrity after rollback, and prevent similar compatibility regressions when upgrading in the future.

HardTechnical
40 practiced

You maintain a widely used CustomResourceDefinition (CRD) and must introduce a breaking schema change. Design a migration strategy across API versions, including CRD versioning, conversion webhooks, a migration controller, data migration steps, and how to coordinate clients and controllers to avoid downtime.

Unlock Full Question Bank

Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.