Kubernetes Architecture, Operations, and Troubleshooting Questions

How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.

MediumSystem Design
75 practiced

Design a multi-tenant strategy for a Kubernetes cluster that will host several internal teams. Discuss the use of namespaces, RBAC roles, resource quotas, network policies, and cost allocation. Provide pros and cons of single-cluster multi-tenant vs multiple clusters per team, and when you'd recommend each approach.

MediumTechnical
57 practiced

HorizontalPodAutoscaler (HPA) is not scaling a deployment even though CPU usage is above the configured target. Provide a troubleshooting plan including which metrics endpoints and API resources to check, how to validate metrics-server or custom metrics adapters, and common misconfigurations that prevent scaling.

HardTechnical
85 practiced

The Kubernetes API server is experiencing increased request latency. What metrics, logs, and traces would you collect to diagnose whether the bottleneck is etcd, admission controllers, or API server CPU/memory? Provide a prioritized triage checklist and remedial actions for each root cause.

EasyTechnical
85 practiced

A pod remains in Pending and the scheduler does not bind it. Describe commands and checks to determine why the pod is unscheduled: inspect resource requests/limits, node allocatable capacity, taints/tolerations, affinity rules, and namespace quotas. Mention concrete kubectl commands you would run.

HardTechnical
47 practiced

You have mixed hardware in the cluster: GPU nodes for machine learning, on-demand nodes for critical services, and spot instances for low-priority batch jobs. Explain how you'd use node labels, taints, tolerations, node selectors or affinity, and PodTopologySpread to ensure correct scheduling and protect critical workloads from being placed on spot instances.

Unlock Full Question Bank

Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.