Kubernetes Architecture, Operations, and Troubleshooting Questions
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Explain the Kubernetes networking model in detail for a DevOps team unfamiliar with it. Describe the IP-per-pod concept, the flat cluster network assumption, how pod-to-pod communication works across nodes, the role of the Container Network Interface (CNI) and kube-proxy, and any common limitations or implicit assumptions operators should be aware of when designing cluster networking.
Sample Answer
Kubernetes gives every Pod its own IP address on one flat, cluster-wide network where any Pod can reach any other Pod's IP directly, without port-mapping or NAT (Network Address Translation) in the middle. That single assumption is the whole model; everything else (Services, kube-proxy, the CNI plugin) exists to make that flat network real and to add load-balancing and service discovery on top of it.
The core assumption: IP-per-Pod, flat and routable
- Every Pod gets its own IP, allocated from the cluster's Pod network, not from the node's own network.
- Containers inside the same Pod share that IP and can reach each other over
localhost. - Any Pod's IP is expected to be reachable from any node in the cluster, as if the whole cluster were one big L3 (Layer 3, meaning IP-address-level) network, even though physically it spans many separate machines.
This is a deliberate simplification: application code never has to deal with port-mapping or think about which node it's running on to reach another Pod. The cost of that simplification is pushed down into the networking layer that has to make the "flat network" illusion actually true.
How Pod-to-pod traffic actually crosses nodes
Two Pods on the same node talk through a local bridge or virtual interface, no different in spirit from two processes on one machine. Two Pods on different nodes need their packets to physically cross the network between those nodes, and that's the job of the CNI (Container Network Interface) plugin, using one of a few common approaches:
- Overlay/encapsulation (for example VXLAN, used by Flannel by default): the CNI wraps the Pod packet inside another packet addressed node-to-node, then unwraps it on arrival. Simple to run, but every packet pays an encapsulation cost.
- Native routing (for example Calico's BGP mode): nodes advertise routes to each other's Pod IP ranges directly, so packets travel as ordinary routed IP traffic with no wrapping.
- In-kernel eBPF forwarding (for example Cilium): programs attached to the kernel's networking hooks forward Pod traffic with less per-packet overhead than either of the above.
Whichever mechanism is used, kube-proxy is not involved in raw Pod-to-Pod traffic; that's entirely the CNI's job. kube-proxy only comes into play for Service traffic, described next.
Where kube-proxy fits in
A Pod's IP is not stable: Pods get rescheduled, restarted, and replaced with new IPs constantly. A Service gives client code one stable address (a ClusterIP) that represents a group of Pods, and kube-proxy is the component that makes connections to that stable address actually land on one of the current, healthy backend Pods. It does this by programming the node's kernel with forwarding rules (commonly iptables or IPVS, an in-kernel load-balancing feature) built from the Service's list of ready backend Pods. Put simply: the CNI plugin makes any Pod reachable by its own IP; kube-proxy makes a stable Service IP resolve to whichever real Pod IP should currently handle the traffic.
Common limitations and implicit assumptions to watch for
- The underlying network must actually support it. Cloud VPC (Virtual Private Cloud) routing tables, security groups, or on-prem firewalls have to allow the CNI's chosen traffic pattern (raw routed IP, VXLAN-encapsulated UDP, or eBPF-forwarded packets) between every pair of nodes, or the "flat network" assumption breaks silently.
- IP address planning matters. The Pod CIDR (the address range Pods are allocated from) needs to be sized for cluster growth and must not overlap with the VPC's own address space, or routing becomes ambiguous.
- NetworkPolicy is opt-in. Without it, any Pod can reach any other Pod by default (east-west traffic is wide open); the flat-network model is about reachability, not isolation, and isolation has to be added deliberately.
- Encapsulation costs latency and throughput. An overlay CNI adds a real, if small, per-packet tax versus native routing or eBPF; this matters more as traffic volume grows.
- kube-proxy itself has a scaling ceiling. iptables-based rule chains get slower to evaluate as the number of Services and endpoints grows; IPVS or an eBPF-based CNI's own Service handling scales better at high counts.
Worked example
A Pod on Node A (IP 10.244.1.7) calling a Service backed by a Pod on Node B (IP 10.244.2.4):
- The calling Pod connects to the Service's ClusterIP, say
10.96.10.20:80. - kube-proxy's node-local rules (built from the Service's endpoint list) DNAT that connection to
10.244.2.4:8080, the real backend Pod. - The CNI plugin now has to deliver a packet addressed to
10.244.2.4, a Pod IP on a different node, across the physical network: an overlay CNI wraps it in a VXLAN packet destined for Node B's real IP; a routed CNI simply routes it there directly using advertised Pod-subnet routes. - Node B unwraps (if needed) and delivers the packet to the backend Pod over its local bridge.
Trade-offs and pitfalls for a team new to this model
- Don't assume Pod IPs are stable enough to hardcode anywhere; always address other workloads through a Service name, never a Pod IP directly.
- Introducing a service mesh or NetworkPolicy changes this baseline model by adding sidecar proxies or enforcement points into the path described above; treat this explanation as the foundation those layers build on, not the final picture.
- Multi-cluster or hybrid-cloud setups often break the single-flat-network assumption outright (two clusters don't share one Pod CIDR space by default), which is exactly why solutions in that space exist to stitch separate flat networks back together.
Describe the end-to-end service discovery and request flow when a Pod resolves a Service DNS name (e.g., 'my-service.default.svc.cluster.local'). Include CoreDNS lookup, how CoreDNS obtains service/endpoints data, kube-proxy behavior, and how traffic is routed to backend pods.
Sample Answer
A Pod's DNS query for my-service.default.svc.cluster.local goes to CoreDNS, which answers from Service and EndpointSlice objects it watches from the API server; kube-proxy separately programs the node's packet-forwarding rules from those same EndpointSlice objects, so the DNS answer and the actual routing path are produced by two independent watchers reading the same underlying data. For a headless Service, CoreDNS skips the middle step entirely and hands back backend Pod IPs directly.
1) The Pod's DNS query
Every Pod's /etc/resolv.conf is set by the kubelet to point at the cluster DNS Service's ClusterIP (CoreDNS), plus a search-domain list (default.svc.cluster.local, svc.cluster.local, cluster.local) that lets short names like my-service resolve without the full suffix. The application's query goes out as a normal UDP (falling back to TCP for large responses) request to that ClusterIP on port 53.
2) CoreDNS: where the answer comes from
CoreDNS runs the kubernetes plugin, which watches the API server for Service and EndpointSlice objects (an EndpointSlice groups the ready backend Pod IPs and ports for a Service; it replaced the older single Endpoints object specifically so that large services don't require rewriting one giant object on every Pod change). Current CoreDNS releases watch EndpointSlices exclusively for this data; they do not fall back to the older Endpoints API, which matters if you're running a pre-1.21-era cluster still on v1beta1 EndpointSlices.
- Normal (ClusterIP) Service: CoreDNS returns a synthetic A record for the Service's stable ClusterIP, plus SRV records for named ports. The Pod's connection always lands on that one virtual IP.
- Headless Service (
clusterIP: None): CoreDNS instead returns one A record per ready backend Pod IP, taken straight from the EndpointSlice. The client picks (or round-robins across) an actual Pod IP and connects to it directly. This is also how a StatefulSet's per-replica DNS names (pod-0.my-service.default.svc.cluster.local) resolve, since StatefulSets are built on a headless Service.
Because CoreDNS reacts to API watch events, a new or removed endpoint typically propagates into DNS within a second or two of the API server processing the change, not on a polling interval. Client-side and CoreDNS-side caching still add their own delay on top of that (see pitfalls below).
3) kube-proxy: turning EndpointSlice data into packet-forwarding rules
kube-proxy runs on every node, watches the same Service and EndpointSlice objects, and programs the node's kernel accordingly. It does not sit in the data path itself; it configures rules the kernel then executes per-packet.
| Mode | Mechanism | Status |
|---|---|---|
| iptables | Chains of DNAT rules, one jump per backend, selected pseudo-randomly | Default today |
| IPVS (IP Virtual Server, a Linux kernel load-balancing feature) | Backends programmed as IPVS "real servers" behind a virtual server, with real scheduling algorithms (round-robin, least-connection, etc.) | Preferred at very large Service/endpoint counts, where iptables' linear rule-chain lookups get expensive |
| nftables | A newer, more efficient kernel packet-classifier replacing the iptables framework underneath | Beta since Kubernetes 1.31, targeted to reach general availability around 1.33; iptables remains the upstream default in the meantime |
| userspace | Proxied connections through a kube-proxy userspace process | Removed entirely in Kubernetes 1.26 after a multi-release deprecation; not available on any current cluster |
Regardless of mode, the rule shape is the same idea: "packets to Service ClusterIP:port get DNAT'd (Destination Network Address Translation, rewriting the destination address) to one of the ready backend Pod IP:port pairs."
4) Runtime packet path
sequenceDiagram
participant API as API Server
participant DNS as CoreDNS
participant App as Client Pod
participant KP as Node kernel (kube-proxy rules)
participant BE as Backend Pod
API-->>DNS: watch Service + EndpointSlice (continuous)
API-->>KP: watch Service + EndpointSlice (continuous)
App->>DNS: query my-service.default.svc.cluster.local
DNS-->>App: ClusterIP A record (or Pod IPs if headless)
App->>KP: connect to ClusterIP:port
KP->>BE: DNAT to a ready endpoint Pod IP:port
BE-->>App: response
- If the chosen backend Pod is on the same node, the packet is delivered locally through the CNI's (Container Network Interface, the plugin that wires up Pod networking) bridge or veth pair without leaving the host.
- If it's on another node, the packet is routed (or encapsulated, depending on the CNI) across the node network, and some CNI/kube-proxy combinations perform source NAT on cross-node traffic so return packets route back correctly.
- For a headless Service, there's no DNAT step at all: CoreDNS already gave the client a real Pod IP, so kube-proxy is not involved in that connection.
- Only Pods that are Ready (passing their readiness probe) appear in the EndpointSlice at all, so an unready Pod is invisible to both DNS and kube-proxy simultaneously, not resolvable and not routed to.
Worked example: representative output
$ kubectl get endpointslices -l kubernetes.io/service-name=my-service
NAME ADDRESSTYPE PORTS ENDPOINTS
my-service-x7f2q IPv4 8080 10.244.1.12,10.244.2.9
$ dig +short my-service.default.svc.cluster.local
10.96.140.201
10.96.140.201 is the stable ClusterIP; 10.244.1.12 / 10.244.2.9 are the actual Pod IPs kube-proxy's rules will DNAT to. If the same Service were headless, the dig output would show 10.244.1.12 and 10.244.2.9 directly instead of the ClusterIP.
Trade-offs and pitfalls
- DNS record churn at scale: a Service with rapidly changing backends (frequent rollouts, aggressive autoscaling) generates a steady stream of EndpointSlice updates. CoreDNS itself answers from its watch cache correctly, but application-level and node-level DNS caching (and any fixed TTL a client library applies) can serve a stale backend IP after that Pod is gone; keep client-side DNS caching TTLs short for volatile Services rather than assuming CoreDNS's freshness is the only factor.
- Conntrack (connection tracking) table exhaustion or stale entries on a node can cause traffic to silently drop even though both DNS and the iptables/IPVS rules are correct; this is a distinct failure mode from DNS or kube-proxy misconfiguration and needs node-level conntrack metrics to diagnose.
- A Pod that resolves DNS successfully but still can't reach the backend usually means the DNS layer and the kube-proxy/CNI layer are fine but something else (NetworkPolicy, a firewall, or the backend Pod itself) is blocking the connection; treat DNS success as ruling out one layer, not all of them.
Design a simple CI/CD workflow that builds a container image, runs tests, pushes to a registry, and deploys to Kubernetes. Compare an imperative pipeline that calls kubectl apply versus a GitOps approach that updates a Git repo and lets a controller (e.g., ArgoCD/Flux) reconcile the cluster.
Sample Answer
An imperative pipeline (CI runs kubectl set image or kubectl apply straight against the cluster) is the fastest thing to stand up and fine for a single low-stakes environment, but it hands CI broad, standing write credentials to the cluster and leaves the cluster's actual state defined by whatever CI last did rather than by anything reviewable. A GitOps controller (Argo CD or Flux) inverts that: CI's job stops at pushing an image and updating a manifest in a Git repository, and an in-cluster controller with its own credentials pulls that repository and reconciles the cluster to match it, so Git becomes the single source of truth and CI never touches the cluster directly.
Shared pipeline stages
Both approaches share the same build side:
- Checkout the repo on a merged change.
- Build a container image tagged immutably by commit SHA, for example
registry.example.com/myapp:sha-abc123(never reuse a mutable tag likelatestfor a deployable artifact). - Run unit and integration tests against that image.
- Push the image to the registry.
They diverge at the deploy step.
Two deploy paths
flowchart LR
A[Merge PR] --> B[CI: build image sha-abc123]
B --> C[CI: run tests]
C --> D[CI: push image to registry]
D --> E{Deploy path}
E -->|Imperative| F[CI: kubectl set image]
F --> G[Cluster updated immediately]
E -->|GitOps| H[CI: bump tag in manifest repo]
H --> I[Argo CD / Flux controller]
I --> J[Controller reconciles cluster to match Git]
| Dimension | Imperative (kubectl apply/set image) | GitOps (Argo CD / Flux) |
|---|---|---|
| Who holds cluster-write credentials | CI runner, directly | Only the in-cluster controller; CI only needs Git and registry access |
| Source of truth for "what's deployed" | Whatever CI last ran, reconstructed from pipeline logs | The Git repository's current commit |
| Drift detection | None built in; a manual kubectl edit is invisible until someone notices | Automatic; the controller continuously diffs cluster state against Git and can auto-heal or alert |
| Rollback | kubectl rollout undo, bounded by the Deployment's retained ReplicaSet history | git revert the offending commit; the controller reconciles the cluster back automatically |
| Audit trail | Pipeline logs plus whatever the cluster's audit log captured | Git history and, for manifest changes, the pull-request review trail |
| Latency to apply | Immediate | Bounded by the controller's poll interval or webhook trigger, not instant |
Why this is an operator-pattern question, not just a CI/CD question
Argo CD's Application custom resource and Flux's GitRepository/Kustomization custom resources are themselves Kubernetes Custom Resource Definitions (CRDs, the mechanism for extending the Kubernetes API with new object types), reconciled by a controller running the same control loop every built-in Kubernetes controller runs: read desired state, read actual state, act to close the gap, repeat. The "desired state" here just happens to be a Git commit instead of a domain object like a database instance. That is why this sits in the same conceptual bucket as writing an operator, not in general CI/CD tooling: the deploy mechanism is a Kubernetes-native reconciliation loop, not a script that mutates the cluster once and exits.
Worked example: one change, two ways
A pull request bumps myapp to sha-abc123 and merges.
- Imperative: CI's final step runs
kubectl set image deployment/myapp myapp=registry.example.com/myapp:sha-abc123 -n prod, followed bykubectl rollout status deployment/myapp -n prodto confirm the rollout finished. If it needs to be undone,kubectl rollout undo deployment/myapp -n prodreturns to the previous ReplicaSet, but only as far back asrevisionHistoryLimitretains history, and there is no record in Git of what "previous" actually was. - GitOps: CI's final step is a commit to the manifest repository changing the image tag field to
sha-abc123and opening (or auto-merging, depending on policy) a pull request. Argo CD or Flux notices the new commit on its next sync, computes the diff against the live cluster, and applies it. Undoing the change isgit revert <commit>on the manifest repo; the controller reconciles the cluster back to the prior tag on its own, and the revert itself is a reviewable, timestamped Git object.
Trade-offs and pitfalls
- GitOps does not remove the security question, it relocates it: whoever can merge to the manifest repository can now change the cluster, so branch protection and required review on that repo are doing the job role-based access control (RBAC, governing who or what can perform which actions against which resources) used to do for direct
kubectlaccess. - Mixing the two models is the most common real-world mistake: an engineer runs a manual
kubectl applyorkubectl editfor a hotfix while a GitOps controller is also watching the same resources. Depending on the controller'sselfHeal/prune settings, the controller will silently revert the manual change on its next sync, which looks like a flaky rollback bug but is actually GitOps working exactly as designed against an out-of-band change. - GitOps reconciliation is eventually consistent by design; a team expecting
kubectl apply-style immediacy for an urgent hotfix needs either a fast webhook-triggered sync or an explicit, audited "break glass" imperative path, not silence about the delay. - Imperative pipelines are still the right choice for a genuinely disposable environment (a short-lived preview namespace per pull request) where the overhead of a Git-mediated reconciliation loop buys little.
A control plane upgrade introduced API incompatibility with a CRD-backed controller and caused mass pod failures. Explain how you would roll back the control plane safely, mitigate the failing controller to stop further damage, validate cluster integrity after rollback, and prevent similar compatibility regressions when upgrading in the future.
Sample Answer
Treat this as an incident, not a debugging session: stop the failing controller from doing further damage first, restore a known-good control plane second, validate correctness third, and only then work backward on prevention. The part worth being precise about is the rollback step itself: "roll back the control plane" is not a single universal command, and what it actually means depends heavily on whether the cluster is managed or self-run.
flowchart TD
A[Detect: mass pod failures after upgrade] --> B[Mitigate: stop the failing controller]
B --> C[Roll back control plane]
C --> D[Validate cluster integrity]
D --> E[Root-cause + prevent recurrence]
1. Stop the bleeding
Scale the misbehaving controller to zero so it stops acting on objects while you work:
kubectl -n <ns> scale deploy/<controller> --replicas=0
If the controller registered an admission webhook that's now rejecting or mutating requests based on the new (incompatible) API shape, remove or disable that webhook configuration rather than leaving it live while the controller itself is down, since a stale webhook can independently block unrelated traffic. If pods are stuck Terminating because the controller's own finalizers can no longer complete their cleanup logic against the new API, that's a sign to investigate before force-removing finalizers, which is a last resort that can leave orphaned cloud resources behind.
2. Roll back the control plane, correctly
This is the step where an overly generic plan is dangerous. What "rollback" actually means splits by environment:
- Self-managed (kubeadm or similar). There is no rollback command. etcd (the cluster's backing datastore) does not support a minor-version downgrade once components have written data in the new version's format, so the only real path is: restore etcd from a snapshot taken before the upgrade, then reinstall the previous kubelet/kubeadm/control-plane component versions. If no pre-upgrade etcd snapshot exists, there is no clean rollback, only forward fixes; this is why taking that snapshot has to happen before every upgrade attempt, not after something breaks.
- Managed control planes vary and are changing quickly. Amazon EKS added a native version-rollback capability in 2026 that lets you revert to the immediately previous minor version, but only within a limited window (documented as 7 days) after the upgrade; past that window it is not available and you're in the same "restore from backup" position as self-managed. Google Kubernetes Engine (GKE) is the opposite case from EKS: outside one narrow exception, GKE does not let you downgrade a cluster's control plane to a previous minor version at all, and attempting it is rejected outright with an explicit error that the requested version is not newer than the current one. GKE does support patch-level downgrades within the same minor version, and it has a Preview-stage rollback that only works mid-way through a two-step control-plane minor upgrade: after the first step (the binary upgrade) and only until the second step (the emulated-version upgrade) completes; once that second step finishes, the previous minor version is no longer reachable either. Don't assume your specific managed provider or cluster version has this capability; check it explicitly rather than treating "it's managed, so it rolls back" as a given, because that assumption is exactly what a previous, now-corrected version of this plan got wrong.
3. Validate cluster integrity after rollback
- API server health:
kubectl get --raw /livez?verboseand/readyz?verbose(prefer these overkubectl get componentstatuses, which is deprecated and unreliable on any control plane that isn't a single static node, since it hard-codes checks that don't reflect real high-availability setups). - Workload state:
kubectl get nodes,pods,deployments --all-namespaces, looking specifically forCrashLoopBackOffor stuckPendingpods that predate the rollback and shouldn't be blamed on it. - The CRD-backed controller itself: once rolled back, restart it at low replica count first, watch its logs and a sample of its managed custom resources'
statusfields for a few reconcile cycles before restoring full replicas. - Run an actual smoke test against the CRD: create/update/delete a throwaway custom resource and confirm the controller reconciles it correctly, rather than only checking that the controller process is running.
4. Prevent recurrence
- Require a pre-upgrade compatibility gate in the upgrade runbook: check the Kubernetes deprecated/removed API list for the target version against every CRD and controller in the cluster, not just built-in workload types.
- Maintain a staging cluster that mirrors production's actual CRD and controller versions (not a clean, controller-free cluster) and run the real upgrade against it first, including the CRD-backed controller's own test suite.
- Take an etcd snapshot immediately before every control-plane upgrade, unconditionally, regardless of whether the provider claims to support rollback.
- Add automated post-upgrade health gates (a scripted smoke test against critical CRDs and workloads) that must pass before an upgrade is considered complete, rather than relying on someone noticing pod failures.
Trade-offs and pitfalls
- Force-removing finalizers to unstick Terminating pods is fast but can leave real orphaned resources (cloud disks, external DNS records) behind; only do it once you've confirmed what that specific controller's finalizer was supposed to clean up.
- Restoring an etcd snapshot rolls back everything written since the snapshot, not just the CRD schema; any legitimate workload changes made in that window are lost too, which is the real cost of not having a native, narrower rollback available.
- The instinct to "wait and see if it recovers" after a bad upgrade is usually wrong for a CRD-backed controller actively causing mass pod failures; the controller should be stopped first, investigated second.
You maintain a widely used CustomResourceDefinition (CRD) and must introduce a breaking schema change. Design a migration strategy across API versions, including CRD versioning, conversion webhooks, a migration controller, data migration steps, and how to coordinate clients and controllers to avoid downtime.
Sample Answer
A breaking CRD (CustomResourceDefinition, the mechanism that registers a new custom object type into the Kubernetes API) schema change is handled by running two versions side by side rather than cutting over: add the new version as served alongside the old one, put a conversion webhook between them so every client keeps working regardless of which version it speaks, run a migration controller that rewrites persisted objects onto the new shape, and only flip the storage version (and eventually drop the old one) once nothing in the fleet still depends on it. The property that makes this safe is that served and storage are independent per-version flags: a version can stay served for old clients long after it has stopped being how objects are written to etcd.
1. CRD versioning
Kubernetes CRDs are apiextensions.k8s.io/v1 (the only supported API version; v1beta1 was removed in Kubernetes 1.22, so pin your tooling to v1 if you're touching anything older). Each entry in spec.versions can independently be served (clients can request it) and storage (this is the shape actually persisted in etcd; exactly one version must have storage: true).
spec:
versions:
- name: v1
served: true
storage: true # existing shape, still the write-path today
- name: v2
served: true
storage: false # new shape, readable/writable via conversion
conversion:
strategy: Webhook
webhook:
clientConfig:
service: {name: crd-converter, namespace: platform, path: /convert}
conversionReviewVersions: ["v1"]
Use conversion.strategy: Webhook (not None) whenever the change is genuinely breaking. None only works when the two versions are structurally identical except for a rename, since it does no real transformation.
2. Conversion webhook
A small, stateless HTTPS service implementing the ConversionReview API: given an object in version A, return it in version B. It runs on every read, list, and watch that crosses versions, so it sits on the request hot path and needs its own availability (multiple replicas, readiness probes, low latency) since a webhook outage makes the CRD's non-storage versions unreadable. The detail most people miss: conversion must be lossless for fields the other version doesn't know about. If v2 adds a field that v1 has no place for, converting v2 to v1 has to stash that field somewhere recoverable (commonly an annotation) rather than silently dropping it, or a client that reads it back as v2 loses data it never touched.
3. Migration controller
The webhook converts on the fly; it does not rewrite what's stored in etcd. A separate migration controller (or the community kube-storage-version-migrator add-on, API group storagemigration.k8s.io/v1beta1, which requires a 1.30+ API server with the StorageVersionMigrator feature gate) performs the one-time job of listing every existing object and issuing a no-op update, which forces the API server to re-serialize it in whatever version is currently marked storage: true. Until that rewrite runs, old objects sit in etcd in the old format and are converted up to v2 on read; new writes go straight to v2. Run the migration rate-limited and namespace-by-namespace so a large object count doesn't overload the API server, and track per-object migration state so a partial run can resume.
4. Coordinating clients and controllers
| Phase | storage version | served versions | who can be running |
|---|---|---|---|
| A: introduce v2 | v1 | v1, v2 | old clients/controllers unaffected; new code can start targeting v2 |
| B: bulk migrate | v1 | v1, v2 | migration controller rewrites objects; both old and new controllers reconcile correctly via the webhook |
| C: flip storage | v2 | v1, v2 | all writes land as v2 natively; v1 reads still work via conversion |
| D: retire v1 | v2 | v2 only | only after telemetry shows zero v1 requests |
Multiple controller versions can run against the same objects during B and C because every read comes back in whatever version the controller asked for; the risk is two controllers racing to write the same object in different versions and stomping each other's fields, which is why a canary-then-bulk migration (not "flip everything at once") matters.
5. Rollback and safety
Take an etcd snapshot before flipping the storage version. Rollback from phase C is just flipping storage back to v1 and leaving the webhook running; rollback from D (after v1 is dropped) is much harder, which is why D should only happen once request-metrics or audit logs confirm nothing is still asking for v1.
Worked scenario
For a CRD with roughly 40 known consumers, a realistic sequencing looks like: week 0, add v2 as served (storage stays v1), start the migration controller against a single low-traffic namespace as a canary and watch its metrics; week 1, if the canary is clean, run the migration across the rest of the objects in rate-limited batches; week 2, once 100% of objects report migrated: true in their status and the conversion webhook's request counts show v1 traffic has stopped growing, flip storage to v2; week 3+, after confirming no client has requested v1 in that window, drop v1 from served and retire the webhook. This is a schedule, not a measured result, and is meant to illustrate the ordering, not a fixed timeline for every CRD.
Trade-offs and pitfalls
- The conversion webhook is a new single point of failure on the CRD's read path; under-provisioning or forgetting its own rollout strategy turns a schema migration into an availability incident.
- Flipping the storage version before validating the bulk migration is fast but risky: if the new schema has a subtle bug, everything written after the flip is in the new (possibly wrong) shape.
- For a CRD with very few consumers, all of this machinery can be overkill: a coordinated breaking change with a short deprecation window and direct communication to the small client list may be cheaper than building conversion + migration infrastructure. Reach for the full multi-version path when you can't enumerate or coordinate every consumer directly.
Unlock Full Question Bank
Get access to all Kubernetes Architecture, Operations, and Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.