Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
Design a zero-downtime secret rotation process for thousands of running containers that avoids secret leakage and guarantees minimal exposure to old secrets. Explain how you would distribute new secrets, revoke old secrets, coordinate dependent services (databases, third-party APIs), and audit the rotation. Discuss use of short-lived tokens, TLS, key-version headers, and potential race conditions.
Sample Answer
Direct answer
Zero-downtime rotation rests on one rule: there must always be a moment when both the old and the new secret are accepted, and the old one is retired only after you have evidence nobody still uses it. So every rotation runs in four ordered phases: make the new secret acceptable to whoever verifies it, distribute it to every consumer, confirm the consumers have switched, and only then revoke the old one. I would push most of the fleet onto short-lived, per-instance credentials issued by a secrets manager (such as HashiCorp Vault) so that "rotation" becomes routine expiry, and run the four-phase process only for the long-lived secrets that cannot be made dynamic (third-party API keys, signing keys, a certificate authority or CA).
Terms used below
- Secret: anything that grants access: a database password, an API key, a TLS private key, a token-signing key.
- Short-lived (dynamic) credential: a credential minted on request for one workload, with a built-in expiry (its lease or TTL, time-to-live). Vault's database secrets engine, for example, creates a new database user per lease and drops it when the lease ends.
- Dual-valid window: the period when both old and new secrets work. Its length is the core design parameter.
- Key version header: a label sent alongside a signed or encrypted message that says which key produced it, such as the
kid(key ID) header in a JSON Web Token (JWT). It lets a verifier hold several keys at once and pick the right one.
The rotation state machine
stateDiagram-v2
[*] --> Staged: generate v2, store encrypted
Staged --> Accepted: verifier trusts v1 AND v2
Accepted --> Distributed: consumers fetch v2
Distributed --> Verified: telemetry shows no v1 use
Verified --> Revoked: v1 disabled at verifier
Revoked --> [*]
Distributed --> Accepted: rollback (consumers back to v1)
Each arrow is a gate with a check, not a timer alone. Rollback is possible right up to Revoked, and never after it, which is why the evidence gate before revocation matters most.
1. Distributing new secrets
- Pull, never bake. The core idea is that the workload's own platform identity, something the platform already hands it for free, doubles as its credential to the secrets store, so nothing has to be pre-shared. Two common implementations: a Kubernetes service-account token (a signed token the cluster issues to each pod) presented through Vault's Kubernetes auth method (Vault's plugin that verifies that token with the cluster), or a cloud workload identity (the equivalent for a VM or serverless task). Containers authenticate this way and fetch secrets at runtime. Nothing lives in the image or in environment variables baked into a deployment spec.
- A local agent per pod. A sidecar (a second container that runs alongside the application container in the same pod, sharing its network and storage) or init container (a container that runs once before the main container starts, then exits), such as Vault Agent, the Secrets Store CSI Driver (a Kubernetes storage plugin that mounts secrets from an external store as files), or a mesh agent (the local proxy a service mesh, a layer that manages service-to-service networking and certificates for every workload, installs next to each one) for certificates, fetches, renews and writes secrets to an in-memory volume (
tmpfs, so they never touch disk) and re-renders the file when a new version appears. The application watches the file and reloads. - Environment variables are the trap. A process reads its environment once at start. Kubernetes updates Secrets mounted as volumes eventually, but does not update values injected as environment variables or files mounted with
subPath. Those require a restart, which turns every rotation into a fleet-wide rolling restart. - Stagger the fetch. Consumers refresh on their own schedule with jitter (random offset) rather than all at once when a new version lands.
2. Coordinating dependent services
Order matters, and the rule is: the side that verifies changes first, the side that presents changes second.
| Dependency | How to get a dual-valid window | Revocation |
|---|---|---|
| Database you control | Preferred: dynamic per-lease users, so there is nothing shared to rotate. Otherwise two users (app_a, app_b): rotate the password of the one not in use, switch consumers to it, then rotate the other next time. | Drop or lock the old user and terminate its sessions: in PostgreSQL, changing a password does not disconnect sessions already logged in. |
| Third-party API key | Most providers allow two active keys. Create key 2, distribute, watch the provider's usage report or your own egress logs until key 1 is idle, delete key 1. | Delete at the provider. If the provider allows only one key, the window is zero: schedule a short maintenance window or put a proxy in front that holds the key. |
| Token-signing key (JWT) | Publish the new public key in the JSON Web Key Set (JWKS, the published list of keys verifiers trust) first, wait for verifiers' caches to refresh, then start signing with the new kid. | Remove the old key from the set only after the longest token lifetime has passed. |
| TLS certificates | Short-lived certs (hours to days) from an internal certificate authority, issued automatically by a mesh or an agent and hot-reloaded. | Expiry does the revoking. For a CA rotation, distribute a trust bundle (the set of CA, certificate authority, certificates a verifier is configured to trust) holding both CAs first, then reissue leaf certs (the end-entity certificates issued to individual services, as opposed to the CA's own certificate) from the new CA, then remove the old CA from the bundle. |
3. Guaranteeing minimal exposure to old secrets
"Exposure" means the time an old secret remains usable after we decided to replace it. The window must cover the slowest consumer, and nothing more:
Wdual=Trefresh+Tconn+Tmarginwhere T_refresh is the longest time for a consumer to pick up the new version, T_conn is the longest a pooled connection opened with the old secret stays open, and T_margin is a buffer added on top of the other two so the revocation decision does not hinge on borderline telemetry.
Key-version headers in practice
Anything signed or encrypted should say which key produced it, so verifiers and decryptors can hold several keys during a rotation:
- Tokens: the
kidin the JWT header selects the verification key from the key set. - Encrypted data at rest: store the key version next to each ciphertext (for example a
v2:prefix), so records written before the rotation still decrypt with v1 while new writes use v2, and a background job re-encrypts old records before v1 is retired. Vault's transit engine formats its ciphertexts this way. - Webhooks and service-to-service signatures: send a header naming the key version, and let the receiver accept the current and previous version during the window.
Without a version label, a verifier has to try every key it holds, and it cannot tell you from logs who is still using the old one, which removes the evidence the revocation gate depends on.
4. Auditing the rotation
- Store-side audit: every read of every secret version, by which identity, from where. In Vault, audit devices (file, syslog, socket) log every request; Vault refuses to serve requests if it cannot write to at least one enabled audit device, so auditing cannot be silently lost.
- Usage-side evidence: the verifier's logs of which key version was presented (DB login user, API key ID,
kidof verified tokens). This is the gate for revocation: "zero v1 uses in the last N minutes across all verifiers". - Rotation record: who or what triggered it, version numbers, the time each phase gate passed, and whether it rolled back. Alert when a secret passes its maximum age without rotating.
- Never log the value, only version IDs or a short hash.
5. Race conditions and how each is closed
- Consumer ahead of verifier. A pod fetches v2 before the database accepts it and fails to log in. Closed by ordering: the verifier accepts v2 before v2 is readable by consumers (the
StagedtoAcceptedgate). - Two rotators at once. A scheduled rotation and a manual one both generate a "v2", one overwrites the other, and half the fleet holds a secret that no longer exists. Closed by a compare-and-set write (a write that only succeeds if the stored value still matches what was last read, so two writers racing to change it cannot silently clobber each other) on the version number (write v3 only if the current version is still v2) or a lock.
- Revoke before drain. v1 is disabled while pooled connections or cached tokens still use it. Closed by the telemetry gate plus the connection-lifetime term in the window formula.
- Rollback re-enables a revoked secret. Rolling back after revocation would resurrect a secret you may be rotating because it leaked. Closed by making rollback only possible before
Revoked; after it, roll forward to v3. - Restart storm. A new version triggers every pod to restart or re-fetch at once, overloading the secrets store. Closed by jittered refresh and by avoiding restart-based delivery.
Worked example: 5,000 containers on one PostgreSQL cluster
Inputs (illustrative, chosen to show the arithmetic): consumers poll for a new version every 5 minutes; the connection pool recycles connections at most every 30 minutes; margin 10 minutes.
Wdual=5+30+10=45 minutesSo the old password is revoked no earlier than 45 minutes after distribution, and only if the database's login telemetry shows zero v1 logins for the last 10 minutes. If you shortened the pool's maximum connection lifetime to 10 minutes, the window drops to 25 minutes: connection lifetime, not secret distribution, is usually the dominant term.
Load on the secrets store if the fleet moves to dynamic credentials with a 1-hour lease renewed at half-life (a common convention: renew once half the lease's time remains, so one missed renewal still leaves margin before expiry), that is, every 30 minutes, 1,800 s:
1800 s5000≈2.8 requests/sThat is trivial. The risk is the burst: if all 5,000 pods restart within 60 seconds after an outage, that is 5000 / 60 ≈ 83 requests/s of new credential creation, and each creation is also a CREATE ROLE on the database. Plan for the burst: rate-limit creation, jitter start-up, and size the database's connection limits for the extra users briefly alive during overlap.
Trade-offs and pitfalls
- Dynamic credentials vs a shared rotated secret. Dynamic per-instance credentials give the smallest blast radius (how much is exposed when one thing goes wrong) (one leaked credential is one pod for one lease) and make revocation per instance, but they make the secrets store a hard dependency for every start-up. Mitigate with long enough leases to ride out a short store outage and a highly available store. I would still choose them for anything we control.
- Short TTLs are not free. A 5-minute TTL means a 5-minute store outage breaks everyone. Pick TTLs that exceed your store's realistic recovery time.
- In-place rotation (one credential, changed in place) always has a gap, because the change and the consumers' reload cannot be simultaneous. Use two credentials or dynamic ones.
- "Revocation" that only stops new logins is incomplete. Kill existing sessions and invalidate caches, or a leaked credential keeps working through an open connection.
- Emergency rotation is a different mode. When a secret has leaked, you shrink the dual-valid window deliberately and accept some errors, because leaving the old secret alive is worse than a brief outage. Decide that policy before the incident.
That is every published Service Discovery and Configuration Management question for Cybersecurity Engineer so far. Browse the other topics in this category, or practice this one interactively.