Distributed Systems Security and Trust Questions
Security problems that exist because a system is distributed: keeping trust state correct while it propagates across many services, clusters and regions. Covers credential, token and certificate revocation under eventual consistency and network partitions; fleet-wide rotation of signing keys, secrets and trust anchors without outages (key rollover and grace windows, canary rotation, recovery from a compromised root); distributed authorization (replicated policy decision points, cached decisions, fail-open versus fail-closed when an auth dependency degrades); propagating caller identity and permissions through service call chains; tamper-evident audit trails across services and regions (hash chains, Merkle proofs, ordering events with imperfect clocks); Byzantine and partially trusted participants; cross-cluster and cross-organization trust federation; securing shared distributed components such as caches and message brokers against injection, replay and cross-tenant access; protecting data in transit across region boundaries; and tenant isolation as a security blast-radius boundary. Steady-state mTLS, service-mesh identity and network segmentation mechanics are covered by zero-trust service-to-service security; single-system cryptography and KMS basics by applied cryptography.
Your platform needs a secrets and API key management system for services sitting behind an API gateway. It must support zero-downtime key rotation, immediate revocation, audit logging, and minimal blast radius if a key leaks. Explain the storage model, how keys get distributed to gateways and services, how you would orchestrate rotation, and the emergency revocation flow.
Sample Answer
Direct answer
Split the problem into two kinds of secret. Client API keys (presented by callers to the gateway) are stored only as hashes in a central key registry and pushed to gateways as a versioned, in-memory snapshot plus a change stream (a continuous feed of create, scope-change and revoke events, so a gateway only has to apply what changed rather than re-fetch everything), so a gateway can validate a key locally and a revocation reaches every gateway in seconds. Service secrets (database passwords, third-party credentials) live in a secrets manager, encrypted under keys held in a KMS (key management service) or HSM (hardware security module), and are fetched at runtime by workloads that authenticate with their platform identity, never with a secret baked into an image. Rotation always runs through a window where old and new are both valid; revocation skips that window and is pushed, acknowledged and audited.
Terms
- Blast radius: how much one leaked key can do. Shrink it with per-client, per-environment, narrowly scoped, short-lived credentials.
- Envelope encryption: each secret is encrypted with its own data key, and that data key is encrypted with a master key that never leaves the KMS/HSM.
- Secret zero: the first credential a brand-new workload needs in order to fetch all its other secrets. How you bootstrap it decides whether the whole design is sound.
Storage model
flowchart LR
ADM[Key admin API] --> REG[(Key registry: hashes + metadata)]
REG --> STR[Change stream]
STR --> G1[Gateway region A]
STR --> G2[Gateway region B]
SM[(Secrets manager)] --> KMS[KMS or HSM master key]
W[Service workload] -->|platform identity| SM
REG --> AUD[(Append-only audit log)]
SM --> AUD
Client API keys. Generate 256 random bits, format as pk_live_<keyid>_<secret>. Store keyid, SHA-256(secret), owner, scopes, environment, status, created and expiry times. Never store the key itself: it is shown once at creation. A fast hash is correct here (unlike passwords): passwords are hashed with deliberately slow algorithms like bcrypt because people choose low-entropy, guessable values, so the hash must be expensive enough to make guessing infeasible. A 256-bit random secret has no such weakness: it cannot be brute-forced regardless of hash speed, so bcrypt's deliberate slowness buys nothing and would add latency on every request. The keyid prefix lets gateways look up the record in one step and lets secret-scanning tools (automated scanners that watch code repositories for accidentally committed credentials) recognise a leaked key in a public repository.
Service secrets. Stored in the secrets manager under envelope encryption, organised by path per service and environment (prod/orders/db), with a policy granting each workload identity exactly its own paths. Prefer dynamic secrets where the backend supports them: the secrets manager creates a unique database user per workload with a lease (a grant that expires automatically unless the workload renews it before the deadline) of, say, 1 hour, so there is nothing long-lived to leak and each credential is attributable to one workload.
Sizing the gateway snapshot. With 2,000,000 active client keys at about 32 bytes of hash plus 100 bytes of metadata each: 2,000,000 x 132 = 264,000,000 bytes, about 252 MiB. That fits in gateway memory, so every validation is a local lookup with no network call. If it did not fit, gateways would cache on demand with a short TTL (time-to-live) and fall back to the registry on a miss.
Distribution to gateways and services
- Gateways load a full snapshot at start-up (tagged with a version number), then apply a change stream of
created / scope_changed / revokedevents in order. Each gateway reports the version it has applied, so the control plane (the system that manages and coordinates gateway configuration, as distinct from the data plane that carries live request traffic) always knows the laggards, the gateways that have fallen behind the latest version. If the stream breaks, the gateway re-fetches a snapshot; if it cannot, it keeps serving from its last snapshot and alerts, because refusing all traffic is worse than a snapshot a few minutes old, except for revocations (see below). - Services fetch their secrets at start-up through a local agent (a sidecar: a helper process deployed alongside the application container that does this fetching and renewal on the app's behalf), which renews leases before expiry and writes secrets to an in-memory file the app re-reads, so rotation does not require a redeploy.
Secret zero: bootstrapping a brand-new workload
The failure pattern is putting a long-lived token in CI (continuous integration) variables or the container image so the new container can log in to the secrets manager. Instead, let the platform vouch for the workload:
- CI/CD pipeline: the CI provider issues a short-lived OIDC (OpenID Connect) token describing the exact repository, branch and job; the secrets manager or the cloud provider's security token service (STS) trusts that issuer and exchanges the token for credentials scoped to that pipeline. No static secret exists in CI at all.
- Ephemeral containers: Kubernetes projects a service-account token that is audience-bound (it can only be redeemed by the exact service it was issued to, not replayed elsewhere) and expires; the agent presents it, the secrets manager validates it against the cluster and returns a short-lived token for that service's paths. On a cloud VM the equivalent is the instance identity document: a document signed by the cloud provider itself, attesting which specific VM instance is making the request.
- When no platform identity exists, use a single-use, short-lived wrapped token delivered by the deploy system (Vault calls this response wrapping: the secrets manager seals the real credential in an envelope that can be opened, or "unwrapped," exactly once): if an attacker used it first, the legitimate workload's unwrap fails and that failure is itself the alarm.
Zero-downtime rotation orchestration
For client API keys, rotation is the client's action, so the platform must allow two active keys per client:
- Client (or automation) creates key B; A and B are both valid.
- Client deploys B; the gateway emits
last_usedper key ID, so both sides can see traffic migrate from A to B. - When A has had zero traffic for a set period (for example 7 days), A is revoked. Keys also carry an expiry (for example 1 year), with warnings at 30 and 7 days.
For service secrets:
- Create the new credential in the backend (new DB password or user) while the old one stays valid.
- Publish it to the secrets manager as a new version; agents pick it up on their next refresh.
- Wait until every consumer reports the new version (or until one full refresh interval plus margin has passed).
- Disable the old credential, then delete it after a cooldown. If error rates rise after step 4, re-enabling the old credential is the rollback, which is why you disable before deleting.
Emergency revocation flow
Leak detected (secret scanner hit, anomaly alert or customer report):
- Revoke in the registry: status
revoked, with reason and actor recorded in the audit log. - Push: the revocation event goes on the change stream with priority; gateways apply it and acknowledge by version. Target: 99% of gateways within 5 seconds.
- Bound the tail: a gateway that has not acknowledged within the deadline is marked unhealthy and removed from the load balancer until it catches up. This turns "eventually revoked" into "revoked, or not serving traffic".
- Scope the damage: query the audit log for every request made with that key ID since the suspected leak time.
- Replace: issue a new key to the owner; for a service secret, rotate immediately without the normal overlap window, accepting a brief error spike.
Audit logging
Every key creation, scope change, revocation, secret read and failed authentication is written to an append-only store in a separate account that application operators cannot delete from. Log key IDs, never key material. Alert on reads from an unexpected identity, a spike in failed validations for one key ID (someone guessing), and use of a key from a new network or country.
Trade-offs and pitfalls
- Local validation vs instant revocation: local snapshots make validation fast and resilient but make revocation a distributed-propagation problem. The acknowledgement-plus-eviction step is what closes that gap; without it you only have "probably revoked".
- Rotation that requires a restart is rotation nobody runs. Hot reload of secrets is a prerequisite.
- Scoping is the cheapest blast-radius control: a key that can only read
ordersinstagingis a far smaller incident than a master key. - Common wrong turn: encrypting API keys reversibly "so support can look them up". If the platform can decrypt them, so can whoever compromises the platform.
Design authentication and authorization propagation across hundreds of microservices. Discuss token formats (JWT vs opaque tokens), token exchange patterns for internal services, token refresh and revocation strategies, minimization of token bloat in headers, and patterns for enforcing fine-grained permissions in downstream services (authorization middleware, PDP/PAP, or policy-as-a-service).
Sample Answer
Direct answer
Use opaque tokens at the edge and short-lived, audience-scoped JWTs inside the mesh (the network of proxies that carries and secures service-to-service traffic across the estate), minted by a central token service through token exchange at each trust hop. The edge gateway turns the user's external token into an internal token that carries only who the user is and which service it is meant for, not every permission they hold. Downstream services authenticate the calling service with its workload identity (for example mTLS, mutual TLS where both sides present certificates), validate the internal token locally, and ask a policy decision point for fine-grained decisions, with a local sidecar cache (a small helper process deployed alongside each service instance, sharing its network so it can intercept and answer calls without a remote hop) to keep that off the latency critical path. Revocation is handled by keeping internal tokens to a few minutes and pushing revocation events for the rare cases that cannot wait.
Vocabulary first
- JWT (JSON Web Token): a signed, self-contained token. Any service holding the issuer's public key can verify it without a network call. Downside: once issued it is valid until it expires.
- Opaque token: a random string meaningful only to the issuer. Verifying it means calling the issuer (introspection, RFC 7662). Revocation is instant, but every check is a network call.
- Token exchange (RFC 8693): an OAuth 2.0 grant where a service presents a token it holds and receives a new token for a different audience, typically with narrower scope.
- PEP / PDP / PAP / PIP: the policy enforcement point (code that blocks or allows the request), policy decision point (evaluates the policy), policy administration point (where policies are authored and versioned) and policy information point (the source of attributes such as team membership).
Architecture
flowchart LR
U[Client] -->|opaque token| GW[Edge gateway]
GW -->|exchange| TS[Token service]
TS -->|internal JWT aud=orders| GW
GW -->|JWT + mTLS| O[Orders service]
O -->|exchange for aud=payments| TS
O -->|new JWT + mTLS| P[Payments service]
O -.->|check| SC1[Policy sidecar]
P -.->|check| SC2[Policy sidecar]
PAP[Policy admin and store] -->|bundles| SC1
PAP -->|bundles| SC2
1. Token formats: commit per boundary
| Boundary | Format | Why |
|---|---|---|
| Internet client to edge | Opaque access token (or a JWT validated only at the edge) | Clients are untrusted and tokens leak (logs, browser storage); opaque tokens can be revoked instantly and reveal nothing |
| Edge to internal services | Short-lived JWT (2 to 5 minutes), signed by the internal token service | Hundreds of services each verify locally with a cached public key, so there is no per-hop call to the issuer |
| Service to service | Workload identity (mTLS certificate) plus the propagated user JWT when acting on a user's behalf | mTLS proves which service is calling; the JWT proves on whose behalf |
Introspecting an opaque token on every internal hop does not scale: 20,000 requests per second fanning out through an average of 6 internal hops is 120,000 introspection calls per second against one service, which becomes the platform's single point of failure.
2. Token exchange for internal calls
Forwarding the user's original token unchanged through the call chain is the common anti-pattern: every downstream service receives a token that is valid at every other service (a replayable bearer credential: whoever holds the bytes can use it, with no extra proof that they are the party it was issued to), and a compromised low-value service can replay it against a high-value one.
Instead, each hop exchanges:
- Orders receives a JWT with
aud: orders. - To call Payments it presents that token plus its own workload identity to the token service and requests
aud: payments. - The token service checks policy ("may orders call payments on behalf of a user?") and returns a JWT with
aud: payments, the samesub(the user), and anact(actor) claim recording that orders is the intermediary. RFC 8693 definesactfor exactly this delegation chain.
Payments rejects any token whose aud is not payments, so a token stolen from Orders is useless against Payments. Cost: one extra call per cross-service hop, which you cache: the exchanged token is reusable for its lifetime for the same user and audience.
3. Refresh and revocation
- Internal tokens are not refreshed; they are re-minted. They live 2 to 5 minutes, so the worst-case window in which a revoked user keeps access internally is that lifetime.
- Refresh tokens exist only between the client and the authorization server, are rotated on every use (a reused refresh token signals theft and revokes the whole family), and never enter the mesh.
- Urgent revocation (account takeover, fired employee): the auth server publishes a revocation event (
subor session ID plus a timestamp) on a stream; sidecars and gateways keep a small in-memory deny set and reject any token for that subject issued before the revocation time. The set stays small because entries can be dropped once every token issued before them has expired.
4. Minimizing token bloat
Putting a user's permissions in the token is the fastest way to break the platform. Measured by encoding real JWTs (ES256: an elliptic-curve signing algorithm producing a compact, fixed 64-byte signature, versus RS256, an RSA signature that runs 256 bytes for a 2048-bit key; compact JSON): jwt_len builds the same three parts a real token has (a header naming the algorithm and key, a claim set, and a fixed-length signature), base64url-encodes each, and adds them up, once with a bare claim set and once with a growing list of permission strings appended to it, to measure how token size scales with how many permissions ride inside it.
import base64, json
def b64url(b):
return base64.urlsafe_b64encode(b).rstrip(b"=")
def jwt_len(claims):
header = {"alg": "ES256", "kid": "2026-09-a", "typ": "JWT"}
h = b64url(json.dumps(header, separators=(",", ":")).encode())
p = b64url(json.dumps(claims, separators=(",", ":")).encode())
sig = b64url(bytes(64)) # an ES256 signature is always 64 bytes (r || s)
return len(h) + 1 + len(p) + 1 + len(sig)
base = {"iss": "https://auth.internal", "sub": "user:48213", "aud": "orders-api",
"exp": 1790000000, "iat": 1789999100, "jti": "8f14e45fceea167a"}
print("minimal claims:", jwt_len(base), "bytes")
for n in (50, 200, 500):
fat = dict(base, perms=[f"tenant-{i:04d}:orders:read" for i in range(n)])
print(f"{n} permission strings:", jwt_len(fat), "bytes")
minimal claims: 319 bytes
50 permission strings: 2066 bytes
200 permission strings: 7266 bytes
500 permission strings: 17666 bytes
Each permission string like tenant-0004:orders:read is 23 raw characters; as one element of a JSON array that is 26 bytes (23 plus 2 quote bytes plus a comma), and base64url encoding then expands every 3 raw bytes into 4 encoded characters, a 4/3 blow-up, so each permission costs about 26 x 4/3 ~= 34.7 bytes in the final token. That matches the numbers above exactly: 200 permissions minus 50 is 7266 - 2066 = 5200 bytes over 150 permissions, 5200 / 150 ~= 34.7 bytes each. nginx's default buffer for one large request header line is 8 KB, so a user with a few hundred permissions starts getting unexplained 4xx errors at the proxy, and a 7 KB header repeated across 6 hops is 42 KB of overhead per request. The fix is structural, not compression: keep identity claims (sub, aud, exp, iat, jti, tenant, maybe a coarse role) in the token and resolve permissions at decision time from the PIP. Using ES256 rather than RS256 also helps: its signature is 64 bytes versus 256 bytes for a 2048-bit RSA key.
5. Fine-grained authorization downstream
Three options, and where each fits:
| Pattern | What it is | Use it when |
|---|---|---|
| Authorization middleware in each service | A shared library checks scopes and simple rules in-process | Coarse checks: "does this token have orders:write?" |
| Centralized PDP (policy-as-a-service) | Services call one authorization service per decision | Relationship-heavy rules ("can this user edit this document?") where the data is large and central |
| Distributed PDP: policy sidecar with pushed bundles | Policies authored centrally (PAP), compiled into bundles, pushed to a sidecar next to each service (Open Policy Agent, OPA: a widely used open-source policy engine that evaluates authorization rules written in its own policy language, is the common one) | The default for hundreds of services: sub-millisecond local decisions, no central runtime dependency |
Recommendation: middleware for coarse scope checks, a sidecar PDP for most fine-grained rules, and a central relationship service only for resource-level sharing graphs. Every decision is logged with the policy version that produced it so an auditor can reconstruct why access was granted.
Trade-offs and pitfalls
- JWTs trade revocation for scale. The honest statement of the design is "internal access can outlive revocation by at most the internal token lifetime", and that lifetime is the knob.
- Token exchange adds a dependency. If the token service is down, cross-service calls fail. Run it per region, cache exchanged tokens, and treat it with the same SLO (service-level objective) as the edge.
- Do not let services mint their own internal tokens. Only the token service holds signing keys; services only verify.
- Policy drift: sidecars running stale bundles make inconsistent decisions. Export the bundle version each sidecar is running and alert when any lags the published version beyond a threshold.
- Common wrong turn: using the same token for user identity and service identity. Keep them separate: the certificate says which service, the token says which user.
Design a secure service-to-service authentication and authorization system across multiple clusters and regions. Cover token issuance, short-lived credentials, certificate/key rotation, trust boundaries, least-privilege authorization, fallback when an identity provider is down, and operational practices for key management.
Sample Answer
Direct answer
Give every workload a cryptographic identity issued by its own region's certificate authority after the platform attests what the workload is, and make every credential short-lived so rotation is continuous rather than an event. Authentication happens with mutual TLS (both sides present certificates) using those identities; authorization is a default-deny policy keyed on the caller's identity, evaluated locally. The design principle that answers "what if the identity provider is down" is: verification never needs the issuer online, and issuance is regional with enough credential lifetime left to ride out an outage.
Terms
- Workload identity: a name for a running service, for example the SPIFFE ID
spiffe://prod.eu.example/ns/payments/sa/api(SPIFFE, the Secure Production Identity Framework for Everyone, is an open standard for naming workloads and issuing them certificates, called SVIDs, short for SPIFFE Verifiable Identity Document: the actual certificate or token a workload presents). - Attestation: proving what a workload is from facts the platform controls (its Kubernetes service account, node, image) rather than from something the workload says.
- Certificate chain: a certificate is signed by the authority above it (its issuer); that issuer's own certificate is signed by the authority above it, and so on up to a root. A workload's certificate, at the bottom of that chain, is called a leaf certificate. To trust a leaf, a verifier walks the chain up to a root it already has in its trust bundle; if it has never been told about one of the intermediates in between, it cannot validate anything that intermediate signed, no matter how correctly the leaf itself was issued. That is what "chained to a CA" means, and it is why the rotation order below matters.
- Trust domain / trust boundary: the set of workloads that share a root of trust; crossing a boundary requires explicit federation: the two sides must each deliberately configure trust for specific identities in the other domain, rather than automatically trusting everything the other domain issues.
- Trust bundle: the set of CA (certificate authority) certificates a verifier accepts.
Architecture
flowchart TB
ROOT[Offline root CA] --> ICA1[Intermediate CA: region EU]
ROOT --> ICA2[Intermediate CA: region US]
ICA1 --> AG1[Node agents EU clusters]
ICA2 --> AG2[Node agents US clusters]
AG1 --> W1[Workload certs, 24h]
AG2 --> W2[Workload certs, 24h]
PAP[Policy repo] --> PE1[Local policy check EU]
PAP --> PE2[Local policy check US]
Token and credential issuance
- A node agent (a small process the platform runs on every host, distinct from the application itself) attests the workload using facts it can verify locally: its Kubernetes service account, its namespace (a Kubernetes grouping boundary that scopes and isolates related workloads), and its image digest (a cryptographic hash identifying the exact container image build, so a tampered or swapped image gets a different identity). It then requests a certificate from the regional intermediate CA.
- The workload receives an X.509 certificate (X.509 is the standard format essentially every TLS certificate uses) valid for 24 hours, with its SPIFFE ID in the subject alternative name, a field in the certificate that states which identity or identities it is valid for.
- For calls that also need application-level claims (user on whose behalf, audience), a regional token service issues JWTs (JSON Web Tokens: compact, signed tokens that carry a set of claims, such as who the caller is and who the token is for) valid for 5 minutes with an audience claim, signed with a key whose public half is distributed alongside the trust bundle.
Short-lived credentials and outage tolerance, with numbers
Agents renew a certificate when 50% of its lifetime remains. The worst case for an issuer outage is that it starts at the moment renewal was due, so the tolerance is the remaining lifetime at the renewal point:
| Certificate lifetime | Renew at | Issuer outage the fleet survives | Useful life of a stolen cert |
|---|---|---|---|
| 1 hour | 30 min left | 30 minutes | up to 1 hour |
| 24 hours | 12 h left | 12 hours | up to 24 hours |
| 7 days | 3.5 days left | 3.5 days | up to 7 days |
Commit to 24 hours for certificates. Twelve hours of outage tolerance is long enough to fix a broken CA across a night, and a day is a tolerable exposure window when combined with the other controls. Load is trivial: 50,000 workloads renewing every 12 hours is 50,000 / 43,200 s, about 1.2 issuances per second. The JWT layer stays at 5 minutes because it carries user context, where faster expiry matters more.
Certificate and key rotation
- Leaf certificates rotate continuously by construction.
- Intermediate CAs rotate every few months by overlap: add the new intermediate to trust bundles first, wait until every verifier has the updated bundle, then start issuing from it, and remove the old one only after the last certificate it signed has expired (24 hours plus margin). Reversing that order causes an outage because verifiers reject certificates chained to a CA they have not been told about.
- The root is kept offline in an HSM (hardware security module) and used only to sign intermediates. Root rotation is the same add-then-switch-then-remove sequence stretched over weeks.
- JWT signing keys rotate the same way, publishing the new public key before signing with it.
Trust boundaries
- One trust domain per region (or per environment), not one global domain. A compromised EU intermediate lets an attacker mint EU identities only.
- Cross-region calls use federation: US verifiers are explicitly configured to accept the EU bundle, for specific identities. The US payments service accepts
spiffe://prod.eu.example/ns/checkout/sa/apiand nothing else from EU. - Staging and production never share a root. This is the most common real-world boundary failure.
Least-privilege authorization
Policies are written against identities, not network addresses: "payments accepts calls from checkout on POST /charges; everything else is denied". Policies live in version control, are compiled and pushed to a local enforcement point in every workload (a sidecar, a helper process deployed alongside the application container, or a library linked into the application itself), and are evaluated without a network call. New services start with zero inbound permissions and receive them through review. Every allow and deny is logged with caller identity and policy version.
Fallback when the identity provider is down
Walk through what depends on the issuer:
| Function | Depends on issuer? | Behaviour during outage |
|---|---|---|
| Verifying an existing certificate or token | No: needs only the cached trust bundle and public keys | Unaffected |
| Renewing a workload certificate | Yes | Existing certs keep working for up to 12 hours; alert immediately |
| New pods starting (a pod is the smallest deployable unit in Kubernetes, one or more containers scheduled together on a host) | Yes | Scale-ups and deploys stall; freeze deploys in that region |
| Minting 5-minute JWTs | Yes, regional token service | Fail over to a second token-service instance in the same region; cross-region failover only if the other region's issuer is in the verifiers' federated bundle |
What we deliberately do not do: fall back to a long-lived static shared secret, or disable mTLS. Both turn an availability incident into a security incident. If the outage outlasts the 12-hour tolerance, the controlled response is an emergency extension: issue longer-lived certificates from a standby intermediate held for this purpose, logged and time-boxed.
Operational practices for key management
- Private keys for leaf certs are generated on the host and never leave it; only CSRs (certificate signing requests) travel.
- CA keys live in HSMs or a cloud KMS (key management service); root operations require two people (dual control) and are recorded.
- Monitor: certificate time-to-expiry across the fleet (alert when any workload is under 25% of lifetime), issuance error rates, and bundle version per verifier.
- Rehearse a compromised intermediate: revoke it by removing it from bundles, re-issue the region's workloads from a fresh intermediate, and measure how long that takes. That measured time is your real recovery objective.
Trade-offs and pitfalls
- Short lifetimes replace revocation lists. Certificate revocation lists and OCSP (Online Certificate Status Protocol) checks are rarely reliable at mesh scale (the scale of a service mesh, the network of many services all calling each other over these short-lived certificates); expiry is. The cost is issuer availability, which the table above quantifies.
- Regional isolation costs federation configuration. It is worth it: it bounds a CA compromise to one region.
- Clock skew breaks short-lived credentials. Allow a small skew tolerance (a minute or two) and monitor time sync.
- Common wrong turn: authorizing by IP or namespace label. Those change with every reschedule; identities do not.
You need to build a tamper-evident, globally-consistent audit trail for security events that supports efficient range proofs and legal requests. Requirements: per-region append-only chains, a way to verifiably merge them across regions with proofs, efficient queries for time ranges and per-entity history, and operational tooling for SREs to generate proofs for auditors. Walk through the data structures and storage backends you would use, your indexing strategy, and how you'd manage retention and proof generation.
Sample Answer
Direct answer
Model each region's audit trail as a Merkle-tree transparency log, the structure Certificate Transparency (the public, append-only logging system browsers use to catch mis-issued TLS certificates) uses (RFC 6962): events are appended as leaves (the tree's bottom-level entries, one per event), and every so often the region publishes a signed tree head, a signature over the tree's root hash (a single hash that commits to the contents of every leaf beneath it) and size. Anyone holding a signed head can check two things with a handful of hashes: that a given event is in the log (an inclusion proof) and that today's log is an append-only extension of yesterday's (a consistency proof). To merge regions verifiably, a global log takes each region's signed head at fixed epochs (say every minute) as its own leaves, producing one global root per epoch that commits to all regions at once. Event data lives in write-once object storage, a separate index serves time-range and per-entity queries, and proof generation is a tool SREs run on demand for auditors.
The data structures
Hash chain vs Merkle tree. A hash chain (each record includes the hash of the previous one) is tamper-evident, but proving that event #5,000,000 is in the log means handing over everything after it. A Merkle tree hashes events in pairs, then pairs of pairs, up to one root. Proving one event is included takes one sibling hash per level, about ⌈log2n⌉ hashes:
| Log size | Sibling hashes in an inclusion proof | Proof size (SHA-256) |
|---|---|---|
| 1,000 | 10 | 320 bytes |
| 1,000,000 | 20 | 640 bytes |
| 8.64 billion (one day at 100k events/s) | 34 | 1,088 bytes |
Per-region append-only log. Each region runs its own log (append-only by construction: the log server only appends leaves, and consistency proofs let anyone detect a rewrite). Leaf = canonical serialisation of the event (encoding it into bytes in one fixed, unambiguous order, so hashing the same event twice always produces the same hash; here: event ID, entity ID, actor, action, timestamp, payload hash).
Verifiable cross-region merge. Regions do not share one total order in real time; forcing one would make every write wait on a cross-region round trip. Instead, at each epoch boundary, each region's signed head becomes a leaf in the global epoch tree, in a fixed region order. The global root for epoch 1042 therefore commits to exactly which events every region had at that moment. A proof for any event is a pair: inclusion in its regional tree, plus inclusion of that regional head in the global tree.
flowchart TB
subgraph US[us region log]
U1[events] --> UR[signed head us]
end
subgraph EU[eu region log]
E1[events] --> ER[signed head eu]
end
subgraph AP[ap region log]
A1[events] --> AR[signed head ap]
end
UR & ER & AR --> G[Global epoch tree, root per minute]
G --> W[(WORM archive + external witnesses)]
Runnable demonstration (Python 3, standard library)
This follows RFC 6962's hashing: leaves are hashed with a 0x00 prefix and interior nodes with 0x01, so a leaf can never be passed off as an interior node.
import hashlib
def H(b): return hashlib.sha256(b).digest()
def leaf(d): return H(b"\x00" + d) # RFC 6962 leaf hash
def node(l, r): return H(b"\x01" + l + r) # RFC 6962 interior hash
def split(n): # largest power of 2 < n
k = 1
while k * 2 < n: k *= 2
return k
def root(items):
if len(items) == 1: return leaf(items[0])
k = split(len(items))
return node(root(items[:k]), root(items[k:]))
def prove(items, i): # audit path for item i
if len(items) == 1: return []
k = split(len(items))
if i < k: return prove(items[:k], i) + [("R", root(items[k:]))]
return prove(items[k:], i - k) + [("L", root(items[:k]))]
def verify(item, path, expected_root):
h = leaf(item)
for side, sib in path:
h = node(h, sib) if side == "R" else node(sib, h)
return h == expected_root
# Regional logs for one epoch (a minute of events per region)
regions = {
"us": [f"us|{i}|user:42|login".encode() for i in range(5)],
"eu": [f"eu|{i}|user:7|export".encode() for i in range(3)],
"ap": [f"ap|{i}|svc:billing|keyread".encode() for i in range(6)],
}
region_roots = {r: root(ev) for r, ev in regions.items()}
# Global epoch tree: leaves are (epoch, region, regional root), fixed region order
epoch = 1042
g_leaves = [f"{epoch}|{r}|".encode() + region_roots[r] for r in sorted(regions)]
global_root = root(g_leaves)
# Proof that eu event #2 is in the globally committed epoch
ev = regions["eu"][2]
p1 = prove(regions["eu"], 2)
gi = sorted(regions).index("eu")
p2 = prove(g_leaves, gi)
ok1 = verify(ev, p1, region_roots["eu"])
ok2 = verify(g_leaves[gi], p2, global_root)
print("regional inclusion:", ok1, "| proof hashes:", len(p1))
print("global inclusion: ", ok2, "| proof hashes:", len(p2))
# Tamper: someone edits the stored event after the root was signed
tampered = b"eu|2|user:7|view"
print("tampered event verifies:", verify(tampered, p1, region_roots["eu"]))
# Proof size grows with log2(n), not n
import math
for n in (1_000, 1_000_000, 8_640_000_000):
print(f"n={n:>13,} -> about {math.ceil(math.log2(n))} sibling hashes, {math.ceil(math.log2(n))*32} bytes")
Output:
regional inclusion: True | proof hashes: 1
global inclusion: True | proof hashes: 2
tampered event verifies: False
n= 1,000 -> about 10 sibling hashes, 320 bytes
n= 1,000,000 -> about 20 sibling hashes, 640 bytes
n=8,640,000,000 -> about 34 sibling hashes, 1088 bytes
The demo omits signatures to stay short; production signs every published root with a key in an HSM or KMS (below). It does not omit consistency proofs.
Walking through the proof, by hand
split(n) finds the largest power of two strictly less than n. That is how RFC 6962 shapes an unbalanced tree for any size: the left side is always a perfect subtree of size k, and the right side (n - k items) is built the same way, recursively. For the eu region above, 3 events, split(3) = 2, so the tree is node( node(L0, L1), L2 ): leaves 0 and 1 pair up first, and leaf 2 hangs directly off the root.
That is why prove(regions["eu"], 2) (the proof for event index 2) returns exactly one sibling instead of the ceil(log2(3)) = 2 you might expect from a balanced tree: at the top level, index 2 is not less than k = 2, so the function takes the ("L", root(items[:2])) branch and recurses into prove(items[2:], 0), a single-item list, which returns []. The whole proof is that one ("L", ...) entry: the combined hash of leaves 0 and 1, labelled L because it sits to the left of leaf 2 when they are joined.
verify() starts from leaf(item) and folds in each sibling in the order the proof lists them: "R" means the sibling is combined on the right (node(h, sib), used when the proved leaf was in the left half at that level), "L" means it is combined on the left (node(sib, h), used here). One fold with the leaves-0-1 hash on the left reproduces the root, so verify returns True.
Consistency proof and range proof, worked small
Both are missing from the demo above; here they are, small enough to read in full and to run on their own (repeating the leaf/node/split/root helpers from the first demo).
import hashlib
def H(b): return hashlib.sha256(b).digest()
def leaf(d): return H(b"\x00" + d)
def node(l, r): return H(b"\x01" + l + r)
def split(n):
k = 1
while k * 2 < n: k *= 2
return k
def root(items):
if len(items) == 1: return leaf(items[0])
k = split(len(items))
return node(root(items[:k]), root(items[k:]))
def consistency_proof(items, m): # proves size-m tree is a prefix of size-n tree
n = len(items)
return _subproof(items, m, n, True)
def _subproof(items, m, n, b):
if m == n:
return [] if b else [root(items)]
k = split(n)
if m <= k:
return _subproof(items[:k], m, k, b) + [root(items[k:])]
return _subproof(items[k:], m - k, n - k, False) + [root(items[:k])]
def verify_consistency(m, n, proof, old_root, new_root):
proof = list(proof)
def rec(m, n, b, claimed_old):
if m == n:
if b:
return claimed_old, claimed_old
h = proof.pop()
return h, h
k = split(n)
if m <= k:
right = proof.pop()
oh, nh_left = rec(m, k, b, claimed_old)
return oh, node(nh_left, right)
left = proof.pop()
oh, nh_right = rec(m - k, n - k, False, claimed_old)
return node(left, oh), node(left, nh_right)
oh, nh = rec(m, n, True, old_root)
return oh == old_root and nh == new_root
def range_proof(items, start, end): # boundary hashes for the contiguous range [start, end]
n = len(items)
needed = []
def walk(lo, hi):
if hi <= start or lo > end:
needed.append(root(items[lo:hi]))
elif lo >= start and hi - 1 <= end:
return
else:
k = split(hi - lo)
walk(lo, lo + k); walk(lo + k, hi)
walk(0, n)
return needed
def verify_range(range_items, start, end, n, needed, expected_root):
needed = list(needed)
def rebuild(lo, hi):
if hi <= start or lo > end:
return needed.pop(0)
if lo >= start and hi - 1 <= end:
return _root_sub(range_items[lo - start:hi - start])
k = split(hi - lo)
return node(rebuild(lo, lo + k), rebuild(lo + k, hi))
def _root_sub(sub):
if len(sub) == 1: return leaf(sub[0])
k = split(len(sub))
return node(_root_sub(sub[:k]), _root_sub(sub[k:]))
return rebuild(0, n) == expected_root
# 1. Consistency proof: the eu region's log grows from 2 events to 3.
eu_old = [f"eu|{i}|user:7|export".encode() for i in range(2)]
eu_new = [f"eu|{i}|user:7|export".encode() for i in range(3)]
old_root, new_root = root(eu_old), root(eu_new)
cproof = consistency_proof(eu_new, 2)
print("consistency proof, 2 events -> 3 events: hashes needed =", len(cproof))
print("verifies (append-only holds):", verify_consistency(2, 3, cproof, old_root, new_root))
rewritten = [b"eu|0|user:7|EDITED-LATER", eu_new[1], eu_new[2]] # attacker edits event 0, still grows to 3
fake_new_root = root(rewritten)
print("same proof against a silently rewritten history:", verify_consistency(2, 3, cproof, old_root, fake_new_root))
# 2. Range proof: prove events 1..3 of a 5-event log, nothing omitted from the middle.
five = [f"us|{i}|svc:billing|keyread".encode() for i in range(5)]
five_root = root(five)
rproof = range_proof(five, 1, 3)
print()
print("range proof for events[1..3] of 5: boundary hashes needed =", len(rproof))
print("verifies with the real 3 events:", verify_range(five[1:4], 1, 3, 5, rproof, five_root))
tampered_middle = [five[1], b"us|2|svc:billing|EDITED", five[3]]
print("verifies if the middle event is edited:", verify_range(tampered_middle, 1, 3, 5, rproof, five_root))
dropped_middle = [five[1], five[3]] # attacker returns only 2 of the 3 claimed events
try:
ok = verify_range(dropped_middle, 1, 3, 5, rproof, five_root)
except IndexError:
ok = False
print("verifies if the middle event is silently dropped:", ok)
Output:
consistency proof, 2 events -> 3 events: hashes needed = 1
verifies (append-only holds): True
same proof against a silently rewritten history: False
range proof for events[1..3] of 5: boundary hashes needed = 2
verifies with the real 3 events: True
verifies if the middle event is edited: False
verifies if the middle event is silently dropped: False
For 3 leaves, proving events 0 and 1 haven't been disturbed while a third is appended needs only 1 hash: the new leaf hangs directly off the old root (the same unbalanced shape explained above), so that old root itself is the one sibling the consistency proof supplies. For the range proof, the 5-event log splits as node(node(node(L0,L1),node(L2,L3)),L4); proving events 1 to 3 needs the hash of L0 (the left edge of the first proved leaf) and the hash of L4 (everything to the right of the range), 2 boundary hashes. Handing over 2 hashes instead of leaf 2's own separate inclusion proof is what proves the auditor's own copy of events 1, 2 and 3 is complete: drop or edit any one of the three and the reconstructed root stops matching, exactly as the last two lines show.
Storage backends
| Layer | Backend | Holds |
|---|---|---|
| Log server | A transparency-log implementation (Google's open-source Trillian is the established one) or a small custom service backed by a relational database | Leaf hashes, tree state, signed heads |
| Event payloads | Object storage with WORM (write once, read many) retention, for example S3 Object Lock in compliance mode, where no user including the account root can delete or shorten retention on a locked version | Batched event files (Parquet, a compressed columnar file format well suited to large batches of structured records), keyed by region/epoch |
| Tree tiles | Same WORM storage | Precomputed subtree hashes so proofs are generated without recomputing the tree |
| Query index | Columnar store (a database that stores each column of data together rather than each row, which is fast for aggregate and filter queries over specific fields) or search engine (ClickHouse, OpenSearch) | Pointers from (entity, time) to (region, epoch, leaf index) |
Indexing strategy
The index is for finding events; the log is for proving them. The index never needs to be trusted, because every result can be proved against a signed head.
- Time ranges: leaves are appended in time order within a region, so a range maps to a contiguous leaf-index range per region. Store
epoch -> (first_leaf, last_leaf)per region; a time-range query is a lookup plus a scan of that slice. - Per-entity history: a secondary index
(entity_id, timestamp) -> (region, leaf_index). For each hit, generate an inclusion proof. Across regions, order by timestamp and then region, and state that cross-region order within one epoch is by timestamp, not causality. - Range proofs: for "all events in [t1, t2]", deliver the contiguous run of leaves plus the sibling hashes along the left edge of the first leaf and the right edge of the last leaf (worked small, with real code and output, in "Consistency proof and range proof, worked small" above); the auditor hashes the delivered events together with those boundary siblings and must reproduce the signed root exactly. That proves nothing was omitted from the middle of the range, which a list of separate inclusion proofs cannot.
Retention, legal hold and compliance
- Defensible retention policy: each event class has a documented retention period (for example seven years for security events), applied as a default retention on the WORM bucket. Deletion happens only by automated expiry after the lock lapses, and each expiry batch is itself an audit event.
- Legal holds: a legal hold on a locked object keeps it immutable beyond its retention date until explicitly removed. Legal-hold search runs on the index ("all events for entity X or matter Y"), then places holds on the matching object versions, and records the hold list as a signed manifest so you can later prove what was held and when.
- Deletion vs immutability: privacy law may require erasing a person's data from a log you promised never to change. Resolve it with crypto-shredding: encrypt each entity's payloads with its own data key and put only the ciphertext hash in the tree. Destroy the key and the content is unreadable, while every proof still verifies.
- Integrity proof via key management: signed heads are signed by a key in a KMS or HSM (key management service or hardware security module) the log operators cannot export; the KMS's own access log shows every signing call. For an external audit you hand over the public key, the signed heads, and the KMS log showing the key was used only by the log service. That is what turns "our logs are intact" into evidence.
- Witnessing: publish each global root to independent witnesses (another team's account, a partner, or a public timestamping service). A log operator who rewrites history would need to fork every witnessed head, which consistency checks expose.
Operational tooling for SREs
audit-proof event <event_id>: returns the event, its regional inclusion proof, the regional signed head, the global inclusion proof and the global signed head, packaged with a standalone verifier script the auditor runs themselves.audit-proof range --entity X --from t1 --to t2: the events plus range proof.- A continuous monitor that fetches every new signed head, checks consistency with the previous one, and pages if any check fails or a region stops publishing heads (a stalled log can hide events as effectively as a rewritten one).
- Quarterly drill: hand a colleague playing auditor a proof package and have them verify it with nothing but the public key.
Trade-offs and pitfalls
- One global log is simpler to prove but puts a cross-region round trip on every append; per-region logs with an epoch merge keep writes local and delay only the global commitment by one epoch.
- Trusting the index is the common mistake; every answer to an auditor must be backed by a proof, not a query result.
- Signing key loss breaks nothing retroactively but stops new heads; keep a documented key rotation that signs the transition with both keys, and retain old public keys for the full retention period.
You operate a multi-tenant control plane where API keys and control APIs are currently shared across all tenants. Design an architecture that minimizes the blast radius if a single tenant is compromised, and explain how you would migrate off the shared-key model and what that isolation costs you versus what it buys you.
Sample Answer
Direct answer
With one shared API key, the blast radius (everything a single compromised credential can reach) of any tenant compromise is every tenant, because the control plane cannot tell tenants apart. The fix is to make the tenant a property of the credential, not of the request: issue per-tenant, short-lived, narrowly scoped credentials; derive the tenant ID from the authenticated identity and enforce it on every control-plane call; encrypt each tenant's data under its own key; and group tenants into cells (independent copies of the control plane, each serving a subset of tenants) so that even an infrastructure-level compromise reaches one cell, not everyone. I would migrate by running per-tenant credentials alongside the shared key, attributing and shadow-enforcing first (running the new tenant check and logging what it would have blocked, without actually blocking anything yet), moving tenants in waves, and revoking the shared key last. The cost is more keys, more deployment units and more operational work; what it buys is turning "one leak exposes all tenants" into "one leak exposes one tenant".
Why the shared model is so dangerous
Today, a request says "act on tenant 42" and carries the shared key. The control plane checks the key (valid) and trusts the tenant ID in the request. So:
- Anyone holding the key can act as any tenant by changing one parameter.
- A compromise of tenant 42's automation (a leaked CI (continuous integration) pipeline secret, a malicious insider at that customer) is a compromise of all tenants.
- You cannot revoke access for one tenant without breaking every tenant, so in an incident you are forced to choose between leaving the attacker in and a platform-wide outage.
- Audit logs cannot say which tenant actually made a call, which makes incident scoping guesswork.
Target architecture
flowchart LR
T[Tenant workload] -->|per-tenant short-lived token| GW[API gateway]
GW -->|tenant ID from token| CP[Control plane in cell 3]
CP --> AZ{Tenant check<br/>resource.tenant == token.tenant}
AZ -->|match| DS[(Tenant data<br/>encrypted with tenant key)]
AZ -->|mismatch| DENY[Deny and alert]
CP --> KMS[Key service<br/>per-tenant keys]
Layer 1: identity per tenant
- Each tenant gets its own credentials, ideally exchanged for short-lived tokens (for example 15-minute tokens from a token service) so a stolen token expires quickly. Long-lived API keys, where unavoidable, are per tenant, hashed at rest, scoped to specific actions, and rotatable without touching other tenants.
- Scopes per credential: a CI deploy key can deploy, but cannot read secrets or manage users.
Layer 2: authorization bound to the tenant
- The control plane takes the tenant ID from the verified token only. Any tenant ID in a URL or body must match it, or the call is denied and logged as a cross-tenant attempt.
- Every data-access path filters by tenant at the lowest layer possible (row-level security in the database, meaning the database itself filters rows by tenant, or per-tenant schemas) so an application bug that forgets the filter still cannot return another tenant's rows.
Layer 3: per-tenant encryption keys
- Envelope encryption: each tenant's data is encrypted with data keys that are themselves encrypted by a tenant-specific master key in the key service. A policy on each master key allows its use only by that tenant's workloads.
- A bug that returns the wrong ciphertext, or a stolen backup, does not yield another tenant's plaintext, and crypto-shredding (deleting the tenant's master key) is a clean offboarding.
Layer 4: cells and namespace isolation
- Cell-based architecture: split the control plane into N independent cells, each with its own database, queues and credentials to the underlying infrastructure. A small routing layer maps tenant to cell. A compromise of a cell's service account (a non-human identity that infrastructure code authenticates as, distinct from a human user's login), or a bad deploy, is limited to that cell's tenants.
- Tenant workloads run in their own namespaces (Kubernetes' isolated groupings of resources within one cluster; or full separate accounts, for high-value tenants) with network policies (Kubernetes rules that control which pods may talk to which, enforced by the cluster network itself) that deny cross-namespace traffic by default.
- This is the classic tenancy spectrum: pool (all tenants share infrastructure, separated by identity and data filters), silo (each tenant gets dedicated infrastructure) and bridge (a mix, for example pooled compute with siloed data). I would use pool within a cell for most tenants and silo for the few tenants whose contracts or regulators require it.
Layer 5: detection and containment
- Alert on cross-tenant denials (any occurrence is suspicious), unusual call volumes per tenant, and tokens used from new networks.
- Automated quarantine: a single action revokes all of one tenant's tokens and keys and freezes its control-plane writes, without affecting any other tenant. This is the capability the shared key made impossible.
Migration plan
- Inventory and attribute. Log every call made with the shared key with enough context (source network, user agent (the identifying string a client sends describing what software made the request), workload identity, the tenant ID it claimed) to know who uses it and for what. Typical finding: a few internal tools legitimately act across tenants; these become separately scoped operator identities with their own audit.
- Issue per-tenant credentials in parallel. Both the shared key and the new credentials work. SDKs (the client libraries customers install to call your API) and docs switch to the new model.
- Shadow enforcement. For calls with new credentials, evaluate the tenant check and log "would deny" without blocking. Fix false positives (usually legitimate cross-tenant operator workflows).
- Enforce per tenant, in waves. Once a tenant is fully on its own credentials, turn on enforcement for that tenant and disable the shared key for that tenant ID. Start with internal and low-risk tenants, then the rest.
- Deadline and cutoff. Publish a date; the remaining users get direct help. On the date, stop accepting the shared key, then rotate and destroy it (assume it has already leaked; it has been in many hands).
- Introduce cells afterwards as a separate project. Moving tenants between cells is a data migration; doing it together with the credential change doubles the risk in each step.
What it costs versus what it buys
Worked example with 500 tenants split into 10 cells of 50:
| Question | Shared key | Per-tenant credentials | Per-tenant credentials + 10 cells |
|---|---|---|---|
| Tenants exposed by one tenant's leaked credential | 500 | 1 | 1 |
| Tenants exposed by one compromised control-plane service account | 500 | 500 | 50 |
| Tenants affected by a bad deploy | 500 | 500 | 50 (if deploys roll cell by cell) |
| Credentials to manage | 1 | 500+ | 500+ |
| Control-plane deployments to operate | 1 | 1 | 10 |
Costs: credential lifecycle for hundreds of tenants (issuance, rotation, revocation, support tickets when a customer loses a key); per-tenant keys in the key service (per-key and per-request charges scale with tenant count); a fixed baseline cost for every cell (each cell's database and queues cost money even when mostly idle); deploy pipelines that must roll across cells; cross-cell features (global search, billing) get harder. Benefits: a single customer's compromise is contained to that customer; containment is a one-tenant action instead of a platform-wide outage; audit logs can scope an incident precisely; enterprise customers' security questionnaires get concrete answers.
My recommendation is per-tenant identity, tenant-bound authorization and per-tenant keys for everyone now, since they are cheap relative to the risk they remove, and cells once the tenant count or the value of the largest tenants makes a platform-wide incident unaffordable. Full silo per tenant only where a contract or regulator requires it, because at 500 tenants the fixed per-silo cost multiplies by 500.
Trade-offs and pitfalls
- Per-tenant keys with a shared admin path (one super-credential used by support tooling) recreates the original problem. Operator access must be scoped, time-bound and approved.
- Trusting the tenant ID in the request body after the migration is the most common way this design quietly fails.
- Cells without cell-aware deploys give no deploy-safety benefit; roll changes one cell at a time.
- Noisy neighbours (tenants that consume so much shared capacity, CPU, bandwidth, connections, that they degrade service for others on the same infrastructure) are an isolation problem too: without per-tenant rate limits, one compromised tenant can still degrade others even if it cannot read their data.
Unlock Full Question Bank
Get access to all 12 Distributed Systems Security and Trust interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.