Distributed Systems Security and Trust Questions

Security problems that exist because a system is distributed: keeping trust state correct while it propagates across many services, clusters and regions. Covers credential, token and certificate revocation under eventual consistency and network partitions; fleet-wide rotation of signing keys, secrets and trust anchors without outages (key rollover and grace windows, canary rotation, recovery from a compromised root); distributed authorization (replicated policy decision points, cached decisions, fail-open versus fail-closed when an auth dependency degrades); propagating caller identity and permissions through service call chains; tamper-evident audit trails across services and regions (hash chains, Merkle proofs, ordering events with imperfect clocks); Byzantine and partially trusted participants; cross-cluster and cross-organization trust federation; securing shared distributed components such as caches and message brokers against injection, replay and cross-tenant access; protecting data in transit across region boundaries; and tenant isolation as a security blast-radius boundary. Steady-state mTLS, service-mesh identity and network segmentation mechanics are covered by zero-trust service-to-service security; single-system cryptography and KMS basics by applied cryptography.

MediumTechnical
38 practiced

Stateless JWTs are widely used on your platform. You need a way to immediately invalidate a compromised token, at scale, without giving up the performance benefits that made you choose JWTs in the first place. Design a revocation system that balances correctness and performance, and explain how SREs would operate, monitor, and scale it.

HardSystem Design
36 practiced

You need to build a tamper-evident, globally-consistent audit trail for security events that supports efficient range proofs and legal requests. Requirements: per-region append-only chains, a way to verifiably merge them across regions with proofs, efficient queries for time ranges and per-entity history, and operational tooling for SREs to generate proofs for auditors. Walk through the data structures and storage backends you would use, your indexing strategy, and how you'd manage retention and proof generation.

HardTechnical
40 practiced

Compare opaque tokens that require introspection with self-contained JWTs in a global microservices environment that experiences intermittent network partitions. Discuss revocation complexity, cacheability, introspection latency, consistency of revocation decisions, and proposed architectures for both approaches. Provide guidelines for SREs on when to prefer opaque tokens vs JWTs based on trust boundaries and SLOs.

HardTechnical
35 practiced

A root signing key used to mint service identity certificates was discovered to be compromised. You are the SRE lead. Produce an incident response and remediation plan: immediate containment and revocation steps, how to remove the compromised trust anchor, re-issue a new CA and rotate workload certificates at scale (millions of instances), communicate with external partners, and minimize downtime while making the system secure again. Discuss cryptographic constraints, rollback strategies, and monitoring to verify success.

MediumTechnical
41 practiced

Implement a simplified JWT verifier in Python that supports RS256 and key rotation. Requirements: function verify_jwt(token: str, jwks: List[Dict]) -> Dict where jwks is a list of key dicts like {"kid": "abc", "pem": "-----BEGIN PUBLIC KEY-----..."}. The verifier should: 1) parse the token header to find kid; 2) verify the signature against the matching public key (or try all keys if kid is missing); 3) validate exp and nbf claims; 4) return the claims as a dict or raise an error. Include a brief note on how you'd handle a missing kid during rotation in production.

Unlock Full Question Bank

Get access to all 17 Distributed Systems Security and Trust interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.