Distributed Systems Security and Trust Questions

Security problems that exist because a system is distributed: keeping trust state correct while it propagates across many services, clusters and regions. Covers credential, token and certificate revocation under eventual consistency and network partitions; fleet-wide rotation of signing keys, secrets and trust anchors without outages (key rollover and grace windows, canary rotation, recovery from a compromised root); distributed authorization (replicated policy decision points, cached decisions, fail-open versus fail-closed when an auth dependency degrades); propagating caller identity and permissions through service call chains; tamper-evident audit trails across services and regions (hash chains, Merkle proofs, ordering events with imperfect clocks); Byzantine and partially trusted participants; cross-cluster and cross-organization trust federation; securing shared distributed components such as caches and message brokers against injection, replay and cross-tenant access; protecting data in transit across region boundaries; and tenant isolation as a security blast-radius boundary. Steady-state mTLS, service-mesh identity and network segmentation mechanics are covered by zero-trust service-to-service security; single-system cryptography and KMS basics by applied cryptography.

HardSystem Design
45 practiced

Your platform needs a secrets and API key management system for services sitting behind an API gateway. It must support zero-downtime key rotation, immediate revocation, audit logging, and minimal blast radius if a key leaks. Explain the storage model, how keys get distributed to gateways and services, how you would orchestrate rotation, and the emergency revocation flow.

HardSystem Design
33 practiced

Design authentication and authorization propagation across hundreds of microservices. Discuss token formats (JWT vs opaque tokens), token exchange patterns for internal services, token refresh and revocation strategies, minimization of token bloat in headers, and patterns for enforcing fine-grained permissions in downstream services (authorization middleware, PDP/PAP, or policy-as-a-service).

HardTechnical
45 practiced

Design a secure service-to-service authentication and authorization system across multiple clusters and regions. Cover token issuance, short-lived credentials, certificate/key rotation, trust boundaries, least-privilege authorization, fallback when an identity provider is down, and operational practices for key management.

HardSystem Design
36 practiced

You need to build a tamper-evident, globally-consistent audit trail for security events that supports efficient range proofs and legal requests. Requirements: per-region append-only chains, a way to verifiably merge them across regions with proofs, efficient queries for time ranges and per-entity history, and operational tooling for SREs to generate proofs for auditors. Walk through the data structures and storage backends you would use, your indexing strategy, and how you'd manage retention and proof generation.

HardTechnical
34 practiced

You operate a multi-tenant control plane where API keys and control APIs are currently shared across all tenants. Design an architecture that minimizes the blast radius if a single tenant is compromised, and explain how you would migrate off the shared-key model and what that isolation costs you versus what it buys you.

Unlock Full Question Bank

Get access to all 12 Distributed Systems Security and Trust interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.