Multi-Region and Geo-Distributed Systems Questions
Running a system across regions and continents: multi-region replication, data residency and sovereignty, geo-routing and CDN edge distribution, cross-region consistency and quorum placement, and conflict resolution when two regions accept writes. Covers regional failover and split-brain prevention, recovery objectives (RTO/RPO), region-by-region rollout and blast-radius containment, and the latency, cost, and consistency tradeoffs of going global. Global distribution strategy across the service and data tiers.
Design cross-border encryption key management for an international SaaS product: customers in EU, US, and APAC. Requirements: customer data must remain within region, keys must rotate regularly, revocation must be auditable, and customers may request key deletion. Propose KMS topology, HSM usage, key replication or isolation strategy, access control, and legal implications.
Sample Answer
Direct answer
Run three independent, region-isolated key management service (KMS) instances, one each for the EU, US, and APAC customer bases, with no cross-region replication of key material by default, since replicating the keys would undercut the same region-lock guarantee being made about the data they protect.
Framework
- KMS topology: one KMS instance per jurisdiction, each managing keys only for that jurisdiction's customer data. A service in one region can request encryption or decryption only against that region's KMS instance, enforced by the same identity and network scoping used for the data itself.
- HSM usage: back each regional KMS with hardware security module (HSM) backed key storage so raw key material is never extractable, even by the provider's own operators, which is typically the strongest, most auditable claim available about key protection.
- Rotation: automate rotation on a fixed schedule, for example every 90 days, using envelope encryption, meaning data is encrypted with a data key that is itself encrypted by the rotating master key, so rotating the master key only requires re-wrapping the much smaller data keys, not re-encrypting the underlying data.
- Auditable revocation: every key operation, use, rotation, revocation, writes an immutable, append-only audit log entry, and revocation is tested to confirm it actually blocks decryption immediately rather than only preventing new encryption operations.
- Customer-requested deletion, crypto-shredding: honor a deletion request by destroying that
customer's specific data encryption key rather than hunting down and deleting every copy of their
data. Once the key is gone, and gone from every recoverable key backup, the ciphertext is
permanently unreadable, satisfying "delete my data" far faster than a data-hunting process, though
it must be paired with actually purging the now-useless ciphertext on a normal retention schedule
afterward. - Emergency key recovery, and what it must deliberately exclude: for the rare case a key needs
recovering after an operational error, not a customer deletion, be precise about what recovery can
mean, because the non-extractability claim above forbids the obvious answer. No operator reads key
material out of the hardware security module; what exists instead is a restore of an encrypted HSM
backup into another HSM, released under dual control, a fixed minimum number of authorized people
out of a larger set independently approving, and logged with the same rigor as any other key
operation. That backup is also the thing that can quietly falsify the deletion story: if a
customer's data encryption key is destroyed from the live HSM while a copy still sits inside a
restorable backup, the ciphertext is not permanently unreadable, it is merely inconvenient to read,
and a regulator asking "can anyone in your company still decrypt this" gets a different answer than
the one you gave the customer. Keys eligible for crypto-shredding therefore have to be excluded from
the recoverable backup set by design, and demonstrating that exclusion, not just describing it, is
what an auditor should be asked to check. - Authorized non-local access: for a legitimate cross-border need, for example a valid legal hold requiring a non-EU legal team to access EU data, route it through an explicit, time-boxed exception process, legal sign-off plus dual-control approval, granting narrowly-scoped, logged access rather than a standing cross-region credential. The exception itself becomes part of the audit trail, distinct from and far rarer than routine operations.
Worked example
An EU customer requests deletion. The system destroys their specific data encryption key, immediately rendering their ciphertext unreadable, and the underlying storage rows are purged on the next scheduled cleanup within the contractual retention window. Separately, a legal hold for a US customer's data is routed through the exception process, requiring both legal sign-off and two independent engineers' dual-control approval before a time-boxed, logged grant is issued, distinct from the automated systems handling normal traffic.
Trade-offs and pitfalls
Crypto-shredding satisfies "delete my data" quickly, but customers or regulators sometimes expect literal row deletion too, so it needs pairing with actual storage cleanup rather than treating key destruction as the entire answer. A break-glass process with weak dual control, where one person can approve it alone in practice even if policy says otherwise, is a common audit finding, so the dual-control rule needs to be a technical control, not just a documented expectation.
Design identity federation and authorization for a multi-tenant SaaS spanning regions with local regulatory constraints. Include token issuance models, central vs regional identity providers, cross-region token validation, key rotation, privacy considerations, and approaches to minimize authentication latency.
Sample Answer
Direct answer
Split identity into two planes. A small, globally-replicated control plane holds
tenant metadata (which region owns each tenant's user data, which identity provider (IdP)
config to use) and public signing keys. Each region runs its own data plane: the actual
user records, credentials, and multi-factor state, kept local to the region that satisfies
that tenant's regulatory residency requirement (an EU tenant's users live and authenticate
in an EU region, a healthcare tenant's users might need to stay in a specific country).
Tokens are short-lived, self-contained, and signed locally, so any region can verify them
without a network call back to the issuing region. That combination is what keeps both
compliance and latency intact at the same time.
How the pieces fit together
Central vs. regional identity provider (IdP). Don't run one global IdP that stores
every tenant's credentials in one place: that creates a single point of both latency and
regulatory failure (a login from Frankfurt round-tripping to us-east-1, or EU personal data
sitting on US soil). Instead, run a regional IdP per data-residency zone, using a standard identity protocol such as OpenID
Connect (OIDC) or SAML (Security Assertion Markup Language, an older but still common protocol that does
the same job), and a thin central directory that maps tenant_id
to "which region's IdP owns this identity." That directory is the one thing worth
replicating globally, because it is small (millions of rows of tenant_id -> region, not
user PII) and read-heavy: a global key-value store with fast eventually-consistent reads
(a global secondary index, or a CDN-fronted lookup service) is enough.
Keep that directory tenant-scoped, and resist the obvious extension of putting user_id in
it, because a user_id -> region table is a different object legally, not just a bigger one.
A user identifier the operator can re-attach to a person through its own regional identity
store is pseudonymised data, and pseudonymised data is still personal data under EU rules
rather than anonymous. Replicating that table everywhere therefore puts a copy of EU personal
data in every region you operate in, which is precisely the outcome the regional data plane
exists to prevent, and it would be the first thing an auditor pulled on. If a login flow
genuinely needs user-level routing, because the same email address can exist under two
tenants, get it without a global user index: carry a tenant hint in the login URL or
subdomain, or let the nearest regional IdP answer "not mine" and hand the flow off. Either
costs one extra hop on a rare login. A global user index costs a permanent, replicated copy
of who your users are. When a login request
lands at any edge, the gateway does one cheap lookup, then redirects the login flow to the
correct regional IdP.
Token issuance. The regional IdP that owns the user issues a signed token (an OAuth2
access token or OIDC ID token, typically a JWT: a compact, digitally signed JSON payload)
after authentication. The token carries claims (user id, tenant id, roles, expiry) but
deliberately minimal PII, since the token itself will travel to whichever region actually
serves the request, which may not be the user's home region.
Cross-region token validation. Because the token is self-contained and signed, a
service in any region can verify it locally: fetch the issuing region's public key from a
JWKS endpoint (JSON Web Key Set, a small published document of public verification keys),
cache it, and check the signature and expiry. No call back to the issuing region is needed
per request, which is what keeps steady-state latency regional. Only the public keys cross
regions, never the private signing key or the underlying credential store.
Key rotation. Each region owns and rotates its own signing key pair independently.
Publish the new public key to the shared JWKS well before the old key is retired (an
overlap window, commonly 24 to 72 hours), so tokens signed just before rotation still
validate everywhere. Retire the old key only after that window passes and after the
longest-lived token issued under it has expired. Never let one region's key compromise
force a rotation of every other region's keys: keys are per-region, isolated blast radius.
Privacy. Two separate residency concerns exist: where the user record lives (must
stay in-region, enforced by the data-plane split above) and what the token carries as it
crosses regions (must be minimal: an identifier and role claims, not name, email, or other
attributes a downstream region has no legal basis to hold). Treat token content
minimization as its own privacy control, separate from where the account record sits.
Minimizing authentication latency. Login (issuance) is rare and can tolerate one
cross-region hop to the home IdP. Token validation (checked on every API call) is the hot
path, and it is local: signature check against a cached public key, no network round trip.
Design so the expensive part happens once per session, not once per request. Be careful about
what that actually removes, though: it takes authentication out of the per-request cross-region
budget, it does not take data access out of it. Where the tenant's records live is a separate
placement decision, and for a residency-pinned tenant it is a decision you are not free to
make.
Worked example
A tenant based in Germany has users authenticating against the eu-west IdP. A user opens
the SaaS from a hotel in Singapore; the request lands at the nearest edge (ap-southeast).
The gateway looks up tenant_id -> eu-west in the global directory (a few milliseconds),
redirects the login to the eu-west IdP, and the user authenticates there. The eu-west
IdP issues a JWT signed with its current key. Every subsequent API call from that session is
authenticated wherever it lands: ap-southeast services fetch (and cache) eu-west's public key
from the shared JWKS once, then verify the token locally on every call, with no callback to
eu-west to authenticate anything. State the win precisely, because it is easy to overclaim. The
authentication hop is gone from the request path. The request is not therefore local: this tenant's
records are pinned to eu-west by the same residency rule that put its IdP there, so any call that
actually reads or writes German tenant data still crosses to eu-west and pays that round trip.
What ap-southeast can genuinely serve on its own is the work that touches no residency-bound data,
static assets, decisions derivable from the token's own claims, region-agnostic reference data,
anything already cached at the edge. The user's PII record never leaves eu-west; the token in
transit carries only user_id, tenant_id, and role claims.
Trade-offs and pitfalls
- The tenant-to-region directory becomes a hard dependency. If it is unavailable, no
region knows where to route a fresh login. Make it read-cacheable at the edge with a
generous time-to-live (TTL), since tenant-to-region mapping almost never changes. - Token replay across regions is a real risk once tokens are globally valid. Keep
expiry short (minutes, with refresh tokens handled by the home region only) and consider
binding tokens to a specific audience or region claim if a tenant's compliance posture
requires that a token issued for one jurisdiction cannot be replayed to serve data in
another. - When you present this design to executives or legal stakeholders, frame it as a
risk-reduction trade, not a pure engineering choice: "the regional data plane is what
lets us tell a regulator that EU user data never leaves the EU; the central directory is
a small, non-sensitive index that makes that possible without adding a network hop to
every request." That framing is usually what actually gets the design approved, because
it answers "what does this cost us in risk" before "how does it work." - Don't over-centralize "just for convenience": a single global IdP is simpler to operate
on day one, but it forecloses the regional data-residency story you will need the moment
a large regulated customer asks where their users' credentials are stored.
Describe cross-region key management and encryption options: a single global KMS, regional KMS with federation, customer-managed keys with HSM, and bring-your-own-key. Discuss rotation, access controls, cross-region replication (if allowed), performance impact, and compliance trade-offs.
Sample Answer
Direct answer
The four options trade off operational simplicity against who actually controls the key and how well the design survives a regional outage or a customer's own security requirement. Most real systems end up with regional key management plus federation rather than either extreme.
Comparison
| Option | Who controls the key | Cross-region behavior | Typical use |
|---|---|---|---|
| Single global key management service (KMS) | Provider, one logical service | Simple, but becomes a single dependency every region relies on | Small systems without residency constraints |
| Regional KMS with federation | Provider, per-region instances with a trust relationship | Independent key material per region; federation lets an authorized identity in one region request access in another under policy | Most residency-sensitive multi-region systems |
| Customer-managed keys with hardware security module (HSM) | Customer, keys generated and stored in dedicated hardware | Customer decides if and how keys replicate; provider never sees raw key material | Enterprise customers with their own compliance mandate |
| Bring-your-own-key (BYOK) | Customer, imports their own key material into the provider's KMS | No replication unless the customer explicitly imports the same key elsewhere | Customers who need to prove they, not the vendor, control the key |
Elaboration
- Rotation: regional KMS with federation supports centrally-scheduled rotation per region without touching customer processes. BYOK rotation is the customer's responsibility, so the system needs to treat "the customer forgot to rotate" or "the customer revoked the key" as expected states, not exceptions.
- Access controls, including cross-account access: in a federated or multi-account setup, a service in one account or region that needs a key in another does so via an explicit cross-account role assumption scoped to exactly that key and audited per use, never a blanket "trust this whole account" grant. This is what makes federation auditable instead of merely convenient.
- Cross-region replication of keys, if allowed: replicating key material itself increases blast radius, since a compromise in one region can then decrypt data protected in another, and can itself violate a residency rule about key material. Most designs prefer per-region independent keys with an access-federation layer over replicating the keys themselves.
- Performance impact: an encrypt or decrypt call to a local regional KMS typically costs single-digit milliseconds. A cross-region KMS call, needed when federation routes to another region's key, adds a full round trip, roughly 80-150ms depending on the regions, which matters if it sits on a hot request path rather than an infrequent operation.
- Compliance trade-offs, including BYOK: BYOK gives a customer the strongest evidentiary claim of control, useful for their own audits and for "we can make our data permanently unreadable by revoking our key" guarantees, but it also means the customer can break their own service by mismanaging the key, and typically shifts key-loss risk explicitly onto the customer by contract. Customer-managed HSM keys are similar but the vendor usually still operates the hardware.
Worked example
A customer requires proof they can revoke access to their data at will. The platform offers BYOK,
the customer imports their own key, and that key is used to wrap the data keys that actually encrypt
their records (envelope encryption: the imported key never touches the records themselves, it only
encrypts the short-lived keys that do). Revoking it means the platform can no longer unwrap any of
those data keys, so the ciphertext stops being readable. The honest version of the timing has two
parts, and it is worth stating both to the customer rather than promising instant: the key
management service starts refusing unwrap requests immediately, while any data key already unwrapped
and sitting in a service's memory keeps working until that cache entry expires. So the guarantee
bites after the data-key cache time-to-live, typically minutes, and that TTL, not the key rotation
schedule, is the number that belongs in the contract. It is also genuinely testable, which is the
real value to the customer's security team: revoke, wait out the cache window, confirm reads fail,
and keep the evidence. Compare that to a global-KMS design where "revoke access" depends on the
vendor's internal process instead of a customer-controlled action.
Trade-offs and pitfalls
Offering BYOK without a clear contractual statement of what happens if the customer loses their key, their data becomes unrecoverable by design, leads to a very difficult support conversation later. Replicating raw key material for convenience undermines the entire reason federation exists.
That is every published Multi-Region and Geo-Distributed Systems question for Security Architect so far. Browse the other topics in this category, or practice this one interactively.