Multi-Tenancy and Isolation Questions
Serving many tenants from shared infrastructure: tenancy models (silo, pool, bridge), data isolation, per-tenant data residency, noisy-neighbor mitigation, per-tenant limits, and security boundaries between tenants. Covers the cost, isolation, and blast-radius tradeoffs of shared versus dedicated resources, and business continuity: per-tenant backup, disaster recovery, and compliant tenant offboarding and deletion. The architecture layer specific to SaaS and platform products.
Design a multi-region replication and routing strategy where some tenants must keep data in one region and others can be globally replicated. Explain how to represent tenant residency policy in metadata, how to route reads and writes, how to replicate while honoring residency, and enforcement points to prevent accidental cross-border writes.
Sample Answer
Direct answer
I would make residency a versioned policy record per tenant in a global tenant directory: a home region that alone accepts writes, and an explicit set of regions where copies may exist. Pinned tenants have a set of one; globally replicated tenants have several. Every component that moves data (the router, each regional service, the replication pipeline, backup jobs) reads that record and makes the same decision from it, so writes always go to the home region, reads go to the nearest permitted copy, and replication only sends a tenant's changes to regions in its allowed set. The policy is enforced where data lands, not only where requests start, because routing bugs happen.
Representing residency policy as metadata
| Field | Meaning | Example |
|---|---|---|
tenant_id | Stable identifier | acme-de |
home_region | The single region that accepts writes for this tenant (this single-writer design, only one region may ever accept writes for a given tenant, avoids conflicting updates across regions) | eu-central-1 |
allowed_regions | Every region a copy (replica, cache, backup, search index) may live in; always includes the home region | {eu-central-1, eu-west-1} |
policy_version | Incremented on every change, stamped on replicated events | 3 |
data_classes (optional) | Finer rules per kind of data, for example "billing records may be global, documents may not" | {documents: pinned, usage_metrics: global} |
Design choices:
- Allowed regions are an explicit allow-list, not a "global: true" flag. "Global" quietly grows when you open a new region; an allow-list forces a decision.
- The directory itself contains no customer data, only these fields, so it can be replicated to every region and cached in every router and service (with a short cache lifetime and invalidation on version change).
- Changes are workflows, not edits. Narrowing
allowed_regionsmeans deleting copies in the removed regions and proving it; widening it means seeding replicas. The version number lets every component detect it is working from a stale policy.
Routing reads and writes
- Writes: always to
home_region, whatever region the user is in. A user in Singapore writing for an EU-pinned tenant pays the round trip to Frankfurt. - Reads: to the caller's own region if it is in
allowed_regions, otherwise to the home region. Reads from replicas can be slightly stale (asynchronous replication lag), so read-your-own-writes paths (the page a user sees right after saving) should read from home or wait for the replica to catch up to the write's position (a position is a marker, typically a sequence number or log offset, for one specific point in the ongoing stream of database changes; "catch up to the write's position" means the replica has applied every change through and including that particular write, not just that some time has passed).
Replicating while honouring residency
- Replicate at the tenant level, not the database level. A regional database holds many tenants with different policies, so whole-database replication to another region would copy pinned tenants too. Use change-data-capture (CDC: streaming each committed row change from the database log) filtered by tenant, or keep pinned and global tenants in separate databases so that database-level replication only ever carries global tenants.
- The replication consumer in the target region (also called a replication sink: the destination that receives and applies a stream of replicated changes) checks the policy again before applying a change, and stamps the policy version on each one. A change stamped with a stale policy version is refused outright, not reinterpreted against whatever the policy is now: reinterpreting an old decision under new rules risks applying a change whose other, now-outdated assumptions (which data class it belongs to, for example) no longer hold. A refused change is not lost: it stays in the source's replication log exactly like any change the sink has not yet acknowledged, so it is redelivered on the next replication pass and checked again, this time against whichever policy is current by then. Checking at the sink matters because it is the last point before data lands.
- Backups follow the same allow-list: a backup copy rule is generated from
allowed_regions, never hand-written.
Worked example (executed)
A minimal model of the directory, the router and the two enforcement checks (standard library only):
from dataclasses import dataclass
@dataclass(frozen=True)
class ResidencyPolicy:
tenant_id: str
home_region: str # the single region that accepts writes
allowed_regions: frozenset # every region a copy may live in
policy_version: int
DIRECTORY = {
"acme-de": ResidencyPolicy("acme-de", "eu-central-1", frozenset({"eu-central-1", "eu-west-1"}), 3),
"kiwi-au": ResidencyPolicy("kiwi-au", "ap-southeast-2", frozenset({"ap-southeast-2"}), 1),
"globex": ResidencyPolicy("globex", "us-east-1",
frozenset({"us-east-1", "eu-west-1", "ap-southeast-2"}), 7),
}
class ResidencyViolation(Exception):
pass
def route(tenant_id, op, caller_region):
p = DIRECTORY[tenant_id]
if op == "write":
return p.home_region
# reads: nearest permitted copy, else the home region
return caller_region if caller_region in p.allowed_regions else p.home_region
def check_local(tenant_id, op, this_region):
"""Runs inside every regional service before it touches storage."""
p = DIRECTORY[tenant_id]
ok = this_region == p.home_region if op == "write" else this_region in p.allowed_regions
if not ok:
raise ResidencyViolation(f"{op} for {tenant_id} refused in {this_region}")
def replicate_allowed(tenant_id, target_region, record_policy_version):
p = DIRECTORY[tenant_id]
return target_region in p.allowed_regions and record_policy_version == p.policy_version
for tenant, op, caller in [("acme-de", "read", "eu-west-1"), ("acme-de", "read", "us-east-1"),
("acme-de", "write", "eu-west-1"), ("kiwi-au", "read", "eu-west-1"),
("globex", "read", "ap-southeast-2"), ("globex", "write", "ap-southeast-2")]:
print(f"route {tenant:8} {op:5} from {caller:15} -> {route(tenant, op, caller)}")
for tenant, op, region in [("kiwi-au", "write", "ap-southeast-2"), ("kiwi-au", "write", "eu-west-1"),
("acme-de", "read", "eu-west-1")]:
try:
check_local(tenant, op, region)
print(f"check {tenant} {op} in {region}: allowed")
except ResidencyViolation as e:
print(f"check: REFUSED ({e})")
print("replicate kiwi-au -> eu-west-1:", replicate_allowed("kiwi-au", "eu-west-1", 1))
print("replicate globex -> eu-west-1:", replicate_allowed("globex", "eu-west-1", 7))
print("replicate globex -> eu-west-1 (stale v6 event):", replicate_allowed("globex", "eu-west-1", 6))
Output:
route acme-de read from eu-west-1 -> eu-west-1
route acme-de read from us-east-1 -> eu-central-1
route acme-de write from eu-west-1 -> eu-central-1
route kiwi-au read from eu-west-1 -> ap-southeast-2
route globex read from ap-southeast-2 -> ap-southeast-2
route globex write from ap-southeast-2 -> us-east-1
check kiwi-au write in ap-southeast-2: allowed
check: REFUSED (write for kiwi-au refused in eu-west-1)
check acme-de read in eu-west-1: allowed
replicate kiwi-au -> eu-west-1: False
replicate globex -> eu-west-1: True
replicate globex -> eu-west-1 (stale v6 event): False
Reading it:
acme-demay be read from Ireland (in its allowed set) but a read fromus-east-1is sent to Frankfurt, and its writes always go to Frankfurt.kiwi-auis pinned to Sydney: a European user's read is routed to Sydney, and if a buggy service ineu-west-1tries to write for it, the local check refuses before storage is touched.globexis globally replicated: reads are local in Sydney, writes still go to its home inus-east-1.- Replication to
eu-west-1is refused for the pinned tenant, allowed for the global one, and refused for a change stamped with a stale policy version.
Enforcement points against accidental cross-border writes
- Router: computes the destination from the policy (above).
- Regional service middleware:
check_localbefore any storage call, so a mis-routed or internally forwarded request fails loudly. - Replication sink: re-checks the allow-list and policy version before applying, refusing (not reinterpreting) a stale-version change as described above.
- Storage layer: per-region credentials and per-region encryption keys, so a service in one region cannot write to, or decrypt, another region's stores even if the code tries.
- Pipelines and backups: generated from the same policy record; a CI (continuous integration) check fails any hand-written cross-region rule.
- Detection: a periodic job lists tenant IDs present in each region's stores and alerts on any tenant found outside its allowed set.
Trade-offs and pitfalls
- Write latency for remote users of pinned tenants is the price of residency; there is no way around it other than moving the home region.
- Single-writer per tenant keeps replication conflict-free. Multi-region writes for global tenants are possible but bring conflict resolution (rules for deciding which value wins when two regions each accept a write to the same record at nearly the same time); I would only add them for a proven latency need.
- Pitfalls: a boolean "global" flag instead of an allow-list; replicating whole databases that mix pinned and global tenants; checking policy only in the router; caching the policy forever so a narrowed policy keeps replicating to a removed region; and forgetting derived copies (search indexes, caches, analytics extracts), which are copies too.
Design a secure multi-tenant microservice platform where tenant isolation, data protection, and noisy-neighbor mitigation matter. Compare isolation options (namespaces, containers, VMs), per-tenant quotas and throttling, encryption choices, monitoring requirements, and the operational trade-offs between cost and isolation strength.
Sample Answer
Direct answer
I would treat isolation as a ladder and put each tenant on the lowest rung its threat allows. If tenants only use our code and send us data, per-tenant Kubernetes namespaces with quotas, network policies and per-tenant encryption keys are enough and cost little. If tenants run their own code on our platform, containers alone are not a security boundary (they share the host's kernel), so I would move them into a sandboxed runtime or small virtual machines. Noisy-neighbour protection is a separate problem from security: it is solved with per-tenant quotas and throttling at every shared resource, and it is proved with per-tenant monitoring.
A few terms used throughout
The kernel is the core program that mediates every process's access to CPU, memory, files and network, by handling system calls (the requests a running program makes to it, like "open this file" or "send this data"); containers on the same machine all share one kernel. A container escape, sometimes via a kernel exploit (a bug in the kernel itself), is when code inside a container breaks out of its boundary and reaches the host, and from there potentially every OTHER tenant's containers on that same machine, not just its own. A hypervisor is the layer that lets one physical machine run several virtual machines, each with its own emulated hardware, which is why a VM-level boundary survives a kernel exploit that a container-level one does not. Attack surface is the total set of ways in an attacker could exploit, so a smaller attack surface means fewer such ways in. The control plane is the part of a cluster that manages it (schedules workloads, stores configuration) rather than running the tenant workloads themselves.
First, name the threat
- Accident and noise: a tenant's workload uses too much CPU, memory or I/O, or a bug in our code crosses tenants. Our code is trusted; tenants are not attackers.
- Malicious or untrusted tenant code: a tenant uploads a plugin, script or container. Now a kernel exploit in that code is a cross-tenant breach.
The first threat is mostly about fairness and correctness; the second is about hard security boundaries. Mixing them up leads either to overspending (VMs for everyone) or to a real breach (containers for untrusted code).
Isolation options compared
One ambiguity to clear up: "namespaces" can mean Linux namespaces (the kernel feature that gives a process its own view of processes, network and filesystem, which is what a container is built from) or Kubernetes namespaces (a named partition inside one cluster for organising objects, quotas and permissions). Below, "namespace" means the Kubernetes kind.
| Option | What is shared | Boundary strength | Cost and density | Good for |
|---|---|---|---|---|
| Kubernetes namespace per tenant (pods from many tenants on the same nodes) | Nodes, kernel, cluster control plane | Policy-level: role-based access control (RBAC), network policies, quotas. A kernel or container escape crosses tenants | Highest density, lowest cost | Trusted code, noise and accident isolation |
| Containers on dedicated nodes per tenant | Cluster control plane only | A container escape reaches only that tenant's nodes | Lower density: idle node capacity per tenant | Large tenants, strong noise isolation |
| Sandboxed containers (gVisor, which puts a second, unprivileged kernel-like layer between the container and the real kernel, so even code that finds a way to misbehave is intercepted by that layer instead of reaching the host's actual kernel) or microVMs (Firecracker, Kata Containers: a tiny virtual machine per pod) | Host hardware, hypervisor | Much smaller attack surface than a shared kernel | Some per-pod overhead in memory and startup | Untrusted tenant code at scale |
| Full VMs or separate clusters per tenant | Nothing but hardware or nothing at all | Strongest | Highest cost, most to operate | Regulated or contractual isolation |
Per-tenant quotas and throttling
Every shared resource needs its own limit, because a noisy tenant will find the one you forgot.
- Admission (API edge): a per-tenant token bucket: each tenant earns request tokens at a fixed rate up to a burst size; with no token, the request gets HTTP 429 and a retry hint. Plus a per-tenant cap on concurrent in-flight requests, which catches slow expensive requests that a rate limit misses.
- Compute: a
ResourceQuotaper namespace (caps total CPU, memory and object counts a tenant can request) and aLimitRange(default and maximum per container, so nothing runs unbounded). Priority classes (labels that tell the scheduler which workloads to keep and which to kill first when a node runs short on resources) decide who is evicted first under pressure. - Data stores: per-tenant connection limits in the connection pooler (a component that shares a small, fixed number of real database connections across many requests, since the database itself can only hold so many open at once), statement timeouts, and per-tenant partitions or weighted fair scheduling (giving each tenant a guaranteed proportional share of throughput, instead of pure first-come-first-served) on shared queues.
- Network: network policies default-deny (block all traffic unless a rule explicitly allows it) traffic between tenant namespaces; egress limits stop one tenant saturating shared bandwidth.
Encryption choices
- In transit: TLS at the edge, and mTLS (mutual TLS: both sides present certificates, so each service proves its identity) between services, typically from a service mesh (an infrastructure layer that runs alongside every service to manage service-to-service traffic, including issuing and rotating the certificates mTLS needs, without changing application code). mTLS also gives every call an authenticated workload identity you can check against the tenant it claims to act for.
- At rest, per tenant: envelope encryption. Each tenant has its own key-encryption key in a key management service (KMS); data is encrypted with data keys that are themselves encrypted by the tenant's key. Benefits: a leaked backup is useless without that tenant's key, you can rotate one tenant's key, and deleting a tenant's key makes all its data unreadable (crypto-shredding), which helps with deletion requests.
- Customer-managed keys (the tenant holds the key in its own account and can revoke it) as a premium option for regulated tenants. The trade-off: if they revoke it, their service stops, and your support team must understand that.
Monitoring requirements
You cannot manage noisy neighbours you cannot see. Every metric, log line and trace carries tenant_id, and you alert on per-tenant saturation (share of quota used, throttled request rate, p99 latency by tenant: p99 is the 99th percentile, the latency 99% of requests beat).
The cardinality trap, with numbers. A request-latency histogram with 12 buckets exports 14 time series (12 buckets plus a sum and a count). Labelled by 20 endpoints and 10,000 tenants:
14×20×10,000=2,800,000 time seriesThat is from one metric. The fix: keep full-detail metrics per endpoint without the tenant label, emit per-tenant metrics only for a few coarse signals (requests, errors, throttles, CPU seconds), and put per-tenant detail in logs or traces, which are sampled and queried on demand.
Worked example: choosing rungs for a real platform
A workflow-automation platform: 3,000 tenants, 2,950 using only built-in steps, 50 enterprise tenants, and a new feature letting tenants run custom JavaScript steps.
- Built-in steps (trusted code): shared nodes, namespace per tenant, quotas, default-deny network policy, per-tenant KMS key.
- Custom JavaScript steps (untrusted code): run in a separate node pool using a sandboxed runtime or microVMs, with no network access except a proxy that enforces the tenant's allowlist, and hard CPU-time and memory caps per execution.
- 50 enterprise tenants: dedicated node pools for predictable performance and an easy story for their security reviews.
The design choice that matters: the untrusted-code pool is separated by what runs, not by who the customer is. A small tenant's custom script is as dangerous as a large one's.
Operational trade-offs between cost and isolation strength
- Each rung up the ladder lowers density: dedicated nodes waste the capacity a tenant does not use, and a cluster per tenant multiplies control-plane cost and upgrade work.
- Stronger isolation also slows operations: patching 3,000 VMs is harder than patching 40 shared node pools.
- The cheapest real improvement is usually not a stronger rung but a missing quota; most noisy-neighbour incidents come from an unbounded shared resource (connections, a queue, a cache), not from weak compute isolation.
Pitfalls
- Treating Kubernetes namespaces as a security boundary for untrusted code.
- Rate-limiting requests but not concurrency, so one tenant's slow reports exhaust a worker pool.
- Per-tenant encryption keys with a shared application role that can decrypt everything anyway. The key boundary only holds if key access is scoped per tenant too.
- What would change my recommendation: if every tenant ran custom code (a hosting or compute platform), sandboxing becomes the default rung for everyone, not an exception.
Explain encryption at rest and encryption in transit in the context of multi-tenant services. Describe a high-level design for per-tenant encryption keys using a cloud KMS: key isolation, envelope encryption for blobs, key access policies, and operational considerations for key rotation and emergency key revocation.
Sample Answer
Direct answer
Encryption in transit protects data while it moves over a network (TLS, Transport Layer Security, the standard protocol for encrypting a network connection, between clients and the service, and mutual TLS, where both sides present certificates, between internal services). Encryption at rest protects data stored on disks, databases, object storage and backups. In a multi-tenant service both are baseline, but neither separates tenants from each other: a disk encrypted with one platform key is readable by any process allowed to read the disk. Tenant isolation through encryption comes from per-tenant keys held in a cloud KMS (key management service), applied with envelope encryption (encrypting each object with its own one-time data key, then encrypting that small data key, not the object, with the tenant's KMS key; the mechanics are below), with key policies that only let the right tenant's requests use the right key.
Three tiers of key granularity
| Model | How it works | Isolation | Cost and effort |
|---|---|---|---|
| Single platform key | One key encrypts everything (typical default disk or storage encryption) | Protects against stolen disks, not against a bug that reads the wrong tenant's rows | Near zero |
| Per-tenant key | Each tenant has its own KMS key; all of that tenant's data keys are wrapped by it | A leaked data key exposes one object; a disabled tenant key makes that tenant's data unreadable without touching others | One KMS key per tenant (AWS KMS lists 1 US dollar per key per month) |
| Per-entity key | Each record, file or user inside a tenant has its own data key | Finest-grained deletion: crypto-shredding means destroying the key itself, so every copy of the data it protects, including backups, becomes permanently unreadable without touching the data at all ("crypto-shred one customer's record") | Many more keys to store and manage |
The common answer: a KMS key per tenant, plus a fresh data key per object or per file under it. Per-entity keys are added only where you need to delete or revoke individual records cryptographically.
Envelope encryption for blobs
Calling KMS to encrypt every byte would be slow and costly, and KMS limits how much data one call can take. Instead:
- Ask KMS for a new data key (DEK, data encryption key) under the tenant's key. KMS returns it twice: in plaintext and encrypted ("wrapped") by the tenant's key-encryption key (KEK), which never leaves KMS.
- Encrypt the blob locally with the plaintext DEK (AES-256-GCM, a standard, widely used symmetric encryption algorithm).
- Store the ciphertext together with the wrapped DEK; discard the plaintext DEK from memory.
- To read: send the wrapped DEK to KMS, which checks that the caller may use that tenant's key, returns the plaintext DEK, and the service decrypts locally.
sequenceDiagram
participant App
participant KMS
participant Store
App->>KMS: Generate data key (tenant 42 key)
KMS-->>App: plaintext DEK + wrapped DEK
App->>App: encrypt blob with DEK
App->>Store: ciphertext + wrapped DEK
App->>KMS: Decrypt wrapped DEK (context tenant 42)
KMS-->>App: plaintext DEK (if policy allows)
Worked example: tenant 42 uploads a 50 MB file. One KMS call produces a DEK; the 50 MB is encrypted locally; the stored object is 50 MB of ciphertext plus a few hundred bytes of wrapped key. Reading it costs one KMS call, not one per block.
Key access policies
- The tenant's key policy allows only the service role (the identity the application itself, not a person, assumes to call KMS) that handles that tenant, never a wildcard.
- Bind each operation to the tenant with encryption context (non-secret key-value pairs such as
tenant_id=42that must match on decrypt; AWS KMS supports conditions on them withkms:EncryptionContext:condition keys). A request that tries to decrypt tenant 42's data key while claiming tenant 7 fails. - Every KMS call is logged (AWS CloudTrail on AWS), which gives an audit trail per tenant key.
Operational considerations
Rotation. Rotating a KMS key generates new key material for new encryptions while keeping the old material to decrypt old data, so nothing needs to be re-encrypted. The benefit is forward-looking, not retroactive: if one specific key version is later suspected of compromise, only the data keys wrapped under that version are exposed, not every data key the tenant has ever had, and every new object from that point wraps under the fresh version. Many compliance frameworks also require rotation on a fixed schedule regardless of suspicion. AWS documents default automatic rotation every 365 days, with a custom period available, and notes that rotation does not re-encrypt data or rotate the data keys already generated. If a data key itself is suspected compromised, rotation does not help: that data must be re-encrypted with a new DEK.
Emergency revocation. Disabling a tenant's key makes every later decrypt of that tenant's data fail immediately at KMS, which is the fastest containment action. Two caveats: data keys already decrypted and cached in application memory keep working until the cache expires (so keep cache lifetimes short, for example minutes), and scheduling key deletion (AWS enforces a waiting period of 7 to 30 days) is irreversible, so it is used for deliberate crypto-shredding at offboarding, not during an incident.
Cost and scale. 5,000 tenants at 1 dollar per key per month is 5,000 dollars a month before request charges, and each automatically rotated key adds 1 dollar per month for each of its first two rotations. Caching decrypted DEKs briefly keeps request volume and latency down.
Trade-offs and pitfalls
- Assuming "encrypted at rest" means tenants are isolated from each other; it does not.
- One key for all tenants means revoking a single tenant is impossible without re-encrypting everyone.
- Long-lived DEK caches quietly defeat emergency revocation.
- Per-entity keys everywhere add management cost with little benefit unless you need per-record deletion.
Explain a safe tenant onboarding and offboarding process for a SaaS product: steps required to provision isolation resources (databases, IAM), apply tenant-specific encryption keys, seed or migrate data, validate access controls, and securely erase and validate data deletion on offboarding to satisfy audits and compliance.
Sample Answer
Direct answer
Tenant onboarding and offboarding should be a single automated, idempotent workflow (safe to re-run after a failure without creating duplicates), not a runbook of manual steps. Onboarding provisions isolation resources from code, creates the tenant's encryption key, loads data with verified counts, and refuses to activate the tenant until automated cross-tenant access tests pass. Offboarding is the mirror image: export, hold for the contractual grace period, delete live data, destroy the tenant's key so any remaining copies (including backups) become unreadable, and produce a signed deletion record an auditor can check.
The lifecycle as a state machine
Model the tenant as a record with an explicit state, advanced by a workflow engine (for example AWS Step Functions, a managed service that runs multi-step workflows with retries, or Temporal, an open-source workflow engine with the same job):
REQUESTED → PROVISIONING → SEEDING → VALIDATING → ACTIVE → OFFBOARD_REQUESTED → EXPORTED → GRACE → DELETING → KEY_DESTROYED → CLOSED
Every step writes its result back to the tenant record, so a failure halfway leaves a known state that the workflow resumes from, and an auditor can read the history.
Onboarding
1. Provision isolation resources
The resources depend on the tenancy model (silo: dedicated resources per tenant; pool: shared resources with a tenant_id on every row; bridge: a mix, such as shared compute with a dedicated database):
| Resource | Pool tier | Silo tier |
|---|---|---|
| Database | Tenant row in the tenant registry; RLS (row-level security) policies already cover it | Dedicated database or cluster created from an infrastructure-as-code (IaC) template |
| Storage | Tenant prefix s3://data/tenants/{tenant_id}/ | Dedicated bucket |
| Identity (IAM, AWS Identity and Access Management) | Shared service roles; access restricted by a session tag (a label attached to a temporary credential when it is issued, here tenant_id) checked in policy conditions (ABAC, attribute-based access control: the IAM rule only allows an action when the tag matches the resource being accessed) | Dedicated role per tenant scoped to its own resources |
| Network | Shared | Optional dedicated VPC (Virtual Private Cloud) or private endpoint |
Everything is created from versioned templates, never by hand, so tenant 4,000 is configured identically to tenant 4.
2. Tenant-specific encryption keys
Create a per-tenant key (or, for the long tail, a per-tenant data key wrapped by a shared master key, encrypted under it, the standard "envelope encryption" pattern explained next) in a key management service such as AWS KMS (Key Management Service). Use envelope encryption: data is encrypted with a data key, and the data key is stored encrypted under the tenant's master key. Tag the key with the tenant ID, restrict its key policy so only that tenant's roles can use it, and record the key identifier in the tenant registry. This one decision is what makes offboarding provable later.
3. Seed or migrate data
- Seed default configuration, roles and an admin user from templates.
- Migrate customer data (from a previous system or a trial environment) through a staged pipeline: land in a quarantine area encrypted with the tenant's key, transform, load, then reconcile: row counts and checksums per table between source and target must match before the step succeeds.
- Every row written carries the tenant ID; the loader refuses a row whose tenant ID does not match the job's tenant.
4. Validate access controls (the activation gate)
The tenant is not ACTIVE until an automated suite passes:
- A token for the new tenant can read its own seeded objects.
- The same token cannot read a fixed canary tenant's (a permanent test tenant seeded with known data, used only to verify that isolation checks catch a leak) objects by ID, search, or export (expect 404, not 403, so existence is not revealed).
- A canary tenant's token cannot read the new tenant's objects.
- The IAM role for the tenant is denied access to another tenant's storage prefix and key (checked with a policy simulator, a tool that evaluates what an IAM policy would allow without actually performing the action, or a live attempt).
Offboarding
1. Export and grace period
Deliver a complete export in a documented format, encrypted, with a checksum manifest. Hold the data for the contractual grace period (commonly 30 days) in a suspended state: no logins, no processing, but restorable if the customer changes their mind.
2. Delete live data
- Silo: destroy the database, bucket and roles via IaC.
- Pool: delete by tenant ID in batches (to avoid long locks), across every store: primary database, search indexes, caches, queues, analytics copies, file storage. Driving this from a data inventory (a registry of every store that holds tenant data) is what prevents a forgotten store.
3. Destroy the key (crypto-shredding)
Backups and replicas are the hard part: you usually cannot surgically delete rows from a snapshot. If the tenant's data in those copies was encrypted with the tenant's key, scheduling the key for deletion renders them unreadable. AWS KMS enforces a waiting period of 7 to 30 days before a scheduled key is actually deleted, which doubles as a safety window. For pooled tenants, only what was actually encrypted with their data key becomes unreadable this way, typically file attachments and exports written through the application under the tenant's own storage prefix. Pooled database snapshots hold the bulk relational rows, which are not encrypted per tenant (they rely on RLS while live, not a per-tenant key) and simply age out under the backup retention policy; the deletion record states that date separately.
4. Validate deletion and produce evidence
- Query every store in the data inventory for the tenant ID and record a count of zero for each.
- Record the key deletion event from the provider's audit log (for AWS, CloudTrail, its API audit log).
- Generate a deletion certificate: tenant ID, stores checked with counts, key ID and deletion date, backup expiry date, workflow run ID, signed off by the system and a named owner.
Worked example
A pooled-tier customer with 2.1 million rows across 14 tables offboards on 1 March with a 30-day grace period:
- 1 March: export delivered, tenant suspended.
- 31 March: live deletion runs in 5,000-row batches, so 2,100,000/5,000=420 batches; the inventory check shows 0 rows in all 14 tables, 0 documents in search, 0 objects under the storage prefix.
- 31 March: tenant data key scheduled for deletion with the maximum 30-day window; deleted 30 April.
- Backups: daily snapshots with 35-day retention, so the last snapshot containing live data (taken 30 March) expires 4 May. This tenant's file attachments were encrypted with its data key; its relational rows in the shared pooled database were not, so the certificate's two dates track two different mechanisms. The certificate lists 4 May as the date after which no copy exists in any form (the last unencrypted database snapshot containing this tenant's rows has aged out of the 35-day retention window), and 30 April as the date after which no encrypted copy can be read (the data key is gone, so any surviving copy of an attachment becomes unreadable).
Trade-offs and pitfalls
- A dedicated master key per tenant vs per-tenant data keys under a shared master key. Dedicated keys give the cleanest audit story and per-tenant revocation but add key cost and API quota pressure at tens of thousands of tenants. Recommendation: dedicated keys for silo and regulated tenants; per-tenant data keys (stored in the tenant registry) for the pool.
- Pitfall: crypto-shredding a key that encrypted more than one tenant. Verify the key-to-tenant mapping is one-to-one before scheduling deletion.
- Pitfall: forgetting derived data. Analytics warehouses, ML (machine learning) feature stores and support tickets are where offboarded tenants live on. The data inventory must include them.
- Pitfall: activating before validating. Provisioning that "succeeded" with a misconfigured policy is how a new tenant starts life able to see an old one.
Architect a multi-tenant microservices platform serving 10,000 tenants worldwide that requires per-tenant isolation (compute and data), per-tenant cross-region failover, and cost transparency to tenants. Discuss tenant isolation models (logical vs physical), deployment strategies (shared cluster vs dedicated cluster), noise isolation, billing implications, and operational tooling needed.
Sample Answer
Direct answer
At 10,000 tenants worldwide I would build a cell-based architecture: the platform is stamped out as many independent, identical "cells" (each a Kubernetes cluster plus its own databases, hosting a few hundred tenants), placed in several regions. Small tenants share a cell with per-tenant namespaces, quotas and a per-tenant database; large tenants get a dedicated node pool (a set of machines reserved just for their workloads) or a dedicated cell. Every tenant has a home cell and a paired standby cell in another region, and a global tenant directory says which is active, so failing over one tenant means replicating its data and flipping one directory entry, not moving a whole region. Per-tenant metering runs from day one because cost transparency is a product feature here, not an internal report.
Requirements and the forces in tension
- Isolation of compute and data per tenant. Strong isolation pushes towards dedicated clusters; 10,000 dedicated clusters is operationally impossible.
- Per-tenant cross-region failover. A tenant (not just the whole platform) must be movable to another region, with a stated RPO (recovery point objective: how much recent data you may lose) and RTO (recovery time objective: how long until service is back).
- Cost transparency to tenants. Each tenant sees what it consumed and what it is billed, which means every CPU-second, byte and request needs a tenant label.
- Worldwide. Latency and data-residency rules decide the home region.
Isolation model: logical vs physical, and where each applies
| Tier | Share of tenants | Compute | Data | Why |
|---|---|---|---|---|
| Small | ~98% | Namespace per tenant in a shared cell, with CPU and memory quotas | Own logical database (or schema) on the cell's shared database servers, own encryption key | Cheap, fast onboarding, a real per-tenant boundary for data |
| Medium | ~1.8% | Dedicated node pool inside a shared cell (their pods run only on their nodes) | Own database server | Removes CPU and cache contention without a whole cluster to run |
| Large | ~0.2% | Dedicated cell | Dedicated database cluster | Contractual isolation, very large load |
A Kubernetes namespace is a named partition of one cluster; with quotas, network policies and per-tenant service accounts (a non-human identity a tenant's own workloads authenticate as, scoped so one tenant's pods cannot use another tenant's permissions) it isolates well against accidents but still shares the node's kernel, so untrusted tenant code would need a stronger sandbox. Physical isolation (dedicated nodes or clusters) is what you sell to tenants who need it.
Deployment strategy: shared cluster vs dedicated cluster
I would not run one giant shared cluster: a bad upgrade, an overloaded API server (the Kubernetes control-plane component every cluster-management request goes through) or a misconfigured network policy would hit all 10,000 tenants. Cells cap the blast radius (how many tenants one failure reaches).
Sizing, with the assumptions stated: 9,800 small and 180 medium tenants share cells, 20 large tenants get their own. At 250 tenants per shared cell:
⌈2509,800+180⌉=⌈39.92⌉=40 shared cellsSpread over 4 regions, that is 10 shared cells per region, so one bad cell touches at most 250 tenants (2.5% of the base).
flowchart TB
U[Tenant request] --> E[Global edge + tenant directory]
E -->|active| A[Home cell, region A]
E -.->|on failover| B[Standby cell, region B]
A -->|per-tenant async replication| B
A --> M[Metering pipeline]
B --> M
M --> BL[Per-tenant usage + bill]
Per-tenant cross-region failover
- Placement: the directory stores each tenant's home cell, standby cell and state (
ACTIVE,FAILING_OVER,ACTIVE_ON_STANDBY). - Data: each tenant's database replicates asynchronously to its standby cell. Asynchronous means the RPO is the replication lag (typically seconds); synchronous cross-region replication would add a cross-region round trip (tens of milliseconds between nearby regions, well over 100 ms between continents) to every write, which I would offer only as a premium tier.
- Compute: the standby cell keeps the tenant's deployment defined but scaled low; failover scales it up.
- Procedure: fence the old primary so it cannot accept a write once failover has started. Two ways to do that: revoke its write credentials at the database, or bump the tenant's epoch, a version number stored in the tenant directory that increments on every failover. Every write is required to carry the current epoch, and the datastore or proxy checks that number against the directory's latest value before accepting the write; once the epoch is bumped, any write still in flight from the old primary carries the now-stale epoch and is rejected, even though the old primary itself never learns the failover happened. Then promote the replica (make the standby's database copy the new primary and start accepting writes), flip the directory entry, and let DNS or the edge route pick it up. Because it is per tenant, the same mechanism doubles as a region-move tool for residency changes.
Capacity arithmetic. If a whole region fails and its tenants spread over the other 3 regions, each remaining region takes on one third of a region's load, reaching 4/3 of its normal load. For that to fit, normal utilisation must be at most 3/4 = 75% of capacity per region.
Replication bandwidth. If the average tenant writes 2 GB a day: 10,000 x 2 GB = 20 TB a day, and 20 x 10^12 bytes / 86,400 s = about 231 MB/s, or about 1.85 Gbit/s of cross-region traffic in total, which is also a line item on the bill (inter-region transfer is charged per GB on the major clouds).
Noise isolation
- Admission: per-tenant rate limits and concurrency limits at the edge, sized by plan.
- Compute: namespace resource quotas (caps on total CPU and memory a namespace may request) and default limits for every pod, so no tenant can schedule unbounded work.
- Data: per-tenant connection caps and statement timeouts; heavy tenants moved to their own database server when their share of a server crosses a threshold.
- Shared queues: per-tenant partitions or weighted fair queuing (giving each tenant a guaranteed proportional share of a shared queue's throughput, instead of pure first-come-first-served) so a tenant's backlog does not delay others.
Billing implications and cost transparency
Tenants see showback (a usage and cost breakdown) on a dashboard, and billing on the invoice. To make these match:
- Label every workload with
tenant_idat deploy time; the metering pipeline rejects unlabelled resources. - Measure per tenant: CPU and memory reserved versus used per pod, storage bytes per database, requests and egress bytes at the gateway, replication bytes.
- Charge dedicated resources directly; split shared ones (control plane, cell overhead, idle capacity) by a published rule, for example proportional to reserved CPU.
- Reconcile monthly: the sum of all tenant charges plus the platform's own share must equal the cloud bill.
Two billing consequences of this design: the standby is not free (a warm standby, kept running at reduced scale so it can take over quickly, as described above, still has storage and replication traffic, and that belongs on the tenant's bill, which is why cross-region failover is usually a paid tier), and dedicated tiers carry a visible minimum charge because their idle capacity cannot be shared.
Operational tooling needed
- Tenant directory and control plane (the small set of services that make decisions about tenants, placement, routing, failover, as opposed to the data plane, which actually carries each tenant's traffic) for onboarding, placement, tier changes and failover.
- Cell factory: infrastructure-as-code to stamp identical cells, plus a staged rollout that upgrades one cell at a time.
- Per-tenant observability: dashboards and alerts filtered by tenant, with care for label cardinality (the number of distinct time series; 10,000 tenant labels on every metric is expensive, so keep per-tenant metrics coarse).
- Failover runbooks and game days (scheduled practice failure drills, on purpose, not waiting for a real incident): fail over a test tenant every week, measure its RTO and RPO, and publish them.
- Tenant migration tooling: move a tenant between cells without downtime (copy, catch up, fence, flip).
Trade-offs and pitfalls
- Cells multiply fixed cost. 40 control planes (one management layer per cell), 40 sets of monitoring. The alternative (one huge cluster) is cheaper and fails for everyone at once. I would accept the overhead.
- Global dependencies defeat cells. If the tenant directory or identity service lives in one region, a regional outage still takes everyone down. Replicate the directory to every region and let cells run on a cached copy.
- Failover without fencing causes split brain (two copies both accepting writes and diverging). Always fence first.
- What would change the design: if tenants were mostly large enterprises, I would make "dedicated cell" the default and focus on cell automation instead of packing density.
Unlock Full Question Bank
Get access to all 16 Multi-Tenancy and Isolation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.