Multi-Tenancy and Isolation Questions
Serving many tenants from shared infrastructure: tenancy models (silo, pool, bridge), data isolation, per-tenant data residency, noisy-neighbor mitigation, per-tenant limits, and security boundaries between tenants. Covers the cost, isolation, and blast-radius tradeoffs of shared versus dedicated resources, and business continuity: per-tenant backup, disaster recovery, and compliant tenant offboarding and deletion. The architecture layer specific to SaaS and platform products.
Implement (in Python) a tenant-aware partition key generator function that maps (tenant_id: str, message_id: str, timestamp: int) -> partition_key such that (1) events for the same tenant mostly map to the same set of partitions, (2) hot tenants are spread across more partitions, and (3) the function is deterministic. Show example outputs for tenant A and tenant B.
Sample Answer
Direct answer
Give every tenant a deterministic, ranked list of partitions (computed by hashing the tenant with each partition number), and let a tenant use the first k entries of its list, where k is 1 for normal tenants and larger for hot tenants. Within that set, pick the slot by hashing the message ID. The same inputs always give the same partition, a normal tenant stays on one partition, a hot tenant spreads over k, and raising k keeps the tenant's existing partitions because the ranking never changes. The timestamp selects which version of the "who is hot" configuration applies, so replaying old events routes them exactly as they were routed originally.
Terms used below
- Partition: one of a fixed number of independent append-only logs in a message system such as Kafka (a widely used distributed message queue). A consumer reads a partition in order, and order is only guaranteed within one partition.
- Partition key: the value that decides which partition a message goes to. Here the function returns the partition index directly (a producer can send to an explicit partition), which avoids a second hash scattering things.
- Hot tenant: a tenant producing far more traffic than others; if all of it lands on one partition, that partition's consumer falls behind (lag) while others sit idle.
- Deterministic: same inputs, same output, on every machine and every process restart.
- Rendezvous hashing (highest random weight): for a key, score every candidate bucket with
hash(key, bucket)and rank buckets by score. The ranking for a key never changes, so taking "the top k" and later "the top k+1" only ever adds buckets.
Approach
- Stable hash. Use
hashlib.blake2b, never Python's built-inhash(), which is randomly salted per process for strings. Two producers usinghash()would disagree. - Spread per tenant, versioned by time. A small config says how many partitions each hot tenant may use, with an
effective_fromtimestamp.spread_for(tenant, ts)picks the entry that was live at the event's timestamp. - Tenant's partition set. Rank all partitions by
hash(tenant, partition)and take the first k. Normal tenants (k = 1) get one partition; different tenants get different ones, spreading tenants evenly. - Slot within the set.
hash(tenant, message_id) % kpicks one of the k partitions. Hashing the message ID (not a random number) means a retried message goes to the same partition, which keeps duplicate detection simple.
Code
import hashlib
from collections import Counter
NUM_PARTITIONS = 32
# Spread config: how many partitions each hot tenant may use.
# Each entry is (effective_from_epoch_seconds, {tenant_id: spread}).
# Versioning by time keeps the function deterministic on replay:
# an event is always routed by the config that was live at its timestamp.
SPREAD_CONFIG = [
(0, {"tenant-A": 4}),
(1_700_000_000, {"tenant-A": 8}), # tenant-A got hotter; widened to 8
]
DEFAULT_SPREAD = 1
def _h(*parts: str) -> int:
"""Stable 64-bit hash. Never use built-in hash(): it is salted per process."""
data = "\x1f".join(parts).encode()
return int.from_bytes(hashlib.blake2b(data, digest_size=8).digest(), "big")
def spread_for(tenant_id: str, timestamp: int) -> int:
spread = DEFAULT_SPREAD
for effective_from, table in SPREAD_CONFIG:
if timestamp >= effective_from:
spread = table.get(tenant_id, DEFAULT_SPREAD)
return max(1, min(spread, NUM_PARTITIONS))
def tenant_partitions(tenant_id: str) -> list[int]:
"""Rendezvous ranking: every partition gets a score for this tenant,
highest first. A tenant with spread k uses the first k entries, so
widening k from 4 to 8 keeps the original 4 partitions."""
return sorted(range(NUM_PARTITIONS),
key=lambda p: _h(tenant_id, str(p)), reverse=True)
def partition_key(tenant_id: str, message_id: str, timestamp: int) -> int:
k = spread_for(tenant_id, timestamp)
candidates = tenant_partitions(tenant_id)[:k]
slot = _h(tenant_id, message_id) % k
return candidates[slot]
if __name__ == "__main__":
T_OLD, T_NEW = 1_690_000_000, 1_710_000_000
for tenant in ("tenant-A", "tenant-B"):
print(tenant, "allowed set (new config):",
sorted(tenant_partitions(tenant)[:spread_for(tenant, T_NEW)]))
for i in range(5):
mid = f"msg-{i}"
print(f" {mid}: old={partition_key(tenant, mid, T_OLD):2d}"
f" new={partition_key(tenant, mid, T_NEW):2d}")
# Load shape over 10,000 messages per tenant under the new config.
for tenant in ("tenant-A", "tenant-B"):
c = Counter(partition_key(tenant, f"m{i}", T_NEW) for i in range(10_000))
print(tenant, "partitions used:", len(c),
"max share: %.1f%%" % (100 * max(c.values()) / 10_000))
# Widening 4 -> 8 keeps the old set: old partitions are a subset of new.
old = set(tenant_partitions("tenant-A")[:spread_for("tenant-A", T_OLD)])
new = set(tenant_partitions("tenant-A")[:spread_for("tenant-A", T_NEW)])
print("old set:", sorted(old), "subset of new:", old <= new)
moved = sum(partition_key("tenant-A", f"m{i}", T_OLD) !=
partition_key("tenant-A", f"m{i}", T_NEW) for i in range(10_000))
print("messages whose partition changed at widening:", moved, "of 10000")
# Determinism: the same inputs always give the same answer.
print("deterministic:", all(
partition_key("tenant-A", f"m{i}", T_NEW) ==
partition_key("tenant-A", f"m{i}", T_NEW) for i in range(1000)))
Output (from running the file above; identical when re-run with different PYTHONHASHSEED values)
tenant-A allowed set (new config): [0, 7, 14, 16, 22, 23, 25, 29]
msg-0: old= 0 new=16
msg-1: old= 7 new= 7
msg-2: old=29 new=29
msg-3: old=25 new=25
msg-4: old=25 new=25
tenant-B allowed set (new config): [29]
msg-0: old=29 new=29
msg-1: old=29 new=29
msg-2: old=29 new=29
msg-3: old=29 new=29
msg-4: old=29 new=29
tenant-A partitions used: 8 max share: 13.1%
tenant-B partitions used: 1 max share: 100.0%
old set: [0, 7, 25, 29] subset of new: True
messages whose partition changed at widening: 4957 of 10000
deterministic: True
Key points, read from the output
- Requirement 1 (locality): tenant-B, a normal tenant, puts every message on partition 29. Tenant-A's traffic is confined to 8 of the 32 partitions, not all of them, so its consumers and caches deal with a known subset.
- Requirement 2 (hot spread): tenant-A uses all 8 partitions in its set, and the busiest one takes 13.1% of its messages, close to the ideal 1/8 = 12.5%.
- Requirement 3 (determinism): repeated calls agree, and the whole output is byte-identical across processes with different hash salts.
- Widening is incremental: tenant-A's old set
[0, 7, 25, 29]is contained in the new set. But 4,957 of 10,000 message IDs changed partition, about half, which is what the arithmetic predicts. The old slot ish % 4and the new slot ish % 8for the same hash h, and since both index the same ranked list, a message stays put exactly whenh % 8is below 4, which is half of all hash values. - A collision to notice: tenant-B's single partition (29) is also in tenant-A's set. Small tenants hashed onto a hot tenant's partition share its lag. Mitigation below.
Complexity
- Each call ranks all P partitions: O(P log P) time, O(P) memory. With P = 32 that is trivial, but on a hot path you would cache
tenant_partitions(tenant)(one list per tenant, invalidated only when P changes), making a call O(1) hashing work. - Config lookup is O(number of config versions); keep it small or index it by time.
Edge cases
- Ordering: a hot tenant's events are no longer totally ordered, because they sit on k partitions. If per-entity order matters (all events for one order or one user in sequence), hash the entity ID instead of the message ID for the slot, so each entity stays on one partition.
- The widening moment: at the config's
effective_fromtime, roughly half of a tenant's keys move partition (measured above). Events for one entity near the boundary can be consumed out of order. Schedule changes, and let consumers tolerate a short reorder window or drain before switching. - Changing P (the partition count) reshuffles every tenant's ranking. Treat it as a migration, not a config tweak.
- Unknown tenant or spread larger than P: clamped to the range 1 to P by
spread_for. - Small tenants sharing a hot tenant's partition: either reserve a block of partitions for hot tenants and rank small tenants only over the rest, or have consumers process per tenant fairly within a partition.
- Empty or very long IDs: hashing handles both; the
\x1fseparator stops("ab","c")and("a","bc")from hashing the same.
Trade-offs
Using k = 1 for everyone gives the best locality and strict per-tenant order but lets one tenant saturate a partition. Plain hash(message_id) % P spreads load perfectly but destroys tenant locality and makes per-tenant lag invisible. The ranked-set design sits between them and lets you move along that spectrum per tenant, with a config change rather than a code change.
Design a multi-tenant API gateway and extensibility model for a SaaS platform that must support tenant-specific rate limits, custom authentication, and plugin-based request transformations. Describe the components, a tenant-config data model, your scaling strategy, and how you'd evolve shared API contracts safely without breaking individual tenants.
Sample Answer
Direct answer
I would build a gateway data plane of stateless nodes that serves every tenant, driven by a control plane that stores each tenant's configuration and pushes versioned snapshots to the nodes. Each request is resolved to a tenant first, then passes a fixed pipeline: tenant-specific authentication, rate limiting, then the tenant's ordered list of transformation plugins, then the backend. Custom authentication is configuration wherever possible (which identity provider, which key format), and custom transformations run as sandboxed WebAssembly plugins with strict CPU, memory and time budgets. Shared API contracts evolve by additive changes, explicit versions tenants can pin, and a pre-release check that replays every tenant's plugins against the new contract.
Terms used below
- API gateway: the front door that every API request passes through; it handles cross-cutting concerns (auth, limits, routing) so backends do not.
- Data plane / control plane: the data plane handles live traffic; the control plane manages configuration and is off the request path.
- Rate limit / token bucket: a limit like "100 requests per second, bursts up to 200". A token bucket refills at the rate and each request spends a token.
- Plugin: tenant-supplied logic that modifies a request or response (rename a field, add a header, map a legacy format).
- WebAssembly (Wasm): a portable bytecode that runs in a sandbox with no access to the host unless explicitly granted, which makes it suitable for running untrusted code.
- OIDC (OpenID Connect) and JWT (JSON Web Token): a standard login protocol and the signed token it issues; the gateway verifies the signature with the issuer's public keys (a JWKS, JSON Web Key Set).
- Feature flag: a runtime switch turning behaviour on for specific tenants without a deploy.
Components
flowchart LR
REQ[Request] --> TR[Tenant resolver]
TR --> AU[Auth stage]
AU --> RL[Rate limiter]
RL --> PL[Plugin chain<br/>Wasm sandbox]
PL --> BE[Backend services]
CP[Control plane<br/>config store + validator] -->|versioned snapshots| TR
RL <--> RC[(Shared counters)]
- Tenant resolver: maps hostname (
acme.api.example.com), API key prefix, or token claim (a field inside a verified token, such as a JWT'stenant_idfield, set by the issuer and not editable by the caller) to a tenant ID. Unknown tenant: reject before any other work. - Auth stage: per-tenant method from config: OIDC with the tenant's issuer, API keys (stored hashed), or mutual TLS (both sides present certificates). The verified identity, not a header, becomes the tenant context passed downstream.
- Rate limiter: per-tenant, per-key and per-route limits. Each node keeps a local token bucket and synchronises with a shared counter store every short interval, so the global limit is approximately enforced without a remote call per request. Concretely, for a tenant limited to 500 rps / burst 1000 behind, say, 10 gateway nodes, each node's local bucket gets an even static share: 50 rps and a burst of 100. A node enforces only its own share regardless of what other nodes see, so the platform-wide total can never exceed the configured 500 rps / 1000 burst: the "approximate" part is fairness, not the ceiling. If the tenant's traffic happens to land unevenly (all of it hitting 2 of the 10 nodes instead of spreading out), those 2 nodes throttle it well before 500 rps platform-wide while the other 8 nodes' shares sit unused; the periodic sync is what lets the counter store notice that skew and rebalance the local shares at the next interval, not what caps the total.
- Plugin chain: ordered list of Wasm modules for this tenant, each with a budget (for example 5 ms CPU and 16 MB memory), no network access, and a declared failure mode (fail closed: reject the request; fail open: skip the plugin).
- Control plane: API and UI for tenants to edit config, a validator (schema checks, plugin compilation, limit sanity), and a publisher that distributes signed, versioned snapshots.
Tenant-config data model
{
"tenant_id": "t-4812",
"version": 57,
"hosts": ["acme.api.example.com"],
"auth": {
"type": "oidc",
"issuer": "https://login.acme.com",
"audience": "acme-api",
"required_scopes": ["orders:read"]
},
"rate_limits": [
{"scope": "tenant", "rps": 500, "burst": 1000},
{"scope": "api_key", "rps": 50, "burst": 100},
{"scope": "route", "route": "POST /orders", "rps": 100, "burst": 100}
],
"plugins": [
{"id": "legacy-date-format", "module_sha256": "<hash of reviewed module>", "stage": "response",
"cpu_ms": 5, "memory_mb": 16, "on_error": "fail_open"}
],
"api_version_pin": "2025-11-01",
"flags": {"orders_v2_fields": true},
"schema_extensions": {"order": {"x_acme_cost_center": "string"}}
}
Key design choices: the whole document is versioned, so a bad change rolls back to version 56 in one step; plugins are referenced by content hash so the running code is exactly the reviewed code; custom schema extensions live in a tenant namespace (x_acme_...) so they can never collide with future platform fields.
Scaling strategy
- Stateless nodes, config in memory. Illustrative sizing: 20,000 tenants x about 4 KB of config each = 80 MB, which every node can hold. Nodes pull whole snapshots or, to save bandwidth, deltas (only the tenants whose config changed since the last version, not the whole 80 MB document), and swap atomically (the node switches every lookup to the new version in one step, so no request is ever served against half-old, half-new config); a node that cannot fetch config keeps serving the last good version.
- Horizontal scale behind a load balancer; the rate-limit counter store is sharded by tenant ID (its data is split across multiple physical stores by tenant ID, so a given tenant's counters always live on the same one store and no single store has to hold everyone's counters) so one tenant's counters live on one shard.
- Hot tenants: a tenant large enough to dominate a node pool gets its own pool (routing by hostname), which is a config change, not a code change.
- Plugin cost control: modules are compiled once and cached per node; the per-plugin CPU budget bounds the worst case. With a 5 ms budget and 3 plugins, a tenant can add at most 15 ms to its own request's processing time. That bound caps the delay added to the request that triggered it, but the CPU still runs on a node shared by other tenants, so it is not free for them: enough concurrent requests from one tenant still compete for the same node's CPU as everyone else's requests. What actually bounds a tenant's total draw on a node is the rate limiter above (it caps how many of that tenant's requests can be in flight at once); the plugin budget alone only bounds the worst case per request, not the aggregate load one tenant can put on a shared node.
Evolving shared contracts without breaking tenants
- Additive by default: new optional fields and new endpoints are safe; removing or renaming fields, tightening validation, or changing a field's meaning are breaking.
- Explicit versions with pins: breaking changes ship as a new dated version; each tenant's
api_version_pinkeeps them on the old behaviour until they opt in. The gateway translates between versions for as long as the old one is supported. - Tolerant readers: clients and plugins are required to ignore unknown fields rather than error on them. Document it, and test it.
- Replay tenant plugins before release: for every tenant plugin, run its recorded sample traffic through the new contract in a staging gateway and diff the outputs. A plugin that relied on a field that moved fails here, not in production.
- Feature flags for rollout: new behaviour turns on per tenant (internal tenants, then a small cohort, then everyone), with instant rollback.
- Deprecation with telemetry: the gateway counts calls per tenant per deprecated field; you only remove a version when that count is zero or the announced deadline (for example 12 months) has passed.
Trade-offs and pitfalls
- Wasm plugins vs declarative transforms: declarative mapping rules (for example JSON path renames: move or rename a field at a given address inside the JSON document, such as
$.customer.email, plus header rules) cover most needs, are easier to validate, and cannot loop. Offer them first and allow Wasm only when rules cannot express the need. - Global vs local rate limiting: exact global counting needs a remote call per request; local buckets with periodic sync allow brief overshoot of the limit. For fairness between tenants the overshoot is acceptable; for billing-grade quotas, reconcile from logs.
- Pitfall: letting custom auth plugins decide the tenant. The tenant must be established by the platform's resolver and auth stage, not by tenant code that could claim to be anyone.
- Pitfall: config pushes as an outage source. Validate, canary (send the new config to a small subset of nodes first and watch for errors before the rest of the fleet gets it) to a few nodes, and keep the last-good snapshot, because one malformed tenant config must not take down the fleet.
Design a multi-tenant ingestion API in Java that enforces per-tenant rate limits and resource isolation, supporting 1M tenants and 100,000 requests/second globally. Each tenant needs its own burst allowance on top of a base quota, and the limiter itself must not become a bottleneck or a single point of failure at that scale. Walk through your architecture end to end, including how you'd keep per-tenant counters consistent under high concurrency.
Sample Answer
Direct answer
At 100,000 requests per second (RPS) the limiter must not add a network round trip per request to a central counter, and it must keep working when its backing store is down. I would make every rate-limit decision locally inside each Java gateway node, from an in-memory token bucket per active tenant, and use a sharded quota service only to hand out the tenant's allowance in small leases (a lease is a time-bounded batch of tokens the node may spend locally without asking again until it runs low or the lease expires) and to keep the global total honest. Per-tenant counters stay correct under concurrency with a lock-free compare-and-set loop on an immutable bucket state (lock-free means no thread ever blocks waiting for another; a thread whose update collides with someone else's simply retries instead of waiting). If the quota service is unreachable, nodes keep deciding from a conservative local share, so the limiter degrades to "approximately right" instead of becoming an outage.
Requirements and numbers
- 1,000,000 tenants; 100,000 RPS globally; each tenant has a base rate plus a burst allowance.
- The limiter must not be a bottleneck or a single point of failure (one component whose failure stops everything).
- Resource isolation: one tenant's traffic must not consume capacity others need.
Assume 20 gateway nodes behind a load balancer: 100,000 / 20 = 5,000 RPS per node, well within a single JVM's reach. Tenant configuration (plan, base rate, burst) at about 64 bytes each is 1,000,000 × 64 bytes, about 61 MiB: small enough to cache in every node. Most of the million tenants are idle at any moment; if 50,000 are active and a bucket entry costs about 200 bytes, the live buckets need about 9.5 MiB per node.
Architecture
flowchart LR
T[Tenant clients] --> LB[Load balancer]
LB --> G1[Gateway node 1: local buckets]
LB --> G2[Gateway node N: local buckets]
G1 -. lease tokens, report usage .-> Q[Quota service, sharded by tenant]
G2 -. lease tokens, report usage .-> Q
Q --> S[(Replicated counter store)]
C[(Tenant config)] --> G1
C --> G2
G1 --> I[Ingestion queue, partitioned by tenant]
G2 --> I
I --> W[Workers with per-tenant concurrency caps]
1. Gateway nodes (Java). Each request is authenticated, the tenant ID comes from the credential, and tryAcquire(tenant) runs against a local bucket. Accepted requests go onto a durable ingestion queue (e.g. Kafka, a distributed log that many independent consumers can read from, which keeps messages in order within each partition), partitioned by tenant so a tenant's data stays in order. Rejected requests get HTTP 429 with a Retry-After header.
2. Leasing instead of per-request coordination. A tenant's bucket on a node is refilled not from the clock alone but from leases granted by the quota service: "you may spend up to 50 tokens for tenant X over the next second" (tenant X has a base rate of 500 RPS here; the lease is sized at about 10% of a second's allowance, the same sizing rule used below for the cross-node overshoot bound). A node asks for a new lease when its local allowance runs low, and reports what it used. This cuts coordination traffic from 100,000 calls per second to a few per active tenant per node per second, and only for tenants actually sending traffic.
3. Quota service, sharded by tenant. Tenants are assigned to shards by consistent hashing (a scheme where adding a shard moves only a small fraction of tenants), so each tenant's global counter has exactly one owner. That removes cross-shard coordination. The owner keeps the tenant's global token bucket and grants leases out of it. Its state is replicated (e.g. a Redis primary with a replica, Redis being an in-memory key-value store often used for counters and caches because it is fast and supports atomic operations, or a replicated store) so a shard failure fails over rather than losing the counters.
4. Burst on top of base. The bucket's capacity is base plus burst and its refill rate is the base rate: a tenant that has been quiet accumulates up to the burst, then is held to base. The same shape applies at both levels (global bucket in the quota service, local bucket on the node).
5. Resource isolation beyond request counts. Rate limits cap arrivals; they do not stop a tenant with expensive requests from hogging workers. Behind the queue, workers enforce a per-tenant concurrency cap, and the queue consumer serves tenants fairly (round robin across tenant partitions) rather than in arrival order.
Keeping per-tenant counters consistent under high concurrency
Within a node, many request threads hit the same tenant's bucket simultaneously. A read-modify-write without coordination lets two threads both see "1 token left" and both proceed. The fix below keeps the bucket's (tokens, last refill time) pair in one immutable object and swaps it with an atomic compare-and-set (CAS): the write only succeeds if nobody changed the state since it was read; otherwise the thread re-reads and retries. No locks, so a hot tenant never blocks threads serving other tenants.
import java.util.concurrent.*;
import java.util.concurrent.atomic.*;
import java.util.function.LongSupplier;
public class TenantLimiter {
// Immutable snapshot of one tenant's bucket; swapped atomically with CAS.
record State(double tokens, long lastNanos) {}
record Plan(double ratePerSec, double capacity) {} // capacity = base + burst
static final class Bucket {
final Plan plan;
final AtomicReference<State> state;
Bucket(Plan plan, long now) {
this.plan = plan;
this.state = new AtomicReference<>(new State(plan.capacity(), now));
}
boolean tryAcquire(long now) {
while (true) {
State cur = state.get();
double elapsed = Math.max(0, now - cur.lastNanos()) / 1e9;
double refilled = Math.min(plan.capacity(), cur.tokens() + elapsed * plan.ratePerSec());
if (refilled < 1.0) {
// no token: still record refill progress, then reject
if (state.compareAndSet(cur, new State(refilled, now))) return false;
continue;
}
if (state.compareAndSet(cur, new State(refilled - 1.0, now))) return true;
// another thread won the race: re-read and retry
}
}
}
private final ConcurrentHashMap<String, Bucket> buckets = new ConcurrentHashMap<>();
private final LongSupplier clock;
private final java.util.function.Function<String, Plan> plans;
TenantLimiter(LongSupplier clock, java.util.function.Function<String, Plan> plans) {
this.clock = clock;
this.plans = plans;
}
boolean tryAcquire(String tenantId) {
long now = clock.getAsLong();
return buckets.computeIfAbsent(tenantId, id -> new Bucket(plans.apply(id), now))
.tryAcquire(now);
}
public static void main(String[] args) throws Exception {
AtomicLong fakeClock = new AtomicLong(0); // frozen time: no refill during the race
TenantLimiter limiter = new TenantLimiter(fakeClock::get,
id -> new Plan(100.0, 1_000.0)); // 100 req/s base, 1,000 burst capacity
int threads = 16, attemptsPerThread = 10_000;
ExecutorService pool = Executors.newFixedThreadPool(threads);
LongAdder granted = new LongAdder(); // LongAdder: a counter built for many threads to increment concurrently with less contention than AtomicLong
CountDownLatch start = new CountDownLatch(1); // CountDownLatch: blocks every thread here until start.countDown() releases them all at once, for a fair race
for (int t = 0; t < threads; t++) {
pool.submit(() -> {
start.await();
for (int i = 0; i < attemptsPerThread; i++)
if (limiter.tryAcquire("tenant-42")) granted.increment();
return null;
});
}
start.countDown();
pool.shutdown();
pool.awaitTermination(1, TimeUnit.MINUTES);
System.out.println("attempts=" + threads * attemptsPerThread + " granted=" + granted.sum());
fakeClock.addAndGet(2_000_000_000L); // advance 2 s: expect 200 new tokens
int more = 0;
for (int i = 0; i < 1_000; i++) if (limiter.tryAcquire("tenant-42")) more++;
System.out.println("granted after 2s refill=" + more);
System.out.println("other tenant unaffected=" + limiter.tryAcquire("tenant-7"));
}
}
Run with java TenantLimiter.java (JDK 21, single-file launch). Output:
attempts=160000 granted=1000
granted after 2s refill=200
other tenant unaffected=true
16 threads made 160,000 attempts against a full bucket of 1,000 tokens with the clock frozen, and exactly 1,000 were granted: no double-spending. Advancing the clock 2 seconds at 100 tokens per second granted exactly 200 more, and a different tenant was untouched. The clock is injected so the test is deterministic; in production it is System::nanoTime. In production the map also needs eviction of idle tenants (a bounded cache), otherwise a million tenants' buckets accumulate.
Where the lease plugs into tryAcquire. The listing above fixes plan.ratePerSec() at construction so it can isolate the concurrency question (does the CAS loop stop double-spending) from the lease question, which is orthogonal. In production, Plan is not a constant: each Bucket holds an AtomicReference<Plan>, and tryAcquire reads the current one on every call. Each time a node's lease response arrives ("50 tokens for the next second"), the lease handler does planRef.set(new Plan(leaseTokens / leaseWindowSeconds, plan.capacity())), so the refill rate tryAcquire uses becomes whatever the most recent lease allows, with no change to the CAS loop or to the test above.
Across nodes, consistency is a deliberate approximation. A tenant's global allowance can be overspent by at most the unspent leases outstanding: with 20 nodes each holding a lease of up to 10 tokens for a tenant, the worst-case overshoot is 20 × 10 = 200 tokens for that tenant. Shrinking lease size tightens the bound and increases coordination traffic; for a tenant with a base of 100 RPS I would size leases at about 10% of a second's allowance and only lease to nodes that actually see that tenant's traffic.
Failure modes
| Failure | Behaviour |
|---|---|
| Quota-service shard down | Nodes keep serving the tenants on that shard from a fallback local rate of base / number of nodes (fail-conservative), until failover completes |
| A gateway node dies | Its unspent leases expire; the global bucket recovers them after the lease TTL (time-to-live) |
| One tenant sends 50,000 RPS | Its local buckets reject at the edge on every node; no coordination traffic beyond its leases, and other tenants' buckets are separate objects |
| Load balancer skews a tenant onto one node | That node leases more often; the global bucket still enforces the total |
Trade-offs and pitfalls
- A central Redis call per request (an atomic Lua script per request: Redis runs the whole check-and-decrement as one indivisible step, so two concurrent requests can't race each other) is simpler and exact, and at 100,000 RPS a Redis cluster can carry it. I prefer leasing here because it removes the per-request network hop from the latency path and keeps working when Redis does not; I would switch to the per-request design if tenants needed exact enforcement (for example, billing by the request against a hard cap).
- Fail-open vs fail-closed: fail-open means letting requests through uncounted when the quota check can't be made, which during a quota outage risks overload; fail-closed means blocking requests when the check can't be made, which turns a limiter outage into a full outage. Fail to a conservative local share instead (the fallback in the failure-modes table above), which is neither: it keeps limiting, just from the last-known-good numbers.
- Locking per tenant (
synchronizedon the bucket) is correct too, but under a hot tenant it queues threads; CAS keeps threads moving and retries are cheap because the critical section is tiny. - Unbounded per-tenant maps are a memory leak at a million tenants; evict a bucket only after its tenant has been idle at least capacity / rate seconds (10 s for the 1,000-token, 100-per-second plan above). By then the bucket would have refilled anyway, so recreating it full changes no decision.
Explain how PostgreSQL Row Level Security (RLS) can be used to implement tenant isolation for a shared-schema database. Describe how the policies are defined, the performance implications, and the gotchas that catch teams out in production. Give examples of when RLS is the right choice and when it isn't.
Sample Answer
Direct answer
PostgreSQL Row Level Security (RLS) lets you attach a policy to a table, a boolean condition such as tenant_id = <current tenant>, that the database silently adds to every query from ordinary roles. In a shared-schema design (all tenants in the same tables, each row tagged with tenant_id), that turns "every developer must remember the WHERE tenant_id = ? clause" into "the database refuses to show or write other tenants' rows". The application sets the current tenant per transaction, and the policy reads it. RLS is a strong safety net, not a complete isolation strategy: superusers, table owners and some views bypass it, and connection pooling can leak the tenant setting unless it is scoped correctly.
How the policy is defined
Three pieces:
- Turn it on for the table:
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;With RLS enabled and no policy, ordinary roles see zero rows (default deny). - Write the policy.
USINGfilters which existing rows a query can see, update or delete.WITH CHECKvalidates rows being inserted or updated, so a tenant cannot write a row stamped with another tenant's ID. - Tell the database who the tenant is. The usual pattern is a custom setting: the application runs
SET LOCAL app.tenant_id = '42'at the start of each transaction, and the policy reads it withcurrent_setting('app.tenant_id', true). Thetruemeans "return NULL instead of erroring if the setting does not exist".
Worked example (executed)
Run against postgres:16-alpine with docker run -d --name pg-demo -e POSTGRES_PASSWORD=pw postgres:16-alpine, then docker exec -i pg-demo psql -U postgres -q < rls.sql:
-- run as the superuser "postgres" (the table owner in this demo)
CREATE TABLE invoices (
tenant_id int NOT NULL,
id int NOT NULL,
amount numeric NOT NULL,
PRIMARY KEY (tenant_id, id)
);
INSERT INTO invoices VALUES (1,1,100),(1,2,250),(2,1,999);
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON invoices
USING (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::int)
WITH CHECK (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::int);
CREATE ROLE app_user LOGIN;
GRANT SELECT, INSERT, UPDATE, DELETE ON invoices TO app_user;
CREATE VIEW invoice_totals AS SELECT tenant_id, sum(amount) AS total FROM invoices GROUP BY tenant_id;
GRANT SELECT ON invoice_totals TO app_user;
\echo '--- 1. superuser/owner, no tenant set: policy does not apply'
SELECT count(*) FROM invoices;
SET ROLE app_user;
\echo '--- 2. app_user, no tenant set: fails closed (NULL matches nothing)'
SELECT count(*) FROM invoices;
\echo '--- 3. app_user as tenant 1'
BEGIN;
SET LOCAL app.tenant_id = '1';
SELECT tenant_id, id, amount FROM invoices ORDER BY id;
\echo '--- 4. tenant 1 tries to write a tenant 2 row'
-- SAVEPOINT marks a point to roll back to. Without it, the failed INSERT below
-- would abort this whole transaction and every later statement in it would also
-- error with "current transaction is aborted"; rolling back to the savepoint
-- undoes just the failed insert so the script can keep running and COMMIT.
SAVEPOINT s;
INSERT INTO invoices VALUES (2, 50, 1);
ROLLBACK TO SAVEPOINT s;
COMMIT;
\echo '--- 5. SET LOCAL ended with the transaction: next request on this pooled connection sees'
SELECT count(*) FROM invoices;
\echo '--- 6. plain SET (session scope) leaks into the next request on the same connection'
SET app.tenant_id = '2';
SELECT 'request A (tenant 2)' AS who, count(*) FROM invoices;
SELECT 'request B (forgot to set tenant)' AS who, tenant_id, amount FROM invoices;
RESET app.tenant_id;
\echo '--- 7. a view owned by the superuser bypasses the policy'
SET app.tenant_id = '1';
SELECT * FROM invoice_totals ORDER BY tenant_id;
RESET ROLE;
ALTER VIEW invoice_totals SET (security_invoker = true);
SET ROLE app_user;
\echo '--- 8. same view with security_invoker = true'
SELECT * FROM invoice_totals ORDER BY tenant_id;
RESET ROLE;
Output:
--- 1. superuser/owner, no tenant set: policy does not apply
count
-------
3
(1 row)
--- 2. app_user, no tenant set: fails closed (NULL matches nothing)
count
-------
0
(1 row)
--- 3. app_user as tenant 1
tenant_id | id | amount
-----------+----+--------
1 | 1 | 100
1 | 2 | 250
(2 rows)
--- 4. tenant 1 tries to write a tenant 2 row
ERROR: new row violates row-level security policy for table "invoices"
--- 5. SET LOCAL ended with the transaction: next request on this pooled connection sees
count
-------
0
(1 row)
--- 6. plain SET (session scope) leaks into the next request on the same connection
who | count
----------------------+-------
request A (tenant 2) | 1
(1 row)
who | tenant_id | amount
----------------------------------+-----------+--------
request B (forgot to set tenant) | 2 | 999
(1 row)
--- 7. a view owned by the superuser bypasses the policy
tenant_id | total
-----------+-------
1 | 350
2 | 999
(2 rows)
--- 8. same view with security_invoker = true
tenant_id | total
-----------+-------
1 | 350
(1 row)
What each step shows:
- Steps 1 and 2: the superuser sees all 3 rows; the application role with no tenant set sees 0. Failing closed (showing nothing) is the behaviour you want when context is missing.
- Step 3: tenant 1 sees only its own 2 rows, with no
WHEREclause in the query. - Step 4:
WITH CHECKblocks tenant 1 from inserting a row labelled tenant 2. TheSAVEPOINTaround it is not part of the isolation mechanism; it only lets this script recover from the expected error and keep running the remaining steps in one transaction. - Step 5:
SET LOCALvanished atCOMMIT, so the next request on the same pooled connection starts with no tenant and sees nothing. - Step 6: a plain
SETlasts for the whole session, so request B, which forgot to set a tenant, silently read tenant 2's 999 invoice. - Steps 7 and 8: a view runs with its owner's rights by default, and the owner here bypasses RLS, so the view leaked tenant 2's total. PostgreSQL 15 added
security_invoker = truefor views, which applies the caller's policies.
The NULLIF(..., '') in the policy is there for a reason found while running this: once a transaction has used SET LOCAL app.tenant_id, later reads of the setting in that session return an empty string rather than NULL, and without NULLIF the cast ''::int raises invalid input syntax for type integer on the next request. That fails closed, but as an error on every pooled connection's next query rather than an empty result.
Performance implications
- The policy predicate is added to the query like an ordinary
WHEREclause, so the cost is whatever that filter costs. With an index that leads withtenant_id, it is usually an index lookup. The follow-on script below (run in the same container afterrls.sql, withdocker exec -i pg-demo psql -U postgres -q < rls2.sql) loads 200 tenants with 501 invoices each and checks the plan, and also checks the table-owner bypass covered under gotchas:
CREATE ROLE app_owner LOGIN;
CREATE TABLE notes (tenant_id int NOT NULL, body text);
ALTER TABLE notes OWNER TO app_owner;
INSERT INTO notes VALUES (1,'a'),(2,'b');
ALTER TABLE notes ENABLE ROW LEVEL SECURITY;
CREATE POLICY p ON notes USING (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::int);
SET ROLE app_owner;
SET app.tenant_id = '1';
\echo 'owner, ENABLE only'
SELECT count(*) FROM notes;
RESET ROLE;
ALTER TABLE notes FORCE ROW LEVEL SECURITY;
SET ROLE app_owner;
\echo 'owner, FORCE'
SELECT count(*) FROM notes;
RESET ROLE;
-- index use (loads 200 tenants x 501 invoices)
INSERT INTO invoices SELECT t, g, 1 FROM generate_series(1,200) t, generate_series(1000,1500) g;
ANALYZE invoices;
SET ROLE app_user;
SET app.tenant_id = '7';
EXPLAIN (COSTS OFF) SELECT sum(amount) FROM invoices;
EXPLAIN (COSTS OFF) SELECT * FROM invoices WHERE id = 1200;
owner, ENABLE only
count
-------
2
(1 row)
owner, FORCE
count
-------
1
(1 row)
QUERY PLAN
-------------------------------------------------------------------------------------------------------------------
Aggregate
-> Bitmap Heap Scan on invoices
Recheck Cond: (tenant_id = (NULLIF(current_setting('app.tenant_id'::text, true), ''::text))::integer)
-> Bitmap Index Scan on invoices_pkey
Index Cond: (tenant_id = (NULLIF(current_setting('app.tenant_id'::text, true), ''::text))::integer)
(5 rows)
QUERY PLAN
-------------------------------------------------------------------------------------------------------------------------
Index Scan using invoices_pkey on invoices
Index Cond: ((tenant_id = (NULLIF(current_setting('app.tenant_id'::text, true), ''::text))::integer) AND (id = 1200))
(2 rows)
Reading the two shapes: the first query (sum(amount), no row filter beyond tenant) matches many rows, tenant 7's whole slice of the table, so the planner picks a Bitmap Index Scan: walk the index once and build an in-memory list ("bitmap") of the matching rows' physical locations, without fetching them yet. That feeds a Bitmap Heap Scan, which fetches exactly those locations from the table in one pass ordered by physical position, cheaper than jumping to each row individually. Index Cond under the bitmap index scan is the condition used to build that list; Recheck Cond on the heap scan is PostgreSQL re-checking that same condition against each fetched row, something bitmap scans always do because the in-memory bitmap can be a lossy approximation when a great many rows match. The second query adds id = 1200. Combined with the policy's tenant condition, that pins down both columns of the (tenant_id, id) primary key, so at most one row can match, and the planner instead uses a plain Index Scan: walk the index and fetch that one matching row immediately, which beats building a whole bitmap when only a single row is expected. Either way, the policy's tenant condition becomes part of the Index Cond on the (tenant_id, id) primary key, so the database reads only tenant 7's slice of the index.
- Without a
tenant_id-leading index, every query scans all tenants' rows and then filters, which gets worse as tenants grow. - Policies that call subqueries or functions (for example "tenant is in the user's membership table") run that logic for queries on every protected table; keep the policy a simple comparison against a setting.
- The query planner (the part of PostgreSQL that chooses the plan, the kind of plan shown by
EXPLAINabove) will not push (reorder to run earlier than) user-supplied conditions that use non-leakproof functions ahead of the policy check. A function is leakproof if PostgreSQL can prove it never reveals anything about its input, not through its result, an error message, a log line, or even how long it takes to run; plain comparisons like=are leakproof, but an arbitrary user-defined function might, for example, raise an error that quotes the value it was checking, revealing a row's content even though the row was never returned. PostgreSQL cannot verify this for every function, so a non-leakproof user condition is never allowed to run before the tenant policy check, only after it. That ordering (policy check first, always) can occasionally produce a worse plan than letting a cheap-looking user condition run first. CheckEXPLAINon hot queries after enabling RLS.
Production gotchas
| Gotcha | What happens | Fix |
|---|---|---|
Superuser and BYPASSRLS roles | Always bypass policies | The application must never connect as a superuser or a role with BYPASSRLS |
| Table owner | The owner bypasses RLS unless the table is set to FORCE ROW LEVEL SECURITY (in the script above, the non-superuser owner app_owner saw 2 rows before FORCE and 1 after) | Have migrations run as the owner, the application as a separate role, and use FORCE as a backstop |
| Connection pooling | SET persists on the pooled connection into the next request (step 6); PgBouncer (a widely used PostgreSQL connection pooler) in transaction mode hands the same physical connection to a different client at each transaction boundary, so a plain session SET left behind by one client can still be active on that connection when a completely different client's next transaction begins, silently applying the wrong tenant to their queries | Use SET LOCAL (or set_config(..., true)) inside the transaction, never session SET |
| Views | Run as the view owner, bypassing the caller's policies (step 7) | security_invoker = true on PostgreSQL 15+, or have views owned by a non-bypassing role |
| Tenant ID from untrusted input | If the tenant value comes from a request header or, worse, is built into SQL by string concatenation, an attacker can switch tenant or inject SQL | Take the tenant from the authenticated token, pass it as a bound parameter to set_config('app.tenant_id', $1, true) |
| Unique constraints and foreign keys | These are checked across all rows regardless of RLS, so "email already exists" can reveal another tenant's data | Make uniqueness per tenant: UNIQUE (tenant_id, email) |
| Missing context | A policy that errors or matches everything when the setting is absent | Fail closed, as with the NULLIF form above |
When RLS is the right choice, and when it is not
Right choice: a shared-schema SaaS (software as a service) product with many small tenants on PostgreSQL, where the risk you most want to remove is a developer forgetting a tenant filter; internal analytics databases where many teams share tables but must see only their own rows.
Not the right choice (or not enough on its own):
- Tenants that need a separate encryption key, separate backup and restore, or proof of physical separation: use a database or infrastructure per tenant.
- Workloads where one tenant's queries overload the database: RLS isolates data, not performance.
- Heavy reporting through views, materialized views or ETL (extract, transform, load) jobs that run as privileged roles: those paths bypass RLS and need their own tenant filtering.
- Very complex policies (hierarchical sharing across tenants): the per-query cost and the difficulty of reasoning about them usually argue for authorization in the application plus a simple RLS tenant check underneath.
Design a tenant-aware backpressure and queueing system for a multi-tenant SaaS product with bursty workloads. Explain fair-share algorithms, per-tenant quotas, circuit-breakers per tenant, and how to avoid noisy-neighbor problems without over-provisioning.
Sample Answer
Direct answer
The core change is to stop putting all tenants' work in one first-in first-out (FIFO) queue. I would give each tenant its own bounded queue, have workers pull from those queues with a weighted fair-share scheduler (deficit round robin), cap each tenant's in-flight work with a quota, and push back on a tenant whose queue is full with a retryable rejection instead of letting its backlog grow. Per-tenant circuit breakers stop one tenant's failing dependency from tying up shared workers. Together these let one tenant's burst use idle capacity, while making it impossible for that burst to delay everyone else, so you do not need to provision for every tenant's peak at once.
Why a shared FIFO queue fails
With one queue, a tenant that drops 500 jobs at once puts every other tenant's next job behind all 500. Nobody did anything wrong; the queue discipline simply gives the earliest arrivals all the capacity. Adding workers only shortens the delay in proportion to the extra spend, which is the over-provisioning we want to avoid.
Backpressure is the general answer: when a system is full, it signals producers to slow down instead of accepting unbounded work. In a multi-tenant system, the signal must be per tenant, so that only the tenant causing the pressure is slowed.
Architecture
flowchart LR
C[Tenant clients] --> A[Admission: auth, tenant ID, quota check]
A -->|queue full| R[429 + Retry-After]
A --> Q1[Queue tenant A]
A --> Q2[Queue tenant B]
A --> Q3[Queue tenant C]
Q1 --> S[Fair-share scheduler]
Q2 --> S
Q3 --> S
S --> W[Shared worker pool]
W --> CB[Per-tenant circuit breakers]
CB --> D[Downstream dependencies]
1. Admission. Identify the tenant from its credentials (never from a client-supplied field), check its rate quota (for example a token bucket: a per-tenant allowance that refills at a fixed rate), and check its queue depth. If the queue is at its bound, reject with HTTP 429 and a Retry-After header. This is where backpressure becomes visible to the producer.
2. Per-tenant bounded queues. Physically these can be one table or stream partitioned by tenant, or a set of logical queues in the scheduler. The bound per tenant is set by tier (e.g. 200 queued jobs for a standard tenant, 2,000 for enterprise), which is what caps any one tenant's claim on memory and latency.
3. Fair-share scheduler. Deficit round robin (DRR) visits each non-empty tenant queue in turn and gives it a "quantum" of work credit per round (its weight); a tenant spends credit to dispatch jobs, and unspent credit carries over only while it has jobs waiting. With equal weights it is plain round robin; with weights 1:2:4 for free, standard and enterprise tiers, capacity divides in that ratio when everyone is busy. Crucially, it is work-conserving: if only one tenant has work, that tenant gets all the workers. That is how one tenant's burst uses idle capacity without taking capacity other tenants need. If job costs vary a lot, charge the quantum in estimated cost units (seconds of work) rather than job count.
4. Per-tenant concurrency quota. Independently of the scheduler, cap how many jobs a tenant can have running at once (e.g. 20% of workers for a standard tenant). This protects against a tenant whose jobs are individually slow, which fair ordering alone does not.
5. Per-tenant circuit breakers. A circuit breaker watches calls to a dependency and, once the failure rate crosses a threshold, fails further calls immediately for a cool-down period instead of letting them hang. Keyed per tenant (and per dependency), it handles the case where one tenant's jobs call something tenant-specific that is broken (its webhook endpoint, its connected database, its bad configuration). Without it, that tenant's jobs sit on shared workers waiting on timeouts. With it, the tenant's jobs fail fast (and are retried later or dead-lettered, i.e. moved to a DLQ, a dead-letter queue for jobs that keep failing) and the workers go back to everyone else.
6. Load shedding under global overload. If the whole system is saturated, shed in order: free-tier background work first, then retries, then new standard work; never an enterprise tenant's in-flight job.
Worked example (executed)
A pool completes 10 jobs per tick. Tenant a submits a burst of 500 jobs at tick 0; tenants b, c, d, e each submit 2 jobs per tick, steadily. Compare one FIFO queue with per-tenant queues (bound 200) served round robin.
from collections import deque
CAPACITY = 10 # jobs the worker pool finishes per tick
TICKS = 100
SMALL = ["b", "c", "d", "e"] # steady tenants: 2 jobs per tick each
BURST = 500 # tenant "a" dumps 500 jobs at tick 0
QUEUE_CAP = 200 # per-tenant queue bound (used by fair mode only)
def arrivals(t):
jobs = [("a", t)] * BURST if t == 0 else []
for s in SMALL:
jobs += [(s, t)] * 2
return jobs
def pct(values, p):
v = sorted(values)
return v[min(len(v) - 1, int(p / 100 * len(v)))]
def run_fifo():
q, waits, rejected = deque(), {}, 0
for t in range(TICKS):
q.extend(arrivals(t))
for _ in range(min(CAPACITY, len(q))):
tenant, arrived = q.popleft()
waits.setdefault(tenant, []).append(t - arrived)
return waits, rejected, len(q)
def run_fair():
queues = {x: deque() for x in ["a"] + SMALL}
waits, rejected = {}, 0
for t in range(TICKS):
for tenant, arrived in arrivals(t):
if len(queues[tenant]) >= QUEUE_CAP:
rejected += 1 # backpressure: caller gets 429 + Retry-After
else:
queues[tenant].append((tenant, arrived))
served = 0
while served < CAPACITY and any(queues.values()):
for tenant in queues: # round-robin, quantum = 1 job per visit
if served < CAPACITY and queues[tenant]:
_, arrived = queues[tenant].popleft()
waits.setdefault(tenant, []).append(t - arrived)
served += 1
return waits, rejected, sum(len(q) for q in queues.values())
for name, fn in [("FIFO", run_fifo), ("fair-share", run_fair)]:
waits, rejected, backlog = fn()
small = [w for s in SMALL for w in waits.get(s, [])]
print(f"{name:10s} small-tenant p99 wait={pct(small, 99)} ticks, "
f"small jobs done={len(small)}, burst jobs done={len(waits.get('a', []))}, "
f"rejected={rejected}, backlog={backlog}")
Output (Python 3):
FIFO small-tenant p99 wait=50 ticks, small jobs done=500, burst jobs done=500, rejected=0, backlog=300
fair-share small-tenant p99 wait=0 ticks, small jobs done=800, burst jobs done=200, rejected=300, backlog=0
Reading it: under FIFO the four well-behaved tenants waited up to 50 ticks at the 99th percentile (p99) and had 300 jobs still stuck at the end, because the burst went first. Under fair share their p99 wait was 0 ticks and all 800 of their jobs finished. Tenant a still got the 2 spare slots per tick (10 capacity minus 8 from the steady tenants), finishing 200 jobs, and its other 300 were rejected at admission and must be retried later. Same worker count, no extra capacity: the only thing that changed is who absorbs the burst. Tenant a does.
Avoiding over-provisioning
Provision for the sum of guaranteed shares plus a modest burst buffer, not the sum of peaks. Because DRR is work-conserving, idle capacity is still lent to whoever is bursting, so utilisation stays high. Scale out the worker pool on the metric that means "tenants inside their quota are waiting" (queue age of in-quota work), not on total queue length, which a single bursting tenant can inflate on its own and which would make you buy capacity for that tenant's backlog.
Trade-offs and pitfalls
- Rejecting work moves the problem to the client. Clients must honour
Retry-Afterwith backoff and jitter (random spread), or the rejections come straight back as a retry storm. Offer an asynchronous submit-and-poll path for tenants who legitimately send large batches. - Thousands of tenants and a naive round robin scans many empty queues; keep an "active tenants" list containing only non-empty queues.
- Fairness by job count is unfair when job sizes differ. Charge cost units, or one tenant with 10-minute jobs still dominates.
- Circuit breakers keyed globally punish every tenant for one tenant's broken webhook; key them by tenant and dependency.
- Priority tiers without a floor starve the free tier. Give every tier a minimum weight so low-tier work always progresses.
Unlock Full Question Bank
Get access to all 8 Multi-Tenancy and Isolation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.