Multi-Tenancy and Isolation Questions
Serving many tenants from shared infrastructure: tenancy models (silo, pool, bridge), data isolation, per-tenant data residency, noisy-neighbor mitigation, per-tenant limits, and security boundaries between tenants. Covers the cost, isolation, and blast-radius tradeoffs of shared versus dedicated resources, and business continuity: per-tenant backup, disaster recovery, and compliant tenant offboarding and deletion. The architecture layer specific to SaaS and platform products.
A major client requests a capability that would weaken multi-tenant isolation for a shared cloud service in order to meet their performance needs. Propose architectural alternatives and contractual constraints that protect platform multi-tenancy and your other customers while still addressing the client's underlying need.
Sample Answer
Direct answer
I would not ship the isolation-weakening capability on the shared platform, because the cost of it lands on every other customer who never agreed to it. Instead I would find the performance need underneath the request (usually predictable throughput or lower tail latency, the slow end of the response-time distribution: the small share of requests that take far longer than a typical one, which is what a p99 figure measures), and meet it with capacity that is reserved or dedicated to that client: a higher committed tier inside the shared pool, or a dedicated cell running the same code. Then I would put the limits in the contract so the arrangement cannot quietly drift back into the shared pool.
Terms used below
- Multi-tenancy: many customers (tenants) run on the same software and infrastructure, separated by software controls such as quotas, rate limits and access checks.
- Isolation: the guarantee that one tenant cannot read another tenant's data, and cannot degrade another tenant's performance beyond an agreed bound.
- Noisy neighbour: a tenant whose load consumes shared resources (CPU, connections, disk I/O) so that other tenants slow down.
- Blast radius: how many tenants are affected when something goes wrong.
- Cell: a complete, independent copy of the service stack (compute, cache, database) serving a subset of tenants. A problem in one cell does not spread to others.
- Silo / pool / bridge: three tenancy models. Pool means everyone shares everything; silo means a tenant gets its own dedicated stack; bridge mixes them (for example shared web tier, dedicated database).
- SLA (service-level agreement): the contractual promise (for example 99.9% availability), usually with service credits if missed. p99 (99th percentile) latency is the response time 99% of requests beat.
- RPS: requests per second.
Step 1: turn the request into the need behind it
Typical "weaken isolation" asks and what they really mean:
| What the client asks for | What they usually need | Why the literal ask is dangerous |
|---|---|---|
| "Remove our rate limit" | Sustained throughput above their tier, or bursts | Their burst consumes headroom (spare capacity kept in reserve above normal load) that other tenants' latency depends on |
| "Let us query the database directly" | Faster bulk reads or reporting | Bypasses the tenant filter in the application layer; one bad query locks shared tables |
| "Pin us to the fastest nodes" / "turn off fair scheduling for us" (fair scheduling: the mechanism that shares capacity across tenants instead of serving requests in strict arrival order) | Lower p99 latency | Moves the queueing delay onto everyone else |
| "Share a cache across our org and a partner's org" | Fewer cold reads (reads that miss the cache and fall through to the slower backing store) | Cross-tenant cache keys are a classic data-leak path |
In the room I ask three questions: what workload, at what rate, with what latency target, and what happens to their business if that target is missed. Those numbers decide which alternative fits.
Step 2: architectural alternatives, cheapest first
- Optimise inside their own quota. Batch APIs, bulk export jobs, async endpoints, a regional endpoint closer to their users, client-side caching. Often the latency complaint is round trips, not server capacity.
- A committed-capacity tier in the pool. They buy a higher limit backed by capacity the platform actually adds (more nodes in their cell), not by borrowing others' headroom. The limit still exists; it is just larger and paid for.
- A dedicated cell (bridge model). Same code, same deploy pipeline, but their traffic lands on hardware nobody else uses. Inside that cell they can have relaxed limits, because the only tenant they can hurt is themselves.
- A full silo. Separate account, network and database. Justified only when they also need it for compliance, custom keys or change-window control.
Feature flags (switches that turn a behaviour on for specific tenants without a separate deploy) are how you ship "relaxed limits" safely: the flag is only allowed to evaluate to "on" for tenants whose placement is a dedicated cell. A flag that relaxes limits inside a shared cell is exactly the isolation weakening we are refusing.
flowchart LR
C[Client traffic] --> R{Tenant router}
R -->|standard tenants| P1[Shared cell 1]
R -->|standard tenants| P2[Shared cell 2]
R -->|this client| D[Dedicated cell<br/>relaxed limits allowed]
P1 --> G[Per-tenant limits<br/>always enforced]
P2 --> G
Step 3: contractual constraints
The contract is what stops tomorrow's escalation from reopening the shared-pool question.
- Committed capacity, stated in numbers: "up to 30,000 RPS sustained on the dedicated cell, p99 under 50 ms at that rate." Anything above that is throttled, and the throttling is not an SLA breach.
- Minimum commitment and term: a dedicated cell has fixed cost whether or not it is used, so pricing includes a floor and a term long enough to recover setup cost.
- Platform protection clause: the provider may throttle, shed (refuse or drop low-priority work outright under overload, rather than queueing it) or isolate any workload that threatens other customers, and the client may not require disabling tenant isolation controls (rate limits on shared components, tenant filtering, encryption boundaries).
- Shared-dependency carve-out: SLA credits cover the dedicated cell, not global dependencies such as identity or DNS, which stay shared.
- Change control: changes to their cell's configuration go through the same review as platform changes; no direct production access.
- Exit and downgrade path: if the commitment lapses, the tenant moves back to the pooled tier with its standard limits, with notice.
- Security obligations: penetration testing of their cell only, with notice; no testing that touches shared components.
Worked example
Assumptions (illustrative, not measured): a shared cell has 20 nodes, each serving 2,000 RPS within the latency target, so capacity is 20 x 2,000 = 40,000 RPS. Normal load from 300 tenants is 28,000 RPS (70% utilisation). The client wants bursts of 30,000 RPS with p99 under 50 ms.
- Literal ask (remove their limit in the shared cell): demand becomes 28,000 + 30,000 = 58,000 RPS against 40,000 capacity, which is 145%: work arrives 45% faster than the cell can finish it. That extra 45% cannot just vanish, so it piles onto the backlog every second instead of draining, which is what "queues grow without bound" means. Queues grow without bound during the burst, and all 300 tenants miss their latency target, not just the requester.
- Dedicated cell: size for 30,000 RPS at 60% target utilisation (leaving headroom so bursts do not queue): 30,000 / 0.6 = 50,000 RPS of capacity, and 50,000 / 2,000 = 25 nodes. That is the number that goes into the price and into the "committed capacity" clause.
- Result: the client's burst can now only degrade their own cell. The other 300 tenants' blast radius from this client is zero.
If step 1 showed their real issue was 400 ms of round trips from another continent, a regional endpoint would fix it for a fraction of the cost of 25 nodes, which is why step 1 comes first.
Trade-offs and pitfalls
- Dedicated cells cost money and operational attention. Each one is another thing to patch, monitor and capacity-plan. Keep them on the same automated pipeline; a hand-tuned snowflake cell is where drift and outages come from.
- What would flip the recommendation: if many clients ask for the same relaxation, it is a product gap, not an exception. Raise capacity or redesign the limit for everyone rather than growing a zoo of special cells.
- Pitfall: "just this once" config overrides in the shared pool. They are rarely removed, and the next incident review finds them.
- Pitfall: selling an SLA the architecture cannot keep. A 50 ms p99 promise that still depends on a shared database is a promise other tenants are paying for.
- Pitfall: treating it as purely technical. The account team needs a clear, positive story ("dedicated performance tier"), not "security said no". Framing the offer as an upgrade is what gets it accepted.
Create an operational runbook to detect, contain, and remediate a suspected cross-tenant data leakage incident. Include detection signals, immediate containment steps, steps for forensic evidence collection, customer notification timelines, regulatory reporting considerations, and long-term remediation to prevent recurrence.
Sample Answer
Direct answer
A cross-tenant leak runbook has one organizing idea: every hour you work out two things faster, which tenants' data went where, and whether it is still flowing. The runbook moves through detection (automated tenant-mismatch signals plus customer reports), containment in minutes (kill switches, cache flush, token revocation), evidence preservation before anything is changed further, an exposure matrix that says which source tenant's data reached which recipient tenant, customer and regulator notification on legally driven clocks (GDPR's 72 hours for controllers is the one that sets the pace), and remediation that closes the whole class of bug, not just the instance.
Roles first
A leak is a security incident and a customer incident at the same time. Assign at declaration:
| Role | Owns |
|---|---|
| Incident commander (IC) | Decisions, timeline, severity |
| Technical lead | Containment and root cause |
| Forensics lead | Evidence preservation, chain of custody |
| Privacy/legal lead | Whether it is a personal-data breach, regulator and contract clocks |
| Customer communications lead | Tenant notices, support scripts |
| Scribe | Timestamped log of every action and decision |
Legal is paged at declaration, not after the investigation, because notification clocks start at awareness, and because counsel may want the investigation run under legal privilege (communications made for the purpose of legal advice that cannot be compelled into evidence later, sometimes used to keep an incident investigation's internal findings confidential).
Phase 1: Detection signals
- Tenant-mismatch assertions in the data access layer. After every query, the repository checks that every returned row's
tenant_idequals the request's tenant context. A mismatch is blocked and alerted. This is the single most valuable signal because it fires on the first leaked row. - RLS denials or unexpected spikes in rows filtered by row-level security (RLS: the database silently filtering rows by a tenant policy) from a code path that normally sees none.
- Cache-key anomalies: a cache entry read under one tenant context that was written under another (store the writer's tenant ID inside the cached value and compare on read).
- Customer reports: "I can see an invoice that isn't ours." Support must have a one-click path to page security, with a script that says do not ask the customer to send screenshots containing the other tenant's data by email.
- Canary tenants: synthetic tenants whose unique marker strings should never appear in another tenant's responses, exports, or search results.
- Audit-log anomalies: an API key or user fetching object IDs belonging to many tenants (an IDOR pattern, insecure direct object reference: fetching an object by guessing its ID without an ownership check).
Phase 2: Immediate containment (target: first 30 minutes)
Containment beats diagnosis. In order:
- Stop the flow. Flip the feature flag (a runtime switch that turns a code path off without a deploy) for the suspected endpoint, or roll back the last deploy if the timing correlates. If neither is clear, put the affected surface into read-only or maintenance mode for the implicated tenants.
- Purge poisoned shared state. First export a snapshot of the affected cache keys and the tenant ID each one was storing, so Phase 3 can still size the exposure after the flush; then flush the relevant cache namespaces, invalidate CDN (content delivery network) objects for affected routes, and pause export or email jobs that might be sending leaked data out.
- Revoke what may have been exposed. Rotate any credentials, API keys or session tokens that could have crossed the boundary.
- Do not delete. Containment must not destroy evidence: disable, do not drop.
Phase 3: Forensic evidence collection
Before anyone "cleans up":
- Snapshot affected databases, and export the relevant logs, traces, CDN logs and audit logs to a write-once location (for example object storage with a retention lock, a setting that blocks deletion or modification for a fixed period, even by an administrator) with restricted access.
- Record hashes of every exported artifact and who collected it, when (chain of custody: an unbroken record of who handled the evidence and when, so it can later be shown that nothing was altered along the way).
- Freeze log retention for the incident window so rotation does not erase it.
- Capture the deployed build identifiers and configuration at the time of the incident.
Then build the exposure matrix, the artifact everything downstream depends on. Worked example of sizing it: the leak was a cache key missing the tenant prefix on /invoices/{id}, live for 108 minutes. Access logs show 1,940 requests to that route in the window; joining the request's tenant (from the access log) against the tenant that value was cached under (from the pre-flush cache export made during containment) finds 37 responses served cross-tenant, touching data of 6 source tenants seen by 9 recipient tenants. Those 6 source tenants are the ones owed a breach notice; the 9 recipients may be asked to confirm deletion of what they received.
Of those 37 leaked responses, 6 went from Tenant A to Tenant C; since each response is one invoice and an invoice can hold many line items, those 6 invoices held about 120 line items between them. The matrix records that pair like this:
| Source tenant (data owner) | Recipient tenant (saw it) | Data categories | Records | Window | Evidence |
|---|---|---|---|---|---|
| Tenant A | Tenant C | invoice line items (names, amounts) | 6 responses, about 120 line items total | 09:14 to 11:02 UTC | access log query #4 |
("Records" here counts individual line items inside a leaked response, not the response count itself, since one leaked invoice can carry many line items.)
Phase 4: Notification timelines
The SaaS provider is usually a processor (acting on the customer's instructions) and the tenant is the controller (who decides why and how personal data is processed). That sets who tells whom.
| Obligation | Clock | Who |
|---|---|---|
| GDPR Art. 33: processor to controller | "Without undue delay" after becoming aware | You, to each affected source tenant |
| GDPR Art. 33: controller to supervisory authority (a country's data protection regulator, for example Ireland's DPC for many EU-based tech companies) | Within 72 hours of the controller becoming aware, where the breach is likely to pose a risk to individuals | The tenant, supported by your report |
| GDPR Art. 34: controller to individuals | Without undue delay, when high risk | The tenant |
| HIPAA: business associate (a vendor, like you, processing health data on a covered entity's behalf) to covered entity (a healthcare provider, insurer or clearinghouse) | Without unreasonable delay, no later than 60 days from discovery | You, if the data is protected health information |
| Contracts (DPA, data processing agreement) | Often 24 to 72 hours, whichever the DPA says | You; check each affected tenant's contract |
| US state breach laws, sector regulators | Vary by state and sector | Counsel decides |
Practical rule: aim to send an initial notice within 24 hours to affected source tenants even with incomplete facts, then follow up as the exposure matrix firms up. A late but complete notice is worse than an early partial one, because the tenant's own 72-hour clock may already be running. Recipient tenants get a separate, carefully worded request to delete and confirm, which must not reveal the source tenant's identity.
Phase 5: Long-term remediation
Fix the class, not the line:
- Make tenant scoping structural. A tenant-scoped repository or ORM (object-relational mapping) layer that refuses unscoped queries; RLS in the database as the backstop.
- Tenant in every key. Cache keys, object storage paths, search index filters and queue message envelopes all carry the tenant ID, enforced by shared libraries rather than convention.
- Automated cross-tenant tests in CI: a suite that authenticates as tenant A and attempts to read every resource type belonging to tenant B, expecting 404.
- Keep the tenant-mismatch assertion permanently in production.
- Blameless post-incident review within two weeks, with owners and dates for each action, and a customer-facing summary for affected tenants.
Trade-offs and pitfalls
- Rolling back vs disabling a flag. Rollback is fast but may also revert an unrelated security fix; a flag is surgical but only exists if someone built it. Build kill switches for every tenant-data-serving path in advance.
- Pitfall: cleaning up before snapshotting. Flushing the cache without first recording which keys held which tenant's data destroys your ability to scope the exposure matrix.
- Pitfall: notifying recipients before sources. Recipients learn a competitor's data existed before the owner does.
- Pitfall: over-notifying. Telling all 4,000 tenants "there may have been an issue" when the matrix shows 6 affected creates panic and regulator questions. Scope first, fast.
Explain single-tenant and multi-tenant models for SaaS products. Describe differences in operational cost, customization capability, security isolation, upgrade cadence, and typical customer preferences across verticals and company sizes.
Sample Answer
Direct answer
Single-tenant means each customer gets its own copy of the software and its own database, like each family owning a detached house. Multi-tenant means all customers share one running system and one set of infrastructure, with each customer's data kept separate by the software, like families in one apartment building with locked flats. Multi-tenant is much cheaper to run and easier to keep up to date; single-tenant gives stronger separation and more room for customization, which is why large regulated customers often ask for it and small businesses rarely care.
Key terms in plain words
- Tenant: one customer (usually a company) and all of its users and data.
- Isolation: how strongly one customer's data and performance are protected from the others.
- Noisy neighbour: in a shared system, one very busy customer slowing everyone else down, like one flat running every tap at once and dropping the building's water pressure.
The five differences
| Single-tenant | Multi-tenant | |
|---|---|---|
| Operational cost | High: every customer has its own servers and database, mostly idle, plus its own monitoring, backups and patching | Low: one set of infrastructure is shared, so capacity is used efficiently and fixed costs are spread |
| Customization | High: the vendor can change configuration, integrations, even versions per customer (risky if taken too far) | Configuration only: settings, branding, custom fields and feature switches, but the same code for everyone |
| Security isolation | Strong: separate database and often separate network; one breach or bug affects one customer | Relies on software controls; a bug in them could expose several customers, so the vendor must invest heavily in testing and in database-level safeguards |
| Upgrade cadence | Slow and uneven: each copy upgraded separately, often on the customer's schedule, so customers drift onto different versions | Fast and uniform: one upgrade reaches everyone, often weekly or continuously |
| Typical buyer | Large enterprises, banks, healthcare, government | Startups, small and mid-sized businesses, most enterprises for non-sensitive tools |
Customer preferences by vertical and company size
- Small and mid-sized businesses prefer multi-tenant: lower price, instant sign-up, no upgrade projects. They rarely ask where their data physically sits.
- Large enterprises often ask for dedicated options or strong isolation guarantees in security reviews, plus custom contracts and change control (for example, a say in when upgrades happen).
- Financial services and healthcare push towards single-tenant or dedicated data stores because of regulators and auditors, and often want their own encryption keys. Note that regulations such as HIPAA (the US health-privacy law) do not strictly require a separate system; these customers ask for it because it is easier to evidence to an auditor.
- Government and public sector frequently need data in a specific country or an accredited environment, which usually means a dedicated or specially certified deployment.
- Tech-forward companies of any size tend to value the fast feature cadence of multi-tenant over customization.
Worked example
A vendor has 100 customers.
- Single-tenant: suppose the smallest workable stack for one customer (app servers, database, backups, monitoring) costs about $2,000 a month (an illustrative assumption). 100 x $2,000 = $200,000 a month, and every release is rolled out 100 times.
- Multi-tenant: suppose one shared, well-sized stack for all 100 costs $30,000 a month. That is $30,000 / 100 = $300 per customer per month, and a release goes out once.
Here the shared system costs about 15% of the dedicated one ($30,000 vs $200,000). That gap is why most SaaS vendors are multi-tenant by default and sell dedicated deployments as a premium tier at a much higher price.
The common middle ground
Many vendors run hybrid: everyone on the shared system by default, with a dedicated deployment (or at least a dedicated database) for customers who need it and pay for it. The discipline that makes this work is keeping one codebase: dedicated customers run the same software, just on separate infrastructure, so they still get upgrades.
Pitfalls
- Assuming single-tenant is automatically "more secure": a dedicated copy that is three versions behind on patches can be less secure than a well-run shared system.
- Promising deep per-customer customization in a single-tenant deal, then being unable to upgrade that customer.
- Treating the choice as all-or-nothing, when the hybrid model covers most real customer bases.
Design a secure multi-tenant microservice platform where tenant isolation, data protection, and noisy-neighbor mitigation matter. Compare isolation options (namespaces, containers, VMs), per-tenant quotas and throttling, encryption choices, monitoring requirements, and the operational trade-offs between cost and isolation strength.
Sample Answer
Direct answer
I would treat isolation as a ladder and put each tenant on the lowest rung its threat allows. If tenants only use our code and send us data, per-tenant Kubernetes namespaces with quotas, network policies and per-tenant encryption keys are enough and cost little. If tenants run their own code on our platform, containers alone are not a security boundary (they share the host's kernel), so I would move them into a sandboxed runtime or small virtual machines. Noisy-neighbour protection is a separate problem from security: it is solved with per-tenant quotas and throttling at every shared resource, and it is proved with per-tenant monitoring.
A few terms used throughout
The kernel is the core program that mediates every process's access to CPU, memory, files and network, by handling system calls (the requests a running program makes to it, like "open this file" or "send this data"); containers on the same machine all share one kernel. A container escape, sometimes via a kernel exploit (a bug in the kernel itself), is when code inside a container breaks out of its boundary and reaches the host, and from there potentially every OTHER tenant's containers on that same machine, not just its own. A hypervisor is the layer that lets one physical machine run several virtual machines, each with its own emulated hardware, which is why a VM-level boundary survives a kernel exploit that a container-level one does not. Attack surface is the total set of ways in an attacker could exploit, so a smaller attack surface means fewer such ways in. The control plane is the part of a cluster that manages it (schedules workloads, stores configuration) rather than running the tenant workloads themselves.
First, name the threat
- Accident and noise: a tenant's workload uses too much CPU, memory or I/O, or a bug in our code crosses tenants. Our code is trusted; tenants are not attackers.
- Malicious or untrusted tenant code: a tenant uploads a plugin, script or container. Now a kernel exploit in that code is a cross-tenant breach.
The first threat is mostly about fairness and correctness; the second is about hard security boundaries. Mixing them up leads either to overspending (VMs for everyone) or to a real breach (containers for untrusted code).
Isolation options compared
One ambiguity to clear up: "namespaces" can mean Linux namespaces (the kernel feature that gives a process its own view of processes, network and filesystem, which is what a container is built from) or Kubernetes namespaces (a named partition inside one cluster for organising objects, quotas and permissions). Below, "namespace" means the Kubernetes kind.
| Option | What is shared | Boundary strength | Cost and density | Good for |
|---|---|---|---|---|
| Kubernetes namespace per tenant (pods from many tenants on the same nodes) | Nodes, kernel, cluster control plane | Policy-level: role-based access control (RBAC), network policies, quotas. A kernel or container escape crosses tenants | Highest density, lowest cost | Trusted code, noise and accident isolation |
| Containers on dedicated nodes per tenant | Cluster control plane only | A container escape reaches only that tenant's nodes | Lower density: idle node capacity per tenant | Large tenants, strong noise isolation |
| Sandboxed containers (gVisor, which puts a second, unprivileged kernel-like layer between the container and the real kernel, so even code that finds a way to misbehave is intercepted by that layer instead of reaching the host's actual kernel) or microVMs (Firecracker, Kata Containers: a tiny virtual machine per pod) | Host hardware, hypervisor | Much smaller attack surface than a shared kernel | Some per-pod overhead in memory and startup | Untrusted tenant code at scale |
| Full VMs or separate clusters per tenant | Nothing but hardware or nothing at all | Strongest | Highest cost, most to operate | Regulated or contractual isolation |
Per-tenant quotas and throttling
Every shared resource needs its own limit, because a noisy tenant will find the one you forgot.
- Admission (API edge): a per-tenant token bucket: each tenant earns request tokens at a fixed rate up to a burst size; with no token, the request gets HTTP 429 and a retry hint. Plus a per-tenant cap on concurrent in-flight requests, which catches slow expensive requests that a rate limit misses.
- Compute: a
ResourceQuotaper namespace (caps total CPU, memory and object counts a tenant can request) and aLimitRange(default and maximum per container, so nothing runs unbounded). Priority classes (labels that tell the scheduler which workloads to keep and which to kill first when a node runs short on resources) decide who is evicted first under pressure. - Data stores: per-tenant connection limits in the connection pooler (a component that shares a small, fixed number of real database connections across many requests, since the database itself can only hold so many open at once), statement timeouts, and per-tenant partitions or weighted fair scheduling (giving each tenant a guaranteed proportional share of throughput, instead of pure first-come-first-served) on shared queues.
- Network: network policies default-deny (block all traffic unless a rule explicitly allows it) traffic between tenant namespaces; egress limits stop one tenant saturating shared bandwidth.
Encryption choices
- In transit: TLS at the edge, and mTLS (mutual TLS: both sides present certificates, so each service proves its identity) between services, typically from a service mesh (an infrastructure layer that runs alongside every service to manage service-to-service traffic, including issuing and rotating the certificates mTLS needs, without changing application code). mTLS also gives every call an authenticated workload identity you can check against the tenant it claims to act for.
- At rest, per tenant: envelope encryption. Each tenant has its own key-encryption key in a key management service (KMS); data is encrypted with data keys that are themselves encrypted by the tenant's key. Benefits: a leaked backup is useless without that tenant's key, you can rotate one tenant's key, and deleting a tenant's key makes all its data unreadable (crypto-shredding), which helps with deletion requests.
- Customer-managed keys (the tenant holds the key in its own account and can revoke it) as a premium option for regulated tenants. The trade-off: if they revoke it, their service stops, and your support team must understand that.
Monitoring requirements
You cannot manage noisy neighbours you cannot see. Every metric, log line and trace carries tenant_id, and you alert on per-tenant saturation (share of quota used, throttled request rate, p99 latency by tenant: p99 is the 99th percentile, the latency 99% of requests beat).
The cardinality trap, with numbers. A request-latency histogram with 12 buckets exports 14 time series (12 buckets plus a sum and a count). Labelled by 20 endpoints and 10,000 tenants:
14×20×10,000=2,800,000 time seriesThat is from one metric. The fix: keep full-detail metrics per endpoint without the tenant label, emit per-tenant metrics only for a few coarse signals (requests, errors, throttles, CPU seconds), and put per-tenant detail in logs or traces, which are sampled and queried on demand.
Worked example: choosing rungs for a real platform
A workflow-automation platform: 3,000 tenants, 2,950 using only built-in steps, 50 enterprise tenants, and a new feature letting tenants run custom JavaScript steps.
- Built-in steps (trusted code): shared nodes, namespace per tenant, quotas, default-deny network policy, per-tenant KMS key.
- Custom JavaScript steps (untrusted code): run in a separate node pool using a sandboxed runtime or microVMs, with no network access except a proxy that enforces the tenant's allowlist, and hard CPU-time and memory caps per execution.
- 50 enterprise tenants: dedicated node pools for predictable performance and an easy story for their security reviews.
The design choice that matters: the untrusted-code pool is separated by what runs, not by who the customer is. A small tenant's custom script is as dangerous as a large one's.
Operational trade-offs between cost and isolation strength
- Each rung up the ladder lowers density: dedicated nodes waste the capacity a tenant does not use, and a cluster per tenant multiplies control-plane cost and upgrade work.
- Stronger isolation also slows operations: patching 3,000 VMs is harder than patching 40 shared node pools.
- The cheapest real improvement is usually not a stronger rung but a missing quota; most noisy-neighbour incidents come from an unbounded shared resource (connections, a queue, a cache), not from weak compute isolation.
Pitfalls
- Treating Kubernetes namespaces as a security boundary for untrusted code.
- Rate-limiting requests but not concurrency, so one tenant's slow reports exhaust a worker pool.
- Per-tenant encryption keys with a shared application role that can decrypt everything anyway. The key boundary only holds if key access is scoped per tenant too.
- What would change my recommendation: if every tenant ran custom code (a hosting or compute platform), sandboxing becomes the default rung for everyone, not an exception.
As an Engineering Manager designing SLIs/SLOs for a multi-tenant SaaS product, propose SLIs, SLOs, and isolation strategies to mitigate noisy-tenant impacts. Cover resource quotas, per-tenant rate limits, fairness algorithms, billing/SLA tiers, detection of noisy neighbors, and remediation/automation approaches.
Sample Answer
Direct answer
The mistake to avoid is measuring reliability only across the whole fleet: a fleet-wide 99.95% can hide one customer who has been completely down all month. I would define every SLI (service-level indicator: the measured quantity, such as the fraction of requests that succeed) per tenant, set SLOs (service-level objectives: the target for that SLI) per tier, and add one SLO that is about the fleet's fairness itself: "at least 99% of tenants meet their own SLO each month". Isolation mechanisms (quotas, per-tenant rate limits, fair scheduling) are what make those per-tenant SLOs achievable; detection and automated remediation are what keep them achieved; the contractual SLA (service-level agreement: the promise with financial credits attached) sits below the internal SLO so the team gets warning before money is owed.
SLIs: what to measure, and per whom
| SLI | Definition (per tenant, per window) | Why |
|---|---|---|
| Availability | good requests / valid requests; "good" = non-5xx and answered within a hard ceiling (e.g. 5 s) | The basic "can I use it" measure; excludes 4xx the tenant caused, and excludes 429s that are within the tenant's own contractual quota, but counts 429s issued while the tenant was below quota (those are our fault) |
| Latency | fraction of requests faster than a threshold, per endpoint class (e.g. under 300 ms for interactive reads) | Proportion-under-threshold turns latency into a budget; averages hide tails |
| Freshness / throughput for async work | fraction of jobs started within N seconds of submission | Queued work is where noisy neighbours hurt first |
| Isolation (fleet level) | fraction of active tenants meeting their own availability and latency SLOs | Directly measures whether noisy neighbours are winning |
| Worst-tenant view | the lowest per-tenant availability this week, with its tenant ID | Operational: shows who is hurting now |
The tenant ID has to be attached at the edge (from the authenticated credential) and carried into every metric and log. Metrics labelled by tenant are high-cardinality (many distinct label values, which is expensive in time-series databases), so compute per-tenant SLIs from request logs or a stream processor, and keep only tier-level labels in the real-time metrics system.
SLOs by tier, and the SLA beneath them
| Tier | Availability SLO | Latency SLO (interactive) | SLA to customer | Isolation level |
|---|---|---|---|---|
| Free | 99.5% | 95% under 500 ms | None | Shared pool, lowest scheduling weight |
| Standard | 99.9% | 99% under 300 ms | 99.5% with service credits | Shared pool, guaranteed quota |
| Enterprise | 99.95% | 99% under 200 ms | 99.9% with service credits | Dedicated cell or reserved capacity for the largest tenants |
The SLA is deliberately looser than the SLO, so that an SLO breach is an internal alarm with room to act before credits are due. Billing tiers and isolation tiers must match: a tenant paying for 99.95% cannot sit in the same unprotected pool as the free tier, or the promise is not backed by anything.
Isolation strategies that make per-tenant SLOs possible
- Resource quotas: caps per tenant on concurrent requests, queued jobs, storage and compute (e.g. in Kubernetes, a
ResourceQuotaper tenant namespace; in the application, a concurrency limit per tenant). - Per-tenant rate limits: a token bucket (an allowance that refills at a fixed rate and permits short bursts) keyed by tenant at the API gateway, sized by tier.
- Fairness algorithms: weighted fair queuing (giving each tenant's queue a share of capacity in proportion to its weight, in every round) or deficit round robin (a scheduler that visits each tenant's queue in turn and lets it send an amount proportional to its weight before moving on) across per-tenant queues, so capacity divides by tier weight when busy but idle capacity is still lent out.
- Placement: shard (a partition holding a subset of tenants' data or traffic, typically one database or node) or cell (an independent copy of the stack serving a subset of tenants) assignment that keeps very large tenants off the cells of small ones, so the blast radius (how many tenants one problem can reach) stays small.
Worked example: why fleet-wide SLOs hide the problem
Suppose 1,000,000,000 requests a month across all tenants. One tenant sends 0.05% of that traffic, 500,000 requests, and every single one fails for the month.
failedfleet availability=109×0.0005=500,000=1−500,000/109=99.95%The fleet SLO of 99.9% is met with room to spare, while that customer had 0% availability. The per-tenant SLI flags it within the first hour.
Per-tenant error budgets. An error budget is the amount of unreliability an SLO still allows: the gap between 100% and the SLO target, spent as outage minutes or failed requests. A 99.9% SLO over 30 days allows 0.1% of the window as bad: 30 × 24 × 60 × 0.001 = 43.2 minutes of full outage, or equivalently, for a tenant sending 2,000,000 requests a month, 2,000 failed requests.
Small tenants. A tenant sending 500 requests a month breaches 99.9% with a single failure (1/500 = 0.2% bad). Two fixes: evaluate small tenants over a longer window or pooled by tier, and require a minimum request count before a per-tenant SLO can page anyone. Otherwise the on-call engineer is paged by noise.
Detecting noisy neighbours
- Share of resource: each tenant's share of CPU seconds, database time, queue slots or bytes scanned, computed per minute. Flag a tenant whose share exceeds its tier entitlement while other tenants' latency SLIs are degrading. Both conditions together, because a tenant using idle capacity harms no one.
- Correlation: when a shard's p99 (99th-percentile) latency rises, rank tenants on that shard by the increase in their resource use over the same minutes.
- Burn-rate alerts per tier: alert when the error budget is being spent fast. The sustainable rate is the pace that would use up exactly the whole budget spread evenly across the 30-day window, that is, 1/720 of the budget per hour, where 720 is simply the number of hours in 30 days (30 x 24 = 720). A rate of 14.4 times that (for example 14.4 times the sustainable rate over one hour, which spends 14.4 / 720 = 2% of a 30-day budget in that hour) is fast enough that, sustained, it would use up the whole month's budget in about 720 / 14.4 = 50 hours (roughly two days), which is worth paging on within the hour rather than waiting for a slow drift to show up in a weekly review.
Remediation and automation
Automate the reversible, low-risk actions and keep humans for the rest:
- Automatic: throttle the offending tenant to its tier entitlement (tighten its rate limit or concurrency cap), lower its scheduling weight, and move its heavy jobs to a batch pool. These are reversible and already contractually allowed.
- Automatic with a notification: move the tenant to a different shard or a quarantine cell when throttling is not enough.
- Human decision: permanently upgrading a tenant's tier or dedicated capacity (a commercial conversation), and anything that could lose data.
Every automated action writes an event linked to the tenant, so the account team sees "your workload was throttled at 14:05 because..." rather than an unexplained slowdown.
Trade-offs and pitfalls
- Too many SLOs dilute attention. Keep two per tier (availability, latency) plus the fleet isolation SLO; track the rest as indicators, not objectives.
- Excluding all 429s from availability lets the platform hide its own capacity problems behind rate limiting. Only exclude throttling the tenant brought on itself by exceeding its quota.
- Per-tenant metrics in the real-time metrics store can explode cost; derive them from logs.
- SLA equal to SLO means every internal miss costs money with no warning margin.
- Automation without a kill switch: a bad detector that throttles your largest customer is its own incident. Rate-limit the remediation itself and make it easy to disable.
That is every published Multi-Tenancy and Isolation question for Engineering Manager so far. Browse the other topics in this category, or practice this one interactively.