Multi-Tenancy and Isolation Questions
Serving many tenants from shared infrastructure: tenancy models (silo, pool, bridge), data isolation, per-tenant data residency, noisy-neighbor mitigation, per-tenant limits, and security boundaries between tenants. Covers the cost, isolation, and blast-radius tradeoffs of shared versus dedicated resources, and business continuity: per-tenant backup, disaster recovery, and compliant tenant offboarding and deletion. The architecture layer specific to SaaS and platform products.
Architect a multi-tenant microservices platform serving 10,000 tenants worldwide that requires per-tenant isolation (compute and data), per-tenant cross-region failover, and cost transparency to tenants. Discuss tenant isolation models (logical vs physical), deployment strategies (shared cluster vs dedicated cluster), noise isolation, billing implications, and operational tooling needed.
Sample Answer
Direct answer
At 10,000 tenants worldwide I would build a cell-based architecture: the platform is stamped out as many independent, identical "cells" (each a Kubernetes cluster plus its own databases, hosting a few hundred tenants), placed in several regions. Small tenants share a cell with per-tenant namespaces, quotas and a per-tenant database; large tenants get a dedicated node pool (a set of machines reserved just for their workloads) or a dedicated cell. Every tenant has a home cell and a paired standby cell in another region, and a global tenant directory says which is active, so failing over one tenant means replicating its data and flipping one directory entry, not moving a whole region. Per-tenant metering runs from day one because cost transparency is a product feature here, not an internal report.
Requirements and the forces in tension
- Isolation of compute and data per tenant. Strong isolation pushes towards dedicated clusters; 10,000 dedicated clusters is operationally impossible.
- Per-tenant cross-region failover. A tenant (not just the whole platform) must be movable to another region, with a stated RPO (recovery point objective: how much recent data you may lose) and RTO (recovery time objective: how long until service is back).
- Cost transparency to tenants. Each tenant sees what it consumed and what it is billed, which means every CPU-second, byte and request needs a tenant label.
- Worldwide. Latency and data-residency rules decide the home region.
Isolation model: logical vs physical, and where each applies
| Tier | Share of tenants | Compute | Data | Why |
|---|---|---|---|---|
| Small | ~98% | Namespace per tenant in a shared cell, with CPU and memory quotas | Own logical database (or schema) on the cell's shared database servers, own encryption key | Cheap, fast onboarding, a real per-tenant boundary for data |
| Medium | ~1.8% | Dedicated node pool inside a shared cell (their pods run only on their nodes) | Own database server | Removes CPU and cache contention without a whole cluster to run |
| Large | ~0.2% | Dedicated cell | Dedicated database cluster | Contractual isolation, very large load |
A Kubernetes namespace is a named partition of one cluster; with quotas, network policies and per-tenant service accounts (a non-human identity a tenant's own workloads authenticate as, scoped so one tenant's pods cannot use another tenant's permissions) it isolates well against accidents but still shares the node's kernel, so untrusted tenant code would need a stronger sandbox. Physical isolation (dedicated nodes or clusters) is what you sell to tenants who need it.
Deployment strategy: shared cluster vs dedicated cluster
I would not run one giant shared cluster: a bad upgrade, an overloaded API server (the Kubernetes control-plane component every cluster-management request goes through) or a misconfigured network policy would hit all 10,000 tenants. Cells cap the blast radius (how many tenants one failure reaches).
Sizing, with the assumptions stated: 9,800 small and 180 medium tenants share cells, 20 large tenants get their own. At 250 tenants per shared cell:
⌈2509,800+180⌉=⌈39.92⌉=40 shared cellsSpread over 4 regions, that is 10 shared cells per region, so one bad cell touches at most 250 tenants (2.5% of the base).
flowchart TB
U[Tenant request] --> E[Global edge + tenant directory]
E -->|active| A[Home cell, region A]
E -.->|on failover| B[Standby cell, region B]
A -->|per-tenant async replication| B
A --> M[Metering pipeline]
B --> M
M --> BL[Per-tenant usage + bill]
Per-tenant cross-region failover
- Placement: the directory stores each tenant's home cell, standby cell and state (
ACTIVE,FAILING_OVER,ACTIVE_ON_STANDBY). - Data: each tenant's database replicates asynchronously to its standby cell. Asynchronous means the RPO is the replication lag (typically seconds); synchronous cross-region replication would add a cross-region round trip (tens of milliseconds between nearby regions, well over 100 ms between continents) to every write, which I would offer only as a premium tier.
- Compute: the standby cell keeps the tenant's deployment defined but scaled low; failover scales it up.
- Procedure: fence the old primary so it cannot accept a write once failover has started. Two ways to do that: revoke its write credentials at the database, or bump the tenant's epoch, a version number stored in the tenant directory that increments on every failover. Every write is required to carry the current epoch, and the datastore or proxy checks that number against the directory's latest value before accepting the write; once the epoch is bumped, any write still in flight from the old primary carries the now-stale epoch and is rejected, even though the old primary itself never learns the failover happened. Then promote the replica (make the standby's database copy the new primary and start accepting writes), flip the directory entry, and let DNS or the edge route pick it up. Because it is per tenant, the same mechanism doubles as a region-move tool for residency changes.
Capacity arithmetic. If a whole region fails and its tenants spread over the other 3 regions, each remaining region takes on one third of a region's load, reaching 4/3 of its normal load. For that to fit, normal utilisation must be at most 3/4 = 75% of capacity per region.
Replication bandwidth. If the average tenant writes 2 GB a day: 10,000 x 2 GB = 20 TB a day, and 20 x 10^12 bytes / 86,400 s = about 231 MB/s, or about 1.85 Gbit/s of cross-region traffic in total, which is also a line item on the bill (inter-region transfer is charged per GB on the major clouds).
Noise isolation
- Admission: per-tenant rate limits and concurrency limits at the edge, sized by plan.
- Compute: namespace resource quotas (caps on total CPU and memory a namespace may request) and default limits for every pod, so no tenant can schedule unbounded work.
- Data: per-tenant connection caps and statement timeouts; heavy tenants moved to their own database server when their share of a server crosses a threshold.
- Shared queues: per-tenant partitions or weighted fair queuing (giving each tenant a guaranteed proportional share of a shared queue's throughput, instead of pure first-come-first-served) so a tenant's backlog does not delay others.
Billing implications and cost transparency
Tenants see showback (a usage and cost breakdown) on a dashboard, and billing on the invoice. To make these match:
- Label every workload with
tenant_idat deploy time; the metering pipeline rejects unlabelled resources. - Measure per tenant: CPU and memory reserved versus used per pod, storage bytes per database, requests and egress bytes at the gateway, replication bytes.
- Charge dedicated resources directly; split shared ones (control plane, cell overhead, idle capacity) by a published rule, for example proportional to reserved CPU.
- Reconcile monthly: the sum of all tenant charges plus the platform's own share must equal the cloud bill.
Two billing consequences of this design: the standby is not free (a warm standby, kept running at reduced scale so it can take over quickly, as described above, still has storage and replication traffic, and that belongs on the tenant's bill, which is why cross-region failover is usually a paid tier), and dedicated tiers carry a visible minimum charge because their idle capacity cannot be shared.
Operational tooling needed
- Tenant directory and control plane (the small set of services that make decisions about tenants, placement, routing, failover, as opposed to the data plane, which actually carries each tenant's traffic) for onboarding, placement, tier changes and failover.
- Cell factory: infrastructure-as-code to stamp identical cells, plus a staged rollout that upgrades one cell at a time.
- Per-tenant observability: dashboards and alerts filtered by tenant, with care for label cardinality (the number of distinct time series; 10,000 tenant labels on every metric is expensive, so keep per-tenant metrics coarse).
- Failover runbooks and game days (scheduled practice failure drills, on purpose, not waiting for a real incident): fail over a test tenant every week, measure its RTO and RPO, and publish them.
- Tenant migration tooling: move a tenant between cells without downtime (copy, catch up, fence, flip).
Trade-offs and pitfalls
- Cells multiply fixed cost. 40 control planes (one management layer per cell), 40 sets of monitoring. The alternative (one huge cluster) is cheaper and fails for everyone at once. I would accept the overhead.
- Global dependencies defeat cells. If the tenant directory or identity service lives in one region, a regional outage still takes everyone down. Replicate the directory to every region and let cells run on a cached copy.
- Failover without fencing causes split brain (two copies both accepting writes and diverging). Always fence first.
- What would change the design: if tenants were mostly large enterprises, I would make "dedicated cell" the default and focus on cell automation instead of packing density.
Design a secure multi-tenant cloud environment providing strong tenant isolation across network, compute, storage, IAM and logging. Compare the account-per-tenant model vs shared-VPC/namespace model, including operational overhead, cost, and security trade-offs.
Sample Answer
Direct answer
I would run most tenants in a shared-account model (pooled VPCs, namespaces or schemas, with per-tenant IAM roles, keys and logging) and offer account-per-tenant as a tier for regulated or very large tenants. Account-per-tenant gives the strongest boundary a cloud provider offers (a separate AWS account is a hard wall for IAM, quotas, billing and blast radius) but carries a fixed cost and operational load per tenant that only pays off above a certain tenant size or compliance need. Either way the model is only viable with fully automated provisioning and tenant-level observability and billing built in from the start.
Terms: a VPC (Virtual Private Cloud) is a private network in a cloud account; IAM (AWS Identity and Access Management) controls who can do what; blast radius is how much is affected when something goes wrong. A NAT gateway is what lets resources with no public address reach the internet (for updates, external APIs) without being reachable from it; a VPC typically needs at least one, which is why it is a fixed cost per account below. A security group is a virtual firewall attached to a resource that allows or denies traffic by port and source. A service mesh is a layer that runs alongside every service to manage service-to-service traffic, including issuing the certificates mTLS needs, without changing application code. A principal is the identity making a request, a user or a role, the subject IAM policies are written about.
The two models compared, layer by layer
| Layer | Account-per-tenant | Shared account (shared VPC / namespace) |
|---|---|---|
| Network | Separate VPC per tenant; no route between tenants unless you build one | Shared VPC; separation by subnets, security groups, Kubernetes NetworkPolicy (rules controlling which pods may talk to which), service mesh mTLS (mutual TLS, both sides authenticate) |
| Compute | Tenant's own instances, clusters or functions | Shared clusters with namespaces, quotas, and optional dedicated node pools |
| Storage | Separate buckets and databases in the tenant's account; per-tenant KMS (key management service) keys | Shared buckets with per-tenant prefixes, or shared databases with row-level security (RLS: the database filters rows by tenant) or schema per tenant; per-tenant keys still possible |
| IAM | Account boundary itself is the control; a misconfigured policy in one account cannot grant access to another's resources | Per-tenant roles and attribute-based policies (tags such as tenant_id on principals and resources); one wrong wildcard can cross tenants |
| Logging | Per-account logs, aggregated to a central security account | Shared log pipeline; every record must carry tenant_id and access must be filtered per tenant |
| Quotas and noisy neighbours | Service quotas are per account, so one tenant cannot exhaust another's API rate limits | All tenants share one account's quotas (for example KMS request rates, Lambda concurrency, the number of function invocations allowed to run at once for the whole account) |
| Billing | Consolidated billing (AWS rolling every member account's usage into one payer-account bill) gives an exact per-account bill | Needs cost allocation tags (billing tags that make spend sliceable by a tag's value, like tenant_id) plus metering for shared resources |
| Operational overhead | Hundreds of accounts to patch, baseline and monitor | One environment; changes roll out once |
Worked cost example
Assume 500 tenants and an assumed per-account baseline of about 150 US dollars a month (NAT gateways, per-account security services such as threat detection and config recording, log storage, idle minimums). This baseline is an illustrative assumption; measure your own from a pilot account.
- Account-per-tenant fixed overhead: 500 x 150 = 75,000 dollars a month before any tenant does real work.
- Shared model: that baseline is paid a handful of times (per environment and region), so it is close to negligible per tenant.
If the average tenant pays 300 dollars a month, a 150-dollar baseline is half their revenue, which settles the question for the long tail. If the enterprise tier pays 20,000 dollars a month, 150 dollars is under 1% and the stronger boundary is easy to justify.
Security trade-offs
- Account-per-tenant turns a class of bugs (overly broad IAM policy, wrong bucket prefix) from a cross-tenant breach into a same-tenant bug. It also makes "delete everything for tenant X" as simple as closing an account.
- Shared depends on every layer enforcing the tenant boundary correctly: every query filtered, every policy scoped, every log tagged. Defence in depth is mandatory: tenant context injected at the edge, RLS in the database, per-tenant keys, and automated tests that attempt cross-tenant access.
- Shared control plane risk exists in both: the provisioning pipeline and central security account can reach every tenant, so they need the strictest access and audit.
Provisioning automation
Neither model works by hand. For account-per-tenant, use an account vending pipeline: AWS Organizations to create accounts under an organizational unit (a folder that groups accounts so a policy can apply to all of them at once), service control policies (SCPs, organization-wide guardrails that cap what any role in the account can do) applied automatically, and either Control Tower Account Factory (AWS's built-in service for provisioning a new account with a standard baseline already applied) or Terraform (a general infrastructure-as-code tool: the baseline is written as code and applied the same way every time) to lay down the baseline (VPC, logging to the central account, KMS keys, IAM roles). For the shared model, the same pipeline creates the namespace, database schema or RLS policies, IAM role, key, and quota. Onboarding should be one API call, idempotent, and reversible (offboarding runs the same pipeline backwards).
flowchart LR
REQ[Tenant signup] --> TIER{Tier}
TIER -->|enterprise / regulated| AV[Account vending: Organizations + SCPs + baseline]
TIER -->|standard| NS[Shared account: namespace, schema, role, key, quota]
AV --> REG[(Tenant registry)]
NS --> REG
REG --> OBS[Per-tenant observability + billing]
Tenant-level observability and billing
- Every metric, log and trace carries
tenant_id; dashboards filter per tenant. - In the shared model, tag every resource with
tenant_idand activate it as a cost allocation tag; meter shared resources (CPU-seconds, requests, storage) per tenant and allocate the shared bill proportionally. - In account-per-tenant, the account bill is the tenant's bill, which is simple and exact.
Recommendation and what would flip it
Shared by default, account-per-tenant for regulated, contract-driven, or very large tenants, with the same provisioning API producing both. I would move more tenants to their own accounts if most revenue came from a few hundred large customers, or if regulators or customer contracts demanded a provable hard boundary.
Pitfalls
- Choosing account-per-tenant without automation: the fleet drifts and becomes unpatchable within a year.
- Choosing shared and relying on application code alone for isolation.
- Forgetting shared quotas: one tenant's batch job exhausting an account-level API limit is a noisy-neighbour outage for everyone.
Define multi-tenancy in the context of data platforms. Describe and compare the common isolation models: shared-schema (single table with a tenant_id column), shared database with a separate schema per tenant, a separate database per tenant, and fully isolated infrastructure (separate clusters or accounts). For each model, cover security, operational cost, scalability, backup and restore granularity, schema evolution and index design, and onboarding time for new tenants, and give a short example of when you'd pick it.
Sample Answer
Direct answer
Multi-tenancy means one platform (one codebase, one operations team, often shared hardware) serves many customers, called tenants, while each tenant sees only its own data and gets predictable performance. The isolation models form a spectrum: the more you share, the cheaper and faster to onboard, but the weaker the security boundary and the coarser your control over any single tenant. On a data platform I default to the shared model for the long tail of small tenants and reserve heavier isolation for tenants whose contracts, regulators or workloads demand it.
The spectrum in plain terms
Think of an apartment building. Shared-schema is everyone's mail in one big sorted pile, with a label on each letter. Schema-per-tenant gives each flat its own mailbox in the same lobby. Database-per-tenant gives each flat its own locked room. Separate storage and compute (the fourth tier below) gives each flat its own room and its own lift. Fully isolated infrastructure is a separate building per tenant.
- Shared schema (the "pool" model): every tenant's rows live in the same tables, distinguished by a
tenant_idcolumn. Every query must filter on it. - Shared database, schema per tenant: one database server; each tenant gets its own namespace of tables (
tenant_42.orders). - Database per tenant: each tenant gets its own logical database, possibly on shared servers.
- Separate storage buckets and compute warehouses per tenant: common on modern data platforms that separate storage from compute. Each tenant's files live in its own object-storage bucket (or prefix with its own encryption key), and its queries run on its own compute cluster (a Snowflake "virtual warehouse", for example, is a named compute cluster you can size and bill independently). Control plane and catalog are still shared.
- Fully isolated infrastructure (the "silo" model): a separate cluster, cloud account or subscription per tenant. Nothing but the deployment pipeline is shared.
The literature also names a "bridge" model: a mix, where some layers are shared and others are per tenant. Schema-per-tenant and the storage-plus-warehouse tier are both bridge designs.
Comparison across the six dimensions
| Model | Security boundary | Operational cost | Scalability | Backup and restore granularity | Schema evolution and index design | Onboarding time |
|---|---|---|---|---|---|---|
Shared schema + tenant_id | Weakest: one missing filter leaks data. Harden with database row-level security (RLS, where the database itself appends the tenant filter) | Lowest: one set of tables to run and monitor | Best for many small tenants; shard by tenant_id (split the data across multiple database servers, consistently routing each tenant to one shard by its tenant_id) once a single database fills up | Coarse: backups are whole-database, so restoring one tenant means restoring to a side copy and copying that tenant's rows back | One migration covers everyone. Every index should lead with tenant_id (for example (tenant_id, created_at)) so queries touch one tenant's slice | Seconds: insert a tenant row |
| Schema per tenant | Moderate: database permissions per schema, but same server and same superuser (an account with unrestricted access to every schema on the server, the one boundary schema-per-tenant does not divide) | Medium: migrations and monitoring multiply by tenant count | Catalog bloat (the database's own internal table of tables grows by tenant count times table count, and a very large catalog slows down planning and admin queries) past a few thousand schemas (every table counted per tenant) | Per-schema logical dump (an export of the schema's data and structure as portable files, not a copy of the raw on-disk storage) and restore is practical | Migrations run N times and can half-fail; indexes can be tuned per tenant | Seconds to minutes: create schema and run migrations |
| Database per tenant | Strong logical boundary: separate credentials, can use a per-tenant encryption key | Higher: connection pools, backups and upgrades per database | Tenants can be moved between servers individually | Clean: point-in-time restore (restoring to exactly how the database looked at a chosen moment, using continuous transaction logs, not just the last nightly backup) of one tenant without touching others | Per-database migrations; drift (tenants silently ending up on different schema versions when a migration partially fails or is skipped for one database) between tenants is the main risk | Minutes: provision database, apply schema |
| Separate storage bucket + compute warehouse | Strong for data at rest (bucket policy and key per tenant); shared catalog and control plane (the shared management layer tracking table definitions, permissions and job scheduling across every tenant, as opposed to the data itself) | Pay per tenant for compute, but compute can auto-suspend when idle | Excellent: a heavy tenant's queries run on its own compute, so they cannot slow others | Per bucket: versioning or snapshot per tenant | Shared table definitions in the catalog, per-tenant physical layout (partitioning, clustering keys) | Minutes: bucket, key, warehouse, grants via automation |
| Fully isolated infrastructure | Strongest: separate account or cluster, separate network, blast radius of one tenant | Highest: full stack per tenant, plus a fleet to patch | Scales per tenant, but the fleet itself becomes the scaling problem | Whole environment per tenant, trivially isolated | Every environment upgraded separately; versions drift unless automation is strict | Hours to days unless fully automated |
Blast radius in the table means how many tenants one failure or breach can affect.
Worked example: when each wins
A B2B (business-to-business) analytics product with 2,000 customers:
- Shared schema: 1,900 self-serve customers on a free or starter plan, each storing a few gigabytes. Per-tenant infrastructure would cost more than they pay. Put them in shared tables with RLS, with
tenant_idleading every index. - Schema per tenant: a mid-market segment that wants custom columns or per-tenant index tuning, where the count stays in the hundreds today. Worth sizing the operational cost this segment would hit at scale, though: if one migration takes 3 seconds per schema, migrating a fleet that grew to 2,000 tenant schemas would take 2,000 x 3 = 6,000 seconds, 100 minutes of serial migration, and any failure in the middle leaves the fleet on mixed versions.
- Database per tenant: a healthcare customer that needs its own encryption key and a restore of its data to yesterday at 14:00 without touching anyone else.
- Separate bucket + warehouse: a retail customer running heavy nightly scans that would starve others. Give it its own compute warehouse, billed to it directly, while the catalog stays shared.
- Fully isolated infrastructure: a bank or government agency whose contract requires a separate cloud account, its own network and an audit showing no shared compute.
Trade-offs and pitfalls
- Pick per tier, not per platform. Most real systems run two or three of these models at once and route each tenant by plan. The mistake is choosing one model for everyone.
- Shared schema is only safe with enforcement below the application. An ORM (object-relational mapping library) filter is one forgotten
WHEREaway from a leak; RLS or a tenant-scoped data-access layer makes the filter impossible to skip. - Indexes without
tenant_idfirst force the database to scan every tenant's rows to answer one tenant's query, and a unique constraint onemailalone would stop two tenants from having the same user. - Restore granularity is usually discovered during an incident. Decide up front how you will restore one tenant in the shared model (a side restore plus a scripted copy of that tenant's rows) and rehearse it.
- Moving a tenant between models later (shared to dedicated when it grows) is a data migration. Design
tenant_idinto every table and every key from day one so the move is a copy, not a re-model.
Architect a multi-tenant SaaS platform to support 10K tenants with widely varying usage patterns over 3-5 years. Describe tenant isolation models (shared schema, separate schema, dedicated infra), noisy-neighbor mitigation, autoscaling and operational trade-offs including pricing and support implications.
Sample Answer
Direct answer
For 10,000 tenants with very uneven usage, I would build a tiered, cell-based platform: the long tail of small tenants shares pooled infrastructure (shared schema with a tenant_id column, enforced by database row-level security, RLS: the database itself adds the tenant filter to every query, so application code cannot forget it), mid-size tenants get their own schema or database inside the same cells, and a small number of enterprise tenants get dedicated infrastructure deployed from the same code. Tenant identity is resolved once from the login token and carried through every API call, query, cache key and message. The tier a tenant sits in is a product decision tied to price and SLA (service-level agreement, the contractual uptime and support promise), so pricing, support and operations are designed together with the architecture rather than after it.
Requirements and numbers I am designing to
- Year 1: 10,000 tenants, about 2,000 requests per second (RPS) at peak. Year 5: plan for up to 100,000 tenants and 10,000 RPS at peak.
- Usage follows a power law: a few tenants generate a large share of load, most are small (illustrative shape: the top 1% of tenants might generate 40% of all requests, while the bottom half generate barely 5%).
- Some customers will contractually require dedicated resources, their own encryption keys, or per-tenant restore.
Assumed tier mix (to be replaced by real sales data): 94% self-serve, 5% business, 1% enterprise.
| Year 1 (10,000 tenants) | Year 5 (100,000 tenants) | |
|---|---|---|
| Pooled (self-serve) | 9,400 | 94,000 |
| Bridge (business: own schema or database) | 500 | 5,000 |
| Silo (enterprise: dedicated stack) | 100 | 1,000 |
The last row is the warning in this table: 1,000 dedicated stacks is a fleet-management problem, so the silo tier needs a price floor and full automation, not bespoke setups.
Isolation models and where each tier lives
- Pool (shared schema): all tenants in the same tables, every row tagged with
tenant_id, every index leading withtenant_id, PostgreSQL row-level security (RLS: the database applies the tenant filter to every query) as the safety net. Cheapest per tenant; weakest boundary. - Bridge (separate schema or database per tenant, shared compute cells): per-tenant backup, restore and encryption key, while app servers stay shared.
- Silo (dedicated infrastructure): separate cloud account or cluster per tenant, deployed from the same templates and the same release train (a fixed, recurring deployment schedule every tenant's stack rides, rather than ad hoc per-tenant releases).
Architecture: cells
A cell is a complete, independent copy of the application stack (app servers, database, cache, queues) that serves a fixed set of tenants. A thin global routing layer maps each tenant to its cell. Cells cap the blast radius (how many tenants one failure or bad deploy can hurt) and are the unit of scaling: when cells fill, add a cell rather than growing one giant database.
flowchart LR
U[Client] --> R[Global router: tenant from token]
R --> D[(Tenant directory: tenant to cell, tier, limits)]
R --> C1[Pooled cell 1]
R --> C2[Pooled cell N]
R --> B1[Bridge cell: schema or DB per tenant]
R --> S1[Silo: enterprise tenant stack]
C1 --> O[Metering and cost attribution]
B1 --> O
S1 --> O
Cell sizing arithmetic. Suppose load testing shows one cell handles 1,500 RPS and we run cells at 60% of that for headroom (900 RPS each):
- Year 1: 2,000 / 900 = 2.2, so 3 pooled cells.
- Year 5: 10,000 / 900 = 11.1, so 12 pooled cells.
The 1,500 RPS is an assumption to replace with a real load test; the method is the point.
Tenant context in the API design
- The client authenticates; the signed token carries the tenant ID. The router never accepts a tenant ID from a URL, header or body as authoritative.
- The router looks up the tenant's cell, tier and limits in the tenant directory and forwards the request with a tenant-context header signed by the router.
- Inside the cell, middleware sets the tenant context once per request: the database transaction sets it (
SET LOCAL app.tenant_id, a Postgres setting scoped to the current transaction only;LOCALmatters because it resets automatically when the transaction ends, so it cannot leak into the next request that reuses the same pooled connection for a different tenant, the way a plainSETwould), the cache wrapper prefixes keys with it, queue messages carry it, logs and metrics are tagged with it. - Resource IDs are looked up as
(tenant_id, id), never byidalone, so a guessed ID from another tenant returns 404.
Noisy-neighbour mitigation
- Edge: per-tenant token-bucket rate limits by plan (steady rate plus a burst allowance), returning HTTP 429 with a retry hint.
- Application: per-tenant concurrency caps on expensive endpoints; background jobs in per-tenant queues served fairly instead of one first-in first-out queue.
- Database: per-tenant statement timeouts and connection caps; reporting traffic sent to replicas.
- Placement: a tenant that stays hot for weeks is moved to a less loaded cell, a bridge database or its own cell. Moving tenants between cells has to be a routine, tested operation, because the power law guarantees you will need it: copy the tenant's rows into the new cell's database while the old cell keeps serving it; dual-write (send every write to both the old and new database while the copy catches up) for a large tenant, or, for a smaller one, a brief write freeze (pause that tenant's writes only, for seconds, once the copy is nearly caught up, simpler than dual-write at the cost of a short window where the tenant cannot write); then flip the directory entry (update the tenant's row in the tenant directory to point at the new cell, so the router sends its next request there) once the copy is verified complete.
Autoscaling
- App tier scales horizontally (adds more instances rather than making one instance bigger) inside each cell on CPU and request latency.
- Databases do not autoscale quickly, so capacity is managed by placement: new tenants go to the cell with the most headroom, and a new cell is provisioned when the fleet crosses about 70% of target utilization.
- Silo stacks scale to the enterprise tenant's contracted size and are billed for it.
Backup and restore per tenant
- Pooled: backups are per cell. A single-tenant restore means restoring the cell's backup to a side database and copying that tenant's rows back by
tenant_id. Script and rehearse it quarterly; it is the most common enterprise request. - Bridge and silo: native per-database point-in-time restore.
- Restore capability is sold as a tier feature, because it genuinely costs more to provide.
Upgrades and time-to-patch
Every silo is another thing to patch. Suppose a patch takes 30 minutes per silo and the pipeline patches 10 in parallel:
- 100 silos: 100 / 10 = 10 batches, x 30 minutes = 300 minutes = 5 hours.
- 1,000 silos: 1,000 / 10 = 100 batches, x 30 minutes = 3,000 minutes = 50 hours.
Pooled cells are patched in waves (for example 12 cells in 4 waves with a bake period between them, a pause after each wave to watch for problems before touching the next one), so the pooled estate's patch time grows with the number of cells, not tenants. That is the operational argument for keeping the silo tier small and fully automated: one release train, no per-tenant forks, and contracts that allow maintenance windows.
Pricing and support implications
- Tier price follows isolation cost. A silo has a fixed monthly floor (dedicated database, cluster, network), so its price must start above that floor plus the support cost.
- Metering and billing are part of the architecture. Record requests, storage and compute seconds per tenant from day one; you need them to set prices, produce usage-based invoices, detect noisy tenants and justify upgrades.
- Support tiers map to technical capability: guaranteed restore time, dedicated capacity, customer-managed keys and change windows are only promised on tiers that have them.
Who owns what
| Team | Owns |
|---|---|
| Product | Tier definitions, SLAs, pricing and which features each tier gets |
| Engineering | Tenancy model implementation: tenant context propagation, RLS policies, cell routing, autoscaling, tenant migration tooling |
| SRE / platform (site reliability engineering) | Quota and rate-limit enforcement, cell capacity planning, per-tenant dashboards, incident runbooks (written step-by-step procedures an on-call engineer follows during an incident) for noisy tenants and cell failures |
| Security and compliance | Isolation audits and penetration tests (authorized, simulated attacks used to find exploitable isolation gaps before a real attacker does), encryption key management including per-tenant keys, access reviews for admin tools |
| Support | Communicating SLAs to customers, handling restore requests and quota-increase requests, escalating repeat noisy tenants for tier changes |
Trade-offs and pitfalls
- I commit to pool-first because 94% of tenants cannot pay for dedicated infrastructure. If most revenue came from a few hundred regulated enterprises, I would flip to silo-first with heavy automation.
- Cells add routing and migration complexity but bound the blast radius; one giant shared database is simpler until the day it becomes everyone's outage.
- Pitfalls: tenant ID trusted from the client;
tenant_idmissing from cache keys or indexes; per-tenant custom code branches that make every upgrade bespoke; no rehearsed single-tenant restore; and no metering, which leaves pricing guesswork.
Compare cloud network isolation options to enforce tenant separation: VPC per tenant, VPC peering, PrivateLink/Private Service Connect, and service-mesh-level isolation. For a data platform that serves thousands of tenants, discuss scalability, cost, operational overhead, and security.
Sample Answer
Direct answer
For a data platform with thousands of tenants I would not use a VPC per tenant or peering as the default boundary: both hit hard quota and routing limits long before a few thousand tenants, and they multiply operational work. I would run the multi-tenant data plane in a small number of shared networks, enforce tenant separation inside them with service-mesh identity and authorization (every call carries a cryptographic workload identity, and policy decides which tenant's resources it may reach), expose the platform to tenants' own networks through PrivateLink (AWS) or Private Service Connect (PSC, the Google Cloud equivalent), and reserve a dedicated VPC in a dedicated account for the few tenants whose contracts or regulators require network-level separation.
The four options, defined
- VPC per tenant. A VPC (Virtual Private Cloud) is an isolated private network in a cloud account. Giving each tenant one means their resources share no network at all with other tenants.
- VPC peering. A private, non-transitive link between two VPCs: A peered with B and B peered with C does not let A reach C. Both sides must have non-overlapping IP address ranges.
- PrivateLink / Private Service Connect. The provider publishes a service behind a load balancer; a consumer creates an endpoint (specifically an interface endpoint: a private IP in the consumer's own network) that reaches only that service. Traffic is one-directional (consumer to service), and the two networks' IP ranges may overlap because no routes are exchanged.
- Service-mesh isolation. A service mesh (e.g. Istio) puts a proxy beside every workload (a sidecar) or in the node, which enforces mTLS (mutual TLS: both sides prove their identity with certificates, issued and vouched for by a certificate authority, the trusted party whose signature both sides check) and per-request authorization policies. The boundary is identity-based, at layer 7 (the request level), rather than network-address-based.
- Networking terms used in the comparison below. A security group is a virtual firewall attached to cloud resources that allows or denies traffic by IP address and port, an address-based control, unlike the mesh's identity-based policy. An Availability Zone (AZ) is one of several physically separate data centers inside a cloud Region; spreading a resource across AZs (for example, one endpoint per AZ) protects against a single data center failing, which is why the cost example below multiplies by 3. A hub (in "hub to tenant VPCs") is a central VPC that every tenant VPC peers with individually, rather than tenants peering with each other. A mesh's control plane is the shared component that distributes identity, certificates and policy to every proxy: it is shared infrastructure, so its own capacity, not a per-network quota, is what limits how many workloads it can serve. Egress is traffic leaving a network for a destination outside it, the opposite of ingress. Kubernetes NetworkPolicy is a Kubernetes-native rule that allows or denies traffic between pods by label and namespace, enforced by the cluster's own networking rather than by a sidecar proxy.
Comparison for thousands of tenants
| VPC per tenant | VPC peering (hub to tenant VPCs) | PrivateLink / PSC | Service mesh in shared VPCs | |
|---|---|---|---|---|
| Scalability | AWS defaults to 5 VPCs per Region per account, raisable to hundreds; thousands of tenants means hundreds of accounts (at the default quota: 3,000 / 5 = 600 accounts) | AWS caps active peerings per VPC at 50 by default, 125 maximum, so one hub cannot reach thousands of tenant VPCs; every tenant needs a non-overlapping range | One endpoint service serves many consumers; overlapping tenant IP ranges are fine | Scales with the mesh control plane, not with network quotas |
| Cost | A full network stack per tenant (network address translation (NAT) gateways for outbound traffic, endpoints, load balancers) repeated thousands of times | Peering has no hourly charge (normal data-transfer charges still apply); the real cost is operational | Per endpoint, per Availability Zone, per hour, plus per GB processed (see example) | Proxy CPU and memory overhead on every workload |
| Operational overhead | Very high: thousands of networks to patch, monitor and route | High: IP address planning across thousands of ranges, route tables, non-transitivity workarounds | Moderate: endpoint lifecycle per tenant connection | Moderate to high: certificate rotation, policy management, mesh upgrades |
| Security boundary | Strongest: no shared network | Network-level, but peering exposes whole address ranges unless security groups are tight | Narrow: the consumer can reach exactly one service and nothing else | Strong if policy is default-deny; relies on correct identity and policy, not network separation |
| Tenant-facing use | Dedicated deployments | Rarely a good fit for customer networks (IP overlap, trust) | Best fit for connecting tenants' own networks to the platform privately | Internal only; tenants never see it |
The quota figures above are from AWS's documented VPC and peering quotas; Google Cloud has comparable per-network peering quotas, which is why the same conclusion holds there.
Worked example: 3,000 tenants connecting privately
Suppose 3,000 tenants each connect from their own AWS network through an interface endpoint spread over 3 Availability Zones, and each moves 50 GB a month through it. AWS lists PrivateLink data processing at $0.01 per GB for the first petabyte; the hourly rate per endpoint per Availability Zone is about $0.01 in most Regions (check the current price for yours).
endpoint hours per tenant per monthacross 3,000 tenantsdata processing=0.01×3×730=$21.90=21.90×3000=$65,700=50×0.01×3000=$1,500Two things follow. First, the endpoint is created in the tenant's account, so the tenant's bill carries most of that, which is why private connectivity is usually an enterprise-tier feature. Second, the provider side is one endpoint service, not 3,000 networks: the platform's own cost and operational work grow slowly with tenant count.
Compare VPC-per-tenant for the same 3,000 tenants: at a few hundred VPCs per account (say 300, after requesting the quota increase mentioned above), that is 3,000 / 300 = 10 or more AWS accounts purely to hold networks, each tenant network with its own NAT gateways, security groups (virtual firewalls scoped to a resource, allowing or denying traffic by IP and port), monitoring and change process. Without that increase, at the default of 5 VPCs per account, the same 3,000 tenants need the 600 accounts computed above; either way it is the operational cost, not the account count itself, that makes VPC-per-tenant the wrong default.
The recommended layered design
- Shared data-plane VPCs (a handful, per Region), sized for many tenants' workloads.
- Mesh for tenant separation inside them: default-deny authorization; every workload's identity encodes the tenant it serves (for tenant-dedicated workers) or the platform role (for shared services), and policies say which identities may call which services. Shared services must still check the tenant in the request against the caller's identity; the mesh cannot do that for data inside a shared database.
- Network policy underneath the mesh (Kubernetes NetworkPolicy or security groups) so that a workload that bypasses its proxy (for example, a misconfigured pod that opens a raw connection instead of going through its sidecar, or an attacker who gains code execution on the host directly) still cannot reach arbitrary addresses. Mesh plus network policy is defence in depth: two independent layers, so a hole in one (a missed mesh policy) is still caught by the other (the network-level rule).
- PrivateLink / PSC for ingress from tenant networks, so no tenant's network is ever routed to the platform's and IP overlap is irrelevant.
- Dedicated VPC in a dedicated account for the few tenants who need it, provisioned by the same automation so it is a configuration choice, not a fork.
What would flip the choice: a small number of very large, regulated tenants (tens, not thousands) makes VPC per tenant affordable and simplifies audits; tenants running their own code inside the platform (not just sending data) raises the bar toward separate VPCs or accounts, because a mesh policy is a weaker boundary against hostile code running inside the network than a separate network is.
Trade-offs and pitfalls
- Peering as a tenant-onboarding path fails on the first customer whose IP range overlaps yours, and peering exposes more than one service.
- Mesh-only isolation is only as strong as its policy and its certificate authority; audit policies as code and test that cross-tenant calls are denied.
- Forgetting egress: isolation is also about where tenant workloads can send data out. Restrict outbound traffic per workload identity, or a compromised tenant worker can exfiltrate through the shared NAT.
- Security-group limits: AWS defaults to 60 inbound rules per security group and 5 security groups per network interface; encoding thousands of tenants into IP-based rules breaks these quickly. Identity-based policy avoids it.
That is every published Multi-Tenancy and Isolation question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.