Multi-Tenancy and Isolation Questions
Serving many tenants from shared infrastructure: tenancy models (silo, pool, bridge), data isolation, per-tenant data residency, noisy-neighbor mitigation, per-tenant limits, and security boundaries between tenants. Covers the cost, isolation, and blast-radius tradeoffs of shared versus dedicated resources, and business continuity: per-tenant backup, disaster recovery, and compliant tenant offboarding and deletion. The architecture layer specific to SaaS and platform products.
Explain encryption at rest and encryption in transit in the context of multi-tenant services. Describe a high-level design for per-tenant encryption keys using a cloud KMS: key isolation, envelope encryption for blobs, key access policies, and operational considerations for key rotation and emergency key revocation.
Sample Answer
Direct answer
Encryption in transit protects data while it moves over a network (TLS, Transport Layer Security, the standard protocol for encrypting a network connection, between clients and the service, and mutual TLS, where both sides present certificates, between internal services). Encryption at rest protects data stored on disks, databases, object storage and backups. In a multi-tenant service both are baseline, but neither separates tenants from each other: a disk encrypted with one platform key is readable by any process allowed to read the disk. Tenant isolation through encryption comes from per-tenant keys held in a cloud KMS (key management service), applied with envelope encryption (encrypting each object with its own one-time data key, then encrypting that small data key, not the object, with the tenant's KMS key; the mechanics are below), with key policies that only let the right tenant's requests use the right key.
Three tiers of key granularity
| Model | How it works | Isolation | Cost and effort |
|---|---|---|---|
| Single platform key | One key encrypts everything (typical default disk or storage encryption) | Protects against stolen disks, not against a bug that reads the wrong tenant's rows | Near zero |
| Per-tenant key | Each tenant has its own KMS key; all of that tenant's data keys are wrapped by it | A leaked data key exposes one object; a disabled tenant key makes that tenant's data unreadable without touching others | One KMS key per tenant (AWS KMS lists 1 US dollar per key per month) |
| Per-entity key | Each record, file or user inside a tenant has its own data key | Finest-grained deletion: crypto-shredding means destroying the key itself, so every copy of the data it protects, including backups, becomes permanently unreadable without touching the data at all ("crypto-shred one customer's record") | Many more keys to store and manage |
The common answer: a KMS key per tenant, plus a fresh data key per object or per file under it. Per-entity keys are added only where you need to delete or revoke individual records cryptographically.
Envelope encryption for blobs
Calling KMS to encrypt every byte would be slow and costly, and KMS limits how much data one call can take. Instead:
- Ask KMS for a new data key (DEK, data encryption key) under the tenant's key. KMS returns it twice: in plaintext and encrypted ("wrapped") by the tenant's key-encryption key (KEK), which never leaves KMS.
- Encrypt the blob locally with the plaintext DEK (AES-256-GCM, a standard, widely used symmetric encryption algorithm).
- Store the ciphertext together with the wrapped DEK; discard the plaintext DEK from memory.
- To read: send the wrapped DEK to KMS, which checks that the caller may use that tenant's key, returns the plaintext DEK, and the service decrypts locally.
sequenceDiagram
participant App
participant KMS
participant Store
App->>KMS: Generate data key (tenant 42 key)
KMS-->>App: plaintext DEK + wrapped DEK
App->>App: encrypt blob with DEK
App->>Store: ciphertext + wrapped DEK
App->>KMS: Decrypt wrapped DEK (context tenant 42)
KMS-->>App: plaintext DEK (if policy allows)
Worked example: tenant 42 uploads a 50 MB file. One KMS call produces a DEK; the 50 MB is encrypted locally; the stored object is 50 MB of ciphertext plus a few hundred bytes of wrapped key. Reading it costs one KMS call, not one per block.
Key access policies
- The tenant's key policy allows only the service role (the identity the application itself, not a person, assumes to call KMS) that handles that tenant, never a wildcard.
- Bind each operation to the tenant with encryption context (non-secret key-value pairs such as
tenant_id=42that must match on decrypt; AWS KMS supports conditions on them withkms:EncryptionContext:condition keys). A request that tries to decrypt tenant 42's data key while claiming tenant 7 fails. - Every KMS call is logged (AWS CloudTrail on AWS), which gives an audit trail per tenant key.
Operational considerations
Rotation. Rotating a KMS key generates new key material for new encryptions while keeping the old material to decrypt old data, so nothing needs to be re-encrypted. The benefit is forward-looking, not retroactive: if one specific key version is later suspected of compromise, only the data keys wrapped under that version are exposed, not every data key the tenant has ever had, and every new object from that point wraps under the fresh version. Many compliance frameworks also require rotation on a fixed schedule regardless of suspicion. AWS documents default automatic rotation every 365 days, with a custom period available, and notes that rotation does not re-encrypt data or rotate the data keys already generated. If a data key itself is suspected compromised, rotation does not help: that data must be re-encrypted with a new DEK.
Emergency revocation. Disabling a tenant's key makes every later decrypt of that tenant's data fail immediately at KMS, which is the fastest containment action. Two caveats: data keys already decrypted and cached in application memory keep working until the cache expires (so keep cache lifetimes short, for example minutes), and scheduling key deletion (AWS enforces a waiting period of 7 to 30 days) is irreversible, so it is used for deliberate crypto-shredding at offboarding, not during an incident.
Cost and scale. 5,000 tenants at 1 dollar per key per month is 5,000 dollars a month before request charges, and each automatically rotated key adds 1 dollar per month for each of its first two rotations. Caching decrypted DEKs briefly keeps request volume and latency down.
Trade-offs and pitfalls
- Assuming "encrypted at rest" means tenants are isolated from each other; it does not.
- One key for all tenants means revoking a single tenant is impossible without re-encrypting everyone.
- Long-lived DEK caches quietly defeat emergency revocation.
- Per-entity keys everywhere add management cost with little benefit unless you need per-record deletion.
Design a secure multi-tenant cloud environment providing strong tenant isolation across network, compute, storage, IAM and logging. Compare the account-per-tenant model vs shared-VPC/namespace model, including operational overhead, cost, and security trade-offs.
Sample Answer
Direct answer
I would run most tenants in a shared-account model (pooled VPCs, namespaces or schemas, with per-tenant IAM roles, keys and logging) and offer account-per-tenant as a tier for regulated or very large tenants. Account-per-tenant gives the strongest boundary a cloud provider offers (a separate AWS account is a hard wall for IAM, quotas, billing and blast radius) but carries a fixed cost and operational load per tenant that only pays off above a certain tenant size or compliance need. Either way the model is only viable with fully automated provisioning and tenant-level observability and billing built in from the start.
Terms: a VPC (Virtual Private Cloud) is a private network in a cloud account; IAM (AWS Identity and Access Management) controls who can do what; blast radius is how much is affected when something goes wrong. A NAT gateway is what lets resources with no public address reach the internet (for updates, external APIs) without being reachable from it; a VPC typically needs at least one, which is why it is a fixed cost per account below. A security group is a virtual firewall attached to a resource that allows or denies traffic by port and source. A service mesh is a layer that runs alongside every service to manage service-to-service traffic, including issuing the certificates mTLS needs, without changing application code. A principal is the identity making a request, a user or a role, the subject IAM policies are written about.
The two models compared, layer by layer
| Layer | Account-per-tenant | Shared account (shared VPC / namespace) |
|---|---|---|
| Network | Separate VPC per tenant; no route between tenants unless you build one | Shared VPC; separation by subnets, security groups, Kubernetes NetworkPolicy (rules controlling which pods may talk to which), service mesh mTLS (mutual TLS, both sides authenticate) |
| Compute | Tenant's own instances, clusters or functions | Shared clusters with namespaces, quotas, and optional dedicated node pools |
| Storage | Separate buckets and databases in the tenant's account; per-tenant KMS (key management service) keys | Shared buckets with per-tenant prefixes, or shared databases with row-level security (RLS: the database filters rows by tenant) or schema per tenant; per-tenant keys still possible |
| IAM | Account boundary itself is the control; a misconfigured policy in one account cannot grant access to another's resources | Per-tenant roles and attribute-based policies (tags such as tenant_id on principals and resources); one wrong wildcard can cross tenants |
| Logging | Per-account logs, aggregated to a central security account | Shared log pipeline; every record must carry tenant_id and access must be filtered per tenant |
| Quotas and noisy neighbours | Service quotas are per account, so one tenant cannot exhaust another's API rate limits | All tenants share one account's quotas (for example KMS request rates, Lambda concurrency, the number of function invocations allowed to run at once for the whole account) |
| Billing | Consolidated billing (AWS rolling every member account's usage into one payer-account bill) gives an exact per-account bill | Needs cost allocation tags (billing tags that make spend sliceable by a tag's value, like tenant_id) plus metering for shared resources |
| Operational overhead | Hundreds of accounts to patch, baseline and monitor | One environment; changes roll out once |
Worked cost example
Assume 500 tenants and an assumed per-account baseline of about 150 US dollars a month (NAT gateways, per-account security services such as threat detection and config recording, log storage, idle minimums). This baseline is an illustrative assumption; measure your own from a pilot account.
- Account-per-tenant fixed overhead: 500 x 150 = 75,000 dollars a month before any tenant does real work.
- Shared model: that baseline is paid a handful of times (per environment and region), so it is close to negligible per tenant.
If the average tenant pays 300 dollars a month, a 150-dollar baseline is half their revenue, which settles the question for the long tail. If the enterprise tier pays 20,000 dollars a month, 150 dollars is under 1% and the stronger boundary is easy to justify.
Security trade-offs
- Account-per-tenant turns a class of bugs (overly broad IAM policy, wrong bucket prefix) from a cross-tenant breach into a same-tenant bug. It also makes "delete everything for tenant X" as simple as closing an account.
- Shared depends on every layer enforcing the tenant boundary correctly: every query filtered, every policy scoped, every log tagged. Defence in depth is mandatory: tenant context injected at the edge, RLS in the database, per-tenant keys, and automated tests that attempt cross-tenant access.
- Shared control plane risk exists in both: the provisioning pipeline and central security account can reach every tenant, so they need the strictest access and audit.
Provisioning automation
Neither model works by hand. For account-per-tenant, use an account vending pipeline: AWS Organizations to create accounts under an organizational unit (a folder that groups accounts so a policy can apply to all of them at once), service control policies (SCPs, organization-wide guardrails that cap what any role in the account can do) applied automatically, and either Control Tower Account Factory (AWS's built-in service for provisioning a new account with a standard baseline already applied) or Terraform (a general infrastructure-as-code tool: the baseline is written as code and applied the same way every time) to lay down the baseline (VPC, logging to the central account, KMS keys, IAM roles). For the shared model, the same pipeline creates the namespace, database schema or RLS policies, IAM role, key, and quota. Onboarding should be one API call, idempotent, and reversible (offboarding runs the same pipeline backwards).
flowchart LR
REQ[Tenant signup] --> TIER{Tier}
TIER -->|enterprise / regulated| AV[Account vending: Organizations + SCPs + baseline]
TIER -->|standard| NS[Shared account: namespace, schema, role, key, quota]
AV --> REG[(Tenant registry)]
NS --> REG
REG --> OBS[Per-tenant observability + billing]
Tenant-level observability and billing
- Every metric, log and trace carries
tenant_id; dashboards filter per tenant. - In the shared model, tag every resource with
tenant_idand activate it as a cost allocation tag; meter shared resources (CPU-seconds, requests, storage) per tenant and allocate the shared bill proportionally. - In account-per-tenant, the account bill is the tenant's bill, which is simple and exact.
Recommendation and what would flip it
Shared by default, account-per-tenant for regulated, contract-driven, or very large tenants, with the same provisioning API producing both. I would move more tenants to their own accounts if most revenue came from a few hundred large customers, or if regulators or customer contracts demanded a provable hard boundary.
Pitfalls
- Choosing account-per-tenant without automation: the fleet drifts and becomes unpatchable within a year.
- Choosing shared and relying on application code alone for isolation.
- Forgetting shared quotas: one tenant's batch job exhausting an account-level API limit is a noisy-neighbour outage for everyone.
You are preparing to demonstrate to auditors that tenant isolation guarantees (no data leakage, proper encryption, access logs) are enforced. What artifacts, logs, and test evidence would you present? How would you design systems so artifacts are easy to extract for audits (PCI/GDPR)?
Sample Answer
Direct answer
I would present evidence in three tiers: design evidence (what controls exist and why they isolate tenants), operating evidence (logs and records showing the controls ran every day of the audit period, not only today), and test evidence (results of attempts to break isolation, including penetration tests). Each item maps to a specific control and requirement. To make this cheap to extract, I would design the system so evidence is produced continuously as a byproduct: structured, tenant-tagged, tamper-evident logs, controls declared as code, and automated control tests that store their results.
What auditors actually need
An auditor tests two things: that a control is designed to meet a requirement, and that it operated effectively across the period (for example, the last 12 months). A screenshot of today's configuration proves neither on its own. Evidence therefore has to be dated, complete for the period, and traceable to the control it supports.
1. Design evidence
- Architecture and data-flow diagrams showing where each tenant's data is stored, processed and replicated, and where the tenant boundary is enforced (application filter, database row-level security (RLS, the database adding the tenant filter to every query), separate databases or accounts for isolated tiers).
- Control matrix: each isolation control mapped to the requirement it satisfies, its owner, and how it is tested.
- Encryption design: which keys encrypt which tenant's data. Per-tenant keys in a key management service (KMS, a managed service that stores keys and logs every use) make "tenant A's data cannot be decrypted with tenant B's key" demonstrable.
- Access model: roles, who may access production tenant data (support, engineers), and the approval flow for break-glass access (emergency access granted outside the normal approval flow, for when normal channels are too slow during an incident).
2. Operating evidence (logs)
- Data access logs with who, which tenant, what object, when, and from where, for every read of sensitive data by staff and by the application.
- Key usage logs from the key management service: every decrypt call, with the key id, which ties decryption of tenant A's data to specific callers.
- Change records: every change to isolation-relevant configuration (RLS policies, bucket policies, network rules), linked to a review and approval.
- Access reviews: periodic sign-offs that staff access to production is still needed.
- Alert and incident records: what fired, how it was handled, and customer notifications where required.
3. Test evidence
- Continuous integration (CI) results, meaning the automated checks run on every code change, for isolation tests (one tenant trying to read another's data through every endpoint), per release, kept with the release record.
- Scheduled production canary results (synthetic tenants probing the boundary).
- Penetration test reports and evidence that findings were fixed and retested.
Mapping to PCI DSS and GDPR
PCI DSS (Payment Card Industry Data Security Standard, v4.0.1 is the current version; PCI DSS v4.0 was retired 31 December 2024) has requirements that line up directly:
| Requirement | Evidence I would show |
|---|---|
| Req 3: protect stored account data | Encryption design, key inventory, key rotation records (proof that encryption keys are replaced periodically without breaking access to data already encrypted under the old key) |
| Req 7: restrict access by business need to know | Role definitions, access reviews |
| Req 10: log and monitor all access to system components and cardholder data | Access logs; 10.5.1 requires audit log history to be kept at least 12 months with the most recent three months immediately available |
| Req 11: test security regularly | Penetration test reports |
| Appendix A1 (additional requirements for multi-tenant service providers) | Evidence of logical separation between customers; A1.1.4 requires confirming the effectiveness of that separation by penetration testing at least every six months; A1.2 covers per-customer logging and incident response |
Of the PCI requirements above, Req 10 (logging) and Req 11 together with A1.1.4 (regular testing) are the two an auditor checks hardest for a multi-tenant isolation claim; the rest describe what the logs and access model must already contain.
GDPR (EU General Data Protection Regulation) is principle-based rather than a checklist: Article 5(2) makes the controller (the customer, who decides why and how personal data is processed) responsible for demonstrating compliance (accountability), Article 32 requires appropriate technical and organisational security measures, and Article 30 requires records of processing activities. For a SaaS provider acting as a processor (the vendor that processes data on the controller's behalf and instructions, rather than deciding why it is processed), Article 28 governs the contract with each customer. Of these, Article 32 is the one the isolation evidence above answers directly; Articles 5(2), 28 and 30 describe the surrounding accountability paperwork. Tenant isolation evidence supports Article 32; the data-flow diagrams and residency configuration support the records of processing.
Designing systems so evidence is easy to extract
- Structured audit events with a fixed schema (actor, tenant_id, action, resource, result, timestamp, request id) written by a shared library, so evidence queries are the same across services.
- Tenant id on every event, so you can produce "all access to tenant A's data in Q3" with one query, which is also what a customer's own auditor or a GDPR access request will ask for.
- Tamper-evident storage: ship logs to write-once storage (object storage with a retention lock that prevents deletion until a date), hash-chain batches (group each batch of log entries and store a hash of the batch that also incorporates the previous batch's hash, so editing any past entry changes every hash computed after it, making a silent edit detectable), and restrict deletion rights to no one in engineering.
- Controls as code: RLS policies, bucket policies and network rules live in source control, so the history of the control is the commit history, with reviews attached.
- Automated control tests that save evidence: a daily job that checks, say, "RLS enabled on every tenant table" and "no bucket policy grants cross-tenant access", writing a signed, dated result. Over a year this becomes 365 records of the control operating.
- An evidence catalog: for each control, the query or job that produces its evidence. Audit prep becomes running the catalog for the period.
Worked example
An auditor asks: "Show that no employee accessed tenant 812's cardholder data without a ticket during the last 12 months." With the design above, one query over the audit events returns every staff access to tenant 812 (for example 14 events); each carries a request id that links to a support ticket and approval. The key-usage logs show the same 14 decrypt calls on tenant 812's key from the support role, and no others. The logs' retention-lock setting and hash chain show they could not have been edited. Without the tenant id on events, the same request means weeks of joining raw logs by hand.
Trade-offs and pitfalls
- Logs themselves can leak data. Log identifiers and actions, not card numbers or personal content, or the evidence store becomes the largest compliance risk you have. PCI DSS forbids storing sensitive authentication data (full magnetic-stripe data, card verification codes and PINs: the data used to authenticate a transaction, not just the card number itself) after authorization, which includes logs.
- Retention conflicts: PCI wants logs kept at least a year; GDPR's storage-limitation principle wants personal data kept no longer than needed. Keep audit logs for the defined period with pseudonymous identifiers (an internal id substituted for a person's real data, such as a hashed user id instead of a name or email, so the log stays useful without holding personal data directly) and document the justification.
- Pitfall: point-in-time evidence. Auditors test the whole period; controls that are only checked before the audit fail that test.
- Pitfall: claiming the penetration test covered separation when the test scope excluded cross-tenant scenarios. Scope the test explicitly as "tenant A attempting to reach tenant B".
Compare Role-Based Access Control (RBAC) and Attribute-Based Access Control (ABAC) for enforcing tenant boundaries in a data platform. Include: policy complexity, scalability across many tenants, ease of audits, and suitability for fine-grained data access (e.g., row or column level).
Sample Answer
Direct answer
For tenant boundaries in a data platform I would use both, with different jobs: RBAC for coarse "what kind of user are you" permissions, and an ABAC rule for the tenant boundary itself (the user's tenant attribute must equal the data's tenant attribute) and for row- and column-level filtering. Pure RBAC turns into role explosion once tenants multiply; pure ABAC is harder to audit. The hybrid keeps a small, auditable set of roles and one tenant rule that scales to any number of tenants.
The two models in plain terms
- RBAC (Role-Based Access Control): users are put into roles, and roles are granted permissions. "Alice is in
acme_analyst, andacme_analystcan read theacmeschema." Access is decided by membership lists. - ABAC (Attribute-Based Access Control): access is decided by a rule evaluated over attributes (properties) of the user, the data and the request. "Allow read if
user.tenant_id == table.tenant_idanduser.clearance >= column.sensitivity." The rule is written once; adding a tenant adds attribute values, not rules. - Row-level security (RLS): a filter the database applies automatically so a query only returns rows the caller is allowed to see. Column-level security / masking: hiding or masking specific columns (for example showing only the last 4 digits of a card number).
Comparison on the four asked axes
| Axis | RBAC | ABAC |
|---|---|---|
| Policy complexity | Each rule is trivial ("role X can read Y"). Complexity moves into the number of roles and grants. | Fewer rules, but each is a small program over attributes. Easy to write a rule that is subtly too broad. Depends on attributes being correct and trusted. |
| Scalability across many tenants | Poor. Roles are usually per tenant per job function, so roles grow as tenants x functions, and every new dataset needs grants per role. | Good. One tenant-match rule covers tenant 1 and tenant 10,000. A new tenant needs attribute values (a tag on data, a claim on users), not new policies. |
| Ease of audits | Easy to answer "who can access X?": list the roles with a grant on X and their members. Auditors are familiar with it. | Harder: "who can access X?" means evaluating the rule against every user's attributes. Needs decision logs and tooling to be auditable. The rule itself is easy to review, though, because there is only one. |
| Fine-grained (row/column) access | Awkward. Row-level via RBAC means a role per slice of rows, which multiplies further. | Natural fit. Row filters and column masks are exactly "compare a user attribute to a row or column attribute". |
Worked example
A data platform hosts 500 tenants. Each tenant has four job functions: admin, engineer, analyst, viewer. Analysts may not see the email column, and analysts also see only rows for their own region.
Pure RBAC
- Roles: 500 tenants x 4 functions = 2,000 roles.
- The regional restriction only applies to analysts, so it is analysts, not the other three job functions, that split further. For, say, 5 regions per tenant: each tenant now needs the 3 non-regional roles (admin, engineer, viewer) plus 5 regional analyst roles (one per region) instead of 1 plain analyst role, so 500 x (3 + 5) = 4,000 roles.
- Each new table needs grants to every relevant role. Forgetting one grant is an outage for that tenant; a wrong grant (
acme_analystgranted onglobexdata) is a cross-tenant leak, and with thousands of grants nobody spots it by reading.
Hybrid (recommended)
- 4 roles total (admin, engineer, analyst, viewer), attached to users as today.
- Every user carries a
tenant_idclaim (a piece of information embedded in the user's signed identity token, which the token's issuer vouches for and the user cannot edit) from the identity provider (the system that authenticates the user and issues that token, for example an internal auth service or a provider like Okta); every table (or row) carries atenant_idtag or column. - One tenant rule, in pseudo-policy. (
principalmeans whoever is asking: the authenticated user or service making the request, andprincipal.role,principal.tenant_id,principal.regionare that principal's own attributes.)
permit read on dataset d
when principal.role in d.allowed_roles
and principal.tenant_id == d.tenant_id
row filter: row.tenant_id == principal.tenant_id
and (principal.region is empty or row.region == principal.region)
column mask: email masked when principal.role == "analyst"
# only analysts lose email, matching the requirement above;
# admin, engineer and viewer keep it
Adding tenant 501 means creating its users with tenant_id = 501 and tagging its data. No policy changes.
Concrete implementations of this shape include PostgreSQL row-level security policies that compare a row's tenant_id to a session setting, Snowflake row access policies and masking policies, and AWS Lake Formation tag-based access control (tags on databases, tables and columns, matched against tags granted to principals).
Trade-offs and pitfalls
- Attributes become security-critical data. If a user can set their own
tenant_id(for example from a request header rather than a signed token claim), ABAC is broken. Attributes must come from a trusted source: the identity provider or the platform's tenant registry. - Default deny. A table missing its
tenant_idtag must be unreadable, not readable by everyone. Test this explicitly. - Audit gap in ABAC, and the fix: keep policies in version control (policy-as-code), log every allow and deny decision with the attributes used, and run a periodic job that materialises (computes the answer once and writes it out as a concrete stored list, instead of making every audit request re-evaluate the rule live) "who can see tenant X" by evaluating the rule over the user directory. That gives auditors the list they are used to.
- RBAC is still the right answer when the platform has a handful of large tenants, each with its own dedicated schema and a small admin team. The role count stays small and auditors get the simplest possible story.
- Pitfall: implementing the tenant filter in each application query instead of in the platform (RLS, or the policy engine: the shared piece of infrastructure that evaluates the ABAC rule and returns allow or deny, so no individual application has to reimplement the check). One forgotten
WHERE tenant_id = ...is a breach; a platform-enforced filter cannot be forgotten.
A security review finds that a support engineer investigating one tenant's ticket could, in principle, find another tenant's data sitting in shared application logs, backups, or a third-party monitoring pipeline. Outline the security controls you would put in place across logging, backups, and telemetry to close this off, and how you would verify none of it leaks through a vendor or pipeline you don't fully control.
Sample Answer
Direct answer
The fix is to treat logs, backups and telemetry as data stores in their own right, with the same tenant boundary as the primary database. Concretely: stop sensitive values entering them (allowlisted structured logging), tag every record with the tenant it belongs to so access can be scoped per tenant, give support engineers time-boxed, ticket-scoped access to one tenant at a time, encrypt backups per tenant and restore them only through an audited break-glass path, and scrub telemetry before it leaves your network. Then verify it continuously with canary data (unique marker strings seeded into a test tenant) that you search for in every downstream system, including vendors you do not control.
Why this leak exists at all
In most SaaS products the primary database is protected by tenant-scoped queries or row-level security (RLS: a database feature that filters rows by a policy such as tenant_id = current tenant). The side channels are not. A stack trace logs the SQL statement with a customer's email in it; a debug line dumps a whole request body; a nightly backup contains every tenant in one file; an APM agent (application performance monitoring, for example a hosted tracing product) captures HTTP payloads. A support engineer with read access to the log index investigating tenant A's ticket then sees tenant B's data in the same search results. Nobody broke a rule; the rules simply never covered these stores.
Controls, layer by layer
1. Logging: keep the data out, then scope what remains
| Control | What it does | Why it matters |
|---|---|---|
| Allowlisted structured logging | Logs are JSON with a fixed set of permitted fields (timestamp, level, tenant_id, request_id, route, status, error code). Unknown fields are dropped, not masked | A denylist ("mask fields called ssn") fails the first time someone logs a field called social. An allowlist fails safe |
| Redaction as a second net | Pattern scrubbing (emails, card numbers, tokens) on free-text fields such as error messages | Catches values that sneak in through exception messages |
Mandatory tenant_id on every record | Injected by the logging library from the request context, never by the caller | Without it you cannot scope access or delete a tenant's logs on offboarding |
| Per-tenant access scoping | Either one index per tenant (or per tenant tier), or a shared index where the log platform enforces a document-level filter (it checks the tenant_id on every individual record before returning it, not just at the index level) on tenant_id | The support engineer's query is physically limited to the tenant on the ticket |
| Just-in-time (JIT) support access | Access is granted for one tenant, tied to a ticket ID, expires after a few hours, and every query is audit-logged | Removes standing access; an auditor can match every search to a ticket |
| Short retention for verbose logs | Debug and request-level logs kept days, not years | Less data exposed, and less to erase later |
A minimal allowlisting formatter, run as shown:
import json, re
# Fields a log line may carry. Anything else is dropped, not masked.
ALLOWED = {"ts", "level", "tenant_id", "request_id", "route", "status", "latency_ms", "error_code"}
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+")
def scrub(event: dict) -> dict:
out = {k: v for k, v in event.items() if k in ALLOWED}
dropped = sorted(set(event) - ALLOWED)
if dropped:
out["dropped_fields"] = dropped # names only, never values
for k, v in out.items():
if isinstance(v, str):
out[k] = EMAIL.sub("[email]", v) # second net for free text
return out
raw = [
{"ts": "2026-09-01T10:00:00Z", "level": "error", "tenant_id": "t_acme",
"request_id": "r1", "route": "/invoices/42", "status": 500,
"error_code": "DB_TIMEOUT", "sql": "SELECT * FROM invoices WHERE customer='bob@acme.io'"},
{"ts": "2026-09-01T10:00:01Z", "level": "info", "tenant_id": "t_globex",
"request_id": "r2", "route": "/export", "status": 200, "latency_ms": 812,
"body": {"ssn": "123-45-6789", "note": "CANARY-GLOBEX-7F3A"}},
]
shipped = [json.dumps(scrub(e)) for e in raw]
for line in shipped:
print(line)
# Verification: seeded canary strings must never appear in anything shipped.
canaries = ["CANARY-GLOBEX-7F3A", "bob@acme.io", "123-45-6789"]
leaks = [c for c in canaries if any(c in line for line in shipped)]
print("leaked canaries:", leaks)
Output:
{"ts": "2026-09-01T10:00:00Z", "level": "error", "tenant_id": "t_acme", "request_id": "r1", "route": "/invoices/42", "status": 500, "error_code": "DB_TIMEOUT", "dropped_fields": ["sql"]}
{"ts": "2026-09-01T10:00:01Z", "level": "info", "tenant_id": "t_globex", "request_id": "r2", "route": "/export", "status": 200, "latency_ms": 812, "dropped_fields": ["body"]}
leaked canaries: []
Note what the engineer still gets: the tenant, the route, the error code and the request ID to correlate with. That is enough to debug most tickets; the rare case needing payload data goes through the JIT path against the tenant's own data store.
2. Backups: one tenant's restore must not expose the others
- Encrypt with per-tenant keys where the data is separable (silo databases, per-tenant object prefixes, per-tenant logical exports). With envelope encryption (each backup encrypted by a data key, which is itself encrypted by the tenant's master key in a key management service), access to the backup file alone reveals nothing.
- Pooled database backups cannot be split by key, so they are treated as the most sensitive artifact: stored in a separate backup account, readable only by an automated restore role, never by humans directly.
- Restores go to an isolated environment through a break-glass workflow (a pre-approved emergency path that requires a second approver and is fully logged). The restore job extracts only the target tenant's rows; the full restored copy is destroyed afterwards.
- Backup ACLs (access control lists) deny
readto the support role entirely. Support never touches a backup.
3. Telemetry and third-party pipelines: scrub before egress
- Run all traces, metrics and error reports through a collector you operate (for example an OpenTelemetry Collector, the open-source agent that receives, processes and exports telemetry) and apply attribute allowlists there. Vendors only ever receive the scrubbed stream.
- Turn off payload capture in APM and error-tracking agents (request bodies, SQL parameter values, local variables in stack frames); these defaults are the usual leak.
- Hash or pseudonymize tenant identifiers sent to vendors if the vendor has no need to know the customer's name.
- Metric labels carry tenant tier or a tenant ID, never user-level identifiers (each distinct label value creates its own stored time series, so a label that can take millions of values, like a user ID, multiplies storage and query cost by millions: a high-cardinality cost blow-up).
- Contractually: a data processing agreement (a contract clause binding the vendor to your privacy and security requirements for any data you send them), the vendor on your sub-processor list (the vendor's published list of who else touches your data on their behalf, which your own contracts require you to track), a region pin, a retention setting you have verified in their console, and their independent audit report (for example SOC 2 Type II, a report confirming an auditor actually tested the vendor's security controls over several months, not just reviewed a design on paper) reviewed annually.
Verifying what you do not control
You cannot inspect a vendor's storage, so you test the outputs and the inputs:
- Canary tenant. A synthetic tenant whose records contain unique strings (
CANARY-GLOBEX-7F3A, a fake but valid-format email). A scheduled job exercises it, including error paths that historically log payloads. - Search every sink for canaries daily through each system's API: your log platform, the APM vendor, the error tracker, the data warehouse, the support tool. Any hit is a sev-2 incident (the second-highest urgency tier: high priority, but not an all-hands page) with a known origin because each canary is unique per source path.
- Egress inspection. (Egress is data leaving your network for an outside destination, here a vendor.) Sample the collector's outbound stream and run a DLP (data loss prevention) scanner over it for card numbers, emails and the canaries. This checks the scrubbing before the vendor ever sees data.
- Support-access audit. Weekly reconciliation: every log query by a support engineer must map to a ticket whose tenant matches the query's tenant filter. A mismatch is investigated.
- Contract and configuration drift checks. Vendor retention, region and payload-capture settings read back through their API on a schedule and diffed against the expected values.
Trade-offs and pitfalls
- Per-tenant log indexes vs a shared index with filters. Separate indexes make access control and deletion trivial but cost more (index overhead per tenant) and break cross-tenant operational queries. Recommendation: shared index with enforced document-level filtering for the long tail, dedicated indexes for tenants whose contracts demand it.
- Dropping fields slows debugging. The honest answer is that it does, occasionally. The JIT path to the tenant's own store is the pressure valve; do not reopen the log floodgates.
- Pitfall: redacting only at the log platform. By then the data has already crossed the network and possibly been buffered by an agent on disk. Scrub at the source and at the collector.
- Pitfall: forgetting derived stores. Search indexes, analytics exports and support-tool attachments hold copies too. Canaries find these; architecture diagrams usually do not.
- Pitfall: canaries only in the happy path. Leaks come from error handlers and retries. Exercise failures deliberately.
Unlock Full Question Bank
Get access to all 9 Multi-Tenancy and Isolation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.