Data Protection and Encryption in Practice Questions
Protecting data at rest and in transit across real systems from an engineering rather than pure-cryptography standpoint. Covers encryption strategy and key management for stored and transmitted data, secrets and sensitive-data handling, tokenization and secure elements for payment and sensitive data, and secure data handling in application code. Applied data-protection controls, distinct from cryptographic primitive design and from privacy-regulation compliance.
You must decide between client-side encryption and server-side encryption for a multi-tenant SaaS application that stores customer documents. Build a short threat model for each option and justify which you would choose. Discuss the operational impact on search, analytics, backups, and who is trusted with the plaintext.
Sample Answer
Direct answer
Server-side encryption means the platform (or its cloud provider) holds the keys and does the encrypt and decrypt work, so the platform is trusted with plaintext as part of normal operation. Client-side encryption means data is encrypted before it ever leaves the client, so the server only ever stores and forwards ciphertext, and a server-side breach, insider, or legal compulsion cannot produce plaintext because the server never had the key to begin with. For most multi-tenant SaaS document storage, the right answer is both: server-side encryption as the baseline everywhere, plus client-side or field-level encryption for the specific highest-sensitivity tenant data.
Structured elaboration
Server-side encryption, threat model:
- Defends against: theft of a raw disk or backup, an operator accessing storage media directly without going through the application.
- Does not defend against: a compromised web application or API tier, a malicious insider with elevated database or key-management access, or a subpoena served directly to the provider, since the provider holds the key and can produce plaintext on request.
Client-side encryption, threat model:
- Defends against: exactly the three gaps above. Three concrete attacks it prevents that server-side encryption cannot: (1) an attacker who compromises the multi-tenant database, say via SQL injection, and dumps rows gets ciphertext with no usable key; (2) a cloud provider employee or someone who compels the provider gets the same ciphertext, because the provider's KMS never held the tenant's key; (3) a legal request served on the provider yields ciphertext, since decryption requires the tenant's own key material.
- Does not defend against: a compromised or malicious client device, weak key handling on the client, or an application bug that leaks plaintext before it's ever encrypted.
Operational impact:
- Search and analytics: server-side encryption is transparent to the platform, so full-text search and analytics work normally. Client-side encryption breaks server-side search outright, unless you deliberately add a searchable-encryption design (with the leakage trade-offs that involves), and analytics has to run on data the platform can't read at all.
- Backups: server-side backups are simple to operate. Client-side encryption makes the backup of the data itself trivial (it's already ciphertext), but it moves the critical dependency to the tenant's own key: lose the tenant's client-side key and that tenant's data is unrecoverable, in a way a provider-managed key never would be.
- Two concrete operational costs client-side introduces beyond search: losing the platform's own ability to run features that require reading content, such as built-in virus scanning or content indexing; and the provider can no longer help a tenant recover from their own key loss, which pushes key-escrow and backup discipline entirely onto the tenant.
This trade-off is exactly why regulated tenants, healthcare and fintech customers in particular, often specifically demand "we can never read your data" from a SaaS vendor: it's a contractual and trust requirement as much as a technical one, and it's the scenario where client-side encryption earns its operational cost.
Worked example
Consider a legal-document SaaS storing attorney-client-privileged files for a regulated tenant. On upload, the client generates a random symmetric key, encrypts the document locally, and wraps that key using a key held in the tenant's own KMS key (envelope encryption). Only the ciphertext and the wrapped key ever reach the server; the server's database, backups, and any breach of either yield unreadable blobs. Full-text search across those documents is no longer possible server-side, so the product either drops that feature for this tier of tenant or builds a client-side or locally-decrypted search index instead.
Trade-offs and pitfalls
A common failure mode is a product that markets "client-side encryption" but actually decrypts inside a shared backend service for convenience, silently reverting to server-side trust while keeping the marketing claim. Another is losing search entirely and then quietly building an unencrypted metadata index to compensate, which can leak more than the team realizes about document contents through filenames, tags, or access patterns.
Explain the trade-offs between using environment variables, configuration files, and a dedicated secrets manager for storing application secrets. Cover developer ergonomics, auditability, and the risk of accidental leakage into repositories or logs.
Sample Answer
Direct answer
Environment variables and config files are the ergonomic default (no SDK integration, no network dependency at startup), but neither is encrypted, neither has an access-control model finer than "who can read this host or this file," and neither produces an audit trail of who read the value and when. A dedicated secrets manager trades that ergonomic simplicity for encryption at rest, per-identity access policies, automatic rotation, and a read-by-read audit log, at the cost of an SDK/agent integration and a bootstrap credential the app needs just to authenticate to the secrets manager itself.
Structured elaboration
A secret, for this purpose, is any credential or key whose disclosure grants access or breaks confidentiality: database passwords, API keys, OAuth client secrets, TLS private keys, encryption keys, and service tokens. A feature flag or a non-sensitive config value is not a secret even though it's also "config."
Risk profile of each option:
- Environment variables: readable via
/proc/<pid>/environor a container inspect command by anyone with host/container access, frequently captured by accident in crash dumps or error-reporting tools, and if set via a committed.envfile, permanently leaked into git history. - Config files: the same plaintext-on-disk exposure as environment variables; slightly better because they can be
.gitignored, but that's a process discipline, not a control, and the file is still unencrypted unless separately protected (for example with SOPS or git-crypt). - Dedicated secrets manager: the value is encrypted at rest inside the store's own storage backend and encrypted in transit over TLS to the caller. "Encrypted at rest inside the store" means the persisted copy on disk is ciphertext that only the store's own key hierarchy can decrypt; it does not by itself stop an over-privileged but authorized caller from reading the plaintext value through the normal API, so access policy still matters even after adding a secrets manager.
There are three common integration patterns for getting a secret from the store into an application:
- API call pattern: application code calls the secrets manager's SDK/API directly at runtime whenever it needs the value (typical for AWS Secrets Manager or GCP Secret Manager).
- Sidecar pattern: a sidecar process or container (for example Vault Agent, or a Kubernetes Secrets Store CSI driver, a Container Storage Interface driver that mounts secrets as pod volumes) authenticates to the store on the application's behalf and writes the secret to a local file or in-memory volume the app reads, so the application code never talks to the secrets manager directly.
- Compile-time/build-time injection: the secret is baked into the built artifact during the CI build. This is generally discouraged, because the value ends up embedded in the image or binary with no way to rotate it short of a rebuild, but it occasionally shows up for compiled clients with no other option.
Worked example
A backend service that reads DATABASE_PASSWORD from an environment variable set in its deployment manifest has that value visible to anyone who can kubectl describe pod or docker inspect the container, with no record of who looked. The same service switched to the sidecar pattern instead: a Vault Agent sidecar authenticates using the pod's own Kubernetes service account, fetches the credential, and renders it to a file on a memory-backed volume the app reads at startup; every fetch is now in Vault's audit log, and rotating the password never requires touching the deployment manifest.
Trade-offs and pitfalls
The bootstrap problem is real: the application needs some credential to authenticate to the secrets manager in the first place, so a secrets manager doesn't eliminate the "where does the first credential live" question, it shrinks it down to one bootstrap identity (often a cloud-native identity like an IAM (Identity and Access Management) role or Kubernetes service account, which needs no stored secret at all) instead of many scattered application secrets.
What is format-preserving encryption, and what problem does it solve that standard symmetric encryption modes do not? Name common approaches and explain the security trade-offs involved in preserving format.
Sample Answer
Direct answer
Format-preserving encryption (FPE) is a mode of encryption that produces ciphertext in exactly the same format as the plaintext, the same character set and length, so encrypting a 16-digit credit card number produces another 16-digit number rather than the binary-looking blob standard modes like AES-GCM produce. It solves the problem of encrypting data that has to keep flowing through rigid legacy schemas, database columns, and validation rules that a normal ciphertext blob would break.
Structured elaboration
The problem it solves. Many legacy systems, database columns typed as a fixed-length numeric field, or downstream validators expecting a specific pattern, weren't built to accept an arbitrary-looking ciphertext. FPE lets you encrypt a value in place without changing the schema or touching every consumer of that field, simply because the output still "looks like" valid data of the same shape.
Common approaches. The standardized constructions are FF1 and FF3-1 (both defined in NIST's FPE standard), which build a Feistel-network cipher on top of a standard block cipher, typically AES, restricted to operate over a small alphabet, such as the digits 0 through 9, instead of raw bytes. FF3-1 is a corrected revision of an earlier version, FF3, after a published weakness was found in the original tweak handling, so FF3-1 is the one that remains recommended.
Security trade-offs of preserving format. The value is still protected by the underlying 128-bit AES key, so FPE's reduced security margin does not come from a weaker key. For a fixed key and tweak, FF1 and FF3-1 are keyed bijections (permutations) over the set of valid inputs, so there is no oversized output blob and no output collisions to search for. The weakness comes instead from the size of that input set, the domain. Two things degrade as the domain shrinks. First, the provable-security bounds of a small-domain Feistel construction are weaker than those of a full-width block cipher, because a Feistel network over a tiny alphabet has little room to mix and hide its structure. Second, there are concrete domain-size-dependent attacks: the Durak-Vaudenay attack that broke the original FF3 (and motivated the FF3-1 revision) recovers the permutation with a work factor that scales polynomially in the domain size, so it becomes practical precisely when the domain is small. This is why applying FPE to a short, low-entropy format such as a four-digit PIN is considered particularly weak, while applying it to a long value is comparatively safe.
Worked example
Because the security concern is the size of the domain, not the AES key, the right quantity to compare is how many values an attacker would have to enumerate to defeat the scheme for a given key and tweak. FPE is a bijection over that domain, so an attacker with chosen-plaintext access can query every possible input once and build the complete plaintext-to-ciphertext table, fully inverting the scheme without ever attacking AES. The cost of that codebook attack is just the domain size:
pin4_domain = 10 ** 4 # a 4-digit PIN
card16_domain = 10 ** 16 # a 16-digit card number
# A keyed bijection can be fully mapped with one chosen-plaintext query per value.
print(f"4-digit PIN: {pin4_domain:,} values to enumerate")
print(f"16-digit card: {card16_domain:.0e} values to enumerate")
print(f"The PIN domain is {card16_domain // pin4_domain:.0e}x smaller")
Output:
4-digit PIN: 10,000 values to enumerate
16-digit card: 1e+16 values to enumerate
The PIN domain is 1e+12x smaller
Ten thousand values is trivially exhaustible: an attacker who can obtain ciphertexts for chosen PINs reconstructs the entire mapping in at most 10,000 queries, and even without an oracle a domain that small is vulnerable to frequency analysis and to the domain-scaling recovery attacks that make small-domain FPE weak. The same codebook attack against the 16-digit format would need 10^16 queries, which is infeasible, so the 16-digit case is not the weak one. The lesson is the opposite of a raw key-size comparison: FPE gets riskier as the format gets shorter, so it should never be used to protect a short, low-entropy field like a PIN.
Trade-offs and pitfalls
FPE should be reserved for genuine legacy or format-compatibility constraints, not chosen by default just because it's convenient to bolt onto an existing schema. Where no such constraint exists, a standard mode with its much larger security margin is the better default.
An application needs to support partial or exact-match search on an encrypted field, for example matching the last four digits of a phone number or looking up a record by an exact value, without exposing the full plaintext. Propose secure design options, and explain what information about the underlying data an attacker could still infer from each approach.
Sample Answer
Direct answer
There is no way to search encrypted data for free: every design that lets you query ciphertext trades away some amount of information to get that capability, and the job is choosing the smallest, most deliberate leak that still meets the actual product requirement.
Structured elaboration
| Approach | How it works | What an attacker can still infer |
|---|---|---|
| Deterministic encryption | Same plaintext always produces the same ciphertext, so an indexed equality lookup works directly | Which rows share a value with each other, and which values are most frequent, via simple frequency analysis |
| HMAC-based blind index | Store a keyed hash (HMAC) of the plaintext alongside randomized ciphertext; search by hashing the query term and matching the stored hash | Same equality and frequency leakage as deterministic encryption, since the HMAC is itself deterministic per key, though the hash key can be rotated somewhat more cheaply since it isn't also the decryption key |
| Deliberate partial exposure (for example last-4-digits search) | Store a low-sensitivity slice of the value in a lightly protected or plaintext column, alongside the full value fully randomized-encrypted separately | Exactly the slice you chose to expose, plus whatever additional narrowing is possible by combining it with other exposed data |
| Order-preserving or bucketed schemes | Ciphertext preserves the relative order of the underlying values, enabling range queries | The relative ordering of every value, which for common real-world distributions (salaries, ages, dates) can often reconstruct values close to exactly, once combined with any public information about the distribution |
| Searchable symmetric encryption (encrypted inverted indexes) | Purpose-built cryptographic schemes for keyword search over ciphertext | Reduced leakage compared to the above, but rarely production-ready outside specialized encrypted-search products; usually not worth the engineering cost for a typical application |
Worked example
Take the specific case in the question: matching the last four digits of a phone number. Store the last four digits in a plain or lightly protected column (a phone number's last four digits have only 10,000 possible values, 10^4, so this is a deliberately small, bounded amount of information to expose) and store the full number separately as randomized ciphertext, unrelated to the last-four column. A search for "ends in 4321" runs a plain WHERE last_four = '4321' query, and the application decrypts the full number only for the small set of matched rows before display. The design is honest about what it gives up: an attacker who sees the last-four column learns exactly that slice, nothing more, and the choice to expose it is deliberate and documented rather than an accident of using deterministic encryption on the whole field.
Trade-offs and pitfalls
The temptation to make more and more fields searchable eventually turns a database that looks encrypted into one that is close to plaintext through the accumulation of leakage channels. Order-preserving encryption in particular looks convenient for range queries but leaks close to as much as the plaintext itself for many distributions, and should be avoided for anything but the least sensitive fields. Restrict searchable or partially exposed fields to the minimum the product genuinely requires, and write down the accepted residual leakage for every one you add, rather than treating "it's encrypted" as the end of the analysis.
Design an end-to-end encrypted messaging system that must support group chats, device syncing, limited server-side search, and a lawful-access or eDiscovery request from the server operator's legal team. Explain the client-server responsibilities, the key hierarchy across devices and conversations, and what metadata you would still minimize even though message content is protected.
Sample Answer
Direct answer
End-to-end encryption (E2EE) means only the sending and receiving devices ever hold the keys needed to read a message's content; the server only ever handles ciphertext. The hard part of this design isn't the cryptography itself, well-established protocols already exist, it's building group chat, multi-device sync, and any server-assisted feature without ever giving the server a decryption path, which each of those features would naturally want if you weren't deliberate about it.
Structured elaboration
Client and server responsibilities. The client generates and stores all key material, encrypts every outgoing message, decrypts every incoming one, and performs key exchange with other devices and users directly, with the server only relaying the exchange messages, never reading them. The server relays ciphertext, stores it temporarily for offline delivery until a recipient device comes back online, and handles the metadata needed to route messages, who sent what to whom, when, but must have no code path capable of decrypting content.
Key hierarchy. Each device holds a long-term identity key pair that proves which device belongs to which user. For one-to-one conversations, a continuously re-keying session mechanism (a Double Ratchet-style design) derives a fresh key for every message, so compromising one message's key does not expose past or future messages, a property called forward secrecy. For group chats, a shared group key, or a per-member sender key others also hold, is distributed to every current member through pairwise one-to-one encrypted channels first, avoiding an expensive full pairwise fan-out for every individual group message. When someone joins or leaves the group, the group key is rotated: a new key is generated and redistributed pairwise only to the remaining members, so a removed member's device cannot decrypt future messages even though it still holds whatever history it already received before removal.
Device syncing. Adding a new device gives it its own identity key, and existing devices or contacts must securely deliver the relevant session keys to it, typically via the same pairwise encrypted exchange, verified by an out-of-band or in-app safety-number or QR-code check. That verification step specifically defends against a malicious server quietly inserting a fake device into the key exchange, sometimes called a key-substitution or ghost-device attack.
Limited server-side search. Since the server never has plaintext, search either happens entirely on-device, each client searching its own locally decrypted message store, which works for single-device search but doesn't span devices without extra design, or uses an encrypted keyword-index technique similar to the equality and blind-index approaches used for searchable encrypted database fields generally, with the same equality and frequency leakage caveats applying here too (a blind index is a searchable fingerprint of a value that lets the server match equal values without ever seeing the values themselves; equality leakage means the server can still tell that two entries hold the same value, and frequency leakage means it can see how often each value appears). There is no server-assisted search option that gives up nothing.
Lawful access or eDiscovery from the operator's own legal team. In a properly designed E2EE system, the operator genuinely cannot produce plaintext under a valid legal order, because it was never in a position to have it. What the operator can still produce is whatever metadata it does hold, who talked to whom, when, message sizes, delivery timestamps, which is exactly why minimizing that metadata matters so much: it's the part legal process can actually reach. If a business or regulatory need genuinely requires the operator to produce plaintext under compulsion, that is a fundamentally different, non-E2EE architecture, and the honest answer to a stakeholder asking for "E2EE, but with a way for us to read messages when legally compelled" is that the two requirements are close to mutually exclusive by design, since any mechanism giving the operator that path is also a channel an attacker, or a compromised operator, could use.
Metadata to minimize even though content is protected. Avoid retaining message-derived metadata beyond what delivery actually requires, round or bucket timestamps where feasible rather than keeping exact ones, and avoid keeping full social-graph history longer than the relay function needs, since delivery-timestamp and message-size metadata alone can reveal a surprising amount about conversation patterns even with content fully protected. Offline delivery requires the server to buffer ciphertext until the recipient is next online, which is fine since it's already ciphertext, but how long undelivered messages sit in that queue is itself a retention and privacy decision worth setting deliberately. The server can still compute coarse operational aggregates, total message volume, active user counts, from metadata alone without touching content, but even those aggregates need review so they don't inadvertently correlate with known message-size patterns.
Worked example
flowchart TD
ID[Device identity key<br/>long-term, one per device] --> RATCHET[Session ratchet key<br/>per 1:1 conversation, rotates every message]
ID --> GROUPDIST[Pairwise key distribution<br/>to each group member's device]
GROUPDIST --> GROUPKEY[Group message key<br/>shared by current members]
RATCHET --> MSG1[1:1 message ciphertext]
GROUPKEY --> MSG2[Group message ciphertext]
MSG1 --> SERVER[Server: relays ciphertext only<br/>stores minimal routing metadata]
MSG2 --> SERVER
SERVER -->|offline queue| RECIPIENT[Recipient device, decrypts locally]
MEMBERCHANGE[Membership change: join or leave] --> GROUPDIST
Every arrow into the server carries only ciphertext and routing metadata; the identity key, ratchet key, and group key never leave the devices that hold them, and a membership change flows straight back into pairwise redistribution, not into anything the server can act on.
Trade-offs and pitfalls
A common way E2EE quietly breaks is adding a convenience feature, link previews, spam scanning, cloud backup, that requires handing the server or a backup service plaintext or a copy of a key. Every new feature request against a system like this should be evaluated as "can this be done fully client-side" before assuming a server-side implementation is acceptable, because it usually isn't compatible with the property the system is named for.
Unlock Full Question Bank
Get access to all 37 Data Protection and Encryption in Practice interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.