Data Protection and Encryption in Practice Questions
Protecting data at rest and in transit across real systems from an engineering rather than pure-cryptography standpoint. Covers encryption strategy and key management for stored and transmitted data, secrets and sensitive-data handling, tokenization and secure elements for payment and sensitive data, and secure data handling in application code. Applied data-protection controls, distinct from cryptographic primitive design and from privacy-regulation compliance.
Explain how you would implement OIDC-based authentication so that ephemeral cloud credentials are issued to CI runners instead of long-lived secrets. Describe the trust relationship between the CI provider and the cloud account, the IAM roles involved, and how you would mitigate replay attacks against the OIDC provider endpoint.
Sample Answer
Direct answer
OIDC (OpenID Connect, an identity layer on top of OAuth 2.0) lets the CI provider act as an identity provider: it issues a short-lived, signed JSON Web Token (JWT) for each workflow run, containing claims like the repository, branch, and environment. The cloud account is configured to trust that specific OIDC issuer and to map particular claim values to an IAM (Identity and Access Management) role. The CI job presents the token to the cloud's token service, which verifies the signature and claims against the trust policy and returns temporary, expiring credentials. No long-lived cloud secret is ever stored in CI.
Structured elaboration
The trust relationship, step by step:
- The cloud account registers the CI provider (for example GitHub Actions) as a trusted OIDC identity provider, pinning its issuer URL and public signing keys.
- An IAM role is created with a trust policy whose condition checks specific token claims, most importantly the repository and the branch or environment (for example
repo:my-org/my-repo:environment:production), not a wildcard. - At run time, the CI job requests an OIDC token scoped to the intended cloud audience.
- The job calls the cloud's token-exchange API (AWS's
AssumeRoleWithWebIdentityvia STS, AWS's Security Token Service, or the equivalent Workload Identity Federation call on GCP) presenting that token. - The cloud validates the signature against the registered issuer's public keys, checks the trust-policy condition against the token's claims, and if it matches, returns short-lived credentials, typically expiring within an hour.
Worked example
sequenceDiagram
participant Runner as CI Runner
participant OIDC as CI OIDC Provider
participant STS as Cloud Token Service
Runner->>OIDC: request short-lived signed token (claims: repo, branch, env)
OIDC-->>Runner: signed JWT
Runner->>STS: AssumeRoleWithWebIdentity(JWT)
STS->>STS: verify signature + match trust policy condition
STS-->>Runner: temporary credentials (expire in ~1 hour)
Mitigating replay against the OIDC endpoint: each token carries a short exp (expiry) claim and an aud (audience) claim pinned to the specific cloud account it's meant for, so a token intercepted after issuance is only useful for a narrow window and only against the one consumer it was minted for. The trust policy should also match exact claim values rather than wildcards (an exact branch name, not refs/heads/*), since a wildcard match lets anyone who can open a branch or fork mint a token that satisfies the condition. Monitoring AssumeRoleWithWebIdentity calls for unexpected source repositories or branches adds a detective control on top of the preventive narrowing.
Trade-offs and pitfalls
Two additional edge cases worth naming explicitly. First, multi-environment scoping: separate trust-policy conditions and separate IAM roles per environment claim (environment:staging versus environment:production) so a staging pipeline's token can never assume the production role, even if both use the same OIDC issuer. Second, build-time secrets ending up in the resulting container image: if the OIDC-derived temporary credentials are used mid-build (to pull a private dependency, for example), they must be passed as build-step environment variables, never as a Dockerfile ARG or ENV, since both persist in the image's layer history; use BuildKit secret mounts (--secret) instead, which are not written into any layer.
What secret-scanning approaches would you recommend to catch secrets before they ever reach source control, covering source code, container images, and CI logs? Compare static, regex-based, and machine-learning-based scanners, and explain how you would keep false positives and false negatives manageable in a production scanning pipeline.
Sample Answer
Direct answer
Static, regex-based scanners match known secret formats (an AWS access key's AKIA prefix, a private key's PEM header, a JWT (JSON Web Token)'s three-part structure) and are fast and deterministic, but blind to secrets that don't match a known pattern. Entropy-based checks catch unknown formats by flagging any high-randomness string, at the cost of more false positives on things that merely look random (hashes, UUIDs, test fixtures). Newer machine-learning and live-verification approaches go a step further by actually testing a candidate against the real provider API to confirm it's a working credential, which sharply cuts false positives without needing a hand-written pattern for every secret type.
Structured elaboration
Coverage needs to span three surfaces, not just source code:
- Source code: the most common target, scanned via regex/entropy/verification tooling at multiple points in the lifecycle.
- Container images: a secret can be baked into an image layer even when the source repository is clean (a build argument or a copied config file), so image-layer scanning is a separate, necessary check.
- CI logs: a job can echo a secret at run time even when nothing sensitive was ever committed to source, so scanning the captured logs themselves closes a gap the other two miss.
Managing false positives and false negatives in production: maintain an allowlist or baseline file for known false positives (test fixtures, example keys in documentation) so the same finding doesn't get re-triaged every run; prefer scanners that support live verification over purely pattern-based detection, since a verified hit ("this AWS key authenticates right now") is close to zero false positives; and accept that entropy-only detection needs a tuned threshold specific to the codebase, since a threshold copied from another project either misses real secrets or drowns the team in noise.
Worked example
Where scanning fits in the SDLC (software development lifecycle), across three stages: a pre-commit hook is the earliest and cheapest point to catch a mistake, before it's even recorded in history; a required CI check on every pull request is the broad safety net that catches what a bypassed or missing pre-commit hook let through (a developer running git commit --no-verify, for instance); and a periodic full-history scheduled scan catches anything that slipped past both of the first two checks, including secrets committed before scanning was ever adopted.
That CI check should be wired as a required status check that fails the build on a high-confidence match, blocking the merge until the finding is resolved or explicitly allowlisted with a documented reason, rather than merely reporting a warning that's easy to ignore.
Trade-offs and pitfalls
Scanning source control and CI is necessary but not sufficient: a credential that was valid, got scanned and found clean, and is still sitting unrotated in a running environment months later is a real and separate risk. Continuous validation of already-deployed secrets, tracking each credential's last-used timestamp and periodically confirming it's still needed, catches the stale-but-technically-not-leaked case that point-in-time source scanning cannot.
For Kubernetes workloads, describe secure patterns for managing secrets: native Kubernetes Secrets versus an external secret store integration, how pods authenticate to the store, and how you protect secret material in etcd and node memory.
Sample Answer
Direct answer
Native Kubernetes Secrets are only base64-encoded, not encrypted, and by default persist in plaintext-equivalent form in etcd (the cluster's key-value store); anyone who can read etcd or take a backup of it can recover them unless etcd encryption at rest is explicitly enabled. External secret store integration keeps the actual secret material in Vault or a cloud secrets manager and either mounts it directly into the pod's filesystem or syncs it into a native Secret object, with pods authenticating to the store using their own Kubernetes service account identity.
Structured elaboration
Two common integration shapes:
- Secrets Store CSI Driver: a Container Storage Interface driver mounts the secret directly from Vault, AWS Secrets Manager, or Azure Key Vault into the pod as a volume. The secret material can be kept off etcd entirely if you skip the driver's optional "sync as native Kubernetes Secret" feature.
- External Secrets Operator: watches an external secret and syncs its value into a native Kubernetes Secret object, which is more convenient for apps that already read Secrets the standard way, at the cost of the value now also existing in etcd.
Pod-to-store authentication, concretely with Vault's Kubernetes auth method: the pod presents its own projected service account token (a signed JWT (JSON Web Token) that Kubernetes automatically mounts into the pod) to Vault; Vault validates that token against the Kubernetes API's TokenReview endpoint, confirms which service account and namespace it belongs to, and maps that identity to a Vault policy scoped to that namespace, for example granting read only on secret/data/<namespace>/*. This means two pods in different namespaces authenticate with the same mechanism but land on completely different, non-overlapping policies.
Worked example
A pod in the orders namespace mounts its projected service account token automatically. A Vault Agent sidecar in that pod uses the token to authenticate via Vault's Kubernetes auth method, and Vault's role binding maps system:serviceaccount:orders:orders-app to a policy that only allows reading secret/data/orders/*. If an attacker compromises a pod in a different namespace, its service account token authenticates successfully to Vault but is bound to a different policy, so it cannot read the orders namespace's secrets even though authentication itself succeeded.
Trade-offs and pitfalls
Protecting material in etcd and node memory takes more than picking an integration pattern:
- Enable etcd encryption at rest (an
EncryptionConfigurationbacked by a KMS (Key Management Service) provider) so even a stolen etcd backup is ciphertext. - Restrict etcd network access to control-plane nodes only; it should never be reachable from a worker node running arbitrary workloads.
- Where possible, avoid the CSI driver's "sync as native Secret" mode, keeping the value only in the pod's memory-backed (tmpfs) mount, so it never touches etcd at all.
- Use tmpfs, not a regular disk-backed volume, for any file-based secret mount, so the value doesn't persist to node disk either.
The most common mistake is treating "we use the CSI driver" as sufficient by itself while leaving the sync-to-native-Secret option enabled, which quietly recreates the exact etcd-plaintext exposure the external store was meant to avoid.
Explain what tokenization is and how it differs from encryption. Sketch a typical tokenization architecture including a token vault, describe where tokenization is advantageous for protecting payment or other sensitive data, and name one operational risk it introduces.
Sample Answer
Direct answer
Tokenization replaces a sensitive value with a random surrogate, a token, that has no mathematical relationship to the original; the only way back to the real value is a lookup in a separately protected token vault. Encryption instead transforms data through a reversible mathematical function keyed by a secret, so anyone holding the key can always reverse it directly, with no vault involved at all.
Structured elaboration
Token vault architecture: A token-generation service issues a random token for a given sensitive value. The vault stores the token-to-value mapping, itself protected with its own encryption and strict access control. Every other application service only ever sees and passes around the token; only the small set of services that genuinely need the real value call the vault, under strong authentication and audit logging, to detokenize.
Where tokenization is advantageous: Payment card data is the classic case: replacing a card's primary account number everywhere in order records, logs, analytics, and customer-support tooling with a token means only the vault and the payment processor ever touch the real number, dramatically shrinking how much of the infrastructure is considered "in scope" for handling truly sensitive data. It is also used for other high-sensitivity identifiers, such as a social security number, when many downstream systems need a stable reference to the record without ever needing the real value.
Operational risk it introduces: The vault becomes a single, extremely high-value target and a single point of failure: if it's unavailable, every service that needs a real value is blocked, and if its own protections fail, every token it protects is compromised at once, a concentration-of-risk pattern similar to a master encryption key, except the risk sits in a database rather than a key.
Tokenization versus format-preserving encryption versus masking. Format-preserving encryption is still a reversible, keyed encryption scheme, just constrained to output the same character format as the input, so it needs no vault, only the key, but has a weaker security margin than standard encryption because its output space is smaller. Masking is typically not reversible at all, for example showing only the last four digits of a value for display, and reduces exposure in a user interface or logs, but it is not a storage-level protection mechanism, since the full real value still exists somewhere else in the system and still needs its own protection.
Worked example
A card-present retail system tokenizes a primary account number, 4111111111111111, to an opaque token such as tok_7f2a9c31, stored in the order record. Every downstream system, order history, analytics, support tooling, sees only tok_7f2a9c31, which has no mathematical path back to the real card number. Only the vault, given that token, returns the real number, and logs that specific access.
Trade-offs and pitfalls
Teams sometimes assume tokenizing a field is equivalent to encrypting it more strongly, when the real difference is architectural: tokenization requires a live, highly available vault dependency that encryption alone does not, and that dependency needs its own resilience plan.
Explain what data classification is and why it matters when designing an architecture that has to meet compliance requirements. Name at least three classification tiers you might use and how the controls differ between them.
Sample Answer
Direct answer: Data classification is the practice of labeling data by how sensitive it is, so that the controls protecting it, who can access it, how it's encrypted, how long it's kept, how it's logged, scale with actual risk instead of being applied uniformly everywhere. It matters for compliance-driven architecture because most regulatory and contractual obligations only apply to specific categories of data, so classification is what tells you where the expensive controls actually need to go.
Structured elaboration:
| Tier | Example data | Access | Encryption | Logging |
|---|---|---|---|---|
| Public | Marketing content, published documentation | Anyone | Not required | Minimal |
| Internal | Internal wikis, non-sensitive operational metrics | Employees and contractors | Recommended at rest | Standard |
| Confidential | Customer PII (personally identifiable information, data that can identify a specific individual), payment data, credentials | Named roles on a need-to-know basis | Required at rest and in transit | Detailed: who accessed what and when, retained longer |
A fourth tier, often called Restricted or Regulated, is common in practice for data subject to a specific named framework, healthcare records or payment card data, for example, where the controls aren't simply "more of the same" but a specific set of requirements, additional access reviews, defined retention limits, driven by that framework. Keeping it distinct from the general Confidential tier matters because those controls are prescribed externally, not decided internally.
Why it matters for architecture: once data is classified, each downstream decision, where it's allowed to be stored, which services may read it, whether it can leave a given region, how long it's retained, becomes a lookup against the tier rather than a one-off judgment call per system. That's both faster to design against and easier to audit later, since an auditor can check whether a Confidential-tagged table has the required controls rather than re-litigating its sensitivity from scratch.
Worked example: A payments feature stores a customer's email, Internal to Confidential depending on context, a shipping address, Confidential, and a card token, Restricted, since payment card data carries its own named handling requirements. Classifying these separately, rather than treating the whole payments database as one uniform blob, means the card token gets the strictest controls, the most restrictive access list, the longest audit retention, without forcing that same overhead onto the email field, which doesn't need it.
Trade-offs and pitfalls: Too many tiers, six or more, makes the system hard to apply consistently and people start guessing; too few, just "sensitive" and "not sensitive," loses the ability to right-size controls, which is the entire point of classifying in the first place. Three or four tiers is a practical sweet spot for most organizations.
Unlock Full Question Bank
Get access to all Data Protection and Encryption in Practice interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.