Data Governance, Contracts, and Classification Questions

Governing data at scale: data contracts between producers and consumers, schema evolution/compatibility, data classification and sensitivity tagging, access control, and lineage/cataloging. Covers policy, ownership, and compliance-driven controls over data. The governance layer over the technical stack.

HardTechnical
47 practiced

An analyst wants to join a table containing PII (say, user profiles) with an events table but should never see the raw PII columns unless specifically entitled, and in a multi-tenant warehouse a query should never be able to see another tenant's rows even through an intermediate step. Show how you'd structure the SQL (CTEs, views, column masking, row-level security) so that intermediate query steps can't leak PII or cross-tenant data to someone without the right permissions, and describe how you'd test that the protection actually holds.

MediumTechnical
48 practiced

Design a practical, org-wide strategy to detect and mask PII across all your streaming and batch pipelines, not just the ones someone remembered to flag, covering both raw lake data and curated warehouse tables. What detection approaches would you combine (schema tagging, regex pattern matching, ML-based classifiers), what masking or redaction strategy follows once something is found, and how would you handle the inevitable false positives and legitimate exceptions?

MediumTechnical
39 practiced

You're responsible for classifying the sensitivity of columns in a shared CRM dataset (say, contacts and accounts tables) and setting access policy for several internal roles with different needs. Design a classification scheme, a column-level access policy per role, and a masking or tokenization approach for anyone exporting the data. What's the trade-off between tightening this and keeping the sales and analytics teams productive?

HardSystem Design
40 practiced

Design an approach that offers both reversible pseudonymization (for internal debugging or an authorized audit) and irreversible anonymization (for general analytics) of the same PII fields, including for PII feeding an ML training pipeline. How do you manage the keys so re-identification is possible only through an approved, logged process, and how do retention windows interact with which form of the data is kept where?

HardTechnical
40 practiced

You need to let analysts join records across datasets by a shared identifier without ever exposing raw PII to them. What are your realistic options for doing this safely, and how would you choose between them, say for a one-time internal join between two warehouse tables versus a live join an external, less-trusted partner needs to perform against your infrastructure? Discuss the trade-offs each option makes between performance and how strongly it protects against re-identification.

Unlock Full Question Bank

Get access to all 15 Data Governance, Contracts, and Classification interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.