InterviewStack.io LogoInterviewStack.io

Data Governance, Contracts, and Classification Questions

Governing data at scale: data contracts between producers and consumers, schema evolution/compatibility, data classification and sensitivity tagging, access control, and lineage/cataloging. Covers policy, ownership, and compliance-driven controls over data. The governance layer over the technical stack.

HardSystem Design
39 practiced

Design an approach to unify metadata catalogs across multiple cloud providers (say AWS Glue, GCP Data Catalog, and Azure Purview) into a single place analysts can search regardless of where a dataset actually lives. How do you keep ownership and sensitivity tags synchronized, and resolve conflicts when the same dataset is described differently in two systems?

HardSystem Design
46 practiced

Design a data governance and access-control architecture for an analytics platform operating across multiple regions with different privacy regimes (say, the EU and the US), covering product analytics, ML feature stores, and customer records. Cover data classification, PII detection and masking, role-based access control, lineage, audit trails, and how you'd structure a governance decision-making body, while keeping ML feature freshness and developer agility workable.

MediumTechnical
35 practiced

Design a secure workflow for labeling sensitive enterprise documents that must respect strict tenant isolation: private per-tenant workspaces, role-based access, audit logging, and a way to check labeling quality (inter-rater agreement) without any tenant's data leaking into another's view.

HardSystem Design
43 practiced

Design a policy enforcement architecture for data-lake governance using a policy-as-code engine (in the style of Open Policy Agent). Where in the stack do policies get evaluated (at ingest, in the catalog, at a query gateway), how are policy changes rolled out safely, and how do you audit what was actually enforced versus what was merely written down?

HardSystem Design
42 practiced

Design a data catalog and lineage layer that spans hundreds of datasets across a data lake, warehouse, and ML artifact store for a large consumer product company. Cover the architecture, how metadata and lineage get captured automatically versus curated by hand, how access controls and sensitivity tags plug in, how a data-retention and deletion policy hooks into the same system, and how you'd measure whether teams are actually adopting it rather than just tolerating it.

Unlock Full Question Bank

Get access to all 6 Data Governance, Contracts, and Classification interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.