Data Platform Architecture and Technology Selection Questions
System-level design of an end-to-end data platform: component selection, build-vs-buy, tool trade-offs, and aligning platform architecture with organizational and analytics needs. Covers reasoning about the whole stack (ingestion through serving) and technology-choice justification. The architect-altitude view above any single pipeline.
Design a cost-allocation (showback/chargeback) model to attribute a shared data platform's compute and storage costs to the product teams that use it: tagging strategy, handling of shared resources fairly, reporting cadence, and a dispute-resolution process.
Sample Answer
A cost-allocation model has to answer, defensibly, "how much of this shared bill belongs to your team," which is straightforward for dedicated resources and genuinely hard for shared ones.
Tagging strategy
Require every compute job and every stored dataset to carry a team or product-owner tag at creation time, enforced by policy (a job or table without a valid tag either fails to run or gets flagged), not by asking people to remember. Tag at the finest practical granularity, per query or per job run, not just per dataset, since a shared dataset queried by five teams needs usage-level attribution, not just ownership-level attribution.
Handling shared resources fairly
For genuinely shared infrastructure (a common raw-data lake, a shared orchestration platform), allocate cost using a defensible proxy for actual usage, bytes scanned per team for a shared warehouse, compute-seconds consumed per team for a shared processing cluster, rather than an even split, which systematically overcharges light users and undercharges heavy ones. For genuinely indivisible shared cost (the platform team's own headcount, a base infrastructure license fee that doesn't vary with usage), either allocate it evenly as an acknowledged "platform tax" or pro-rate it by team size or overall usage share, whichever the organization has agreed is fair, but be explicit that this portion is a policy choice, not a measured usage number.
Reporting cadence
Report monthly at a minimum, with a runnable, self-service breakdown available to each team on demand rather than only in a monthly PDF, so a team can investigate a cost spike in the same week it happens rather than a month later when the details are hard to reconstruct. Include trend, not just the current month's number, since a single month's snapshot doesn't show a team whether their cost trajectory is improving or worsening.
Dispute-resolution process
Publish the tagging and allocation methodology itself (not just the resulting numbers) so a team can check the calculation, not just the conclusion. Provide a defined escalation path (a specific person or a lightweight review committee) for a team that believes its allocated cost is wrong, with a documented turnaround time, and track how often disputes are raised and resolved, since a rising dispute rate is itself a signal the methodology needs revisiting, not that teams are being difficult.
Worked example
If a shared raw-data lake costs a fixed amount in storage per month and three teams query it with wildly different frequency, allocate the storage cost by data volume each team's data occupies (a clean, defensible per-team number), and allocate the compute cost of QUERYING that lake by bytes scanned per team (measurable directly from query logs), rather than splitting either cost three ways evenly, which would let the lightest user subsidize the heaviest one indefinitely with no visibility into why.
Trade-offs and pitfalls
The most common mistake is allocating shared cost evenly because it's simple to compute, which quietly subsidizes heavy users at light users' expense and erodes trust in the whole model once someone notices. The second is building a cost model with no visible methodology or dispute path, which turns every allocation disagreement into a political argument instead of a checkable calculation.
You must evaluate whether to build a data-platform component in-house (e.g. a data catalog, an orchestrator) or adopt a managed/vendor solution, for an organization with mixed data maturity. Outline an evaluation framework covering multi-year total cost of ownership, vendor lock-in risk, time-to-value, and team skills, and describe how you would pilot before committing.
Sample Answer
The build-versus-buy question is really an evaluation-framework question: what does it cost to build and maintain this ourselves, what does it cost to adopt a managed solution, and which risk profile fits the team's actual skills and time horizon.
The evaluation framework
Total cost of ownership over 3+ years, not just sticker price: for building, this means engineer salary time for initial build plus ongoing maintenance, infrastructure costs, and the opportunity cost of that engineering time not going toward the product. For buying, this means license or usage-based fees plus the (often underestimated) integration, migration, and customization work.
Vendor lock-in risk: how hard would it be to leave this vendor later. Standard, open interfaces (SQL, open table formats, exportable metadata) are low risk; a proprietary API that "traps" your metadata or code is high risk.
Time-to-value: building takes months before the first real user benefits; a managed solution can often deliver value in weeks, which matters when the business need is urgent.
Team skills and headcount: building and then operating a component like a data catalog or an orchestrator requires ongoing engineering attention, not just an initial build. A team without a dedicated platform engineer is signing up for perpetual part-time maintenance if it builds.
Worked example: build or buy a data catalog
Say a 500-person company with mixed data maturity is deciding whether to build a lightweight internal catalog (a service plus a UI reading table metadata from the warehouse) or adopt a managed catalog product. Building might cost an estimated 2 engineers for 3 months to reach a usable first version (roughly half an engineer-year), then an ongoing 20 to 30 percent of one engineer's time to maintain and extend it as new data sources are added. Buying costs a recurring subscription fee, but reaches a usable state in weeks rather than months, and the vendor absorbs the maintenance burden of keeping up with new source-system integrations. If the company has no dedicated platform team and catalog needs are fairly standard (ownership, lineage, search), buying wins on time-to-value and total cost once you count engineer time honestly. If the company's catalog needs are unusual (deep integration with a proprietary internal system no vendor supports out of the box), building may be the only realistic option regardless of cost.
How to pilot before committing
Run a bounded pilot, a single team or a handful of high-traffic datasets, on the leading option (build a minimal version, or trial the vendor) for a fixed period, with a small number of concrete success criteria defined up front: does search actually return the right table in under a specified number of clicks, does the team's satisfaction improve, does the pilot surface integration gaps the evaluation missed. A pilot without predefined success criteria tends to be judged only by whether it "felt fine," which doesn't generalize to the full rollout decision.
Trade-offs and pitfalls
The single biggest analytical mistake in these evaluations is comparing the vendor's subscription price directly against a build estimate that only counts initial development time, ignoring the ongoing maintenance a build requires. The second is treating "no lock-in" as an absolute good: a small team may genuinely benefit from accepting some vendor dependency in exchange for never having to think about that component again, if the vendor uses reasonably open interfaces underneath.
Tell me about a time you had to convince leadership or stakeholders to adopt a new data architecture or technology decision, such as moving from nightly batch to streaming, or adopting a new platform standard, despite short-term disruption. How did you build the case, and what was the outcome?
Sample Answer
A strong answer here centers on a specific decision, the disruption it caused, and the concrete evidence used to win the argument, not a general description of "being a good communicator."
What a strong story includes
The specific decision and why it faced resistance: name a concrete architecture or technology change (batch to streaming, adopting a new warehouse or platform standard) and the real, legitimate reason people were skeptical, usually short-term disruption, migration risk, or a team's existing investment in the current approach. Skepticism grounded in a real cost is more credible than a strawman objection.
The case actually built: what evidence moved the decision. This is usually a combination of a quantified current pain point (a specific, measured cost or limitation of the status quo) and a bounded, low-risk way to demonstrate the new approach before asking for full commitment, a pilot on one team or one dataset rather than a company-wide bet up front.
Handling pushback: name a specific objection someone raised and how you responded to it, ideally by addressing the underlying concern directly (offering a rollback plan, scoping the pilot smaller, bringing in a skeptic as a collaborator on the pilot) rather than simply repeating the case louder.
The outcome: a specific, verifiable result, the pilot's measured outcome, the decision that followed, and (if enough time has passed) whether the change held up under real usage.
Worked example structure
Situation: the team ran nightly batch reporting, and a growing subset of stakeholders needed same-day answers, but the prevailing view was that streaming was too operationally risky for a team without deep streaming experience.
Task: build the case for adopting a streaming path for the specific use cases that needed it, without a wholesale rip-and-replace of the working batch system.
Action: quantified the actual business cost of the current staleness (a specific recurring decision stakeholders were making on data that was, say, 18 hours old when same-day would have changed the decision), proposed a scoped pilot streaming path for just that one use case rather than the whole platform, and directly addressed the loudest objection (operational risk) by proposing the pilot run alongside the existing batch path rather than replacing it, so a failure would be low-stakes.
Result: the pilot ran for a defined period, the specific business metric it was meant to improve moved measurably, and that evidence, not the original argument alone, is what got broader adoption approved.
Trade-offs and pitfalls
The most common weak version of this story skips straight from "I proposed it" to "it was approved," with no specific pushback and no specific evidence, which reads as either an oversimplified account or a decision that didn't actually face real resistance. The strongest version names the real, legitimate cost the skeptics were worried about and shows the case was won by addressing that cost directly (a pilot, a rollback plan), not by overriding the concern.
Your analytics stack has grown increasingly dependent on a single cloud vendor's proprietary features. Assess the vendor lock-in risk and propose a playbook: short-term mitigations, a long-term migration strategy, and a rough cost estimate for regaining independence.
Sample Answer
Vendor lock-in risk assessment starts with a concrete question: if we had to leave this vendor in six months, what specifically would break, and how long would it take to rebuild.
Assessing the risk
Audit which proprietary features are actually load-bearing versus merely convenient: a proprietary SQL extension used in every transformation query is load-bearing (rewriting it is real work); a UI feature used for occasional debugging is not. Check whether the data itself is stored in an open, portable format (an open table format, standard file formats) or a proprietary storage layer that only that vendor's engine can read; the latter means even "exporting your data" doesn't actually free you, because the export format may not preserve what made the data useful. Check whether your team's operational knowledge (tuning, troubleshooting, cost optimization) has become vendor-specific tribal knowledge that wouldn't transfer.
The playbook
Short-term mitigations: stop adding NEW dependencies on proprietary features going forward, even while continuing to use existing ones, so the lock-in stops growing while you plan the response. Where feasible, start writing new transformation logic in a portable form (standard SQL, or a transformation tool that can target multiple backends) rather than the vendor's proprietary dialect.
Long-term migration strategy: prioritize migrating away from the SPECIFIC features that are hardest to replace, not the easiest ones, since the hard ones are exactly what makes leaving expensive later; tackle them while the total surface area is still smaller than it will be in another year of continued dependency. Where full migration isn't justified yet, invest in an abstraction layer (a transformation framework, a query interface) that could be repointed at a different backend with less rewrite effort than a direct migration would need.
Rough cost estimate for regaining independence
Estimate cost as the sum of: engineering time to rewrite proprietary-feature-dependent logic in a portable form, the cost of running two systems in parallel during validation (a real, often underestimated cost), and the opportunity cost of the engineering time not spent on other roadmap work during the migration window. For an organization with, say, a few dozen pipelines with moderate proprietary-feature dependency, this commonly lands in the range of several engineer-months to a few engineer-years, heavily dependent on how deep the dependency actually runs, which is exactly why the audit step above has to come first: without it, this estimate is a guess rather than a plan.
Trade-offs and pitfalls
The most common mistake is treating "no vendor lock-in" as an absolute goal worth paying for regardless of actual risk; a small team may rationally accept meaningful lock-in in exchange for never having to operate a component themselves, if the switching cost, honestly estimated, is genuinely low relative to the ongoing operational savings. The second mistake is assessing lock-in risk only at adoption time and never revisiting it; dependency accumulates gradually as teams add "just one more" proprietary feature over years, so the actual risk at year three is rarely what it looked like at year one.
Compare a data mesh (federated, domain-oriented data ownership) to a centralized data platform. Discuss ownership, discoverability, governance, latency, cost, and developer velocity, and describe when an organization should favor one approach over the other.
Sample Answer
A data mesh decentralizes data ownership to the domain teams that generate it (marketing, orders, payments), each publishing "data products" with clear contracts, while a centralized platform keeps one team owning ingestion, transformation, and the warehouse for the whole company.
The comparison
Ownership: mesh puts the people closest to the data (who understand it best) in charge of its quality and its contract; centralized puts one platform team in charge of everyone's data, whether or not they understand every domain equally well.
Discoverability: mesh requires a strong shared catalog and standardized metadata across domains, or discovery becomes worse than centralized, since data is now scattered across many owners. Centralized discovery is simpler by construction because everything sits in one place.
Governance: mesh uses federated computational governance, shared standards enforced by tooling (schema validation, SLAs as code) rather than a single team manually reviewing everything; centralized governance is easier to enforce consistently because one team controls the whole pipeline, but that team becomes a bottleneck as the company grows.
Latency (time-to-new-data-product): mesh domain teams can ship a new data product without waiting on a central team's backlog; centralized platforms often become the bottleneck once a company has more than a handful of domains competing for the same platform team's attention.
Cost: mesh usually costs more in tooling and duplicated infrastructure across domains; centralized concentrates cost (and the ability to optimize it) in one place.
Developer velocity: mesh scales velocity horizontally (more domains means more parallel capacity) once the platform is mature; centralized velocity is capped by the size of the platform team, and that cap gets worse as the company scales.
When to favor each
Favor a centralized platform when the organization has fewer than roughly a dozen distinct data domains, when the central platform team is not yet a bottleneck, and when the cost and complexity of federated governance tooling would outweigh its benefit. A 50-person company with three product lines rarely needs a mesh; it needs a competent, well-staffed central team.
Favor a data mesh once a large organization has enough independent domains (order of dozens or more) that a central team has become the bottleneck for every new dataset, and once there is executive appetite to invest in the self-serve platform and governance tooling a mesh requires to avoid becoming an ungoverned mess of inconsistent, undiscoverable domain data. A mesh adopted before that tooling investment is made tends to produce exactly the failure mode critics warn about: each domain reinventing its own formats and quality bar, with no way to reliably query across domains.
Trade-offs and pitfalls
The most common mistake is adopting "data mesh" as an organizational restructuring (just move ownership to domain teams) without investing in the platform and governance tooling that makes federation work; the result is worse than a centralized platform, because now nobody is accountable for cross-domain consistency at all. The second is applying mesh principles to an organization too small to need them, which adds coordination overhead (data contracts, domain-team accountability meetings) that a single central team would have handled faster.
Unlock Full Question Bank
Get access to all 11 Data Platform Architecture and Technology Selection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.