System Design Methodology and Trade-off Analysis Questions
The end-to-end approach to an open-ended design problem and the judgment that resolves it: clarifying scope and constraints, gathering functional and non-functional requirements, capacity and back-of-envelope estimation, and mapping requirements to a high-level architecture, then reasoning explicitly about competing options on cost, complexity, latency, and reliability to defend a choice. Covers driving a design interview from ambiguity to a proposal, trade-off frameworks, decision-making under uncertainty and incomplete information, reversible-versus-irreversible decisions, and defending choices under scrutiny. The process-and-judgment skill underneath every system-design case study.
You need to map a requirements list for a payment-processing subsystem (99.99% availability, sub-200ms p95 authorize latency, PCI-DSS compliance, 7-year data retention, and a fixed monthly budget) onto an actual architecture. How would you structure that mapping, and walk through three example rows: which requirement drove which component, and what you gave up to satisfy it?
Sample Answer
Direct answer
Structure the mapping as a matrix: one row per requirement, columns for the target metric, the component(s) that satisfy it, and what you gave up to get there. Walking three rows for this payment subsystem: 99.99% availability drives multi-availability-zone (multi-AZ) redundancy at the cost of doubled infrastructure and failover complexity; sub-200ms p95 (95th-percentile) authorize latency drives a token cache and dedicated crypto hardware at the cost of extra compute spend; and PCI-DSS (Payment Card Industry Data Security Standard) plus 7-year retention drives tokenization and immutable long-term storage at the cost of losing raw-card analytics fidelity and paying for years of storage.
Structured elaboration
Use a table with these columns for every requirement in the list:
| Column | What it captures |
|---|---|
| Requirement | The stated constraint, in one line |
| Target / metric | The number you're accountable for (99.99%, <200ms p95, 7 years) |
| Component(s) | What actually implements it |
| Metric to instrument | How you'd know if you're meeting it in production |
| Cost impact | Rough $/month or engineering-time delta |
| What you gave up | The trade-off accepted to hit the target |
This format forces every requirement to land on a concrete component and a concrete cost, rather than staying as an aspiration in a requirements document. It also makes conflicts visible: if two rows both compete for the same fixed budget, that surfaces in the table instead of being discovered mid-build.
Worked example
Three rows from the matrix, with the underlying arithmetic shown:
Row 1: 99.99% availability. A 99.99% target permits:
allowed downtime/year=(1−0.9999)×365×24×60 min=52.56 min/year
Component: the authorize API runs multi-AZ with automated failover rather than a single instance. Gave up: roughly double the compute footprint (active-active or hot-standby) plus the operational cost of regularly testing failover, in exchange for that 52.56-minute annual downtime budget instead of the far larger downtime a single-AZ deployment would risk.
Row 2: sub-200ms p95 authorize latency. An illustrative latency budget that sums to the target:
20ms (network)+30ms (tokenize/HSM)+50ms (fraud rules)+20ms (cache read)+60ms (network to processor)+20ms (buffer)=200ms
Component: an in-memory cache for token lookups and a hardware security module (HSM) colocated with the authorize path, rather than a network round trip to a shared crypto service. Gave up: dedicated cache and HSM capacity that sits idle outside peak hours, which is more expensive per request than a shared pool would be.
Row 3: PCI-DSS plus 7-year retention. Assume, as illustrative pinned inputs, 1 million transactions/day and a 2 KB (kilobyte) retained metadata record per transaction (tokenized, not raw card data):
bytes/day=1,000,000×2KB=2,048,000,000 bytes≈2.05 GB/day
total (7yr)=2.05 GB/day×365.25×7 days≈5,236 GB≈5.2 TB
Component: a tokenization service so raw card numbers never enter long-term storage, plus write-once immutable object storage for the 5.2 TB of retained metadata. Gave up: the ability to run ad hoc analytics on raw card attributes, since only tokens and derived fields are retained.
Trade-offs & pitfalls
- The fixed monthly budget row is where the other three collide: if multi-AZ plus dedicated cache/HSM plus 7 years of immutable storage exceeds the budget, something has to re-scope, not silently degrade in production.
- A weak answer lists components without naming what was given up; the "what you gave up" column is the actual trade-off-analysis signal, not the component list itself.
- Treat compliance requirements (PCI-DSS, retention) as filters applied before cost optimization, not something to negotiate down after the architecture is built.
- Revisit the matrix at each design review; a requirement's target or its owning component can shift as the system evolves, and a stale matrix gives false confidence.
For a social feed serving 200M monthly active users and 10k writes/sec, would you fan out a new post to followers' feeds on write, or compute the feed on read? What does each choice cost you, and how would a celebrity account with millions of followers change your answer?
Sample Answer
Direct answer
For 200 million monthly active users (MAU) and 10,000 writes/sec, a pure fan-out-on-write pushes every new post into every follower's feed at write time, buying very low read latency at the cost of massive write amplification and storage. Pure fan-out-on-read defers that work to feed-view time, keeping writes cheap but making every read do more work. A celebrity account with millions of followers breaks the pure push model outright, which is why the practical answer is a hybrid: push for ordinary accounts, pull (or a separate merge step) for very high-fan-out accounts.
Structured elaboration
| Dimension | Fan-out-on-write (push) | Fan-out-on-read (pull) |
|---|---|---|
| Storage | High: one copy of the post lands in every follower's inbox | Low: one canonical copy per post |
| Read latency | Very low: a feed read is a simple lookup | Higher and more variable: must merge recent posts from every followee at read time |
| Write amplification | O(followers) per post; scales with fan-out size | O(1) per post; writes stay cheap regardless of follower count |
| Rebuild after failure | Complex: losing the inbox store means replaying historical writes | Simple: the canonical post store is the source of truth, caches are just recomputed |
| Best fit | Accounts with small-to-medium follower counts | Accounts with very large follower counts (celebrities) |
The decision criterion is the read:write ratio implied by a given account's follower count, not a single global choice: an account followed by 200 people generates trivial fan-out and huge read-latency benefit from push; an account followed by millions generates enormous fan-out for a benefit (marginally faster reads for those followers) that pull-at-read can approximate at read time instead.
Worked example
At 10,000 writes/sec, assume (illustrative, pinned input) an average of 300 followers per post for non-celebrity accounts:
fan-out ops/s=10,000 writes/s×300 avg followers=3,000,000 inbox writes/s
That is the write-amplification cost a pure push model pays continuously just for ordinary accounts.
Now take one celebrity post going to 5 million followers, and assume (illustrative) a fan-out cluster capable of sustaining 500,000 inbox writes/s:
time to fan out one celebrity post=500,000 writes/s cluster capacity5,000,000 followers=10 s
A single celebrity post would take roughly 10 seconds to fully propagate through push fan-out, and that's before accounting for every other post competing for the same fan-out capacity at the same time. This is the concrete reason celebrity accounts change the answer: pushing their posts synchronously into millions of inboxes is not just expensive, it measurably delays delivery to everyone else sharing that fan-out capacity.
Trade-offs & pitfalls
- Treating fan-out-on-write and fan-out-on-read as a single global choice, rather than a per-account decision keyed on follower count, is the most common shallow answer.
- A hybrid design still needs a merge step at read time for celebrity posts, so pull-style merge logic doesn't disappear; it just gets scoped to a small fraction of accounts instead of all of them.
- Async, idempotent fan-out pipelines are required regardless of strategy, because retries and partial failures are certain at this scale; a synchronous fan-out-on-write implementation is a reliability risk independent of the storage trade-off.
- Caching the celebrity's own recent posts aggressively (rather than fanning them out) reduces the read-time merge cost without reintroducing full push fan-out.
How do you decide the right granularity when splitting a system into services? Walk through how coupling versus cohesion, data ownership, and team boundaries change your answer.
Sample Answer
Direct answer
Split along business capability and data ownership, not by technical layer, and treat coupling and cohesion as the actual test: a service boundary is right when it groups things that change together and separates things that don't, and when one team can own its full lifecycle (build, deploy, operate) without waiting on another team to also deploy. Team size and deployment cadence usually decide the timing more than the theory does: a well-modularized monolith can run comfortably until the coordination cost of shared deploys and shared blast radius starts to exceed the operational cost of running the same code as separate services.
Structured elaboration
The criteria, applied together
- Bounded context or business capability: one service per coherent business concept (Orders, Inventory, Billing), not per database table.
- Data ownership: the service that owns a piece of data is its only writer; everyone else goes through its API or its events, never a shared schema.
- Deployment independence: if two "services" cannot be deployed on separate schedules without breaking each other, they are one service wearing two names, a distributed monolith.
- Team boundaries (Conway's Law: a system's structure tends to mirror the structure of the team that builds it): align a service to a team that can own it end to end, so ownership and org chart don't fight each other.
- Transaction boundary: keep operations that need a real ACID (atomicity, consistency, isolation, durability) transaction inside one service; cross-service consistency should default to eventual consistency plus an explicit compensating action, not a distributed transaction.
- Chattiness: if two components exchange many synchronous calls per user request, the network hop between them is pure overhead with no ownership benefit; merge them.
The team-size-driven worked example
Consider an org at 200 people, organized as roughly 20 teams, running a well-modularized monolith with clear internal module boundaries (a modular monolith). Model the shared deploy pipeline as a single server processing one deploy at a time, 30 minutes each, across a 16-hour working day (960 minutes):
deploy capacity/day=30960=32 deploys demand at 20 teams (1 deploy/day each)=20 deploys/day utilization=3220=62.5%At 62.5% utilization there is queueing delay, but the pipeline is stable. Now grow to 500 people, roughly 50 teams, same one-deploy-at-a-time pipeline:
demand at 50 teams=50 deploys/day>32 deploys/day capacityDemand exceeding capacity on a single-server queue means the queue is unstable: it does not just get slower, it grows without bound. That crossing point, not a stylistic preference for microservices, is the concrete signal to start extracting services along the module boundaries the modular monolith already has, so teams stop sharing one serialized deploy pipeline and one shared blast radius.
Anti-patterns that signal you split wrong (or didn't split at all)
- Shared database schema across "separate" services: the clearest sign of a distributed monolith with extra network hops.
- Splitting by technical layer (a UI service, an API service, a database-access service) instead of by capability: nothing can deploy alone, because every user-facing change touches all three.
- A "god" service or shared library that every team depends on for routine changes: it recreates the same coordination bottleneck a monolith had, with worse debugging.
- Over-splitting a capability that still needs real ACID guarantees just because a diagram looks tidier with more boxes.
Trade-offs & pitfalls
- Splitting too early, before the coordination cost above actually bites, buys distributed-systems complexity (network calls, partial failure, eventual consistency) for a coordination problem you didn't have yet.
- Splitting too late means the deploy-pipeline math above turns into a real, measured queue of waiting teams, not a hypothetical.
- The bounded-context choice is the expensive one to get wrong: correcting a wrong service boundary later means a data migration, not just a configuration change.
- Watch for teams treating microservices as a goal instead of a response to a specific coupling problem; the checklist above should produce the boundary, not the other way around.
What's the difference between a high-level architecture (system context and major components) and a component-level design (interfaces, data flows, sequencing)? What would you actually show stakeholders at each level, and what's one decision that only makes sense at the high level?
Sample Answer
Direct answer
A high-level architecture shows the system's scope: the major building blocks (client, API layer, service tier, datastore, cache, external dependencies), how they relate, and the non-functional constraints (scale, availability) that shaped them. A component-level design zooms into one of those blocks and specifies its interfaces, request/response schemas, data flows, and sequencing. You show the high-level view to stakeholders who need to understand what the system is and what it costs or risks; you show component-level design to the people who have to build, test, or integrate against one specific piece.
Structured elaboration
| Dimension | High-level architecture | Component-level design |
|---|---|---|
| Purpose | Scope, responsibilities, external actors, major blocks, non-functional constraints | Internals of one component: interfaces, data formats, control flow, error paths, sequencing |
| Typical diagrams | System context diagram, high-level component diagram, deployment diagram (regions, load balancers, replicas) | Sequence diagram for a specific flow, API contract (request/response schema), data model / entity-relationship diagram |
| Audience | Product managers, other architects, executives, site reliability engineers (SRE), business stakeholders | Backend/frontend engineers, QA, API consumers, integration partners |
| Question it answers | "What is this system, and what are its risk and cost boundaries?" | "How exactly does this one feature work end to end?" |
| Example decision that only lives here | Monolith vs microservices for the whole platform (changes team structure, operational model, and cost) | The exact endpoint shape, schema, and authentication header format for one API |
The reason both layers matter: the high-level view sets the strategy and the constraints everyone else has to work inside; the component-level view is what actually gets implemented, tested, and integrated. A good design doc keeps an explicit mapping from each high-level block down to its component-level detail, so a reviewer can move between the two without re-deriving context.
Worked example
Say you're designing a subscription billing feature. At the high level you'd draw: client apps, an API gateway, a billing service, a payments component, a database, and a message queue for async notifications, with an arrow showing the billing service calls out to a third-party payment processor. The one decision that belongs only at this level: whether billing lives inside the existing monolith or is split into its own service, because that choice affects deployment, on-call ownership, and the blast radius of an incident, not just this one feature.
At the component level, you'd zoom into just the billing service and produce: a sequence diagram for "create subscription" (client → billing service → payments component → processor → database write → event published), the exact request/response schema for the POST /subscriptions endpoint, and an entity-relationship diagram for the subscription and invoice tables. None of that detail belongs on the high-level diagram; it would bury the one decision (monolith vs separate service) that the high-level view exists to surface.
Trade-offs & pitfalls
- Showing component-level detail (full schemas, every retry path) to an executive or product stakeholder buries the one decision they actually need to weigh in on.
- Skipping the high-level view and jumping straight to component design risks locking in a boundary (a shared database, a synchronous call where an event would do) that is expensive to undo later, because it was never surfaced as a decision.
- A common weak answer just says "high-level is the big picture, low-level is the details" without naming a decision that is exclusive to one level; naming that decision is the signal an interviewer is listening for.
- Keep a living link between the two artifacts (a component-level design should reference which high-level block it belongs to) so the documentation doesn't drift apart as the system evolves.
A service is reported to become CPU-bound under heavy load. How would you design an experiment to confirm whether the real bottleneck is CPU, network, or I/O, rather than taking that claim at face value?
Sample Answer
Direct answer
Treat "CPU-bound" as a hypothesis to disprove, not a fact to accept. Instrument all three candidate resources at once (CPU, network, disk/I/O) under the same load, then run bounded isolation experiments that stress one resource at a time to see which one, when constrained, actually reproduces the reported slowdown. A claim of CPU-bound only holds up if CPU utilization is near saturation while the other two are not, and if artificially limiting CPU makes latency worse while limiting network or disk does not.
Structured elaboration
Metrics to collect per candidate resource, under the same load window:
- CPU: percent user and system time, run-queue length (how many processes or threads are ready to run but waiting for a free CPU core, so a growing number means work is piling up faster than the CPU can drain it), per-core utilization, and context switches (how often the CPU swaps between tasks, which adds overhead and can signal contention even when raw CPU percent looks moderate).
- Network: throughput, retransmits, socket queue depth, round-trip time.
- Disk / I/O: percent I/O wait, average read/write latency, queue depth.
- Application: request rate, latency distribution, error rate, so the resource data can be aligned against the actual symptom.
Experiment design:
- Baseline under light load, then reproduce the reported heavy load with a controlled, documented load generator.
- Ramp load stepwise, capturing all resource metrics at each step to see which resource's utilization tracks the load ramp most tightly.
- Isolation tests: constrain one resource at a time (limit CPU to fewer cores, throttle network bandwidth, saturate disk I/O with a separate workload) and observe whether application latency degrades specifically when that resource is constrained.
- If CPU tracks the symptom, use a sampling profiler during the ramp to identify which functions are actually consuming the cycles, this distinguishes genuine compute-bound work from CPU time spent spinning on a lock.
Interpreting the signals, the actual diagnostic logic:
- CPU utilization near saturation (say, 90%+ of user and system time) with a growing run-queue, and latency that worsens specifically when CPU is artificially constrained further, is consistent with CPU-bound.
- Moderate CPU utilization (for example, 50-60%) while latency is still climbing under load is inconsistent with a pure CPU-bound explanation, that pattern points toward lock contention, a downstream dependency, or I/O wait instead, even though the process may show elevated CPU time from spinning.
- High I/O wait with low CPU user/system time, worsened specifically when disk is stressed, points to I/O-bound rather than CPU-bound.
- Retransmits, saturated network interface throughput, or growing socket queues, worsened specifically when bandwidth is throttled, point to network-bound.
Worked example
Consider the diagnostic logic concretely: if profiling shows CPU utilization at 55% during the reported slowdown, with the run-queue not growing, but request latency still rises as concurrency increases, that combination is inconsistent with the CPU-bound claim, because a genuinely CPU-bound service would show utilization tracking toward saturation as latency degrades. The next check, rather than accepting either conclusion on this alone, is the isolation test: artificially cap available CPU further (for example, via a CPU limit or cgroup) and observe whether latency gets meaningfully worse. If it does not, CPU is not the binding constraint regardless of what the initial monitoring dashboard reported, and the same targeted constrain-and-observe check should be repeated for network bandwidth and disk I/O until one of the three actually moves the needle.
Trade-offs & pitfalls
- Trusting a single metric (CPU percent) without run-queue or context-switch context conflates "CPU busy" with "CPU is the bottleneck," a thread spinning on a lock also shows as CPU time but the real fix is a concurrency bug, not more compute.
- Isolation tests using synthetic stress tools can introduce noisy-neighbor effects that don't reflect production traffic patterns, treat isolation results as directional evidence, not a final answer, and validate against production telemetry.
- Skipping the baseline-and-ramp step and jumping straight to isolation tests risks constraining a resource that was never actually near its limit, wasting a change window on a resource that wasn't the constraint.
- Presenting a conclusion without the underlying time-aligned dashboards and the exact isolation test performed makes the finding hard for others to trust or reproduce, keep the experiment scripted and the artifacts (metrics, profiler output) attached to the conclusion.
Unlock Full Question Bank
Get access to all System Design Methodology and Trade-off Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.