System Design Methodology and Trade-off Analysis Questions
The end-to-end approach to an open-ended design problem and the judgment that resolves it: clarifying scope and constraints, gathering functional and non-functional requirements, capacity and back-of-envelope estimation, and mapping requirements to a high-level architecture, then reasoning explicitly about competing options on cost, complexity, latency, and reliability to defend a choice. Covers driving a design interview from ambiguity to a proposal, trade-off frameworks, decision-making under uncertainty and incomplete information, reversible-versus-irreversible decisions, and defending choices under scrutiny. The process-and-judgment skill underneath every system-design case study.
Your request path chains three components in series, each sitting at 99.9% availability on its own. How does that combine into your end-to-end availability, and if the SLA actually requires 99.99%, what would you be willing to spend to close that gap?
Sample Answer
Direct answer
For three components in series, end-to-end availability is the product of the individual availabilities, so 99.9% x 99.9% x 99.9% is well below 99.9% overall, since every additional serial link can only make things worse, never better. Closing the gap to a 99.99% service-level agreement (SLA) target costs money in redundancy, and how much you're willing to spend depends on the business cost of the downtime you're trying to eliminate versus the cost of the additional infrastructure and operational complexity redundancy requires.
Structured elaboration
For independent components A, B, C in series, the system availability is:
Asys=A×B×C
If all three sit at 99.9% (0.999) individually:
Asys=0.9993≈0.997→99.7%
that's a meaningfully worse number than any single component's own 99.9%, which is the core insight: serial dependencies compound failure probability, they don't average it.
To hit 99.99% with identical components and no redundancy, each component alone would need to reach:
p=(0.9999)1/3≈0.999967→99.9967% per component
That's a very high bar for a single component to hit on its own. The more common and often cheaper path is redundancy: for n independent replicas of a component each at availability p:
Acluster=1−(1−p)n
With p=0.999 (the original 99.9% component) and n=2 replicas:
Acluster=1−(0.001)2=0.999999→99.9999%
Replacing all three serial components with a 2-replica redundant cluster each gives:
Asys=(0.999999)3≈0.999997→99.9997%
well above the 99.99% target, using only 2x redundancy per component rather than pushing any single component to a much harder 99.9967% target.
Worked example
Translating these percentages into downtime budgets makes the trade-off concrete. At the original 99.99% target:
downtime at 99.99%/year=(1−0.9999)×525,600 min≈52.6 min/year
At the redundancy-based 99.9997% actually achieved above:
downtime at 99.9997%/year≈(1−0.999997)×525,600 min≈1.6 min/year
So 2x redundancy per component doesn't just meet the 99.99% SLA, it clears it with roughly 51 fewer minutes of annual downtime budget spent. That headroom is what you'd spend, or not spend, depending on the actual cost of an outage minute to the business: if an outage minute costs far less than the added infrastructure and on-call complexity of running every component as an active redundant pair, the cheaper path might be accepting the tighter 99.99% target with only one component redundant, not all three.
Trade-offs & pitfalls
- The multiplicative model assumes independent failures; shared infrastructure (the same power feed, the same network path, the same region) breaks that independence, and correlated failures can quietly undermine a redundancy plan that looks solid on paper.
- Redundancy only helps if failover is fast and reliable; factoring in mean time to recovery (MTTR) alongside mean time between failures (MTBF) matters as much as the raw availability percentage.
- Don't over-invest in the component that's cheapest to make redundant if it isn't the actual constraint; prioritize redundancy where it buys the most availability per dollar, not where it's easiest to implement.
- A common shallow answer stops at "multiply the availabilities" without translating the result into a downtime budget and a cost decision; the interviewer is listening for the willingness-to-spend reasoning, not just the formula.
How do you evaluate build-vs-buy for a core platform capability like authentication or observability? What technical, organizational, and financial criteria drive the decision?
Sample Answer
Direct answer
Build-vs-buy for a core platform capability like authentication or observability comes down to weighing three sets of criteria: technical (does the vendor cover the required functionality without excessive integration work), organizational (does the team have the skills and bandwidth to build and operate it, and is it a genuine differentiator worth owning), and financial (total cost of ownership over several years versus subscription cost, and the opportunity cost of the engineering time either path consumes). Commodity capabilities with real compliance or reliability requirements usually favor buying; capabilities that are a genuine competitive differentiator favor building.
Structured elaboration
Technical criteria: feature coverage against requirements (for authentication: single sign-on/OpenID Connect support, role-based access control; for observability: traces, metrics, logs, retention), integration complexity and API quality, scalability and the vendor's own reliability track record, and how hard it would be to migrate away later (portability, data export).
Organizational criteria: whether the team has the skills and spare capacity to build and operate this well, how urgent time-to-market pressure is, and whether this capability is a genuine long-term differentiator for the product or a commodity everyone needs and nobody differentiates on. Building a commodity capability is usually a distraction from what the team should be differentiating on.
Financial criteria: total cost of ownership (TCO) over a multi-year horizon (engineering time to build and maintain, hosting, licensing), the opportunity cost of the features that don't get built while the team builds this instead, and how predictable vendor pricing is versus the variability of an in-house maintenance burden.
This same framework generalizes beyond auth and observability. Adopting a search-as-a-service offering instead of operating a self-hosted search cluster turns on exactly the same criteria: cost, time-to-market, vendor lock-in, the quality of the vendor's service-level agreements (SLAs), and compliance requirements, the same list that applies to auth-as-a-service.
Worked example
Take authentication for a product with 500,000 monthly active users (MAU), comparing a managed identity provider against building in-house. Buy, at an illustrative $0.05 per MAU per month:
buy: 500,000 MAU×$0.05/MAU-month=$25,000/month⇒$25,000×36=$900,000 over 3 years
Build, assuming two engineers dedicated to building and operating it at an illustrative fully loaded cost of $180,000/year each:
build: 2 engineers×$180,000/year=$360,000/year⇒$360,000×3=$1,080,000 over 3 years
At this scale and time horizon, buying is roughly $180,000 cheaper over three years, before even counting the opportunity cost of the two engineers' time not going toward the product's actual differentiator. That gap would close or reverse at a different MAU count or a different per-MAU vendor price, which is exactly why this needs to be computed per situation rather than assumed.
Trade-offs & pitfalls
- The crossover point between build and buy moves with scale (MAU, request volume): a TCO comparison done once at launch can become wrong as the product grows, so it should be revisited, not treated as permanent.
- Vendor lock-in risk is real but is a cost to be mitigated (contract exit clauses, data portability, thin integration layers), not an automatic reason to build; building in-house has its own lock-in in the form of institutional knowledge walking out the door.
- A common weak answer treats "build" as inherently more control and "buy" as inherently faster, without pricing either side; the financial criterion is the one most often skipped under interview time pressure.
- Compliance requirements can flip the decision entirely regardless of cost: if a vendor can't meet a required certification, buy is off the table no matter how favorable the TCO looks.
You're designing for a messaging app with 1M monthly active users. Midway through, you learn a new feature will increase message throughput by 10x. What changes about your design, and how do you decide what to revisit versus leave alone?
Sample Answer
Direct answer
A 10x jump in message throughput doesn't uniformly stress every part of a messaging app's design; it stresses the components whose load scales directly with message volume (the message broker, delivery workers, database writes for messages) and leaves largely untouched the components whose load scales with something else (user authentication, profile lookups, once-per-session connection setup). Deciding what to revisit versus leave alone comes down to tracing which components' load is actually a function of message throughput.
Structured elaboration
For each system component, ask: does its load scale with message volume, with active user count, or with something independent of both? That answer decides whether the 10x change touches it.
Scales with message throughput, revisit: the message broker/queue (partition count and per-partition throughput), delivery/fan-out workers, database write capacity for message storage, and any per-message monitoring or logging pipeline.
Scales with user count or session activity, mostly leave alone: authentication, user profile storage, push-notification token registration, and connection/session management, none of which get 10x busier just because message volume did.
Needs a fresh look regardless: cost forecasting (10x throughput changes the cost curve even where architecture doesn't change), and operational readiness (on-call load, alerting thresholds, and mean time to detect/restore all need revisiting because incidents become more consequential at higher throughput, even in components that didn't need architectural changes).
Worked example
Assume, as illustrative pinned inputs, 1 million monthly active users (MAU) sending an average of 50 messages/user/day:
baseline total msgs/day=1,000,000 MAU×50 msgs/user/day=50,000,000 msgs/day
baseline avg=86,400 s50,000,000≈579 msgs/s
After the 10x throughput change:
after 10x=579×10≈5,787 msgs/s average
and, using an illustrative 4x peak-to-average ratio for messaging traffic during busy hours:
illustrative peak (4x average)≈5,787×4≈23,148 msgs/s
That rise from roughly 579 to nearly 23,000 msgs/s at peak is what forces a hard look at broker partition count and delivery-worker concurrency. Meanwhile the authentication service, whose load tracks login attempts per MAU rather than messages sent, sees no comparable change and doesn't need re-architecting just because this number moved.
Trade-offs & pitfalls
- The most common mistake is treating a throughput change as a blanket "redesign everything" trigger; tracing each component's actual load driver is what separates urgent work from unaffected components.
- Cost still needs re-forecasting even for unaffected components' surrounding infrastructure (network egress, storage growth), because 10x more messages moving through the system has cost implications beyond the components that need architectural change.
- Don't defer operational readiness (alert thresholds, on-call capacity, incident runbooks) just because it isn't an architectural change; an incident at 10x throughput is a bigger incident even if the design handles the load correctly.
- If the 10x increase is concentrated in a small subset of highly active users rather than spread evenly, the actual bottleneck (a handful of hot conversations or channels) may look different from what a uniform-average calculation like the one above would suggest; validate the assumption behind the average before committing to a fix.
You're designing a user profile service with global, low-latency reads. Fields like email, password, and account status need strong consistency. Fields like display name and profile picture can tolerate eventual consistency. How would you decide, field by field, which guarantee each needs, and how would you defend keeping the split instead of making everything strongly consistent?
Sample Answer
Direct answer
Decide per field with a simple test: what does a user or the business lose if this field is read stale for a few seconds, and does that loss involve authorization, money, or identity? Email, password, and account status gate who can act as whom, so they get a linearizable (single, globally agreed order) read/write path even at a latency cost. Display name and avatar are cosmetic: a stale value for a few seconds costs nothing but a visual blip, so they get eventual, region-local, low-latency writes and reads. Defending the split means showing what making everything strong actually costs on the read path, not just asserting that it is safer.
Structured elaboration
Per-field decision table
| Field | Guarantee | Why | Cost of getting it wrong |
|---|---|---|---|
| Password / auth credentials | Strong (linearizable) | A stale read could let an old, revoked credential keep working | Account takeover window |
| Account status (banned/suspended) | Strong | A stale read lets a banned account keep acting | Abuse, trust and safety failure |
| Email (used for login/recovery) | Strong | Same identity-resolution risk as password | Locked-out or hijacked account |
| Display name | Eventual | Cosmetic; a few seconds of staleness is invisible risk | Momentary visual mismatch only |
| Profile picture | Eventual | Same as display name; also a large binary, cheap to serve from cache or object storage | Momentary visual mismatch only |
| Billing / payment state (extension) | Correctness-critical but not necessarily linearizable | Money is at stake, but the fix is compensating transactions, not blocking global writes | Double charge or missed charge, needing a refund/reversal workflow |
Mechanism
This paragraph is implementation detail, useful to know by name but not required to follow the field-by-field argument made above it. Two logical stores per user: a small, strongly-consistent store (consensus-replicated, for example a Raft-based database, where Raft is an algorithm that gets a cluster of replicas to agree on the same order of writes, or a globally-consistent database) for the identity-critical fields, and a multi-region, eventually-consistent store (Dynamo-style or similar) for everything else. Reads compose a single user object from both stores, so only the strong-store portion pays the cross-region latency cost. Read-after-write for the strong fields comes from routing that specific read to the writer's region or the current leader; monotonic reads (once a client has seen a value, a later read never shows it an older one) for the weak fields come from a session token, not from the strong store.
Extending the framework: billing correctness without going fully strong
Billing state is the case that tempts people into "just make everything strong." Resist it: instead of a synchronous global commit for every billing event, use compensating transactions, an idempotent charge (safe to run the same charge request twice, say after a retry, without actually billing the customer twice) plus a defined reversal or refund path if a downstream step (fraud check, inventory hold) fails after the charge already happened. This gets you correctness (the ledger is right once reconciliation finishes) without paying the linearizable-everything latency tax on a field written far less often than it is read.
Defending the split with a number, not an opinion
The strongest defense against "why not just make it all strong" is quantifying what "all strong" costs on the read path, since profile reads vastly outnumber profile writes.
Worked example
Assume a region-local cache read costs 5 ms, and a linearizable read from the strong store (contacting a majority of replicas across 3 regions, with an illustrative one-way inter-region round-trip time (RTT) of 100 ms) costs roughly two one-way trips:
strong-store read latency≈2×100 ms=200 ms latency multiplier if every read used the strong path=5 ms200 ms=40×If, say, 95% of profile reads only ever touch display-name or avatar fields (illustrative traffic mix, would come from real access logs), forcing all of them through the strong store means 95% of read traffic pays a 40x latency tax for a guarantee only the remaining 5% of fields ever needed. That is the number to put in front of someone asking why you didn't make everything strongly consistent.
Now the revenue-risk quantification (the second absorbed angle): the case for still investing in correctness on the billing fields, even though they don't get the fully linearizable treatment either.
assumed error rate on a race-prone billing path=0.1%=0.001 assumed volume=200,000 billing transactions/day at average value $50 expected daily exposure=200,000×0.001×50=$10,000/dayTen thousand dollars a day of exposure (illustrative; in practice pulled from real incident and error-rate data) is what justifies spending engineering time on compensating transactions for billing.
Trade-offs & pitfalls
- The strong store becomes a small, high-value target: shard it narrowly (identity fields only) so its lower throughput ceiling never becomes the bottleneck.
- Session tokens that carry the last-seen strong-store commit are what give read-your-own-writes on the critical fields without every read hitting the leader; skipping this is a common miss that reintroduces stale-password bugs.
- Pitfall: treating "eventual consistency" as a synonym for "no correctness work needed." The weak store still needs a conflict-resolution rule (last-writer-wins or a merge function), or two concurrent display-name edits silently lose one.
- Pitfall: treating billing as either fully strong or fully eventual instead of reaching for the third option, compensating transactions, which is usually the right cost and correctness balance for money-adjacent but not identity-adjacent fields.
A request path is built from several synchronous cross-service calls, and end-to-end latency is creeping past your SLO. Where would you introduce asynchronous decoupling to bring it back under budget, and what do you give up (immediacy, simpler error handling) to get there?
Sample Answer
Direct answer
Convert the calls whose result the client-facing response does not actually need into async, queue-backed steps, and keep only the calls that determine what you tell the client (an authorization decision, a price, a reservation outcome) on the synchronous path. What you give up is immediacy for the deferred steps (the caller no longer knows they succeeded before the response returns) and simple error handling (you now need retries, idempotency, and a plan for a step that fails after you already told the client it succeeded).
Structured elaboration
How to pick decoupling candidates
For each hop in the chain, ask, in order:
- Does the client's response body or status depend on this call's result? If no, it is a decoupling candidate.
- Is it only on the critical path because of implementation order, not because it is logically required before responding (sending a confirmation email after an order is placed is the classic case)? Decouple it.
- Can the caller tolerate this step failing and retrying later without the user noticing? If yes, move it behind a queue with at-least-once delivery and an idempotency key so a retry cannot double-apply the effect.
- If it must run after you've already told the client the request succeeded, what compensates if it fails? Sagas and compensating-transaction patterns are the standard answer here (a Saga: a sequence of local transactions where each step has a paired undo action that runs if a later step fails, so you get a rollback without a distributed transaction); treat them as a named sibling mechanism rather than re-deriving them.
What you give up, named explicitly
- Immediacy: the client no longer gets confirmation that the deferred step (email sent, loyalty points applied, analytics recorded) actually happened; if the product needs that confirmation, either keep the step synchronous or change the experience to a pending state.
- Simple error handling: a synchronous chain fails loudly and immediately; an async step fails quietly somewhere else, later, and needs monitoring (consumer lag, dead-letter queue depth) to even notice.
- Ordering: once two steps are decoupled, the free ordering guarantee a sequential call chain gave you for nothing is gone; if two async steps can race, an explicit ordering key or a saga is needed to keep them coherent.
The one thing not to decouple just to hit the number
Do not move a correctness-critical write (a payment capture, an inventory decrement, a seat reservation) to async purely to shave latency. That trades correctness for speed: the client sees a fast success response for something that has not actually been secured yet, and overselling or double-charging becomes an incident instead of a design decision.
Worked example: latency-budget arithmetic
Assume today's chain and its 95th-percentile (P95) latencies, all sequential: 10 ms gateway, 50 ms auth check, 120 ms inventory check, 80 ms pricing calculation, 300 ms fulfillment-order creation, 150 ms confirmation-email send, 90 ms audit-log write.
current P95=10+50+120+80+300+150+90=800 msIf the service-level objective (SLO) is P95 at or under 700 ms, that is 100 ms over budget. The email send and the audit-log write are both decoupling candidates by the test above, since the client's response does not need either to have completed:
after decoupling=10+50+120+80+300=560 msThat clears the 700 ms budget with 140 ms of headroom, without touching the correctness-critical inventory check or the fulfillment write.
Worked example: the absorbed booking-system angle
A synchronous seat-booking monolith migrating to event-driven has the same one thing it must not decouple: the seat reservation itself. Keep "reserve the seat" synchronous, using an atomic decrement or a compare-and-swap style check so two concurrent bookings cannot both win the same seat, which is exactly the double-booking risk the migration has to guard against. Move "send the confirmation email," "credit loyalty points," and "sync to the analytics warehouse" behind a queue. If payment fails after the seat was reserved, that is a compensating action (release the hold), not a reason to make the reservation itself asynchronous.
Trade-offs & pitfalls
- New failure mode: a message that fails repeatedly needs a dead-letter queue (DLQ) and an owner who actually looks at it, or side effects silently disappear.
- New monitoring surface: queue and consumer lag become a latency input in their own right; if the queue backs up, "async" steps can end up more stale than the SLO tolerates even though they are off the synchronous critical path.
- Pitfall: decoupling a call because it is slow rather than because its result is unneeded. If the client genuinely needs the answer, moving it to async just hides the latency problem behind a pending-state experience instead of solving it.
Unlock Full Question Bank
Get access to all System Design Methodology and Trade-off Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.