System Design Methodology and Trade-off Analysis Questions
The end-to-end approach to an open-ended design problem and the judgment that resolves it: clarifying scope and constraints, gathering functional and non-functional requirements, capacity and back-of-envelope estimation, and mapping requirements to a high-level architecture, then reasoning explicitly about competing options on cost, complexity, latency, and reliability to defend a choice. Covers driving a design interview from ambiguity to a proposal, trade-off frameworks, decision-making under uncertainty and incomplete information, reversible-versus-irreversible decisions, and defending choices under scrutiny. The process-and-judgment skill underneath every system-design case study.
Compare monolithic and microservices architectures. For each, list the benefits and drawbacks across development velocity, deployment complexity, operational overhead, and testing.
Sample Answer
Direct answer
A monolith is a single deployable unit; microservices split a system into independently deployable services that communicate over the network. The monolith wins on development velocity and simplicity while the team and codebase are small; microservices win on independent scaling, fault isolation, and team autonomy once the organization and traffic have grown enough to need them, but they add real operational cost that a small team pays for even before it needs the benefits.
Structured elaboration
| Dimension | Monolith | Microservices |
|---|---|---|
| Development velocity | Fast at first: one codebase, one build, easy cross-module refactors. Slows as the team grows: everyone contends for the same repository and build queue. | Slower at first: more moving parts and network contracts to define. Stays fast as the org grows: teams change their own service without waiting on others. |
| Deployment complexity | One pipeline, one artifact, predictable rollback, but any change, even a one-line fix, requires redeploying the whole system. | Independent deploys per service shrink blast radius, but many pipelines now need coordinating, and services need versioned, backward-compatible APIs between them. |
| Operational overhead | Low at small scale: one thing to monitor, one thing to scale, coarsely, as a whole. | Higher: service discovery, inter-service network reliability, distributed tracing and logging, and typically a container orchestrator, all needed just to operate. |
| Testing | End-to-end tests run in one process, straightforward to set up; the suite slows and tangles as the codebase grows. | Unit and contract tests per service stay fast and isolated, but full end-to-end behavior now needs integration or contract tests across services, and network-related flakiness becomes real. |
Worked example
A five-person startup with 200 daily active users splits its checkout flow into a separate payments service, an inventory service, and a notifications service on day one. In practice: three CI/CD pipelines to maintain instead of one, a network call, with its own latency and failure modes, added to every checkout in place of a function call, and the same five engineers now also debugging cross-service request tracing for a system with barely any real traffic. None of the microservices benefits, independent team ownership or independent scaling under real load, apply yet, because there is one team and no bottleneck to isolate. A modular monolith, meaning a single deployable codebase with clean internal module boundaries and clear ownership per module, gets the same code-organization benefit without the network and operational cost, and it can be decomposed later once an actual bottleneck, not a hypothetical future one, justifies the split.
flowchart TB
subgraph MONO["Monolith: one deployable unit"]
direction TB
W[Web layer]
B[Business logic]
D[Data access]
end
subgraph MICRO["Microservices: independently deployable, network-connected"]
direction TB
PaySvc[Payments service]
InvSvc[Inventory service]
NotifSvc[Notifications service]
PaySvc <--> InvSvc
InvSvc <--> NotifSvc
end
Trade-offs & pitfalls
- Adopting microservices for resume-driven or "best practice" reasons rather than a named bottleneck.
- Splitting along technical layers (a services layer, a database layer) instead of business capability boundaries, which just moves tight coupling onto the network instead of removing it; this is the organizational mirror named Conway's Law (a system's structure tends to mirror the communication structure of the organization that built it, so splitting along technical layers just recreates the same coordination problems on the network instead of removing them), worth knowing by name without needing to re-derive it here.
- Treating "the codebase feels big" as the signal a split is overdue, instead of a concrete one: one team's deploy regularly breaks or is blocked by another team's unrelated changes.
You need storage for two different use cases in the same system: a low-latency user-session store with automatic expiry, and an append-only audit log. Would you use the same datastore for both, or different ones? Walk through your reasoning.
Sample Answer
Direct answer
Use two different datastores, because the two workloads sit at opposite ends of nearly every axis that matters: the session store needs low-latency point lookups by key with automatic expiry, while the audit log needs durable, ordered, append-only writes with long retention and replay. Forcing one engine to serve both means either paying the audit log's durability and retention cost on every fast session read, or giving up the audit log's append-only, replayable guarantees.
Structured elaboration
| Dimension | Session store | Audit log |
|---|---|---|
| Access pattern | Point lookup and update by key | Append-only write, range/time-based scan |
| Expiry | Automatic (time-to-live, TTL) | None, typically long or indefinite retention |
| Latency need | Very low, on the request's critical path | Can tolerate higher write latency, rarely on the user-facing critical path |
| Durability | Often acceptable to lose on crash (session can be re-created) | Must be durable, it's the record of what happened |
| Natural fit | Redis (an in-memory key-value store) with native TTL | An append-only log or stream (for example, Kafka, a distributed log/streaming platform), or a partitioned relational table |
Why this generalizes beyond sessions and audit logs: the absorbed framing of a machine-learning (ML) metadata store shows the same underlying shape. Continuous integration (CI) pipelines generate a high write rate logging experiment runs, while dashboards and monitoring run read-heavy, scan-heavy queries over historical runs. That's structurally the identical pattern, a write-optimized append path pulling against a read-optimized query path, which is why the "one store or two" question should be re-asked per component rather than defaulting to whatever store is already in the stack.
Worked example
Primary scenario (session + audit). Use Redis for sessions, with native per-key TTL handling expiry automatically, no cleanup job required. Use an append-only log or a partitioned table for the audit trail.
Storage growth check on the audit side: assume, illustratively, 500 audit events per second at 1 KB each.
500 events/s×1KB=500KB/s
500KB/s×86,400s/day=43,200,000KB/day=43.2GB/day (decimal GB, 1 GB = 1,000,000 KB)
43.2GB/day×730days (2 years)=31,536GB≈31.5TB
Thirty-one and a half terabytes of append-only history is exactly the kind of cheap, sequential, long-retention storage a log-oriented system is built for, and exactly the kind of volume that would be prohibitively expensive to hold in an in-memory session store's provisioned capacity.
Second scenario (ML metadata, the absorbed framing). The same reasoning applies with a relational or document store for structured experiment metadata (read-heavy, queried by dashboards across many attributes) paired with an append-only log or object storage for the raw high-volume run logs written continuously by CI, again two engines, chosen for two genuinely different access patterns, not one engine stretched across both.
Trade-offs & pitfalls
- Storing audit events in the session store "for convenience" either evicts old sessions to make room or silently drops the durability guarantee the audit log actually needs, neither is acceptable for an audit trail.
- Using an append-only log system for session storage means reimplementing per-key expiry logic that a key-value store like Redis provides natively, most log/stream systems have topic-level retention, not per-key TTL.
- Two datastores are two systems to operate, monitor, and back up, that overhead is worth paying here because the access patterns are genuinely incompatible; it would not be worth it if the two workloads were similar enough to share one engine.
- Deletion requirements, for example a regulatory "right to be forgotten" request, are harder to satisfy in an append-only log than in a key-value store, this needs a tombstone (a marker record saying "treat this key as deleted") plus a compaction strategy (the background process that rewrites and shrinks the log over time, actually purging tombstoned data) decided up front, not retrofitted after the first request arrives.
flowchart LR
Client --> SessionAPI[Session Read/Write]
Client --> AuditWriter[Audit Event Writer]
SessionAPI --> Redis[(Redis Session Store)]
AuditWriter --> Log[(Append-only Log)]
Log --> ColdStorage[(Long-term Storage)]
Partway through designing a system, you're told to plan for three possible curveballs: a region outage, an upstream schema change that breaks your data pipeline, and a sudden 10x traffic spike. How would you prioritize which to design for first, and how does each change your architecture?
Sample Answer
Direct answer
Prioritize by expected business impact combined with how quickly the failure mode compounds if unaddressed: a region outage first, because it's a full-availability event with no partial-degradation option; a sudden 10x traffic spike second, because it threatens availability but usually has partial mitigations (throttling, degraded modes) available immediately; and an upstream schema change third, because it's typically detectable and containable with fast rollback before it causes user-facing damage, even though it can silently corrupt data if left uncaught.
Structured elaboration
For each curveball, separate the immediate runbook response from the longer-term architectural change it justifies.
Region outage. Immediate: fail over reads and writes to a secondary region using health-checked traffic routing, and pause non-essential batch work to reduce write pressure during the transition. Architectural change: multi-region active-passive (or active-active) replication for the data layer, with regularly rehearsed failover drills; a design that was never built to fail over won't fail over correctly under real pressure, only under a rehearsed one.
Sudden 10x traffic spike. Immediate: autoscale the serving tier, shed or degrade non-critical functionality (serve cached or slightly stale results rather than fail outright), and throttle low-priority background jobs to protect the real-time path. Architectural change: pre-warmed capacity headroom, adaptive rate limiting, and a defined degraded mode that's tested before it's needed, not designed during the incident.
Upstream schema change breaking the data pipeline. Immediate: fail fast on schema-validation errors at ingestion rather than let malformed data propagate, quarantine the bad batch, and roll the downstream transform back to the last known-good schema. Architectural change: enforce a schema contract at the pipeline boundary (a strongly typed serialization format with a compatibility check, such as Avro or Protocol Buffers) so a breaking upstream change is caught at ingestion rather than discovered downstream after it has already corrupted derived data.
Worked example
An illustrative prioritization exercise, scoring each curveball on business impact (1 low to 5 high) and detectability/containability (1 hard to 5 easy) to make the ranking auditable rather than a gut call: region outage scores high impact (5/5: full outage, all users) and moderate containability (3/5: requires a rehearsed failover, not just a code fix); 10x traffic spike scores high impact if unmitigated (4/5) but higher containability (4/5: autoscaling and shedding are standard, fast-acting levers); schema break scores lower immediate user-facing impact (2/5: the pipeline can often keep serving stale-but-correct data while paused) but containability that depends entirely on whether validation exists at the ingestion boundary (2/5 without it), if it doesn't, undetected corruption can silently spread for a long time before anyone notices, which is exactly why validation is the priority architectural investment for that curveball specifically, even though it's ranked last for immediate response.
Trade-offs & pitfalls
- Ranking these purely by which is scariest in the abstract, rather than by business impact and how fast each compounds if left unaddressed, produces a plausible-sounding but ungrounded priority order; tie the ranking to a concrete criterion.
- A schema break that lacks ingestion-time validation is deceptively low-priority in the short term and highest-priority for silent, compounding damage; don't let "least immediately visible" become "least urgent to architect for."
- Building all three mitigations simultaneously from scratch during a single design pass is rarely realistic; sequence the architectural investments and say explicitly which curveball's mitigation ships first and why.
- Rehearsing failure (game days, chaos testing, restore drills) is what turns a runbook from theory into something that actually works under pressure; a runbook that has never been executed is a plan, not a capability.
You're building a stateful, write-heavy service that needs to sustain 10,000 writes per second with low latency. How does that write-heavy profile change your datastore and architecture choices compared to a read-heavy service?
Sample Answer
Direct answer
A sustained 10,000 writes-per-second, low-latency, stateful workload pushes you away from a design tuned for reads (a single write primary, heavy indexing, read replicas) and toward one built for write scaling: a storage engine optimized for sequential writes, a partitioning scheme that spreads writes across many nodes, and a replication model with an explicit, tunable durability-versus-latency trade-off rather than a single write bottleneck.
Structured elaboration
Why a read-optimized design breaks down here. Traditional B-tree storage engines perform random-access writes and update every index on every insert, each additional index roughly adds another write per record. A single-writer relational primary caps total write throughput at whatever one node's disk and CPU can sustain, and read replicas do nothing for write capacity, they only copy the primary's write stream.
What changes for write-heavy:
- Storage engine: log-structured merge (LSM) tree engines (used by databases like Cassandra, HBase, and the storage layer behind DynamoDB-style stores) append writes sequentially and merge them in the background, trading some read amplification (a single logical read may have to check several separate on-disk files before it can answer, since recent and older writes land in different segments) for much higher sustained write throughput than a B-tree.
- Partitioning: writes are sharded across many nodes by a partition key. The key must be chosen for even cardinality, a monotonically increasing key (like a timestamp or auto-increment ID) concentrates all new writes on one shard regardless of how many nodes exist.
- Replication and durability: instead of one primary with no built-in fan-out, use a quorum-based replication scheme, writes are acknowledged once a majority of replicas confirm, giving a tunable point between "acknowledge on one node" (fast, risks data loss) and "acknowledge on all nodes" (safest, slowest).
- Indexing discipline: keep secondary indexes to the minimum the write path can afford, every index is a write, this is the opposite instinct from a read-heavy design where more indexes are usually free wins.
Worked example
Assume, illustratively, that a single write-optimized node sustains 2,000 writes per second at the target latency.
Nodes needed for raw throughput: 10,000/2,000=5 shards.
For durability, replicate each shard three ways (tolerate one node failure without data loss): 5×3=15 total storage nodes.
A quorum write with N=3 replicas and a write quorum of W=2 means the client waits only for the second-fastest replica to acknowledge, not the slowest, bounding tail write latency while still guaranteeing the write survives a single node failure.
Cost contrast, provisioned versus per-operation pricing. At an illustrative $0.00001 per write operation under a consumption-priced managed service:
ops/day=10,000×86,400=864,000,000 writes/day
daily cost=864,000,000×$0.00001=$8,640/day≈$259,200/month
Against 15 provisioned nodes at an illustrative $400/node/month: 15×$400=$6,000/month. At this sustained write rate the per-operation model costs roughly 40 times more, which is why sustained high-volume writes usually favor provisioned or self-managed clusters, and why consumption pricing fits bursty, low-average workloads instead.
Trade-offs & pitfalls
- Carrying over every index from a read-heavy design roughly multiplies write cost by the number of indexes, audit which indexes the write path can actually afford.
- A low-cardinality or monotonically increasing partition key creates a hot shard that caps total throughput no matter how many nodes you add, this is the single most common write-scaling mistake.
- Waiting for all replicas (W=N) is the safest durability setting but the slowest; a majority quorum balances safety and latency, the exact quorum size is itself a trade-off decision, not a default.
- High write concurrency needs connection pooling and write batching, naive one-connection-per-request patterns hit connection limits long before they hit the storage engine's real capacity.
flowchart LR
Client --> Router[Write Router]
Router --> ShardA[Shard A Leader]
Router --> ShardB[Shard B Leader]
Router --> ShardC[Shard C Leader]
ShardA --> ShardARep[Shard A Replicas x2]
ShardB --> ShardBRep[Shard B Replicas x2]
ShardC --> ShardCRep[Shard C Replicas x2]
You're designing a user profile service with global, low-latency reads. Fields like email, password, and account status need strong consistency. Fields like display name and profile picture can tolerate eventual consistency. How would you decide, field by field, which guarantee each needs, and how would you defend keeping the split instead of making everything strongly consistent?
Sample Answer
Direct answer
Decide per field with a simple test: what does a user or the business lose if this field is read stale for a few seconds, and does that loss involve authorization, money, or identity? Email, password, and account status gate who can act as whom, so they get a linearizable (single, globally agreed order) read/write path even at a latency cost. Display name and avatar are cosmetic: a stale value for a few seconds costs nothing but a visual blip, so they get eventual, region-local, low-latency writes and reads. Defending the split means showing what making everything strong actually costs on the read path, not just asserting that it is safer.
Structured elaboration
Per-field decision table
| Field | Guarantee | Why | Cost of getting it wrong |
|---|---|---|---|
| Password / auth credentials | Strong (linearizable) | A stale read could let an old, revoked credential keep working | Account takeover window |
| Account status (banned/suspended) | Strong | A stale read lets a banned account keep acting | Abuse, trust and safety failure |
| Email (used for login/recovery) | Strong | Same identity-resolution risk as password | Locked-out or hijacked account |
| Display name | Eventual | Cosmetic; a few seconds of staleness is invisible risk | Momentary visual mismatch only |
| Profile picture | Eventual | Same as display name; also a large binary, cheap to serve from cache or object storage | Momentary visual mismatch only |
| Billing / payment state (extension) | Correctness-critical but not necessarily linearizable | Money is at stake, but the fix is compensating transactions, not blocking global writes | Double charge or missed charge, needing a refund/reversal workflow |
Mechanism
This paragraph is implementation detail, useful to know by name but not required to follow the field-by-field argument made above it. Two logical stores per user: a small, strongly-consistent store (consensus-replicated, for example a Raft-based database, where Raft is an algorithm that gets a cluster of replicas to agree on the same order of writes, or a globally-consistent database) for the identity-critical fields, and a multi-region, eventually-consistent store (Dynamo-style or similar) for everything else. Reads compose a single user object from both stores, so only the strong-store portion pays the cross-region latency cost. Read-after-write for the strong fields comes from routing that specific read to the writer's region or the current leader; monotonic reads (once a client has seen a value, a later read never shows it an older one) for the weak fields come from a session token, not from the strong store.
Extending the framework: billing correctness without going fully strong
Billing state is the case that tempts people into "just make everything strong." Resist it: instead of a synchronous global commit for every billing event, use compensating transactions, an idempotent charge (safe to run the same charge request twice, say after a retry, without actually billing the customer twice) plus a defined reversal or refund path if a downstream step (fraud check, inventory hold) fails after the charge already happened. This gets you correctness (the ledger is right once reconciliation finishes) without paying the linearizable-everything latency tax on a field written far less often than it is read.
Defending the split with a number, not an opinion
The strongest defense against "why not just make it all strong" is quantifying what "all strong" costs on the read path, since profile reads vastly outnumber profile writes.
Worked example
Assume a region-local cache read costs 5 ms, and a linearizable read from the strong store (contacting a majority of replicas across 3 regions, with an illustrative one-way inter-region round-trip time (RTT) of 100 ms) costs roughly two one-way trips:
strong-store read latency≈2×100 ms=200 ms latency multiplier if every read used the strong path=5 ms200 ms=40×If, say, 95% of profile reads only ever touch display-name or avatar fields (illustrative traffic mix, would come from real access logs), forcing all of them through the strong store means 95% of read traffic pays a 40x latency tax for a guarantee only the remaining 5% of fields ever needed. That is the number to put in front of someone asking why you didn't make everything strongly consistent.
Now the revenue-risk quantification (the second absorbed angle): the case for still investing in correctness on the billing fields, even though they don't get the fully linearizable treatment either.
assumed error rate on a race-prone billing path=0.1%=0.001 assumed volume=200,000 billing transactions/day at average value $50 expected daily exposure=200,000×0.001×50=$10,000/dayTen thousand dollars a day of exposure (illustrative; in practice pulled from real incident and error-rate data) is what justifies spending engineering time on compensating transactions for billing.
Trade-offs & pitfalls
- The strong store becomes a small, high-value target: shard it narrowly (identity fields only) so its lower throughput ceiling never becomes the bottleneck.
- Session tokens that carry the last-seen strong-store commit are what give read-your-own-writes on the critical fields without every read hitting the leader; skipping this is a common miss that reintroduces stale-password bugs.
- Pitfall: treating "eventual consistency" as a synonym for "no correctness work needed." The weak store still needs a conflict-resolution rule (last-writer-wins or a merge function), or two concurrent display-name edits silently lose one.
- Pitfall: treating billing as either fully strong or fully eventual instead of reaching for the third option, compensating transactions, which is usually the right cost and correctness balance for money-adjacent but not identity-adjacent fields.
Unlock Full Question Bank
Get access to all System Design Methodology and Trade-off Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.