Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
You join a team that has resisted formal documentation. How do you introduce decision records so they get used, what format do you pick, and how do you keep them from going stale?
Sample Answer
Direct answer
I start by finding out why they resist, then make decision records small, useful on day one, and attached to work they already do. I pick a one-page markdown format in the repository (context, decision, alternatives, consequences, and a status line), introduce it on one real decision that is happening this week, and keep it alive with a review-by date and a short recurring triage. I do not announce a documentation policy.
Step 1: understand the resistance (first week)
Ask three engineers what happened to past documentation. Typical answers: "it was out of date within a month", "the template took an hour", "nobody read it". Each answer shapes the design. If the pain is staleness, the format must be cheap to keep current; if it is bureaucracy, no approvals.
Step 2: pick the format
An ADR (Architecture Decision Record, a short document capturing one significant decision, what else was considered and its consequences) in the lightweight style popularised by Michael Nygard (who described it in a short blog post as a few sentences each for title, status, context, decision and consequences), extended with a one-line "alternatives considered". Reasons:
- Markdown in the repo (for example
docs/adr/0001-use-postgres.md): no new tool, reviewed in the pull request like code, and history comes free from Git. - One page, about ten minutes to write. Longer formats can come later if wanted.
- Where it lives and who owns the habit: in the repository of the service it governs, with a shared index for cross-team ones. The team's tech lead approves the content; the engineering manager (the person who manages the team, not its technical design) makes sure the habit exists.
Step 3: introduce it on a live decision
Choose an upcoming choice the team is already debating (for example which queue to use). In the design discussion, write the ADR together in 20 minutes on a shared screen. The team experiences the payoff immediately: the argument ends, the reasoning is saved, and the next person will not repeat it.
Step 4: rules of thumb, not mandates
Write one when a decision is hard to reverse, affects another team, or was argued about. Skip it for small choices. This keeps volume low and value high.
Step 5: embed it in existing work
- The pull request template asks "does this change a recorded decision?".
- Code review comments link to the ADR instead of re-arguing.
- The onboarding list (what new hires work through in their first weeks) includes the 5 most important ADRs.
Step 6: keep them from going stale
| Mechanism | How it works |
|---|---|
Status and review_by date | Every ADR states its status and a date by which someone must re-check it (for example review_by: 2027-03-12) |
| Owner | The team that owns the service owns the ADR, not "the architects" |
| 30-minute quarterly triage | List accepted ADRs past their date; confirm, update or supersede each one |
| Supersede, don't edit | When a decision changes, write a new ADR that replaces the old one instead of rewriting it; keeps history trustworthy |
| Link from code | README and config point to the ADR so it is found when the code is touched |
Worked 90-day plan
- Days 1-10: listen; agree the format and folder.
- Days 10-30: write two ADRs on live decisions, with the team.
- Days 30-60: add the pull request template question; new hire reads them.
- Days 60-90: first triage meeting; ask the team whether it helped and drop anything that did not.
Success measures (qualitative first): people cite ADRs in reviews; a new hire finds the reason for a design without asking; the number of repeated debates falls. Counting ADRs alone is a vanity metric (a number that looks impressive but says nothing about whether anyone benefits).
Pitfalls
- Mandating from the top before any value is seen.
- Backfilling years of history; write only the decisions that still matter.
- Turning the ADR into an approval gate so people avoid writing them.
What would change my choice
If the team truly lives in a wiki and never opens the repo docs, I would put ADRs in the wiki with page history, but the repo is preferable because it sits with the code.
Sketch the deployment view for a service that runs across several availability zones or regions with managed data stores. What do you put on it, what do you deliberately leave off, and which reader is it for?
Sample Answer
Direct answer
A deployment view answers "where does each thing actually run, and what survives when a zone or region fails?" I show regions, availability zones (AZ, separate data centres in one region), network tiers, load balancers, compute groups, managed data stores with replication direction, and the traffic and failover paths between them, layered as network, compute and data. I leave off application internals, individual instance IDs and IP ranges, and every security rule. It is written for SREs, security reviewers and capacity planners, not for product or new engineers.
What goes on it
- Network layer: regions, AZs, virtual private cloud (VPC, an isolated private network) with public, private and data subnets (slices of the network: public ones face the internet, private ones hold services, data ones hold databases and accept traffic only from services); the entry path (DNS, which turns names into addresses; a CDN, servers near users that cache content; and a global or regional load balancer that spreads requests across copies).
- Compute layer: each service as a group, with how many copies per AZ (for example "3 copies per AZ") and the platform (container cluster, functions).
- Data layer: managed stores with role and replication: primary in AZ a (takes writes), standby in AZ b (a copy that takes over if the primary fails), read replica in region B (a read-only copy), labeled "async, lag alarm" (the copy trails the primary by some seconds; that delay is replication lag) or "sync" (a write is confirmed only after the copy has it).
- Arrows: protocol, direction, and dashed lines for failover or replication.
- Annotations: RTO/RPO (recovery time and data-loss targets), owner, diagram date.
flowchart TB
U[Users] --> G[DNS + global load balancer]
subgraph RA[Region A, primary]
LA[Load balancer] --> SA1[Service, AZ 1]
LA --> SA2[Service, AZ 2]
SA1 --> DB[(Managed DB primary)]
SA2 --> DB
DB -. sync .-> DBS[(Standby, AZ 2)]
end
subgraph RB[Region B, warm: running at reduced size, ready to scale up]
SB[Service, reduced size] --> RR[(Read replica)]
end
G --> LA
G -. failover .-> SB
DB -. async replication .-> RR
What I deliberately leave off
Class or component internals (that is the component view), individual instance IDs and IPs (put CIDR ranges, the notation for blocks of IP addresses such as 10.0.0.0/16, in a table), full security group rules (firewall rules per resource), CI/CD, monitoring agents, and every minor dependency. If it does not help someone decide "does this survive an AZ loss?", it is noise.
Which reader
Primary: SRE/on-call and cloud architects (failure and capacity), secondarily security reviewers (boundaries) and finance (what is duplicated per region). A product audience gets the container view instead.
Variant: three regions (less commonly asked). Add region roles (one write primary, two read replicas), state the replication lag you accept, show where the event backbone (the shared message stream services publish to, such as a replicated Kafka cluster) sits and how events cross regions, and mark which data must stay in-region.
Variant: Kubernetes (less commonly asked). Kubernetes is a container orchestrator. Zoom into one cluster: namespaces per team or environment, the ingress (the entry point routing outside traffic in), network policies as security boundaries, and node pools per AZ. Keep it as a second diagram rather than crowding the first.
Worked example: a review question the view must answer
The question is "if one AZ dies, can the remaining AZs carry peak traffic?" Take the drawn Region A with 2 AZs and 3 service copies in each, 6 in total. Lose one AZ and 3 of 6 copies remain, half the capacity. If peak traffic needs 4 copies, the view shows it is under-provisioned for that failure. With 3 AZs and 3 copies each (9 total), losing one leaves 6 of 9, two thirds, which covers a peak of 4 with room to spare. Capacity must be sized for the surviving fraction, not the total.
Trade-offs and pitfalls
- One diagram cannot serve network engineers and executives; split by audience.
- Showing replication without its direction and lag makes failover claims unverifiable.
For a message-driven order-processing service, what failure modes and degraded behaviours would you record in its architecture documentation, and how would you write each mitigation so an on-call engineer who did not build it can act on it?
Sample Answer
Direct answer
I would give the service a "Failure modes and degraded behaviour" page that is one table plus one runbook-style entry per row. Each entry says what the on-call engineer will see (the symptom, as a named alert or graph), what customers experience meanwhile (the degraded behaviour), the safest first action, what NOT to do, and when and to whom to escalate. The test of a good entry: someone paged at 3 a.m. who has never opened the code can act on it in the first five minutes. (On-call means the engineer currently carrying the pager for the service; a runbook is the step-by-step operating guide they follow. The queue holds orders waiting for a worker; queue depth is how many are waiting and the age of the oldest message is how long the longest-waiting one has sat. A message broker is the server that stores and hands out those messages.)
The table has eight rows. Rows 1 to 6 (dependency timeout, resource exhaustion, network partition, poison message, duplicate delivery, consumer backlog) apply to almost every message-driven service, so document those first. Rows 7 and 8 (multi-region failover and ML fraud scoring) are advanced extras that only some services need.
The failure modes I would record (order flow assumed: API accepts order, publishes to a message broker, workers consume, write to a database, call a payment provider)
| Failure mode | Symptom on-call sees | Degraded behaviour by design | Safe first action | Escalate |
|---|---|---|---|---|
| Dependency timeout (payment provider slow or down) | Provider-call error rate and latency alarms; queue depth rising | Orders accepted as "payment pending", not rejected | Confirm on provider status page; leave the circuit breaker (a switch that stops calling a failing dependency) open; do not restart workers | Provider account owner if still failing after 30 min |
| Resource exhaustion (worker memory, database connection pool) | Worker restarts, "pool exhausted" errors, rising lag | Slower fulfilment; API still accepts | Scale workers only if the database has connection headroom (spare connections before its limit); otherwise reduce consumer concurrency (how many messages each worker handles in parallel) | Database owner if pool stays saturated |
| Network partition (workers cannot reach broker or database) | Consumers drop to zero, broker healthy for other clients | Orders queue durably; nothing is lost, only delayed | Check network path and security rules changed recently; do not purge queues | Platform/network team |
| Poison message (a message that fails every time) | Same message ID retried; dead-letter queue (DLQ, the parking queue for messages that keep failing) count above zero | Order parked, others proceed | Inspect DLQ message, fix data, replay one message | Service owner for a code fix |
| Duplicate delivery (the broker is at-least-once: it guarantees a message arrives, but may deliver it twice) | Customer double-charge reports | Handlers are idempotent (repeat-safe) via an order-ID key | Check how often the dedupe table (the record of order IDs already processed) rejects a repeat; some hits are normal, but zero hits alongside double-charge reports means the check is not running, so escalate | Owner if any duplicate charge occurred |
| Consumer backlog | Age of oldest message above the target | Delayed confirmations | Check whether a downstream is slow before adding consumers | Owner if age exceeds the SLO (service-level objective, the reliability target) |
| Multi-region stateful failover (advanced) | Replication lag alarm; regional health check red | Secondary region serves; writes since the last replicated point may be lost (the RPO, recovery point objective) | Follow failover checklist; do not fail back without owner approval | Incident commander (the person coordinating a major incident response) |
| ML fraud-scoring dependency returns garbage after a schema change (schema drift: input fields renamed or retyped) | Score distribution shifts; validation-error counter climbs | Fall back to simple rules; flag orders for review | Switch the fraud flag to rules-only mode | Model owner |
Retry amplification (the failure that turns one problem into an outage)
flowchart LR
C[Client, 3 attempts] --> A[API, 3 attempts]
A --> W[Worker, 3 attempts]
W --> P[Payment provider]
W -.->|guard| G[Retry budget + jitter + breaker]
If each layer makes up to 3 attempts in total (the diagram's meaning: 1 original try plus 2 retries), one order can produce 3 x 3 x 3 = 27 calls (if each layer instead retried 3 times after the first try, that is 4 attempts per layer and 4 x 4 x 4 = 64 calls) against a provider that is already struggling. The documented guard: retry at ONE layer only, use exponential backoff with jitter (randomised waits so clients do not retry in lockstep), cap retries with a budget, and open the circuit breaker on sustained failure.
Worked example of one entry, written for a stranger
PAYMENT PROVIDER TIMEOUTS (illustrative thresholds)
Alert: provider_call_p95 (the 95th percentile: 95% of calls finish faster than this) above 5s for 10 min, or DLQ depth above 0
Customer impact: orders show "payment pending"; no orders are lost
Do first: 1) open provider status page 2) confirm breaker state on dashboard
Do NOT: restart workers (loses in-flight state: messages a worker has taken but not yet finished, which then get retried or duplicated), purge queue, disable retries
Recovered when: queue age under 5 min for 15 min
Escalate: 30 min unresolved -> payments owner; any suspected double charge -> immediately
Short entries for the two advanced rows
REGION FAILOVER (illustrative)
Alert: regional health check red for 5 min
Do first: confirm with the cloud status page, then run the failover checklist
Do NOT: fail back without owner approval
Escalate: incident commander immediately
FRAUD SCORING GARBAGE (illustrative)
Alert: validation-error counter above 1% of scored orders
Do first: switch fraud flag to rules-only mode
Do NOT: delete queued orders
Escalate: model owner
Trade-offs and pitfalls
- Write symptoms as alert names and dashboard panels, never as "the system is unhealthy".
- Every mitigation must be safe to run by someone who has not seen the code; put irreversible steps behind an explicit "ask the owner" line.
- Mark thresholds as tuned from the SLO and review them after each incident, otherwise they rot.
- Documenting only outages and skipping degraded modes hides the most useful design decision: what the system deliberately keeps doing when a dependency fails.
What makes an architecture diagram good or bad? What conventions do you insist on for shapes, arrows, layering, legends and labelling, and what would make you send a diagram back?
Sample Answer
Direct answer
A good architecture diagram answers one question for a named audience at one level of abstraction, and a reader can understand it without the author in the room. A bad one mixes levels, has unlabeled arrows and no legend, and shows every component the author knows about. I insist on: a title stating purpose and audience, consistent shapes with a legend, arrows that say what flows and how, sensible layering, and metadata (owner, date, version). I send a diagram back when a reader would have to guess what an arrow means.
Conventions I insist on
- Shapes: one shape per kind of thing (service, data store, queue, external system), explained in a legend. Color never carries meaning alone (colour-blind and print readers).
- Arrows: labeled with a verb plus protocol ("submits order, HTTPS"); direction means "who initiates the call" and I say so on the legend.
- Sync versus async: solid line for a call the caller waits on, dashed for messages and events. Mixing them is a common source of wrong assumptions about failure.
- Layering: users at the top or left, data at the bottom or right; group things inside boundaries (zone, team, trust boundary) rather than scattering them.
- Labelling: names match the code repository and dashboards; every acronym is expanded once.
- Annotations: owner and team, a risk marker on single points of failure (SPOF, one component whose loss takes the system down), a link from each box to its runbook (the operating guide for when it breaks), and the date and version the diagram describes.
In the example below, gRPC is a fast request protocol between services, and the last node is the legend: a note box stating what solid and dashed arrows mean, deliberately connected to nothing so nobody reads it as a data flow. Note the queue-to-worker arrow points from the worker to the queue because, by the convention above, the arrow points at what is called and the worker polls the queue; if the broker pushes messages to the worker, draw it the other way and say so on the legend.
flowchart LR
U[Client] -->|POST order, HTTPS| A[API]
A -->|reserve stock, gRPC| S[Stock service]
A -.->|OrderPlaced event| Q[[Queue]]
W[Fulfilment worker] -.->|polls for events| Q
L[Legend: solid = waits for reply, dashed = event or message]
Structure views versus a sequence view These follow the C4 model (four zoom levels: Context, Container, Component, Code). A container view shows the deployable pieces (apps, services, databases); a component view zooms into the parts inside one container; both show structure, what exists and connects. Each diagram should sit at one such level of abstraction (zoom level). A sequence view (participants across the top, time flowing down) shows order and timing, so use it to show retries, timeouts and failure handling. If the question is "what happens when this fails", structure alone cannot answer it.
Worked example of sending one back
Submitted: a box "Backend" with unlabeled arrows to "DB" and "Cache", cloud icons, no legend, dated last year. My review: name the components ("Order API", "Orders database"); label each arrow; add a legend; mark the cache as a shared dependency; add the owner and a runbook link; remove the dated icons that convey nothing. After the edit a stranger can say what fails when the cache is down. The corrected diagram, as text:
[Order API] --reads/writes orders, SQL--> [Orders database]
[Order API] --reads hot items, Redis--> [Shared cache] (shared dependency, owner: platform)
Legend: solid = synchronous call. Owner: orders team. Runbook: link. Date: 2026-09-01
Send-back triggers
- Unlabeled arrows. 2. Mixed abstraction levels. 3. No legend, or unexplained icons. 4. More elements than a reader can hold, roughly fifteen or more, without grouping. 5. Undated or stale. 6. No indication of sync versus async. 7. Crossing lines that could be untangled by regrouping.
Trade-offs and pitfalls
- Strict conventions help only if the team shares them; agree on a template, not on taste.
- A beautiful diagram that is out of date is worse than an ugly one that is correct, so keep it in the repository next to the code.
Which testing and resilience artifacts belong in architecture documentation so teams can validate the system, and how would you document a failover test plan with hypothesis, steps and expected outcomes?
Sample Answer
Direct answer
Architecture documentation should include the artifacts that let a team prove the design behaves as claimed: a test strategy, a list of failure modes with the expected behaviour, capacity and performance targets with load-test results, recovery objectives with backup-restore evidence, and resilience experiment plans. A failover test plan is written as a hypothesis, a controlled set of steps, expected outcomes with pass and fail criteria, and an abort and rollback path.
Artifacts that belong
| Artifact | Purpose |
|---|---|
| Failure-mode table | For each dependency: what if it fails, expected system behaviour, detection |
| Recovery targets | RTO (recovery time objective: how long an outage may last) and RPO (recovery point objective: how much data loss is tolerable), with the last measured result |
| Load and capacity test plan | Expected peak, headroom, and how it was validated |
| Backup restore test | Date, result, restore time |
| Failover and chaos plans | Below |
| Staging-parity notes | How the test environment differs from production, so results are read with the right caution |
Failover test plan (worked example, illustrative)
Terms used below: the primary is the database instance currently taking writes; a replica is a continuously updated copy that can take over; promoting the replica makes it the new primary; replica lag is how far behind the primary the copy is, in seconds of writes; staging is the production-like test environment; synthetic load is fake but realistic traffic generated by a tool; steady state is the normal measurable behaviour (for example 99% of orders succeed).
Title: Primary database failover, orders service
Hypothesis: If the primary database instance is stopped, the replica is promoted
and orders resume within the 5 minute RTO, with no more than
the 1 minute RPO of writes lost.
Environment: Staging, sized as production-lite (see parity notes)
Preconditions: steady synthetic load; dashboards open; on-call informed;
abort criteria agreed.
Steps:
1. Record baseline: order success rate, latency, replica lag.
2. Stop the primary instance.
3. Observe detection time and promotion.
4. Watch application reconnect behaviour and errors.
5. Restore the old primary as a replica.
Expected: Alert fires within 1 minute; promotion completes; success
rate returns to baseline; no duplicate or lost confirmed orders.
Pass/fail: Recovery within RTO and data loss within RPO, else fail.
Abort: Success rate stays under 50% for 10 minutes: roll back per runbook.
Record: Actual times, surprises, follow-up tickets, date of next test.
Judging the result against the targets (worked numbers, illustrative)
- Primary stopped at 10:00:00. Alert fired 10:00:40. Replica promoted 10:02:30. Order success back to baseline 10:03:00. Recovery took 3 minutes, inside the 5 minute RTO: pass on time.
- Replica lag recorded at the moment of the stop was 20 seconds, so at most 20 seconds of writes could be lost, inside the 1 minute RPO: pass on data.
- Had success only recovered at 10:07:00, recovery would be 7 minutes, above the 5 minute RTO, so the test fails even if no data was lost. Both conditions must hold.
Failure-mode table, one example row
| Dependency | Failure | Expected behaviour | Detection |
|---|---|---|---|
| Orders primary database | Instance stops | Replica promoted, orders resume within 5 minutes, at most 1 minute of writes lost | Alert on failed health checks within 1 minute |
Chaos engineering plan. Chaos engineering means deliberately injecting failure in a controlled way to find weaknesses. Document each experiment as: steady-state definition (the normal measurable behaviour), hypothesis, blast radius (the limit on what could be affected), injected fault, abort conditions, and findings. Start in staging, then move to production only with a small blast radius and abort switches.
Staging parity. List what differs from production (instance sizes, data volume, network topology, third-party stubs). A failover that recovers in staging with a tiny dataset says little about recovery time with production data, so state which conclusions transfer and which do not.
Pitfalls: a plan with no expected outcome (cannot fail); tests never rerun after architecture changes; results not stored beside the design.
Unlock Full Question Bank
Get access to all Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.