Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
You are reviewing an architecture diagram of a payments platform that mixes synchronous authorisation with asynchronous settlement, and the diagram has no reliability annotations. What do you ask the author to add, and which gaps worry you most?
Sample Answer
Direct answer
I would ask the author for four things before I approve: (1) per-arrow reliability annotations (timeout, retry policy, idempotency, and what the caller does on failure); (2) an explicit state machine for a payment (a diagram of every state it can be in and which events move it between them), including an "unknown outcome" state; (3) the reconciliation path that proves money and records agree; (4) trust boundaries and ownership. The gap that worries me most is a synchronous authorisation call that times out with no defined way to find out what actually happened, because that is how you get double charges or lost payments.
What to ask the author to add, grouped by the two halves
Synchronous authorisation (customer waiting)
- Latency budget: if checkout must answer within 5 s, each hop gets a share and timeouts must nest, each outer timeout longer than the inner one plus its retries. For example client 5 s, gateway 4 s, acquirer call 3 s with no retry, so the acquirer times out first and the gateway can still report the outcome. If the acquirer call is retried once, the inner total is 3 + 3 = 6 s, which already exceeds the gateway's 4 s, so the budget must change: for example 1.5 s per attempt with one retry (3 s worst case, ignoring the backoff wait), or no retry at this hop and rely on the status inquiry instead.
- Idempotency (repeat-safe requests): an idempotency key from the client, stored with the result, so a retried request returns the original outcome instead of charging again.
- Behaviour when the acquirer (the bank-side card processor) is slow or down: decline, retry a second processor, or queue? Pick one and draw it.
Asynchronous settlement (money movement later)
- Delivery guarantee on the queue (at-least-once) and consumers that dedupe on payment ID.
- Dead-letter queue (DLQ, where repeatedly failing messages are parked), with an owner and an alert on its depth.
- Reconciliation: a scheduled comparison of our ledger (our record of every payment and amount) against the processor's settlement file (the daily report of what money it actually moved), with an alert on any mismatch. Without it, drift is silent.
- The outbox pattern (write the event and the state change in one database transaction, publish afterwards) so events are never lost between the database and the queue.
Across both
- SLOs (service-level objectives) per path: authorisation availability and latency; settlement completeness within a stated window.
- Trust boundaries and PCI DSS scope (the card-industry security standard): where card data lives, and that only tokens travel elsewhere.
- Ownership per component, runbook links, and RTO/RPO (recovery time objective, how fast we recover, and recovery point objective, how much data we may lose).
The payment state machine I would insist on
stateDiagram-v2
[*] --> Received
Received --> Authorized: acquirer approves
Received --> Declined: acquirer declines
Received --> Unknown: timeout
Unknown --> Authorized: status inquiry says approved
Unknown --> Declined: status inquiry says none
Authorized --> Settled: settlement file matches
Authorized --> Mismatch: settlement file differs
Mismatch --> Settled: manual correction
Worked example: the timeout
The acquirer call has a 3 second timeout and it fires. Without the Unknown state, the service either retries (risking a second authorisation) or fails the payment (risking a hold on the customer's card with no order). With it, the service records Unknown, asks the acquirer "what happened to reference k1?" (a status inquiry), and moves the payment to Authorized or Declined. If the inquiry also fails, a background job keeps asking, and reconciliation is the safety net.
Which gaps worry me most (in order)
- Ambiguous timeout outcome (above).
- No idempotency on settlement events, so duplicates double-post the ledger.
- No reconciliation, so errors are found by customers.
- Unmarked card-data boundary.
My rule for what blocks approval: a gap blocks if, when it happens, money is wrongly moved or nobody can tell. The first three each meet that test (double charge, double posting, silent drift); the card-data boundary is a compliance blocker. Everything else can be fixed after launch.
Extra check, service mesh (less commonly asked). If the diagram includes a service mesh (sidecar proxies handling service-to-service traffic), it must also show: whether retries are configured in the mesh AND in application code (retries multiply), timeouts per route, circuit-breaking settings, whether traffic is encrypted between services (mutual TLS: both sides prove their identity with certificates), and which calls bypass the mesh.
Extra check, shared cache (less commonly asked). If a shared Redis cache sits in a model-serving path, ask: what happens on a cache miss or Redis outage (fallback and load on the backing store)? Time-to-live (TTL, how long an entry lives) versus staleness of features; a stampede (many simultaneous misses all hitting the backing store together); eviction policy (which entries are dropped when memory is full); who owns the keyspace; and whether a single shared instance is a single point of failure (one component whose loss takes everything down). Example: a 60 s TTL on a feature means predictions may use values up to 60 s old, which the model owner must accept in writing.
Trade-offs and pitfalls
- Asking for everything makes the review unusable; mark which asks block approval (timeout outcome, idempotency, reconciliation) and which are follow-ups.
- Adding a second acquirer improves availability but adds routing complexity and reconciliation for two files.
Write the decision, alternatives and consequences sections of an ADR for choosing eventual consistency over strong consistency for a read-heavy user-profile or cart service. Show how you would state the consequences honestly, from the user's point of view, rather than only the benefits.
Sample Answer
Direct answer
The ADR (Architecture Decision Record, a short document capturing one design decision, what else was considered, and what the team accepts as the price) states plainly that we accept eventual consistency, meaning replicas may briefly disagree but converge to the same value if writes stop, for the read-heavy user-profile and cart reads, and keep strong consistency (every read sees the latest acknowledged write) only where being wrong costs money. The consequences section is written as things a user can see, with a mitigation for each, not as a list of architectural virtues.
The ADR excerpt (cart service, illustrative numbers are targets, not measurements)
Terms used in the excerpt: the primary region is the one data-center region that accepts all writes; read replicas are copies of the data in other regions that only serve reads; replication lag (also called convergence lag) is the delay before a change on the primary appears on a replica; a network partition is a break in connectivity between regions; a runbook is a step-by-step guide for handling a specific problem.
ADR-0031: Serve cart and profile reads with eventual consistency
Status: Proposed
## Decision
We will write cart and profile changes to the primary region and serve
reads from regional read replicas and a 30-second cache. Reads are
eventually consistent by default. Checkout, payment and refunds always
read from the primary (strongly consistent). Concurrent edits to the same
cart line are resolved by version number: highest version wins, deletes
are kept as tombstones (markers that say "this item was removed"). Concurrent adds of different items are merged as a union. Losing one of two concurrent quantity changes is an accepted cost.
## Alternatives considered
A. Strong consistency everywhere (single primary for all reads).
Rejected: reads outnumber writes many times over, so every read would
pay cross-region latency and the primary becomes the scaling limit.
B. Eventual consistency everywhere, including checkout.
Rejected: a stale price or stock count at payment is a money error.
C. Read-your-writes only for the writing session (the session sends the
version it last wrote and reads wait for it). Adopted as a refinement
of the decision, not a separate option.
## Consequences
What users can see:
- Add an item on the phone, open the laptop within a few seconds: the
cart may look empty. Mitigation: session token carries the last version
written, so the same device always sees its own change.
Limit: a different device has no such token, so the phone-to-laptop
case is accepted, not solved. It clears once the replica catches up
(2 s target).
- Remove an item on one device while adding on another: the removed item
can reappear. Mitigation: tombstones, and the UI shows "cart updated".
- Price on the cart page may be up to 32 seconds old (2 s replication
target + 30 s cache). Mitigation: price is re-read from the primary at
checkout; a change shows a "price changed" prompt before payment.
- If the primary region is unreachable, users can still browse and edit
carts from a replica, but changes queue on the client and reach the primary later (replicas themselves are read-only).
What operations inherit: convergence lag must be monitored; an alert fires
when replication lag exceeds 5 s for 5 minutes; a runbook covers a stuck
replica; tests below prove the behaviour.
Why this wording is honest
- Every negative is a user-visible scene ("the cart looks empty"), so a product manager can judge it, not just an engineer.
- Each has a named mitigation and its limit, and the 32 s figure is arithmetic the reader can check: 2 + 30 = 32.
- Alternatives include a rejected extreme on both sides.
Conflict resolution and compensation (the rules the Decision section relies on)
This explains the merge rules that the ADR's Decision paragraph names. Version-based "highest version wins" can silently drop one of two concurrent changes. The decision states that cost, and adds a compensation: if two devices change the quantity of the same line within the convergence window, we keep the higher version and show the user the final cart, since losing a quantity tweak is cheap. For adds we prefer a union of items, since losing an added item is worse than a duplicate the user can delete.
Traced with real cart values. The cart starts as {shoes x1 at version 4, hat x1 at version 2}.
- Adds (union): the phone adds socks and the laptop, still on the old copy, adds a scarf. Merge: {shoes, hat, socks, scarf}. Nothing is lost.
- Quantity (higher version wins): the phone sets shoes to 2 and the laptop sets shoes to 3, both starting from version 4. The primary assigns version numbers in arrival order, so the phone's write becomes version 5 and the laptop's becomes version 6. Highest version wins: shoes x3. The phone's change to 2 is lost, and the user sees the final cart with "cart updated".
Payments-ledger variant
For a payments ledger the same ADR would choose the opposite: strongly consistent writes and reads from the primary. The SLO (service-level objective, the reliability target we promise) changes: "no double spend and balance reads are never stale" is the objective, at the cost of higher latency and lower availability during a network partition (the primary must be reachable). The consequences would say that plainly: during a regional outage, payments pause rather than risk an incorrect balance.
Validation entries recorded in the ADR
- Test: two simulated devices edit the same cart within the lag window; assert the final cart equals the defined merge result.
- Monitoring: replication lag, cache hit rate, and a counter of "tombstone resurrections" reported by clients.
- Playbook (a written checklist of steps to follow): how to fail reads over to the primary if lag stays above threshold.
Pitfalls
- Writing "may experience slight delays" hides the anomaly. Name it.
- Do not leave out what would make us reverse the decision. Add a revisit trigger, such as "if support tickets about disappearing items pass an agreed rate, revisit read-your-writes scope".
You have a list of service-dependency pairs. Write a function that produces a Graphviz DOT graph from them, and explain why you might generate diagrams from data rather than draw them by hand.
Sample Answer
Approach
Turn the pairs into a set (removing duplicates), sort nodes and edges so the output is identical on every run, and emit one node line and one edge line per item in the DOT language (Graphviz's plain-text graph format: a -> b; means a directed edge). Escape quotes so an odd service name cannot break the file. The same structured data also produces a markdown summary table, so the diagram and the documentation come from one source.
Code (save as archdocs.py)
import json
import sys
def esc(text):
return text.replace("\\", "\\\\").replace('"', '\\"')
def to_dot(pairs, owners=None):
owners = owners or {}
edges = sorted(set(pairs)) # dedupe + stable order
nodes = sorted({n for pair in edges for n in pair} | set(owners))
lines = ["digraph services {", " rankdir=LR;", " node [shape=box];"]
for n in nodes:
label = esc(n) + ("\\n(" + esc(owners[n]) + ")" if n in owners else "")
lines.append(f' "{esc(n)}" [label="{label}"];')
for src, dst in edges:
lines.append(f' "{esc(src)}" -> "{esc(dst)}";')
lines.append("}")
return "\n".join(lines)
def to_markdown(services):
rows = ["| Service | Owner | SLO | Diagram |", "|---|---|---|---|"]
for name in sorted(services):
s = services[name]
rows.append(f"| {name} | {s['owner']} | {s['slo']} | [{s['diagram']}]({s['diagram']}) |")
return "\n".join(rows)
if __name__ == "__main__":
mode, path = sys.argv[1], sys.argv[2]
with open(path) as f:
data = json.load(f)
if mode == "dot":
owners = {n: s["owner"] for n, s in data["services"].items()}
print(to_dot([tuple(e) for e in data["edges"]], owners))
else:
print(to_markdown(data["services"]))
Sample input, saved as services.json, then two runs:
cat > services.json <<'EOF'
{
"edges": [["web", "auth"], ["web", "catalog"], ["catalog", "db"], ["auth", "db"], ["web", "auth"]],
"services": {
"auth": {"owner": "identity-team", "slo": "99.9% login success", "diagram": "docs/auth.svg"},
"catalog": {"owner": "catalog-team", "slo": "p95 < 300 ms", "diagram": "docs/catalog.svg"}
}
}
EOF
python3 archdocs.py dot services.json
python3 archdocs.py table services.json
Output of the first run (the duplicate web -> auth pair appears once):
digraph services {
rankdir=LR;
node [shape=box];
"auth" [label="auth\n(identity-team)"];
"catalog" [label="catalog\n(catalog-team)"];
"db" [label="db"];
"web" [label="web"];
"auth" -> "db";
"catalog" -> "db";
"web" -> "auth";
"web" -> "catalog";
}
Output of the second run:
| Service | Owner | SLO | Diagram |
|---|---|---|---|
| auth | identity-team | 99.9% login success | [docs/auth.svg](docs/auth.svg) |
| catalog | catalog-team | p95 < 300 ms | [docs/catalog.svg](docs/catalog.svg) |
Reading the code line by line
escputs a backslash before\and", so a name likea"bcannot end the quoted string early.edges = sorted(set(pairs)): thesetdrops the duplicate("web", "auth"), andsortedmakes the order the same on every run.nodes = sorted({n for pair in edges for n in pair} | set(owners)): a set comprehension that loops over each pair and then each name in it, so every service is listed once even if it appears in many pairs; the| set(owners)adds services that have an owner but no dependency edges, which would otherwise vanish from the picture (an isolated service is exactly the kind of gap a landscape diagram should show).rankdir=LR;tells Graphviz to lay the graph out left to right (the default is top to bottom);node [shape=box];makes every node a rectangle unless overridden.- The node loop writes one line per service, adding the owner on a second label line (
\nis a line break inside a DOT label) when one is known; the edge loop writes"src" -> "dst";. to_markdownbuilds one table row per service from the same data; SLO means service-level objective, the reliability target the team commits to (for example "p95 < 300 ms", meaning 95 percent of requests finish under 300 milliseconds).
Pipe the DOT output into dot -Tsvg (from the Graphviz tools) to get an image, or paste it into any online Graphviz viewer.
Why generate diagrams from data instead of drawing them
- Single source of truth: the dependency list already lives in a service catalog or config. A diagram built from it cannot disagree with it.
- Stays current: regenerate on every merge in continuous integration (CI, the automated build), so the picture never lags behind the system.
- Reviewable: a change to the dependency list appears as a small text diff in a pull request, which a screenshot of a hand-drawn box cannot offer.
- Consistent: every diagram gets the same shapes and names, which removes the arguments about icons.
- Cost of hand-drawing: nicer layout and emphasis. I would still hand-draw the executive overview and generate the detailed dependency views.
Complexity
Building the set is O(E) for E pairs and sorting is O(E log E) (nodes are at most 2E plus the number V of owned services), so the cost is O((E + V) log (E + V)), with O(E + V) memory for the output.
Edge cases
- Duplicate pairs are removed by the set.
- Names with quotes or backslashes are escaped (
to_dot([('a"b', 'c')])yields the node"a\"b"). - A service with no owner in the data gets no owner line, which is itself a useful documentation gap to flag.
- A service that is in the service data but appears in no pair is still drawn, as an unconnected box, because its name comes from the owners map.
- Cycles are valid in DOT and drawn as-is, which is helpful because circular dependencies are worth seeing.
- Very large graphs (hundreds of nodes) become unreadable: filter to one domain, or one hop from a chosen service, before drawing.
Walk me through the architecture of a system you built, in five minutes. How do you structure it, what do you leave out, and how do you adapt if I interrupt or cut you short?
Sample Answer
Direct answer
I treat the five minutes as a story with a fixed skeleton: why the system exists, the big picture, one path through it, one or two decisions that shaped it, and what I owned. I deliberately leave out the technology inventory and the edge cases, and I keep a 30-second version in my pocket in case I am cut short. If interrupted, I answer the question directly and then rejoin the story with a signpost.
The five-minute structure (with time budget)
| Time | Part | Content |
|---|---|---|
| 0:00-0:30 | Purpose and scale | What the system did, for whom, and why it was hard. |
| 0:30-1:30 | The big picture | Three to five boxes on the whiteboard: users, entry point, core service(s), data store, external dependencies. |
| 1:30-3:00 | One request end to end | Follow a single realistic request through the boxes, naming what happens at each step. |
| 3:00-4:15 | Two decisions | The trade-offs I made and what I would do differently. |
| 4:15-5:00 | My role and outcome | What I personally built or led, and what happened. |
What I leave out
The full list of technologies, every service, the org history, minor features, and code-level detail. If it does not help the listener understand the one request path or the two decisions, it waits for a question.
Adapting when interrupted or cut short
- Interruption with a question: answer it in one or two sentences, say "I will come back to that at the data step if that is useful", and return to the point I left.
- Interviewer clearly wants depth on one area: ask "Would you rather I go deeper on X than finish the overview?" and follow their lead. The overview is a means, not the goal.
- Cut to 90 seconds: I lead with the punchline first (what it was, one design decision, the result) and drop the middle. That is why the 30-second version exists.
Worked example (illustrative script): a photo-import service for an online marketplace
"Sellers upload up to fifty photos per listing, and slow uploads were costing us listings. I owned the redesign. Sellers upload directly to object storage using a temporary signed link, so our servers never carry the image bytes. An upload event lands on a queue; a resize worker produces thumbnails and stores the results; the listing service reads them. Follow one upload: the browser asks the API for a link, uploads, the queue message triggers the worker, the listing shows a placeholder until the thumbnail is ready. Two decisions: I chose eventual thumbnails over blocking the seller, trading a short placeholder for a faster upload; and I chose a queue rather than direct calls so a resize backlog could not take down listing pages. I would add better retry visibility next time. I designed the flow and wrote the worker; the outcome was that upload complaints dropped."
The script above is deliberately compact: about 150 words, roughly a minute of speech. It is the skeleton, not the full five minutes. At the pace in the time budget, the 'one request end to end' beat (1:30-3:00) would carry the extra detail, for example what the seller sees at each step and what happens if the upload fails halfway, and the two decisions (3:00-4:15) would each get a sentence on the alternative rejected and why. If cut to 90 seconds, the opening two sentences, one decision and the outcome are what survive.
Trade-offs and pitfalls
- Starting with the tech stack: the listener loses the purpose in the first thirty seconds.
- Telling every component with equal weight: it sounds like a diagram being read aloud.
- Inventing precise numbers: give scale and outcome honestly, in words you can defend.
- Ignoring the audience: check by eye or by asking whether to go deeper.
A decision was made when traffic was a tenth of today's, and a new regulation has since appeared. How do you find out that the assumptions behind an old ADR have gone stale, who owns the review, and what happens next?
Sample Answer
Direct answer
Do not rely on anyone remembering. Build freshness into the decision record when it is written: an ADR (architecture decision record: a short document capturing one significant decision, why it was made, and its consequences) lists the assumptions it rests on, each with a measurable tripwire (a pre-agreed condition that, once crossed, triggers a review) and a review-by date. Staleness then surfaces through four channels: metrics crossing a threshold, code or runtime drifting from the decision, external signals such as a new regulation, and the calendar. The team that owns the affected system runs the review, and it ends in one of four recorded outcomes.
1. Write the assumptions down as checkable claims (at decision time)
---
id: ADR-0014
title: Single-region PostgreSQL primary for order storage
status: accepted
date: 2023-02-14
owner: orders-team
review-by: 2024-02-14
tags: [data-residency, capacity]
---
## Assumptions and tripwires
| Assumption | Tripwire |
|----------------------------------------------|-------------------------------------------|
| Peak load stays under 1,000 requests/second | Sustained peak above 800 (80% of ceiling) |
| Customer data may live in one country | Any new rule on where personal data lives |
| A regional outage of a few hours is tolerable| Contract or SLA (service-level agreement) tightens |
Worked numbers: peak load was 500 requests per second when the ADR was accepted, so the ceiling of 1,000 allowed 2x growth. The tripwire at 800 (0.8 x 1,000) fires at 1.6x growth. A system now at 10x (5,000 per second) crossed it long ago. Without the tripwire, nobody looked until something hurt.
2. Freshness signals (mostly a script and a tag, not a product)
- Metric tripwires: an alert on the traffic dashboard opens a review ticket (not a page) titled "Review ADR-0014", with one open ticket per ADR so repeat firings do not spam.
- Code and runtime drift: each ADR lists the paths or services it covers. A check on pull requests (PRs) that touch those paths comments "this area is governed by ADR-0014, do its assumptions still hold?". A second-region resource appearing in infrastructure code while the ADR says single-region is a deviation worth a ticket.
- Regulatory and external signals: ADRs carry tags such as data-residency or retention. When legal or compliance flags a new rule, they search by tag and get the list of decisions to re-check. This is the fix for the new-regulation half of the scenario.
- Calendar: a scheduled job lists ADRs past their review-by date (12 months for critical systems, 24 otherwise) and opens tickets.
Regulation search in practice: with ADRs stored as files in a repository, the search is one command, run here on three sample files (only two carry the tag):
$ mkdir -p adr
$ printf -- '---\nid: ADR-0014\ntags: [data-residency, capacity]\n---\n' > adr/0014-single-region-postgres.md
$ printf -- '---\nid: ADR-0031\ntags: [retention, data-residency]\n---\n' > adr/0031-eu-customer-backups.md
$ printf -- '---\nid: ADR-0022\ntags: [capacity]\n---\n' > adr/0022-queue.md
$ grep -l "tags:.*data-residency" adr/*.md | sort
adr/0014-single-region-postgres.md
adr/0031-eu-customer-backups.md
The legal contact gets a short list of files, and each becomes a review ticket for its owning team.
3. Who owns the review
The owner field names a team, not a person, because people move. That team's tech lead runs the review; an engineering manager makes sure it lands in the backlog with a due date (two sprints is reasonable). Architects are consulted only for critical or cross-team decisions. If the owning team no longer exists, ownership follows the system in the service catalogue (a registry listing every service and the team that owns it); if no team owns the system, the architecture lead assigns one or marks the ADR deprecated.
4. What happens next: four outcomes, each recorded
| Outcome | When | What you write |
|---|---|---|
| Still valid | Assumptions re-checked and hold | A dated "Reviewed" line with evidence; move review-by forward |
| Amend | Choice stands, details change (new ceiling, new tripwire) | A dated amendment section |
| Supersede | The chosen option no longer fits | New ADR; old one set to superseded with links both ways |
| Deprecate | The system or decision no longer applies | Status change and a reason |
In the scenario: the ticket opens in week 0. In week 1 the owner pulls the load graph (5,000 per second against a 1,000 ceiling) and gets a written note from legal on the new rule. A 30-minute outcome meeting in week 2 supersedes the ADR: a new ADR is proposed, then accepted, and the migration becomes a normal tracked project. If the regulation carries a legal deadline, it skips the routine cycle: compliance sets the date, and any interim mitigation is recorded as its own proposed ADR.
Trade-offs and pitfalls
- Too many tripwires cause alert fatigue (people start ignoring alerts because there are too many): keep two to four load-bearing assumptions per ADR.
- Reviews become rubber stamps (approvals with no real check) unless "still valid" requires a link to evidence.
- A stale ADR was not a wrong decision. Leave the old body unedited and say it was right for its time; rewriting history destroys the reason future readers need.
- Very small organisations with a handful of ADRs can skip the automation and use the calendar alone.
Unlock Full Question Bank
Get access to all Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.