Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
You inherit a poorly documented microservice landscape with intermittent cascading failures. How do you audit and map the architecture, decide which diagrams to create first, and prioritise fixes?
Sample Answer
Direct answer
I would treat the cascading failures as the reason and the guide for the documentation work: start from evidence (what actually calls what, and what the failures look like), draw only the diagrams that explain the failure path first, and fix the riskiest amplification mechanisms (missing timeouts, retry storms, no isolation) before any broad cleanup. A cascading failure (one struggling service causing its callers to fail, and their callers) is almost always about how services fail together, so those pictures come first.
Step 1: get a real map, fast (week 1)
- Build the observed dependency graph from distributed traces (records of one request as it passes through services), gateway or service-mesh logs (a service mesh is a layer that proxies and logs every service-to-service call), and network flow data (records of which machines talked to which); do not trust the outdated wiki.
- Read the last 5 to 10 incident postmortems (the written reviews after an outage) and note which services were involved each time.
- Interview on-call engineers for "the services that scare you".
- Record for each service: owner, criticality, call timeouts, retry settings, and whether anything protects it.
Step 2: which diagrams first
| Order | Diagram | Why |
|---|---|---|
| 1 | Landscape or context view: the boundaries, callers and top-level dependencies (about 10 boxes) | Shared vocabulary for everyone |
| 2 | Dependency graph of the request path that failed in the incidents, annotated with timeouts and retries | This is where the cascade travels |
| 3 | Sequence diagram of that request including the failure branch | Shows where waiting and retrying compound |
| 4 | Deployment view for the shared pieces (databases, caches, queues, zones) | Reveals shared fate |
Everything else waits until it is needed.
Step 3: prioritise fixes by blast radius times likelihood
Blast radius means how much breaks if this fails. Score each finding from 1 to 5 on blast radius (how many users or services are affected) and on likelihood (how often it has failed or nearly failed), then multiply. Illustrative scores:
| Finding | Blast radius | Likelihood | Score |
|---|---|---|---|
| Rules service called with no timeout | 5 (every checkout) | 4 (in 3 of the last 5 incidents) | 20 |
| One database shared by 6 services | 5 | 2 (near misses only) | 10 |
| Analytics call retried 3 times | 1 (reports only) | 2 | 2 |
Fix the 20 first, then the 10. Typical fixes, in order (1 and 2 are the week-one core; 3 to 5 come later once the map exists):
- Timeouts on every outbound call, shorter than the caller's own deadline, because a call without a timeout holds threads while the dependency is slow. Timeouts and retries are amplification mechanisms: settings that make one failure bigger as it travels.
- Retry limits with backoff and jitter (waiting longer between attempts, with random variation). Retries multiply at each layer. If A calls B and B calls C, and each caller makes up to 3 attempts (1 call plus 2 retries), then A tries B up to 3 times and each of those B attempts tries C up to 3 times, so C can receive 3 x 3 = 9 attempts for one request. Put one more retrying layer above (four services in a chain, or a retrying client in front) and the deepest service can receive 3 x 3 x 3 = 27.
- Circuit breakers (stop calling a failing dependency for a while) and bulkheads (separate resource pools so one slow dependency cannot use up all threads).
- Load shedding (deliberately rejecting some requests to protect the service) and fallbacks (a simpler substitute response) for non-critical calls, so an analytics outage does not take down checkout.
- Remove shared single points such as one database serving many services.
Worked example
Suppose the incident traces show checkout calls pricing, which calls a rules service, which times out at 30 s, while checkout itself gives up at 5 s. The pricing threads pile up. First fix: set the rules timeout to 500 ms and cap retries at one retry (2 attempts). Along a checkout, pricing, rules chain that changes the worst case for the rules service from 3 x 3 = 9 attempts to 2 x 2 = 4. That single change is cheap, testable in a game day (a planned failure drill), and prevents the most common cascade, before anyone has finished a diagram of the whole estate.
What would change my order
If the incidents share a database or a cloud zone, the deployment view and redundancy jump to the front. If the estate is 20 services, I would just diagram everything in a week.
Pitfalls
Trying to document all services first delays fixes for months. Documenting without an owner per service leaves questions unanswered. Fixing symptoms in one service without agreeing timeout and retry rules across teams just moves the failure elsewhere, so publish the rules as a short standard in the same effort.
How would you capture in a decision record what the decision will cost the people who run the system, and how do you estimate that cost honestly so the decision-maker sees it?
Sample Answer
Direct answer
Add an "operational cost" section to the decision record, written by or with the people who will be on call, that lists what the decision adds to their work (new alerts and pages, runbooks (step-by-step instructions for handling an alert), upgrades, capacity, on-call training) and what it does to the SLO (service-level objective: a reliability target, such as 99.9% of requests succeed). Estimate it by analogy to something the team already runs, give a range rather than a single figure, and state the assumptions so the decision-maker can challenge them.
Template section for the record
## Operational cost
- SLO impact: adds one dependency in series; availability budget change below.
- New alerts/pages: expected pages per month (range) and basis.
- Recurring work (toil: repetitive manual operations work): upgrades, certificate/credential rotation, capacity reviews (hours/month).
- One-off work: runbooks, dashboards, training, game day (a planned failure drill, in hours).
- Rollout timeline and rollback: phases, who is responsible, how to undo.
- Assumptions and confidence: what we measured versus guessed.
- Owner of this cost: which team absorbs it and what they drop to make room.
Estimating honestly
- Anchor on an existing comparable. Suppose the team's current queue service produced 6 pages in a typical month and each page cost about 1.5 hours including follow-up (measured from the on-call log). That is 6 x 1.5 = 9 hours a month.
- Scale with reasoning, not hope. The new component is broadly similar but newer, so you estimate 1x to 2x: about 9 to 18 hours per month, plus a higher figure in the first quarter while the team learns it.
- Add the recurring and one-off work from the template as separate lines. Do not bury them in "maintenance".
- Show the range and the confidence. "9 to 18 hours a month, medium confidence, based on the queue service" is more credible than "12 hours".
- Name what is not known, such as failure modes nobody has seen yet.
Turning SLO impact into a number
An SLO of 99.9% availability over 30 days allows 30 x 24 x 60 x 0.001 = 43.2 minutes of downtime (this allowance is the error budget). Now suppose the decision adds a dependency that is itself 99.9% available and is needed for every request. That is a dependency in series: a request succeeds only if both parts work, so the chances multiply, 0.999 x 0.999 = 0.998001, about 99.8%. Downtime at 99.8% is 30 x 24 x 60 x 0.002 = 86.4 minutes (about 86 with the unrounded figure). If the two parts fail at different times, users see up to 86.4 minutes of failure against a 43.2-minute budget: double. To stay inside 43.2 minutes the two parts would have to share the budget, about 21.6 minutes each, which means each needs 99.95%. Alternatively add a fallback so the dependency's outage does not fail the request. That is a sentence a decision-maker can react to.
A filled-in example (adding a managed queue for order intake)
## Operational cost
- SLO impact: queue is needed for every order; 99.9% x 99.9% is about 99.8%,
so up to 86.4 min of failure a month against a 43.2 min budget.
Mitigation: fall back to a direct database write if the queue is down (to be tested).
- New alerts/pages: 6 to 12 pages a month, at 1.5 h each = 9 to 18 hours a month.
Basis: current queue service, 6 pages last month; new component assumed 1x to 2x.
- Recurring work: upgrades and credential rotation, about 3 hours a month.
- One-off work: runbook 8 h, dashboards 4 h, game day 4 h = 16 hours.
- Rollout and rollback: internal traffic in week 1, 10% of users in week 2,
everyone in week 3. Roll back if error rate stays above 1% for 10 minutes
(illustrative trigger). On-call lead decides.
- Assumptions and confidence: medium; pages measured for the old queue, guessed for the new one.
- Owner: platform team; it defers the dashboard refresh project to make room.
Making the decision-maker see it
- Put the cost next to the benefit on the same page, in the same units (hours per month, budget minutes).
- Say who pays: "this comes out of the platform team's roadmap capacity".
- Give the on-call engineers a review step before approval; they will spot the missing pages.
- Set a checkpoint: after 90 days compare actual pages and hours with the estimate and record the difference.
Rollout and rollback plan
Phase the rollout (internal traffic, small share of users, everyone), write the rollback trigger as a measurable condition (for example error rate above a stated level for a stated period), and record who decides. Include the timeline so on-call schedules can plan for the riskiest weeks.
Pitfalls
- Counting only build effort, and treating running cost as free.
- Precise-looking numbers with no source; a fabricated figure is worse than a labelled range.
- Estimates written by the proposer alone: they are optimistic.
- Recording the cost once and never comparing it with reality.
How would you represent, and keep current, the dependency map between many teams' APIs so that single points of failure are visible? What signals tell you the map has drifted from reality?
Sample Answer
Direct answer
I keep two maps and compare them. The declared map is what teams say they depend on, written in each service's descriptor file and rendered as a graph. The observed map is what traffic really does, built from distributed traces (records that follow one request across services) or from logs of the API gateway (the front door that receives external calls) and of the service mesh (infrastructure that proxies every service-to-service call and records it). Single points of failure are visible on the declared graph by fan-in (how many services depend on a node) and by whether the dependency is hard (no fallback), and drift is the difference between the two maps.
Representing it
A descriptor file is a small file kept in each service's repository that describes the service for tools (name, owner, dependencies). A sketch (illustrative field names):
name: checkout
owner: checkout-team
dependsOn:
- service: payments
contract: payments-api@v2
criticality: hard
- service: recommendations
contract: reco-api@v1
criticality: soft
- service: order-events
contract: OrderCreated@v3
criticality: async
A script reads every such file and turns each dependsOn entry into one directed edge (checkout -> payments, and so on), tagged hard, soft or async. In more detail:
- Each service declares
dependsOnentries with the callee's name, the API contract version, and a criticality flag: hard (requests fail without it), soft (degrades gracefully) or async (event-based, buffered). - A script builds a graph from all descriptors, in the catalog or as generated diagrams per domain.
- SPOF (single point of failure) flags: a node with fan-in of 2 or more, a hard dependency, and no redundancy or fallback recorded.
- Ownership is attached to each node, so a risky node always has a team behind it.
Keeping it current: declared versus observed
Illustrative example:
| Edges | |
|---|---|
| Declared | A to B, A to C, B to D, C to D, C to E |
| Observed in the last 30 days | A to B, B to D, C to D, C to E, A to E, B to F |
Comparing the two sets:
- Observed but undeclared: A to E and B to F. These are the dangerous ones, since nobody documented them and nobody is planning around them.
- Declared but never seen: A to C. Either the dependency is dead code, or it is a rarely used path such as a failover, which the team should confirm.
- Fan-in: D is called by B and C (fan-in 2), and E by C and A (fan-in 2, counting the undeclared edge). If D is a hard dependency with no fallback, it is the first SPOF to review, and the undeclared A to E edge shows E was underrated.
What to fix first. Fan-in of 2 or more only makes a node a candidate, because almost every shared service has that. Rank the candidates: first those where the dependency is hard and there is no fallback, then by how much traffic or revenue passes through, then by how weak the node is (single instance, unowned, poor reliability history). In the example, D (hard, no fallback, called by B and C) goes first; E is next only if the A to E path turns out to be hard.
Signals that the map has drifted
- Undeclared observed edges, and declared edges with zero traffic for 30 days (the set difference above).
- Services in traces or in the deploy system that have no descriptor, or descriptors whose service no longer deploys.
- Descriptors whose last-modified or last-verified date is older than a threshold such as 90 days while the service's code changed.
- A new owner or on-call team that no descriptor mentions.
- An incident where the postmortem (written review after an outage) says "we did not know X depended on Y".
- Contract version mismatch: the caller declares
v1while traffic showsv2.
Process to keep it honest
A weekly job diffs declared against observed and opens a ticket on the owning team for each mismatch. Declarations are updated in the same pull request as a new integration, and a CI (continuous integration, the automated build) check refuses names that do not resolve to a descriptor.
Trade-offs and pitfalls
- Traces are sampled (only a fraction of requests are recorded as traces to save cost), so a rare path may be missing from the observed map; use the mesh or gateway logs for completeness, and treat "not seen" as a question, not proof.
- The observed map cannot say whether an edge is hard or soft; that judgement stays with the owning team, and chaos or failure-injection tests (deliberately switching a dependency off in a controlled way to see whether callers survive) can confirm it.
- If the tool requires heavy manual entry, teams will skip it. Generate as much as possible from traces and treat declarations as the reviewed layer.
You are reviewing an architecture diagram of a payments platform that mixes synchronous authorisation with asynchronous settlement, and the diagram has no reliability annotations. What do you ask the author to add, and which gaps worry you most?
Sample Answer
Direct answer
I would ask the author for four things before I approve: (1) per-arrow reliability annotations (timeout, retry policy, idempotency, and what the caller does on failure); (2) an explicit state machine for a payment (a diagram of every state it can be in and which events move it between them), including an "unknown outcome" state; (3) the reconciliation path that proves money and records agree; (4) trust boundaries and ownership. The gap that worries me most is a synchronous authorisation call that times out with no defined way to find out what actually happened, because that is how you get double charges or lost payments.
What to ask the author to add, grouped by the two halves
Synchronous authorisation (customer waiting)
- Latency budget: if checkout must answer within 5 s, each hop gets a share and timeouts must nest, each outer timeout longer than the inner one plus its retries. For example client 5 s, gateway 4 s, acquirer call 3 s with no retry, so the acquirer times out first and the gateway can still report the outcome. If the acquirer call is retried once, the inner total is 3 + 3 = 6 s, which already exceeds the gateway's 4 s, so the budget must change: for example 1.5 s per attempt with one retry (3 s worst case, ignoring the backoff wait), or no retry at this hop and rely on the status inquiry instead.
- Idempotency (repeat-safe requests): an idempotency key from the client, stored with the result, so a retried request returns the original outcome instead of charging again.
- Behaviour when the acquirer (the bank-side card processor) is slow or down: decline, retry a second processor, or queue? Pick one and draw it.
Asynchronous settlement (money movement later)
- Delivery guarantee on the queue (at-least-once) and consumers that dedupe on payment ID.
- Dead-letter queue (DLQ, where repeatedly failing messages are parked), with an owner and an alert on its depth.
- Reconciliation: a scheduled comparison of our ledger (our record of every payment and amount) against the processor's settlement file (the daily report of what money it actually moved), with an alert on any mismatch. Without it, drift is silent.
- The outbox pattern (write the event and the state change in one database transaction, publish afterwards) so events are never lost between the database and the queue.
Across both
- SLOs (service-level objectives) per path: authorisation availability and latency; settlement completeness within a stated window.
- Trust boundaries and PCI DSS scope (the card-industry security standard): where card data lives, and that only tokens travel elsewhere.
- Ownership per component, runbook links, and RTO/RPO (recovery time objective, how fast we recover, and recovery point objective, how much data we may lose).
The payment state machine I would insist on
stateDiagram-v2
[*] --> Received
Received --> Authorized: acquirer approves
Received --> Declined: acquirer declines
Received --> Unknown: timeout
Unknown --> Authorized: status inquiry says approved
Unknown --> Declined: status inquiry says none
Authorized --> Settled: settlement file matches
Authorized --> Mismatch: settlement file differs
Mismatch --> Settled: manual correction
Worked example: the timeout
The acquirer call has a 3 second timeout and it fires. Without the Unknown state, the service either retries (risking a second authorisation) or fails the payment (risking a hold on the customer's card with no order). With it, the service records Unknown, asks the acquirer "what happened to reference k1?" (a status inquiry), and moves the payment to Authorized or Declined. If the inquiry also fails, a background job keeps asking, and reconciliation is the safety net.
Which gaps worry me most (in order)
- Ambiguous timeout outcome (above).
- No idempotency on settlement events, so duplicates double-post the ledger.
- No reconciliation, so errors are found by customers.
- Unmarked card-data boundary.
My rule for what blocks approval: a gap blocks if, when it happens, money is wrongly moved or nobody can tell. The first three each meet that test (double charge, double posting, silent drift); the card-data boundary is a compliance blocker. Everything else can be fixed after launch.
Extra check, service mesh (less commonly asked). If the diagram includes a service mesh (sidecar proxies handling service-to-service traffic), it must also show: whether retries are configured in the mesh AND in application code (retries multiply), timeouts per route, circuit-breaking settings, whether traffic is encrypted between services (mutual TLS: both sides prove their identity with certificates), and which calls bypass the mesh.
Extra check, shared cache (less commonly asked). If a shared Redis cache sits in a model-serving path, ask: what happens on a cache miss or Redis outage (fallback and load on the backing store)? Time-to-live (TTL, how long an entry lives) versus staleness of features; a stampede (many simultaneous misses all hitting the backing store together); eviction policy (which entries are dropped when memory is full); who owns the keyspace; and whether a single shared instance is a single point of failure (one component whose loss takes everything down). Example: a 60 s TTL on a feature means predictions may use values up to 60 s old, which the model owner must accept in writing.
Trade-offs and pitfalls
- Asking for everything makes the review unusable; mark which asks block approval (timeout outcome, idempotency, reconciliation) and which are follow-ups.
- Adding a second acquirer improves availability but adds routing complexity and reconciliation for two files.
Your organisation keeps architecture diagrams in a mix of drawing tools and every team draws differently. You want a lightweight standard plus automated checks in the pipeline. What would you standardise, what could a machine check versus a human reviewer, and where do diagrams drawn as code fall short of hand-drawn ones?
Sample Answer
Direct answer
Standardise a small set of things that make any diagram readable and traceable, keep the diagram's source in the repository next to the code, and let the pipeline check the mechanical rules while humans judge whether the picture is true and well pitched. Diagrams as code are easy to diff and review but fall short on layout control and polish. Docs-as-code means treating documentation like source code: text files, version control, review and automated checks. C4 is a notation with four zoom levels: context, containers, components and code.
What to standardise (lightweight)
- One notation: C4 levels 1 (context: our system and its users) and 2 (containers: separately running parts such as a web app, an API or a database) are mandatory; component level only where needed.
- One source format, stored in the repo (for example Mermaid, a text language where you write lines like
web --> apiand a tool draws the picture), rendered by the pipeline (the CI job, the automated run that builds and checks files on every change). - A header on every diagram: owner, level, last-reviewed date.
- Every arrow labelled with what flows and the protocol, plus sync or async.
- Element names that match the service catalogue (the registry listing every service, its owner and its links).
- A legend, so colour and line style carry no hidden meaning.
Machine versus human
| A machine can check | A human reviewer must judge |
|---|---|
| Source parses and renders | Whether the diagram matches the running system |
| Header present with owner, level, date | Whether it shows the right level of detail for its audience |
| Every arrow has a label | Whether an important dependency is missing |
| Element names exist in the catalogue | Whether the story is clear and uncluttered |
| Review date not older than the policy (say 90 days, a team choice) | Whether the labels are accurate, not merely present |
A minimal check (run it; the output is exactly what it printed)
import re
REQUIRED = ("owner", "level", "reviewed")
GOOD = """%% owner: payments-team
%% level: container
%% reviewed: 2026-09-01
flowchart LR
web[Web app] -->|"HTTPS, JSON"| api[Orders API]
api -->|"async event via queue"| pay[Payments service]
"""
BAD = """%% owner: payments-team
flowchart LR
web[Web app] --> api[Orders API]
api --> pay[Payments service]
"""
def check(name, text):
problems = []
meta = dict(re.findall(r"^%%\s*(\w+):\s*(.+)$", text, re.M))
for key in REQUIRED:
if key not in meta:
problems.append(f"missing header '{key}'")
for line in text.splitlines():
if "-->" in line and "|" not in line:
problems.append(f"unlabelled arrow: {line.strip()}")
return problems
for name, text in (("good.mmd", GOOD), ("bad.mmd", BAD)):
problems = check(name, text)
print(name, "OK" if not problems else "FAIL")
for p in problems:
print(" -", p)
good.mmd OK
bad.mmd FAIL
- missing header 'level'
- missing header 'reviewed'
- unlabelled arrow: web[Web app] --> api[Orders API]
- unlabelled arrow: api --> pay[Payments service]
How the check works, line by line: lines starting with %% are Mermaid comments, so the team uses them as a header (%% owner: payments-team). The pattern ^%%\s*(\w+):\s*(.+)$ matches a line that starts with %%, then a key word, a colon and a value; re.findall collects each pair into the dictionary meta. The loop then reports any of owner, level, reviewed that is absent. Separately, any line containing --> (an arrow) but no | (the character that wraps an arrow's label) is reported as an unlabelled arrow.
Edge case to know about: the check only matches the plain arrow -->. Mermaid also has dotted arrows (-.->) and thick arrows (==>), and dotted arrows are exactly what teams use for asynchronous calls, so an unlabelled a -.-> b (or a ==> b) passes this check silently (running the same test on those lines reports no problem). A production version should match every arrow form, for example with re.search(r"(-->|-\.->|==>)", line) and then test for a label. The check also does not verify the 'sync or async' wording or the review-date age, which need extra rules.
This only checks structure. The pipeline should also run the Mermaid command-line tool to confirm each file actually renders.
Where diagrams as code fall short
- Automatic layout is hard to tune, so large diagrams tangle and executive-ready visuals need manual polish.
- Limited expressiveness: unusual shapes, annotations and trust boundaries (lines marking where data crosses from a more trusted zone to a less trusted one) are awkward.
- Code drift still happens: a valid diagram file can describe a system that no longer exists.
- Authors need to learn a syntax, so contribution drops. Allow an escape hatch (a permitted exception to the rule): a hand-drawn diagram is acceptable if its editable source file is committed with the same header.
Migrating the mixed legacy set
- Inventory every diagram and its owner. Delete anything nobody claims.
- Convert only the top few that people actually use. Editable files from some tools can be imported; screenshots and slides cannot be converted and must be redrawn.
- Make the source authoritative and the rendered image read-only. Reject manual edits to rendered files (a reviewer sees them in the diff, the line-by-line list of changes), otherwise the two diverge.
- Keep unconverted legacy diagrams with a visible "legacy, unverified" banner and an expiry date.
Unlock Full Question Bank
Get access to all 27 Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.