Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
You join a team that has resisted formal documentation. How do you introduce decision records so they get used, what format do you pick, and how do you keep them from going stale?
Sample Answer
Direct answer
I start by finding out why they resist, then make decision records small, useful on day one, and attached to work they already do. I pick a one-page markdown format in the repository (context, decision, alternatives, consequences, and a status line), introduce it on one real decision that is happening this week, and keep it alive with a review-by date and a short recurring triage. I do not announce a documentation policy.
Step 1: understand the resistance (first week)
Ask three engineers what happened to past documentation. Typical answers: "it was out of date within a month", "the template took an hour", "nobody read it". Each answer shapes the design. If the pain is staleness, the format must be cheap to keep current; if it is bureaucracy, no approvals.
Step 2: pick the format
An ADR (Architecture Decision Record, a short document capturing one significant decision, what else was considered and its consequences) in the lightweight style popularised by Michael Nygard (who described it in a short blog post as a few sentences each for title, status, context, decision and consequences), extended with a one-line "alternatives considered". Reasons:
- Markdown in the repo (for example
docs/adr/0001-use-postgres.md): no new tool, reviewed in the pull request like code, and history comes free from Git. - One page, about ten minutes to write. Longer formats can come later if wanted.
- Where it lives and who owns the habit: in the repository of the service it governs, with a shared index for cross-team ones. The team's tech lead approves the content; the engineering manager (the person who manages the team, not its technical design) makes sure the habit exists.
Step 3: introduce it on a live decision
Choose an upcoming choice the team is already debating (for example which queue to use). In the design discussion, write the ADR together in 20 minutes on a shared screen. The team experiences the payoff immediately: the argument ends, the reasoning is saved, and the next person will not repeat it.
Step 4: rules of thumb, not mandates
Write one when a decision is hard to reverse, affects another team, or was argued about. Skip it for small choices. This keeps volume low and value high.
Step 5: embed it in existing work
- The pull request template asks "does this change a recorded decision?".
- Code review comments link to the ADR instead of re-arguing.
- The onboarding list (what new hires work through in their first weeks) includes the 5 most important ADRs.
Step 6: keep them from going stale
| Mechanism | How it works |
|---|---|
Status and review_by date | Every ADR states its status and a date by which someone must re-check it (for example review_by: 2027-03-12) |
| Owner | The team that owns the service owns the ADR, not "the architects" |
| 30-minute quarterly triage | List accepted ADRs past their date; confirm, update or supersede each one |
| Supersede, don't edit | When a decision changes, write a new ADR that replaces the old one instead of rewriting it; keeps history trustworthy |
| Link from code | README and config point to the ADR so it is found when the code is touched |
Worked 90-day plan
- Days 1-10: listen; agree the format and folder.
- Days 10-30: write two ADRs on live decisions, with the team.
- Days 30-60: add the pull request template question; new hire reads them.
- Days 60-90: first triage meeting; ask the team whether it helped and drop anything that did not.
Success measures (qualitative first): people cite ADRs in reviews; a new hire finds the reason for a design without asking; the number of repeated debates falls. Counting ADRs alone is a vanity metric (a number that looks impressive but says nothing about whether anyone benefits).
Pitfalls
- Mandating from the top before any value is seen.
- Backfilling years of history; write only the decisions that still matter.
- Turning the ADR into an approval gate so people avoid writing them.
What would change my choice
If the team truly lives in a wiki and never opens the repo docs, I would put ADRs in the wiki with page history, but the repo is preferable because it sits with the code.
Why do you record the reasoning behind a decision and not just the decision? What goes wrong later without it, and which elements of the rationale (constraints, alternatives, assumptions, revisit triggers) matter most?
Sample Answer
Direct answer
Because the decision tells people what was chosen, but only the reasoning tells them whether it is still the right choice. Without it, later engineers either avoid changing something they do not understand or change it without knowing what it was protecting. An ADR (Architecture Decision Record, a short document capturing one design decision) should therefore record the constraints, the alternatives, the assumptions and a revisit trigger, not just a title.
What goes wrong without it
- Re-litigation: the same argument recurs every year because nobody remembers it was settled or why.
- Cargo-culting: people copy a pattern into places where its reason does not apply (copying a practice because it worked elsewhere, without understanding why it worked).
- Frozen or reckless changes: "nobody dares touch it" or "let's rip it out" (a fence taken down without asking why it was built, from the idea known as Chesterton's fence: before removing something, find out why it was put there).
- Onboarding and handoff: newcomers and successors see structure but not intent.
- Audit: an auditor or reviewer asking "why did you accept this risk, on what evidence, and who agreed?" gets no answer.
Which elements matter most
- Constraints (time, budget, team skills, regulations, existing systems): the most valuable, because they explain the choice and expire on their own.
- Assumptions (what we believed about traffic, growth, users): these are what silently become false.
- Alternatives, and why each lost: stops people proposing the discarded option again.
- Revisit triggers (a measurable condition that reopens the decision): converts an assumption into a tripwire.
Worked example
Decision: run the app in a single region.
Constraints: launch in 3 months, 4 engineers, users almost all in one country.
Alternatives: multi-region (rejected: doubles operational work for the team we have).
Assumptions: traffic stays in one country; a few hours of regional outage is tolerable.
Revisit trigger: more than 20% of users outside the region, or the availability target is missed twice in a quarter.
Two years later a new team sees one region. Without the record they either spend months "fixing" it or never notice the company expanded. With it, the trigger tells them exactly when to reopen.
Other uses
- Handoff: the new owner gets the reasoning along with the code.
- Audit: a dated trail of who decided what, with which evidence.
- Future review: reviewers can compare current reality with the recorded assumptions.
Reusable structure
Title, status, date; context and constraints; decision; alternatives with reasons rejected; assumptions; consequences; revisit trigger.
ML angle: batch versus online inference
Inference is running a trained model to get predictions. Batch inference computes predictions in bulk on a schedule; online inference computes them per request. A short rationale for it:
| Question | Batch | Online |
|---|---|---|
| Freshness the product needs | Hours are fine | Seconds |
| Latency limit per request | None | Tight |
| Cost and infrastructure | Cheaper, simpler | Always-on serving |
| Failure mode | Stale predictions until next run | Errors visible to users |
Revisit trigger for the ML choice: if the freshness need drops below hourly, or volume grows enough to change the cost picture. That table lets a later reviewer grasp the trade-off in a minute. With numbers (illustrative): a store has 2 million users and a slowly changing catalog. Batch: one nightly job of about 1 hour, so recommendations are at most about 25 hours old: 24 hours between runs plus the hour the job itself takes. Online: at a peak of 500 requests per second, each needing about 20 ms of model compute, the always-on servers need 500 x 0.02 = 10 cores busy at peak, plus headroom. Rationale recorded: 24-hour freshness is acceptable, so batch is cheaper. Revisit trigger: if product needs a click reflected within 5 minutes, batch cannot meet it and the decision reopens.
Pitfalls
Writing only the conclusion, or recording rationale that is a sales pitch with no alternatives. And never editing an old ADR to make past reasoning look better.
Write the decision, alternatives and consequences sections of an ADR for choosing eventual consistency over strong consistency for a read-heavy user-profile or cart service. Show how you would state the consequences honestly, from the user's point of view, rather than only the benefits.
Sample Answer
Direct answer
The ADR (Architecture Decision Record, a short document capturing one design decision, what else was considered, and what the team accepts as the price) states plainly that we accept eventual consistency, meaning replicas may briefly disagree but converge to the same value if writes stop, for the read-heavy user-profile and cart reads, and keep strong consistency (every read sees the latest acknowledged write) only where being wrong costs money. The consequences section is written as things a user can see, with a mitigation for each, not as a list of architectural virtues.
The ADR excerpt (cart service, illustrative numbers are targets, not measurements)
Terms used in the excerpt: the primary region is the one data-center region that accepts all writes; read replicas are copies of the data in other regions that only serve reads; replication lag (also called convergence lag) is the delay before a change on the primary appears on a replica; a network partition is a break in connectivity between regions; a runbook is a step-by-step guide for handling a specific problem.
ADR-0031: Serve cart and profile reads with eventual consistency
Status: Proposed
## Decision
We will write cart and profile changes to the primary region and serve
reads from regional read replicas and a 30-second cache. Reads are
eventually consistent by default. Checkout, payment and refunds always
read from the primary (strongly consistent). Concurrent edits to the same
cart line are resolved by version number: highest version wins, deletes
are kept as tombstones (markers that say "this item was removed"). Concurrent adds of different items are merged as a union. Losing one of two concurrent quantity changes is an accepted cost.
## Alternatives considered
A. Strong consistency everywhere (single primary for all reads).
Rejected: reads outnumber writes many times over, so every read would
pay cross-region latency and the primary becomes the scaling limit.
B. Eventual consistency everywhere, including checkout.
Rejected: a stale price or stock count at payment is a money error.
C. Read-your-writes only for the writing session (the session sends the
version it last wrote and reads wait for it). Adopted as a refinement
of the decision, not a separate option.
## Consequences
What users can see:
- Add an item on the phone, open the laptop within a few seconds: the
cart may look empty. Mitigation: session token carries the last version
written, so the same device always sees its own change.
Limit: a different device has no such token, so the phone-to-laptop
case is accepted, not solved. It clears once the replica catches up
(2 s target).
- Remove an item on one device while adding on another: the removed item
can reappear. Mitigation: tombstones, and the UI shows "cart updated".
- Price on the cart page may be up to 32 seconds old (2 s replication
target + 30 s cache). Mitigation: price is re-read from the primary at
checkout; a change shows a "price changed" prompt before payment.
- If the primary region is unreachable, users can still browse and edit
carts from a replica, but changes queue on the client and reach the primary later (replicas themselves are read-only).
What operations inherit: convergence lag must be monitored; an alert fires
when replication lag exceeds 5 s for 5 minutes; a runbook covers a stuck
replica; tests below prove the behaviour.
Why this wording is honest
- Every negative is a user-visible scene ("the cart looks empty"), so a product manager can judge it, not just an engineer.
- Each has a named mitigation and its limit, and the 32 s figure is arithmetic the reader can check: 2 + 30 = 32.
- Alternatives include a rejected extreme on both sides.
Conflict resolution and compensation (the rules the Decision section relies on)
This explains the merge rules that the ADR's Decision paragraph names. Version-based "highest version wins" can silently drop one of two concurrent changes. The decision states that cost, and adds a compensation: if two devices change the quantity of the same line within the convergence window, we keep the higher version and show the user the final cart, since losing a quantity tweak is cheap. For adds we prefer a union of items, since losing an added item is worse than a duplicate the user can delete.
Traced with real cart values. The cart starts as {shoes x1 at version 4, hat x1 at version 2}.
- Adds (union): the phone adds socks and the laptop, still on the old copy, adds a scarf. Merge: {shoes, hat, socks, scarf}. Nothing is lost.
- Quantity (higher version wins): the phone sets shoes to 2 and the laptop sets shoes to 3, both starting from version 4. The primary assigns version numbers in arrival order, so the phone's write becomes version 5 and the laptop's becomes version 6. Highest version wins: shoes x3. The phone's change to 2 is lost, and the user sees the final cart with "cart updated".
Payments-ledger variant
For a payments ledger the same ADR would choose the opposite: strongly consistent writes and reads from the primary. The SLO (service-level objective, the reliability target we promise) changes: "no double spend and balance reads are never stale" is the objective, at the cost of higher latency and lower availability during a network partition (the primary must be reachable). The consequences would say that plainly: during a regional outage, payments pause rather than risk an incorrect balance.
Validation entries recorded in the ADR
- Test: two simulated devices edit the same cart within the lag window; assert the final cart equals the defined merge result.
- Monitoring: replication lag, cache hit rate, and a counter of "tombstone resurrections" reported by clients.
- Playbook (a written checklist of steps to follow): how to fail reads over to the primary if lag stays above threshold.
Pitfalls
- Writing "may experience slight delays" hides the anomaly. Name it.
- Do not leave out what would make us reverse the decision. Add a revisit trigger, such as "if support tickets about disappearing items pass an agreed rate, revisit read-your-writes scope".
The checkout payment flow could use synchronous RPC or asynchronous messaging. Write the justification you would put in the architecture record, then compress it to a three-sentence summary for a director. What do you keep and what do you drop?
Sample Answer
Direct answer
For checkout, keep the payment authorisation as a synchronous call (the customer is waiting for a yes or no and you must know the answer to continue) and move everything that happens after the charge (receipt, stock update, shipping, analytics) to asynchronous messaging. Record both the choice and the price you are paying. For the director, keep the decision, the reason and the cost, and drop protocol detail.
The justification for the architecture record
# ADR-33: Synchronous authorisation, asynchronous post-payment steps
Status: Accepted | Deciders: payments lead, checkout lead, SRE lead
## Context
Checkout must tell the customer immediately whether payment worked.
After payment, several systems must react (receipt, stock, shipping)
but the customer does not wait for them. Today all are called in one chain,
so a slow shipping service delays or fails checkout.
## Options
A. Synchronous RPC (remote procedure call: the caller waits for the
reply) everywhere.
B. Asynchronous messaging everywhere (payment request goes on a queue).
C. Hybrid (chosen): sync for authorisation, async for what follows.
## Decision
C. Authorisation is a request/response call with a 3 s timeout and an
idempotency key (a unique ID so a retry cannot charge twice). After it
succeeds, the checkout service writes the order and an "order paid" message
in one database transaction, and a relay publishes it. This is the outbox pattern: the message is stored in a table beside the order, so a crash cannot save the order without the message, and a separate relay process then publishes it.
## Consequences
(+) Customer gets an immediate answer; downstream slowness cannot fail checkout.
(-) Consumers must be idempotent (messages can be delivered twice).
(-) Eventual consistency (other systems catch up shortly after, not instantly):
receipt may arrive seconds after the page says "paid". (-) A message broker (the server that holds queued messages until consumers take them) to run and monitor; a dead-letter queue
(where undeliverable messages go) needs an owner.
(-) Debugging spans systems, so we require a trace ID (a unique ID copied onto every message so one order can be followed across systems) on every message.
Why not the extremes
- All-synchronous couples checkout's availability to every downstream: with dependencies in series, availabilities multiply (three at 99.9% give roughly 99.7%).
- All-asynchronous for the payment itself turns a simple "declined" into a status the customer must poll, and complicates the error handling for the most important step.
Optional extension: rolling out a change of protocol on the synchronous hop
This is a side topic to the hybrid decision above, useful only if the synchronous hop also changes technology. If it moves from REST (JSON over HTTP) to gRPC (a Google-originated RPC framework using HTTP/2 and typed protobuf schemas), record the rollout in phases. Protobuf is a compact, strictly typed message format defined in a schema file:
- Define the protobuf contract and add contract tests (automated checks that caller and server both keep to the agreed request and reply shape) run in CI (continuous integration, the automatic build and test run on each change).
- Run the new endpoint alongside the old one; the server supports both.
- Shadow traffic: send a copy of real requests to the new path, compare responses, and do not serve them.
- Canary (try it on a small group first, like a canary warning miners of danger): move a small percentage of live traffic, watch error rate and latency, then step up.
- Retire the old path only after a full release cycle with no rollbacks. Rollback at any stage is a routing flag flip.
Compressing to three sentences for a director
Payment authorisation stays a synchronous call because the customer is waiting for a yes or no. Everything after the charge (receipts, stock, shipping) moves to queued messages so a slow downstream cannot block checkout. Cost: more moving parts to monitor.
That is 256 characters, inside an assumed 300-character limit for an executive summary (a common size for a summary box, not a universal standard; use your own template's limit).
| Keep | Drop |
|---|---|
| The decision and its business reason (customer experience, resilience) | Broker choice, timeout values, protobuf details |
| The one cost the director may be asked to fund or accept | Full option comparison (link to the record) |
| A consequence they can act on (monitoring effort) | Idempotency and outbox mechanics |
Rule of thumb: keep what changes a decision or a budget, drop what only changes an implementation.
Pitfalls
- Summary that promises "no downtime" or "faster" without a mechanism. Say what you know.
- A record that names the winner but not the accepted anomaly (receipt lag), so support is surprised later.
- Async without idempotency: duplicates become double emails or double stock deductions.
Walk me through a time you produced architecture documentation on a real project. Which notation and tools did you choose, why, and what did the choice change for the team?
Sample Answer
Direct answer
I documented a payments-and-orders backend for a team of about twelve engineers that had grown from three services to nine. I chose the C4 model (context, containers, components, code: four zoom levels of diagram, from "the system and its neighbours" down to "the classes inside one part") drawn as diagrams-as-code (diagrams defined in a text file that lives in the repository and renders to pictures), plus short Architecture Decision Records (ADRs: one-page notes recording a single decision and why it was made). The choice changed how the team worked: architecture questions were answered in pull requests instead of in meetings.
Situation and task (STAR: Situation, Task, Action, Result, the four parts of a story answer)
- Situation: onboarding (getting a new engineer productive) took two weeks of asking around. The only diagram was a whiteboard photo that was two years stale, and a wiki page nobody trusted.
- Task: give the team a picture that stays true, that a newcomer can read in one sitting, and that costs almost nothing to update.
Action: what I chose and why
| Option | Strength | Why I passed or picked it |
|---|---|---|
| Whiteboard tools (Miro and Excalidraw are free-form online drawing boards) | Fast, great for live design sessions | Kept for workshops only. Pictures drift because nobody owns the file. |
| Wiki with pasted images | Easy to find | A picture cannot be diffed (compared line by line to show what changed) or reviewed, so it rots. |
| ArchiMate (an enterprise-architecture modelling language with strict notation for business, application and technology layers) | Precise, good for portfolio-wide governance | Rarely needed outside large enterprises, and too heavy for one team. Few of the engineers could read it without training. |
| C4 as diagrams-as-code (Structurizr DSL and Mermaid are two text languages that turn a few lines into a diagram; both are common, kept in the repo) | Small vocabulary, four levels, reviewable diffs | Chosen. |
I stopped at the container level (a container here means a separately running piece such as a service, database or queue) and drew components only for the one service that was confusing. I wrote a legend (a small key) on every diagram: what a box means, what an arrow means, and the date it was last checked. The render step ran in continuous integration (CI, the automated build), so a broken diagram file failed the pull request.
What the artifacts looked like (illustrative)
The container diagram was a text file of about twenty short lines (the Structurizr DSL takes one statement per line, so the view block is written across separate lines), so a change to it appeared as a normal diff in a pull request (a proposed code change that teammates review before it is merged):
workspace {
model {
customer = person "Customer"
shop = softwareSystem "Shop" {
orders = container "Order service" "Owns order state" "Java"
payments = container "Payment service" "Charges cards" "Java"
ordersDb = container "Orders DB" "Stores orders" "PostgreSQL"
}
customer -> orders "Places order"
orders -> payments "Requests charge"
orders -> ordersDb "Reads and writes"
}
views {
container shop "Containers" {
include *
autoLayout lr
}
}
}
Each ADR was a title line plus four short fields (status, context, decision, consequences):
ADR-007: Order service owns payment retries
Status: Accepted
Context: Payment timeouts were retried by three callers, causing double charges.
Decision: Only the order service retries, with an idempotency key on every charge.
Consequences: One place to reason about retries; the payment service must honour the key.
Result
- New engineers were reading the container diagram on day one and could name every service's owner by the end of week one, instead of week two.
- A change to the order service was proposed as a pull request that edited the diagram and added an ADR in the same commit. Reviewers spotted a duplicated call path before any code was written, which avoided a rework that I estimated at about a sprint (an estimate, not a measurement).
- I measured nothing fancier than onboarding time and review comments, and I would say so rather than invent a precise figure.
Trade-offs and pitfalls
- Diagrams-as-code limit layout control; I accepted ugly-but-true over pretty-but-stale.
- A notation only helps if the audience can read it. Pick the one your readers already half-know, and say what would change your mind (for example a regulated client that requires ArchiMate views).
- Do not document everything. One context diagram, one container diagram and the ADRs cover most questions; component and code diagrams are drawn on demand.
Unlock Full Question Bank
Get access to all Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.