Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
How would you represent, and keep current, the dependency map between many teams' APIs so that single points of failure are visible? What signals tell you the map has drifted from reality?
Sample Answer
Direct answer
I keep two maps and compare them. The declared map is what teams say they depend on, written in each service's descriptor file and rendered as a graph. The observed map is what traffic really does, built from distributed traces (records that follow one request across services) or from logs of the API gateway (the front door that receives external calls) and of the service mesh (infrastructure that proxies every service-to-service call and records it). Single points of failure are visible on the declared graph by fan-in (how many services depend on a node) and by whether the dependency is hard (no fallback), and drift is the difference between the two maps.
Representing it
A descriptor file is a small file kept in each service's repository that describes the service for tools (name, owner, dependencies). A sketch (illustrative field names):
name: checkout
owner: checkout-team
dependsOn:
- service: payments
contract: payments-api@v2
criticality: hard
- service: recommendations
contract: reco-api@v1
criticality: soft
- service: order-events
contract: OrderCreated@v3
criticality: async
A script reads every such file and turns each dependsOn entry into one directed edge (checkout -> payments, and so on), tagged hard, soft or async. In more detail:
- Each service declares
dependsOnentries with the callee's name, the API contract version, and a criticality flag: hard (requests fail without it), soft (degrades gracefully) or async (event-based, buffered). - A script builds a graph from all descriptors, in the catalog or as generated diagrams per domain.
- SPOF (single point of failure) flags: a node with fan-in of 2 or more, a hard dependency, and no redundancy or fallback recorded.
- Ownership is attached to each node, so a risky node always has a team behind it.
Keeping it current: declared versus observed
Illustrative example:
| Edges | |
|---|---|
| Declared | A to B, A to C, B to D, C to D, C to E |
| Observed in the last 30 days | A to B, B to D, C to D, C to E, A to E, B to F |
Comparing the two sets:
- Observed but undeclared: A to E and B to F. These are the dangerous ones, since nobody documented them and nobody is planning around them.
- Declared but never seen: A to C. Either the dependency is dead code, or it is a rarely used path such as a failover, which the team should confirm.
- Fan-in: D is called by B and C (fan-in 2), and E by C and A (fan-in 2, counting the undeclared edge). If D is a hard dependency with no fallback, it is the first SPOF to review, and the undeclared A to E edge shows E was underrated.
What to fix first. Fan-in of 2 or more only makes a node a candidate, because almost every shared service has that. Rank the candidates: first those where the dependency is hard and there is no fallback, then by how much traffic or revenue passes through, then by how weak the node is (single instance, unowned, poor reliability history). In the example, D (hard, no fallback, called by B and C) goes first; E is next only if the A to E path turns out to be hard.
Signals that the map has drifted
- Undeclared observed edges, and declared edges with zero traffic for 30 days (the set difference above).
- Services in traces or in the deploy system that have no descriptor, or descriptors whose service no longer deploys.
- Descriptors whose last-modified or last-verified date is older than a threshold such as 90 days while the service's code changed.
- A new owner or on-call team that no descriptor mentions.
- An incident where the postmortem (written review after an outage) says "we did not know X depended on Y".
- Contract version mismatch: the caller declares
v1while traffic showsv2.
Process to keep it honest
A weekly job diffs declared against observed and opens a ticket on the owning team for each mismatch. Declarations are updated in the same pull request as a new integration, and a CI (continuous integration, the automated build) check refuses names that do not resolve to a descriptor.
Trade-offs and pitfalls
- Traces are sampled (only a fraction of requests are recorded as traces to save cost), so a rare path may be missing from the observed map; use the mesh or gateway logs for completeness, and treat "not seen" as a question, not proof.
- The observed map cannot say whether an edge is hard or soft; that judgement stays with the owning team, and chaos or failure-injection tests (deliberately switching a dependency off in a controlled way to see whether callers survive) can confirm it.
- If the tool requires heavy manual entry, teams will skip it. Generate as much as possible from traces and treat declarations as the reviewed layer.
How would you capture in a decision record what the decision will cost the people who run the system, and how do you estimate that cost honestly so the decision-maker sees it?
Sample Answer
Direct answer
Add an "operational cost" section to the decision record, written by or with the people who will be on call, that lists what the decision adds to their work (new alerts and pages, runbooks (step-by-step instructions for handling an alert), upgrades, capacity, on-call training) and what it does to the SLO (service-level objective: a reliability target, such as 99.9% of requests succeed). Estimate it by analogy to something the team already runs, give a range rather than a single figure, and state the assumptions so the decision-maker can challenge them.
Template section for the record
## Operational cost
- SLO impact: adds one dependency in series; availability budget change below.
- New alerts/pages: expected pages per month (range) and basis.
- Recurring work (toil: repetitive manual operations work): upgrades, certificate/credential rotation, capacity reviews (hours/month).
- One-off work: runbooks, dashboards, training, game day (a planned failure drill, in hours).
- Rollout timeline and rollback: phases, who is responsible, how to undo.
- Assumptions and confidence: what we measured versus guessed.
- Owner of this cost: which team absorbs it and what they drop to make room.
Estimating honestly
- Anchor on an existing comparable. Suppose the team's current queue service produced 6 pages in a typical month and each page cost about 1.5 hours including follow-up (measured from the on-call log). That is 6 x 1.5 = 9 hours a month.
- Scale with reasoning, not hope. The new component is broadly similar but newer, so you estimate 1x to 2x: about 9 to 18 hours per month, plus a higher figure in the first quarter while the team learns it.
- Add the recurring and one-off work from the template as separate lines. Do not bury them in "maintenance".
- Show the range and the confidence. "9 to 18 hours a month, medium confidence, based on the queue service" is more credible than "12 hours".
- Name what is not known, such as failure modes nobody has seen yet.
Turning SLO impact into a number
An SLO of 99.9% availability over 30 days allows 30 x 24 x 60 x 0.001 = 43.2 minutes of downtime (this allowance is the error budget). Now suppose the decision adds a dependency that is itself 99.9% available and is needed for every request. That is a dependency in series: a request succeeds only if both parts work, so the chances multiply, 0.999 x 0.999 = 0.998001, about 99.8%. Downtime at 99.8% is 30 x 24 x 60 x 0.002 = 86.4 minutes (about 86 with the unrounded figure). If the two parts fail at different times, users see up to 86.4 minutes of failure against a 43.2-minute budget: double. To stay inside 43.2 minutes the two parts would have to share the budget, about 21.6 minutes each, which means each needs 99.95%. Alternatively add a fallback so the dependency's outage does not fail the request. That is a sentence a decision-maker can react to.
A filled-in example (adding a managed queue for order intake)
## Operational cost
- SLO impact: queue is needed for every order; 99.9% x 99.9% is about 99.8%,
so up to 86.4 min of failure a month against a 43.2 min budget.
Mitigation: fall back to a direct database write if the queue is down (to be tested).
- New alerts/pages: 6 to 12 pages a month, at 1.5 h each = 9 to 18 hours a month.
Basis: current queue service, 6 pages last month; new component assumed 1x to 2x.
- Recurring work: upgrades and credential rotation, about 3 hours a month.
- One-off work: runbook 8 h, dashboards 4 h, game day 4 h = 16 hours.
- Rollout and rollback: internal traffic in week 1, 10% of users in week 2,
everyone in week 3. Roll back if error rate stays above 1% for 10 minutes
(illustrative trigger). On-call lead decides.
- Assumptions and confidence: medium; pages measured for the old queue, guessed for the new one.
- Owner: platform team; it defers the dashboard refresh project to make room.
Making the decision-maker see it
- Put the cost next to the benefit on the same page, in the same units (hours per month, budget minutes).
- Say who pays: "this comes out of the platform team's roadmap capacity".
- Give the on-call engineers a review step before approval; they will spot the missing pages.
- Set a checkpoint: after 90 days compare actual pages and hours with the estimate and record the difference.
Rollout and rollback plan
Phase the rollout (internal traffic, small share of users, everyone), write the rollback trigger as a measurable condition (for example error rate above a stated level for a stated period), and record who decides. Include the timeline so on-call schedules can plan for the riskiest weeks.
Pitfalls
- Counting only build effort, and treating running cost as free.
- Precise-looking numbers with no source; a fabricated figure is worse than a labelled range.
- Estimates written by the proposer alone: they are optimistic.
- Recording the cost once and never comparing it with reality.
How should ADRs relate to runbooks and operational playbooks? Which consequences belong in the decision record, which in the runbook, and how do you keep the two in sync?
Sample Answer
Direct answer
An ADR (architecture decision record: a short, dated, immutable note stating a decision, its context and its consequences) explains why the system is built the way it is. A runbook (a step-by-step guide an on-call engineer follows during a specific situation such as a failover or a full disk) explains what to do when it misbehaves. The ADR records consequences that shape operations; the runbook turns each one into actions. A playbook is the wider, higher-level guide for handling a class of situation (who leads, who is told, what decisions are made, for example during a major outage). A runbook is narrower and technical: the exact steps for one specific alert. A playbook often points to several runbooks. Everything below about runbooks applies to playbooks as well, except that playbooks change less often and carry more coordination content. They stay in sync by linking to each other and by making a decision that changes operations fail review until the runbook is updated.
What goes where
| Content | ADR | Runbook |
|---|---|---|
| Context, options considered, rationale | Yes | No (one-line link only) |
| Operational consequences (new failure mode, new on-call task, new dependency) | Named and summarized | Turned into steps, commands, thresholds |
| Exact commands, dashboards, escalation contacts | No (they change often) | Yes |
| Status (proposed, accepted, superseded) | Yes, and never edited retroactively | "Last verified" date and owner |
| Who to call at 3 a.m. | No | Yes |
The simple test: if the text would still be true in three years, it is a decision and belongs in the ADR. If it changes when a hostname, a dashboard or a person changes, it belongs in the runbook.
Worked example
ADR-014: "Use a managed message queue with dead-letter handling for order events." Consequences recorded: messages that fail three times land in a dead-letter queue (a holding queue for messages that could not be processed); someone must review it; ordering is not guaranteed across partitions (a partition is one parallel lane of the queue; messages in different lanes can be processed in a different order than sent). The linked runbook, "Order events: dead-letter queue is growing", then contains: the alert that fires, the dashboard link, how to inspect a failed message, how to replay after a fix, when to page the owning team, and the safe order of steps.
Excerpt of the ADR (illustrative):
ADR-014: Use a managed queue with dead-letter handling for order events
Status: Accepted (2026-05-12)
Decision: Publish order events to a managed queue; after 3 failed deliveries, move the message to a dead-letter queue.
Consequences: Someone must review the dead-letter queue. Ordering is not guaranteed across partitions. Consumers must be idempotent.
Operational impact: see runbook "Order events: dead-letter queue is growing".
Excerpt of the runbook (illustrative):
Implements ADR-014. Owner: orders team. Last verified: 2026-08-20.
Alert: dead-letter queue depth > 0 for 15 minutes.
1. Open the dashboard (link). 2. Inspect one failed message (command).
3. If the cause is a bad deploy, roll back, then replay. 4. Page the owning team if depth keeps growing.
The ADR says why replay must be idempotent (safe to run twice); the runbook says how to check that before replaying.
Keeping them in sync
- Link both ways. The ADR has "Operational impact: see runbook X"; the runbook header says "Implements ADR-014".
- Pull request checklist: any change accepted as an ADR with an "operational consequences" field must include a runbook change or an explicit "none needed" with reason.
- Supersede, do not edit. When a decision changes, write a new ADR, mark the old one superseded, and update every runbook that links to it. A quick search for the old ADR ID finds them.
- Exercise them. Game-day drills (scheduled practice sessions where the team deliberately triggers a failure to test the runbook) and real incidents test runbooks; a postmortem (written review after an incident) action item that says "runbook was wrong" produces an edit, and if the cause was a changed decision, a new ADR.
- Same repository and review path where possible (docs-as-code: documentation kept as text files in the repository, reviewed like code), with owners named. Superseded means replaced by a newer decision but kept for history.
Pitfalls
- Pasting commands into the ADR: it goes stale while pretending to be immutable.
- Writing runbooks with no link to why: on-call cannot judge when it is safe to deviate.
- Leaving ownership of the link unclear: the on-call team owns the quality of the runbook, and the architect owns making sure each operational consequence of a decision actually reaches it. If neither owns the hand-off, consequences never get written down.
Which architecture diagram would you show to executives, to engineering leadership and to the engineers implementing a change? How do a high-level and a detailed diagram differ in what they must include and avoid?
Sample Answer
Direct answer
Three audiences, three diagrams from the same underlying model. Executives get a system-context picture with a few boxes in business language. Engineering leadership gets a container-level view showing major services, datastores, ownership and risk. The implementing engineers get a detailed diagram of the specific change: components, interfaces, data flow and failure paths. The high-level diagram must show purpose, boundaries and risks and avoid technology detail; the detailed one must show interfaces and behaviour and avoid the whole-system sprawl.
Which diagram for whom
| Audience | Diagram | Must include | Must avoid |
|---|---|---|---|
| Executives | System context (C4 level 1: the system, its users, its neighbours) | Users, the product as one box, external systems, one risk marker, plain labels | Protocols, cloud service names, internal services |
| Engineering leadership | Container diagram (C4 level 2: services, datastores, queues) | Owners, dependencies, regions, where the change lands | Class or library detail, every endpoint |
| Implementing engineers | Component diagram of the affected service (C4 level 3: the parts inside one service) plus a sequence diagram of the changed flow (not a C4 level: it lists the parts across the top and shows the messages between them in time order, top to bottom, for one scenario) | Interfaces, data formats, error paths, what is changing (highlighted) | Unchanged parts drawn in full, marketing wording |
Visual conventions and audience assumptions
- Executives: one colour for "this is what changes", nothing else coloured, and no cloud vocabulary without a gloss. State the assumption on the slide, for example "assumes no cloud background".
- Leadership: colour by owning team, so ownership gaps are visible.
- Engineers: bold or dashed outlines for the changed parts, so review focuses on the diff.
Worked example: ML recommendation service
A recommendation service is a machine learning model that suggests products to a shopper. One diagram, three layers, drawn on the same canvas or as a toggle so the shared names and arrows agree.
- Backend engineers see the request path: API, feature lookup (fetching the stored facts about this shopper, such as recent clicks), model server (the program that runs the trained model and returns suggestions), cache.
- Data scientists see the training loop: raw events, feature store (a database of prepared inputs for the model), training job (the run that fits the model to past data), model registry (the versioned catalogue of trained models).
- Product managers see a simplified strip: shopper, "suggestions", business KPI (key performance indicator: click-through, the share of shown suggestions that get clicked).
Shared across all three layers: the boxes "Shopper", "Recommendations" (the model server plus its model) and "Suggestions". These keep one name each. The implementing engineer's diagram for, say, a change to feature lookup shows the API, feature lookup and feature store as full boxes, with the changed feature lookup drawn in a bold outline and a note "new: reads shopper's last 10 clicks"; the training loop appears as one greyed-out box so it does not compete for attention.
Pitfalls
- Showing engineers the executive picture (no actionable detail) or executives the engineer picture (noise).
- Different names for the same thing across the diagrams.
- No legend, no date, no owner: a diagram nobody can trust is worse than none.
A decision must satisfy data-residency and regulatory constraints across regions. What fields and evidence would your decision-record process require, how do you record the legal constraint versus the technical control, and how would checks be automated?
Sample Answer
Direct answer
I would split every residency decision into two linked records: a legal constraint (what the law or contract demands, owned by legal) and a technical control (how our systems satisfy it, owned by engineering), and make the decision record (an ADR, architecture decision record: a short dated note of what was decided and why) point at both with evidence. Automation then checks the control against the real infrastructure, so compliance is verified continuously, not asserted once in a document.
Required fields
- Identity: id, title, status, date, decision owner, approvers (including legal or the data protection officer, DPO).
- Scope: data classes affected (for example personal data of EU customers, cardholder data for payments), regions, services.
- Legal constraint record: statement in plain words, source (regulation article or contract clause), jurisdiction, owner, review date. Example: GDPR (the EU General Data Protection Regulation) restricts moving personal data outside the EU/EEA without a valid transfer mechanism.
- Technical control record: mechanism, allowed regions, encryption key location, replication and backup rules, who may access, and which constraint id it
satisfies. - Evidence: data-flow map, infrastructure-as-code (IaC) configuration reference, test or scan results with dates, sub-processor list, DPIA (data protection impact assessment) where required, PCI DSS (Payment Card Industry Data Security Standard) scope diagram for cardholder data.
- Exceptions and expiry: any approved deviation, its compensating control and end date.
Plain meanings: a data class is a label for how sensitive a kind of data is (public, personal-eu, cardholder). A transfer mechanism is the legal basis that allows personal data to leave its home region, such as standard contractual clauses signed with the receiver. A sub-processor is a vendor that handles your customers' data on your behalf (for example a hosting or analytics provider), so it must be on the list and in an allowed region. A compensating control is a substitute safeguard that reduces the risk while an exception lasts (for example encrypting the data with a key held only in the EU). The DPO (data protection officer) is the designated person who informs, advises and monitors the organisation on data-protection compliance (GDPR Art. 39); legal responsibility for compliance stays with the controller (the organisation deciding why and how data is processed), not with the DPO. That is why the DPO reviews the constraint record and its reading of the law, while an accountable business owner still approves the decision.
Why keep constraint and control separate
If legal reinterprets a rule, only the constraint changes and every linked control is re-reviewed: search for records whose satisfies field equals LC-7, here TC-12, and re-review each. If engineering swaps a mechanism, the constraint stays put and the control record is superseded. One combined paragraph hides which one changed.
Fragment and a check that actually runs (data is illustrative)
import yaml
adr = yaml.safe_load("""
id: ADR-042
status: accepted
legal_constraint:
id: LC-7
statement: Personal data of EU customers is stored and processed only in EU regions.
source: GDPR transfer rules plus customer contract clause 9.2
owner: data-protection-officer
technical_control:
id: TC-12
satisfies: LC-7
mechanism: Storage and databases for data_class personal-eu are created only in allowed_regions
allowed_regions: [eu-central-1, eu-west-1]
check: inventory-region-scan
review_by: 2027-03-31
""")
inventory = [
{"name": "orders-db", "data_class": "personal-eu", "region": "eu-central-1"},
{"name": "audit-logs", "data_class": "personal-eu", "region": "eu-west-1"},
{"name": "analytics-bk", "data_class": "personal-eu", "region": "us-east-1"},
{"name": "public-cdn", "data_class": "public", "region": "us-east-1"},
]
allowed = adr["technical_control"]["allowed_regions"]
bad = [r for r in inventory if r["data_class"] == "personal-eu" and r["region"] not in allowed]
for r in bad:
print(f"FAIL {adr['id']} {adr['technical_control']['id']}: {r['name']} is in {r['region']}")
print("checked", len(inventory), "resources,", len(bad), "violation(s)")
raise SystemExit(1 if bad else 0)
What the check does: it reads the ADR's allowed regions, then loops over the resource inventory. Any resource whose data class is personal-eu and whose region is not in that allowed list is a violation. analytics-bk fails because it holds personal-eu data in us-east-1, while public-cdn is also in us-east-1 but holds public data, so it passes.
Output when run:
FAIL ADR-042 TC-12: analytics-bk is in us-east-1
checked 4 resources, 1 violation(s)
Automating it (GDPR and PCI evidence)
- CI (continuous integration) validates that every residency ADR has all required fields and a live constraint link.
- A scheduled job compares the ADR's allowed regions with the real resource inventory, as above, and opens a ticket per violation.
- The job stores its result with a timestamp as evidence; auditors read the trail instead of screenshots.
- For PCI, the same pattern checks that systems tagged cardholder data sit only inside the documented scope.
Trade-offs and pitfalls
- Automation proves the technical control, not the legal reading. Keep human review of the constraint.
- Stale evidence is worse than none: expire it and re-run.
Unlock Full Question Bank
Get access to all Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.