Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
How do you record SLOs and error-budget policies for a customer-facing API in its architecture documentation? Give a concrete example for a checkout success rate.
Sample Answer
Direct answer
Record each SLO as a small structured entry: the SLI (how we measure), the target, the window, what counts as good and bad events, the data source, the owner, and an error-budget policy saying what the team does as the budget is consumed. Put it in the reliability section of the architecture document, next to the diagram of the components it measures, and link it to the alerts and the runbook.
Definitions (one line each)
- SLI (service level indicator): a measured ratio, good events divided by valid events.
- SLO (service level objective): the internal target for that ratio over a window.
- SLA (service level agreement): a contractual promise to customers, with consequences (credits, penalties). It should be looser than the SLO so the team is warned before breaching it.
- Error budget: the allowed amount of failure, which is 100% minus the SLO.
Concrete example: checkout success rate
SLI: successful checkouts / valid checkout attempts
Good: order confirmed and payment authorized
Valid: all attempts except client-caused failures (card declined, invalid input)
Target: 99.9% over a rolling 28 days
Source: request logs at the API gateway
Owner: Payments team
Excluding card declines matters: a declined card is the system working correctly, and counting it would let customer behaviour, not service health, drive the number.
Budget arithmetic. Assume 12,000,000 valid attempts in 28 days (illustrative volume).
12,000,000×(1−0.999)=12,000
So 12,000 failed checkouts are allowed in the window. If 3,000 have failed by day 7, the budget used is 3,000 / 12,000 = 25%, and the window elapsed is 7 / 28 = 25%: exactly on pace. If the error rate is a sustained 0.5% against an allowed 0.1%, the burn rate (observed error rate divided by allowed rate) is 5, and a full budget is spent in 28 / 5 = 5.6 days.
Error-budget policy (what changes at each level)
| Budget used | Action |
|---|---|
| Below 50% | Normal feature delivery |
| 50% to 100% | Reliability work is prioritized in planning; risky releases need extra review |
| Over 100% | Feature releases pause except fixes; a postmortem (blameless written review of what happened) is required; resume when back within budget |
Measurement windows and remediation. A rolling 28-day window is smooth; a calendar-month window resets and can hide a bad week. Example: if checkouts fail heavily from the 24th to the 30th, a calendar-month SLO shows a fresh, full budget on the 1st, while a rolling window still counts those days until they slide out 28 days later. In the rolling window each day, the oldest day drops out and the newest is added, so the number moves gradually. Add a fast-burn alert that pages (automatically alerts the on-call engineer, night or day) and a slow-burn alert that opens a ticket. With the numbers above (12,000,000 attempts over 28 days is about 17,900 per hour, and a 28-day window is 672 hours):
- Fast burn: burn rate above 14 for 1 hour. That is an error rate above 14 x 0.1% = 1.4%, about 250 failed checkouts in the hour, and it uses 14 / 672, about 2%, of the whole budget in one hour. At that pace the budget is gone in 2 days, so a human must look now.
- Slow burn: burn rate above 3 for 24 hours. That is an error rate above 0.3%, about 1,290 failures per day (using about 428,600 attempts per day), and about 10.7% of the budget (3 / 28). Nothing is on fire, but the trend needs a ticket.
Both thresholds are illustrative; teams choose them from how much budget they accept losing before someone reacts. "Valid events" are the attempts that count (client-caused failures such as declined cards are removed), and "good events" are the valid ones that succeeded. Link a remediation playbook per alert (a short step-by-step guide, like a runbook): what to check first, how to roll back, when to escalate.
The 50 percent and 100 percent levels are conventions, not laws: 50 percent is an early warning with half the budget and roughly half the window left to react, and 100 percent is the point where the promise is broken. Teams tune them.
Pitfalls: SLO higher than the dependencies can deliver; counting invalid events; no policy consequences, so the number is decorative; the SLA set equal to the SLO.
A decision affects service latency and availability. How do you make sure its ADR carries observability and SLO requirements, and what is the minimum a reader needs recorded so operators can tell later whether the decision is holding up?
Sample Answer
Direct answer
I make observability and SLO (service-level objective, a measurable reliability target such as "99.9% of requests succeed") a required section of the ADR template, checked by the reviewer and by a service owner or SRE (site reliability engineer) before acceptance. The minimum a reader needs is five things: what we promised (the SLO), how we measure it (the indicators), when we will be told it is failing (the alert), what to do (a runbook link: a written step-by-step guide for handling that alert), and when we will look again (a revisit trigger). With those five recorded, an operator a year later can answer "is this decision still holding up?" from data instead of from memory.
The minimum record
| Field | What it must say | Why |
|---|---|---|
| SLO | Target, measurement window, who agreed it | Turns "fast and available" into a testable claim |
| SLIs (service-level indicators) | The exact metrics and where they are measured | Two teams can compute "latency" differently |
| Alert | Threshold, or error-budget burn rate (how fast we spend allowed failures) | Links the promise to a page (an urgent notification to whoever is on call) or a ticket |
| Runbook | Link, owner, first three diagnostic steps | Someone can act at 3 a.m. |
| Revisit trigger | A measured condition, and a review-by date | Tells us when the decision's assumptions have broken |
Making sure it is there (process)
- The ADR template has an "Operational impact" section that cannot be deleted. An empty section is a review blocker.
- The pull request check (or reviewer checklist) asks: does this change latency, availability, or dependencies, and is the SLO impact stated?
- An SRE is a named reviewer on decisions that touch latency or availability, so they see the ADR before acceptance, not during the first incident.
- Post-implementation: the dashboard and alert are linked from the ADR, and the ADR from the runbook, so both directions of search work.
Worked example: moving order processing to a queue
Decision: Order placement writes to a queue; a worker completes the order.
Operational impact
SLO: 95% of orders complete within 60 seconds of placement, measured
over 28 days. Availability of the placement endpoint: 99.9%.
SLIs: (1) age of the oldest unprocessed message (seconds),
(2) queue depth, (3) processing latency per order,
(4) placement endpoint success rate.
Alerts: page if oldest message age exceeds 300 s for 10 minutes;
ticket if queue depth grows for 30 minutes.
Runbook: link; first checks are worker health, poison messages, dependency errors.
Revisit: if the 60-second SLO is missed in 2 of any 6 consecutive
28-day windows, or peak traffic doubles, reopen this ADR.
Two terms in that example: a poison message is one that always fails processing, so it keeps being retried and blocks the messages behind it; queue depth is how many messages are waiting.
Readers can check the arithmetic: with 100,000 orders in a window, a 95% target allows at most 5,000 slower orders; 99.9% availability over 30 days allows 43.2 minutes of downtime, since 30 x 24 x 60 = 43,200 minutes and 0.1% of that is 43.2. Note the example SLO above is measured over 28 days, which allows 28 x 24 x 60 x 0.001 = 40.32 minutes; the 43.2 figure is the 30-day calendar-month version used for the error budget below.
Burn rate example: an error budget is the failure the SLO allows (here 43.2 minutes a month). A burn rate of 1 uses exactly that budget in 30 days. A burn rate of 10 (failing ten times faster than allowed) uses it in 30 / 10 = 3 days, which is why a high burn rate is worth an urgent alert.
Chain from SLO impact to alerts to runbooks
State the SLO impact as numbers ("this adds one hop, so the p99 (99th percentile, the time 99 of 100 requests finish within) latency budget drops by our estimate"). Worked example: checkout has a p99 budget of 300 ms; today's steps use about 240 ms, leaving 60 ms of headroom. The new queue hop is estimated at 25 ms, so headroom becomes 60 - 25 = 35 ms (a rough sum, good enough for budgeting; measure it after launch). Record "headroom: 60 ms before, 35 ms after". Each alert names the SLO it protects, and each alert links to a runbook. An alert with no SLO behind it is noise; an SLO with no alert is a wish.
Telemetry checklist for a network-facing change
Minimum (always record): request rate (requests per second), error rate (share of requests that fail), and latency percentiles (median and 99th percentile).
Add when relevant: saturation (how full a resource is, for example queue depth at 900 of 1,000 or a database connection pool 95% in use); dependency health checks (is the service we call up?); certificate expiry (the date a TLS certificate stops being valid, after which connections fail); and a trace ID (a unique ID stamped on a request and passed along every hop, so one slow request can be followed end to end).
Pitfalls
- Do not paste dashboards into the ADR; they change. Record the metric names and link out.
- A revisit trigger phrased as "if performance is bad" cannot fire. Use a number and a window.
- Set SLOs from what users need, not what the current system happens to deliver.
Which testing and resilience artifacts belong in architecture documentation so teams can validate the system, and how would you document a failover test plan with hypothesis, steps and expected outcomes?
Sample Answer
Direct answer
Architecture documentation should include the artifacts that let a team prove the design behaves as claimed: a test strategy, a list of failure modes with the expected behaviour, capacity and performance targets with load-test results, recovery objectives with backup-restore evidence, and resilience experiment plans. A failover test plan is written as a hypothesis, a controlled set of steps, expected outcomes with pass and fail criteria, and an abort and rollback path.
Artifacts that belong
| Artifact | Purpose |
|---|---|
| Failure-mode table | For each dependency: what if it fails, expected system behaviour, detection |
| Recovery targets | RTO (recovery time objective: how long an outage may last) and RPO (recovery point objective: how much data loss is tolerable), with the last measured result |
| Load and capacity test plan | Expected peak, headroom, and how it was validated |
| Backup restore test | Date, result, restore time |
| Failover and chaos plans | Below |
| Staging-parity notes | How the test environment differs from production, so results are read with the right caution |
Failover test plan (worked example, illustrative)
Terms used below: the primary is the database instance currently taking writes; a replica is a continuously updated copy that can take over; promoting the replica makes it the new primary; replica lag is how far behind the primary the copy is, in seconds of writes; staging is the production-like test environment; synthetic load is fake but realistic traffic generated by a tool; steady state is the normal measurable behaviour (for example 99% of orders succeed).
Title: Primary database failover, orders service
Hypothesis: If the primary database instance is stopped, the replica is promoted
and orders resume within the 5 minute RTO, with no more than
the 1 minute RPO of writes lost.
Environment: Staging, sized as production-lite (see parity notes)
Preconditions: steady synthetic load; dashboards open; on-call informed;
abort criteria agreed.
Steps:
1. Record baseline: order success rate, latency, replica lag.
2. Stop the primary instance.
3. Observe detection time and promotion.
4. Watch application reconnect behaviour and errors.
5. Restore the old primary as a replica.
Expected: Alert fires within 1 minute; promotion completes; success
rate returns to baseline; no duplicate or lost confirmed orders.
Pass/fail: Recovery within RTO and data loss within RPO, else fail.
Abort: Success rate stays under 50% for 10 minutes: roll back per runbook.
Record: Actual times, surprises, follow-up tickets, date of next test.
Judging the result against the targets (worked numbers, illustrative)
- Primary stopped at 10:00:00. Alert fired 10:00:40. Replica promoted 10:02:30. Order success back to baseline 10:03:00. Recovery took 3 minutes, inside the 5 minute RTO: pass on time.
- Replica lag recorded at the moment of the stop was 20 seconds, so at most 20 seconds of writes could be lost, inside the 1 minute RPO: pass on data.
- Had success only recovered at 10:07:00, recovery would be 7 minutes, above the 5 minute RTO, so the test fails even if no data was lost. Both conditions must hold.
Failure-mode table, one example row
| Dependency | Failure | Expected behaviour | Detection |
|---|---|---|---|
| Orders primary database | Instance stops | Replica promoted, orders resume within 5 minutes, at most 1 minute of writes lost | Alert on failed health checks within 1 minute |
Chaos engineering plan. Chaos engineering means deliberately injecting failure in a controlled way to find weaknesses. Document each experiment as: steady-state definition (the normal measurable behaviour), hypothesis, blast radius (the limit on what could be affected), injected fault, abort conditions, and findings. Start in staging, then move to production only with a small blast radius and abort switches.
Staging parity. List what differs from production (instance sizes, data volume, network topology, third-party stubs). A failover that recovers in staging with a tiny dataset says little about recovery time with production data, so state which conclusions transfer and which do not.
Pitfalls: a plan with no expected outcome (cannot fail); tests never rerun after architecture changes; results not stored beside the design.
You own an eventually consistent shopping-cart service that other teams call. How would you document its consistency behaviour so consumers know which anomalies to expect and what to do about them? Give a concrete example of what you would write.
Sample Answer
Direct answer
I document consistency as a contract, not a mood. The cart's page says, in testable terms: what is guaranteed (a client always sees its own writes), what is only eventual (other clients, caches, read models), the numeric time window in which they converge, the named anomalies a consumer can meet, and what to do about each. "Eventually consistent" (replicas may briefly disagree but converge if writes stop) with no window and no anomaly list is unusable, so I never write it alone.
Concrete example: the Cart API "Consistency behaviour" page (illustrative API and targets)
Guarantees
- Read-your-writes: pass the `cart-version` returned by a write on your next
read and you will see that write or a newer one.
- Other readers: converge within 2 s at p99 (99th percentile, the value 99% of
requests stay under) in normal operation. Alert threshold 5 s.
Not guaranteed during a regional failure.
- Checkout: use `consistency=strong`; it reads the primary region (the one region and database copy that accepts writes; replicas are read-only copies of it in other places).
Anomalies you may see
1 Stale read Cart looks older than what you just did in another client.
Cause: replica lag or cache. Do: re-read with cart-version.
2 Duplicate add A retried add appears twice. Cause: retry after timeout.
Do: send an Idempotency-Key on every write (a unique ID the client generates per action and sends as a header; the server stores it with the result and, if the same key arrives again, returns the stored result instead of adding the item twice).
3 Resurrected A removed item reappears after a concurrent add.
Cause: merge of two concurrent edits. Do: show
"cart updated" and let the user remove it again.
4 Lost quantity Two concurrent quantity changes: highest version wins.
Do: treat the response as truth, never compute locally.
5 Stale price Cart price may be up to 32 s old (2 s lag + 30 s cache).
Do: never charge from cart data; checkout re-prices.
The 32 s is arithmetic the consumer can check: 2 s replication target plus the 30 s cache TTL (time-to-live, how long a cached copy is served). It is a p99-based bound, not a hard maximum: the 2 s convergence target is a p99 and the alert only fires at 5 s, so a rare read can be staler.
Operation-to-consistency table
| Operation | Consistency | Notes |
|---|---|---|
| Add / update / remove line | Acknowledged after durable write on primary | Returns new cart-version |
| Get cart (default) | Eventual, via replica or cache | Up to 32 s stale |
Get cart with cart-version | Read-your-writes | May wait for replica catch-up |
| Checkout read | Strong | Slower, primary only |
Showing it on diagrams (notation)
Use one convention across all diagrams: solid arrows are synchronous, dashed arrows are asynchronous and carry a label with the lag target, and cache boxes carry their TTL. In the Mermaid diagram below the dashed arrow with an open head (--)) is the asynchronous replication; dashed arrows with a filled head (-->>) are just the replies to synchronous calls, which Mermaid draws dashed by convention.
Read the diagram top to bottom: the stale read appears at the step where the client asks the replica without a cart-version and gets version 6, even though the primary already acknowledged version 7. The last two steps show the fix.
sequenceDiagram
participant C as Client
participant P as Primary
participant R as Replica
C->>P: add item
P-->>C: ok, cart-version 7
P--)R: async replicate (lag target 2 s)
C->>R: get cart (no version)
R-->>C: version 6 (stale window)
C->>R: get cart (cart-version 7)
R-->>C: version 7
Stale-window annotation on a cache tier
Label each read path with "stale up to X": replica read 2 s, cache 30 s, combined worst case 32 s. This tells a consumer which path to choose for their need.
ML feature store variant
A feature store (a system that serves precomputed inputs to models), a kind of read model (a copy of data reshaped and kept separately for reading), may hold the cart's item count refreshed hourly. Document "feature freshness up to 1 hour" and the consequence: recommendations may ignore items added in the last hour. The model owner decides whether that is acceptable.
Reconciliation job
A reconciliation job is a scheduled repair task that compares two copies of the data and fixes differences. Here a nightly job compares primary and replica for carts changed in the last 24 hours, repairs divergence from the primary, and reports the count fixed. Document its schedule and that consumers must not depend on it for correctness; a rising repair count is a warning that replication is degrading.
Making teams not misread guarantees
Write the page with the same discipline as an ADR (Architecture Decision Record, a short document recording a decision and its reasons): state guarantees with numbers and a measurement point; ban bare words like "real time" and "eventually"; name anomalies; state the failure-mode behaviour. Add a 30-minute onboarding walkthrough that replays the anomalies with a demo cart. Contract tests (automated checks that replay each anomaly) keep the page honest.
Pitfalls
Promising a window you do not monitor; describing the design instead of what the consumer can rely on; leaving out what to do.
Design a one-page service page or repository README that lives with the code: which fields does it carry, who owns each, and how do you stop it from going stale?
Sample Answer
Direct answer
One page, kept in the same repository as the code, with about 15 to 20 short items (mostly a link or one line each, grouped under five headings in the skeleton below) that an on-call engineer or a new teammate needs, where each field has a named owner (mostly the owning team) and as many fields as possible are generated or checked instead of typed. It stays fresh because it changes in the same pull request as the code, has an automatic check in CI (continuous integration, the build that runs on every change), and shows a "last reviewed" date.
The page (skeleton)
# billing-invoice-api
Purpose: creates and stores invoices for completed orders.
Owner: billing team, on-call channel #billing-oncall
Tier / lifecycle: tier 1, production
Last reviewed: 2026-09-01
## Architecture
Diagram source: docs/architecture.mmd (rendered in the repo). Upstream: order-service.
Downstream: tax-provider (external), ledger-db.
## Interfaces
API spec: openapi.yaml. Events published: InvoiceCreated.
## Reliability
SLO (service-level objective, the reliability target): 99.9% of requests succeed each month.
Dashboard: <link>. Alerts: <link>. Runbook (step-by-step guide for handling problems): docs/runbook.md
## Operating it
Deploy: merge to main, pipeline deploys. Rollback: redeploy the previous tag.
Config and secrets: names and where stored (never the values).
Data: stores invoices, contains personal data, retained 7 years.
## Local development
Three commands to run it and its tests.
Reading the skeleton
- Tier is the business-criticality class a company assigns each service, for example tier 1 means an outage hurts customers or revenue directly and gets the fastest response and strictest checks, tier 3 means an internal tool that can wait until morning. The lifecycle value says whether it is experimental, production or being retired. Use your organisation's own definitions.
docs/architecture.mmdis a Mermaid file: a few lines of text that a renderer turns into a diagram, so the diagram is reviewed like code instead of being a stale screenshot.- On-call means the engineer currently responsible for responding to alerts, according to a paging schedule (a rota kept in a tool such as PagerDuty or Opsgenie). A CI check can look up the
Ownerteam in that schedule export and fail with, for example, "Owner 'billing team' has no on-call rotation" if it does not exist. - A monorepo is one repository holding many services; there the README lives in each service's folder and the check runs per folder.
Who owns each field
| Field | Owner | Freshness mechanism |
|---|---|---|
| Purpose, architecture, runbook, deploy and rollback | Owning team | Changed in the same pull request as code; reviewers required by the repo's CODEOWNERS file (a file naming who must approve changes to which paths) |
| Owner and on-call | Team lead | Checked in CI against the on-call schedule; a missing or unknown team fails the build |
| SLOs, dashboards | SRE or the team | Generated from monitoring configuration where possible |
| Dependencies, API spec | Team | Generated or checked from code and the spec file |
| Last reviewed | Team lead | CI opens an issue when it is older than 180 days |
Stopping it from going stale
- Docs-as-code (treating documentation like source code: plain text files, reviewed in pull requests and checked by CI): the page lives beside the code, so a change to behaviour and to the page are one review.
- Pull request template with one checkbox: "Does this change alter deploy, rollback, dependencies or the runbook? Updated the README?"
- CI checks: required headings and fields exist and links are not broken (these fail the build), and a "last reviewed" date older than 180 days opens an issue rather than failing it (see the last pitfall).
- Generate what is generated: dependency lists and SLO tables from source data, not by hand.
- Cap it at one page. Long pages are not maintained; deep material links out.
ML service variant
Add: model purpose and intended use, training data source, link to the model card (the standard summary of a model's intended use, data and evaluation), evaluation metrics and the threshold for shipping, the feature pipeline and its owner, the serving endpoint, how to roll back to the previous model version, and the monitor that watches for input drift (the live data changing away from the training data). Runbook link, SLOs, diagrams and owners stay the same as above.
Trade-offs and pitfalls
- A template with 30 required fields is filled with placeholders; keep the required set small.
- An owner that is a person becomes wrong when they change teams; use a team.
- Auto-failing builds on a stale date annoys people; prefer a ticket first, and fail the build only for missing owner or runbook.
Unlock Full Question Bank
Get access to all 23 Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.