Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
A cross-region system can lose connectivity between regions or lose its primary database. How do you show those scenarios on architecture diagrams and in operational docs so that both the user-visible degradation and the recovery path are clear?
Sample Answer
Direct answer
I would draw the normal state once, then give each failure its own diagram overlay and its own runbook page, with the same four headings every time: what the user sees, what data is at risk, how we recover, and who decides. A single "resilient architecture" picture hides the difference between losing the link between regions and losing the primary database, and those are different problems with different recoveries.
The diagram (deployment view: which software runs where, in which region, with the two scenarios marked)
flowchart LR
U[Users] --> G[Global traffic routing]
G --> A[Region A: app tier]
G -.->|"only if A is declared down"| B[Region B: app tier]
A --> P[(Primary database)]
P ==>|"async replication, lag shown on dashboard"| R[(Replica in Region B)]
B -.-> R
X1{{"Scenario 1: link between regions is cut (strikes the replication path)"}} -.-> R
X2{{"Scenario 2: primary database is lost"}} -.-> P
How to read it: solid arrows are the normal path; the dotted arrow from global routing to Region B is used only after Region A is declared down (a named person or an agreed rule, for example health checks failing for 3 minutes, formally decides A is out); the thick arrow is data replication; each hexagon marks where a scenario strikes: scenario 1 touches the replica end of the thick replication arrow (Mermaid cannot point at an arrow itself), scenario 2 touches the primary. Alert on lag during scenario 1, because the lag is the data you would lose.
Beside the picture, always state: the consistency model (what readers are promised to see after a write; here replication is asynchronous, meaning the primary confirms a write before the replica has it, so the replica trails), the RPO (recovery point objective: the most data we accept losing, measured in time) and the RTO (recovery time objective: how long recovery may take). Put real values on the page, for example RPO 5 seconds and RTO 15 minutes. Measured lag of 4 seconds fits inside the 5 second RPO; lag of 30 seconds would not, and the dashboard should alert before that.
Scenario pages
| 1. Link between regions is cut | 2. Primary database is lost | |
|---|---|---|
| User-visible | Region A keeps serving. Data that changed in A is not visible in B. Any users or jobs already reading from B (for example a reporting job or users pinned to B) see older data; global routing does not send new traffic to B while A is healthy. | Writes fail or queue until a new primary is promoted. Reads may continue from the replica. |
| Data at risk | Nothing lost yet, but replication lag grows for as long as the link is cut, so the RPO clock is running: if Region A is then lost before the link returns, everything written since the cut is gone. Also divergence risk if both sides accept writes. | Writes not yet replicated. Worst case equals the replication lag at the moment of failure (if lag was 4 seconds, up to 4 seconds of committed writes). |
| Recovery | Wait and let replication catch up; do not promote B. | Promote the replica after the decision rule is met; repoint the app; rebuild a new replica. |
| Decision | On-call, no failover needed. | Named authority, using the rule on the page. |
A runbook is the step-by-step operating procedure an on-call engineer follows during an incident. A filled scenario 2 page, for example: User sees: checkout fails for up to 15 minutes. Data at risk: up to 4 seconds of orders (current lag). Recover: confirm the primary is truly gone, promote the Region B replica, repoint the app, build a new replica. Decides: the incident commander (the person in charge during the incident), after 3 minutes of failed health checks.
Split-brain and flapping
Split-brain means two regions both believe they are the primary and both accept writes, so their data diverges. The docs must state the single-writer rule: only the holder of the current write lease (an expiring right to write) may accept writes, and each promotion increments an epoch number (a counter that only goes up) so an old primary is rejected when it returns. Traced: Region A is primary in epoch 7. A goes dark and B's replica is promoted in epoch 8. When A returns it still thinks it is epoch 7, but the storage layer and gateway only accept writes stamped epoch 8 or higher, so A's writes are rejected and A must resync as a replica. Flapping means leadership bouncing back and forth on a shaky link. The runbook adds a hold-down time (a minimum wait before another promotion) and forbids automatic promotion on a partition alone.
Make lag visible
The topology diagram links to a live dashboard showing replication lag, because the recovery decision ("how much would we lose?") is answered by that number.
Trade-offs and pitfalls
- Drawn failovers that skip fail-back (returning to the original primary after the outage) leave teams stuck with the wrong primary. Document the return path.
- Do not promise a zero RPO with asynchronous replication.
You have to define the minimum useful architecture documentation set for a critical, globally distributed service. Which documents and diagrams would you insist on, who is each for, and what would you refuse to write?
Sample Answer
Direct answer
For a critical, globally distributed service I would insist on a small set: a context diagram, a container diagram (one box per separately running service, database or queue) with a data-flow view, a deployment (topology) view (which parts run in which region and how they connect), Architecture Decision Records (ADRs, one-page records of individual decisions), an operations runbook (a step-by-step page an on-call engineer follows when an alert fires), a failure-mode/disaster-recovery document, and a short quality-goals-and-SLOs page: seven items in all, matching the table below. Each has a named reader. I would refuse to write anything whose only reader is "future auditors" or that duplicates what code and infrastructure files already state. For the frame I would start with C4 (four zoom levels of diagram, learnable in an hour) for the pictures, and treat arc42 (a free template with twelve numbered sections for architecture documentation) as a checklist of headings to borrow from, not a second thing to fill in. A team new to this should pick C4 first.
The minimum set, with reader and purpose
| Document | Reader | Question it answers |
|---|---|---|
| Context diagram + one-paragraph purpose | Executives, new hires, client procurement | What is this, who uses it, what does it depend on? |
| Container diagram (services, datastores, queues) | All engineers | What runs, and who owns each piece? |
| Deployment view: regions, failover paths (failover means switching traffic to a healthy region), data residency boundaries (legal limits on which country's servers may hold customer data) | SREs (site reliability engineers), security, cloud architects | Where does it run, and what happens when a region fails? |
| ADRs | Future maintainers | Why is it built this way, what did we trade away? |
| Runbook: alerts, dashboards, rollback, escalation | On-call engineers at 3 a.m. | What do I do right now? |
| Failure modes and recovery targets (RTO: recovery time objective; RPO: recovery point objective, how much data loss is tolerable) | SREs, engineering leadership | What breaks, how fast do we recover, how much data can we lose? |
| Quality goals and SLOs (service-level objectives: measurable reliability targets) | Product, SREs | What promises have we made? |
For a client proposal, the technical appendix is the container and deployment diagrams plus the quality goals, with notation stated on the page and detail pitched at engineers, while the front section stays at business language (availability, compliance, cost).
Worked example: sizing "globally distributed"
A service in three regions with a 99.95% SLO allows 0.05% of a 30-day month (43,200 minutes) as downtime: 43,200 x 0.0005 = 21.6 minutes. That single derived number drives the failure-mode document (this 21.6 minutes is the error budget, the total downtime the promise tolerates): if regional failover takes 15 minutes, one failover uses most of the month, so the document must say who decides to fail over and how it is rehearsed. A failover plan written from it would read: "the on-call engineer may trigger failover after 5 minutes of regional failure; target recovery 15 minutes; rehearsed quarterly". Each of those lines follows from the budget: a decision that waits half an hour would exhaust it.
The first lines of the runbook page for that alert look like this (illustrative):
Alert: RegionHealthCheckFailing (region eu-west)
Impact: EU customers see errors or slow checkout.
First step: open the regional dashboard and confirm errors exceed 5% for 5 minutes.
Then: run the failover procedure, step 1 of 6 (shift traffic to the standby region).
Escalate: page the incident commander if not recovered in 15 minutes.
What I would refuse to write
- A component or class diagram for every service (generate it from code if anyone needs it).
- A prose copy of the configuration (the infrastructure-as-code repository, where servers and networks are defined in text files, is the source of truth).
- A 60-page architecture "bible" that no one owns; every document needs an owner and a last-reviewed date or it gets deleted.
- Status-of-project material: that belongs in a tracker.
Trade-offs and pitfalls
The risk of too little is tribal knowledge (facts that live only in a few people's heads); of too much, stale pages that mislead on-call engineers. My rule: keep a document only if a named reader would suffer without it, and put it beside the code (docs-as-code, meaning docs versioned and reviewed like code) so it changes in the same pull request.
Write a small script that scans a repository for ADR files, reads their metadata (id, status, title, date), and fails a CI check if an ADR is missing required fields or has an invalid status. Then explain what else you would lint.
Sample Answer
Approach
Treat each ADR (architecture decision record: a short document capturing one decision) as a Markdown file that starts with a small metadata block between --- lines (this block is called front matter). The script walks the folder for files named ADR-*.md, parses the block, collects every problem instead of stopping at the first one, prints the problems, and exits with code 1 (an exit code is the number a program returns when it finishes: 0 means success, anything else means failure) so CI (continuous integration: checks that run on every pull request) fails. It also writes the ADR relationships as a JSON graph. It uses only the Python standard library and a deliberately flat key: value parser, so there is nothing to install. Save it as adr_lint.py.
import json, re, sys, tempfile
from datetime import date
from pathlib import Path
REQUIRED = ["id", "title", "status", "date"]
STATUSES = {"proposed", "accepted", "deprecated", "superseded", "rejected"}
ID_RE = re.compile(r"^ADR-\d{4}$")
DATE_RE = re.compile(r"\d{4}-\d{2}-\d{2}")
def parse_front_matter(text):
"""Return dict of key: value from a leading --- block, or None if absent."""
lines = text.splitlines()
if not lines or lines[0].strip() != "---":
return None
meta = {}
for line in lines[1:]:
if line.strip() == "---":
return meta
if ":" in line:
key, value = line.split(":", 1)
meta[key.strip().lower()] = value.strip()
return None # block never closed
def split_ids(value):
return [v.strip() for v in value.split(",") if v.strip()]
def lint(root):
errors, adrs = [], {}
for path in sorted(Path(root).rglob("*.md")):
if not path.name.startswith("ADR-"):
continue
meta = parse_front_matter(path.read_text(encoding="utf-8"))
if meta is None:
errors.append(f"{path.name}: missing or unclosed front matter")
continue
for field in REQUIRED:
if not meta.get(field):
errors.append(f"{path.name}: missing required field '{field}'")
adr_id = meta.get("id", "")
if adr_id and not ID_RE.match(adr_id):
errors.append(f"{path.name}: id '{adr_id}' must look like ADR-0007")
if adr_id and not path.name.startswith(adr_id):
errors.append(f"{path.name}: filename does not start with id '{adr_id}'")
if adr_id in adrs:
errors.append(f"{path.name}: duplicate id '{adr_id}'")
status = meta.get("status", "").lower()
if status and status not in STATUSES:
errors.append(f"{path.name}: invalid status '{status}'")
if meta.get("date"):
try:
if not DATE_RE.fullmatch(meta["date"]):
raise ValueError("not YYYY-MM-DD")
date.fromisoformat(meta["date"])
except ValueError:
errors.append(f"{path.name}: date '{meta['date']}' is not YYYY-MM-DD")
if status == "superseded" and not meta.get("superseded-by"):
errors.append(f"{path.name}: status superseded needs 'superseded-by'")
if adr_id:
adrs[adr_id] = (path.name, meta)
nodes = sorted(adrs)
edges = []
for adr_id, (name, meta) in sorted(adrs.items()):
for kind in ("supersedes", "depends-on"):
for target in split_ids(meta.get(kind, "")):
if target not in adrs:
errors.append(f"{name}: {kind} points at unknown {target}")
else:
edges.append({"from": adr_id, "to": target, "type": kind})
for adr_id, (name, meta) in sorted(adrs.items()):
for target in split_ids(meta.get("supersedes", "")):
if target in adrs and adrs[target][1].get("status", "").lower() != "superseded":
errors.append(f"{name}: supersedes {target}, but {target} is not marked superseded")
for target in split_ids(meta.get("superseded-by", "")):
if target not in adrs:
errors.append(f"{name}: superseded-by points at unknown {target}")
return errors, {"nodes": nodes, "edges": edges}
def make_demo(root):
docs = {
"ADR-0001-cache-aside.md": "---\nid: ADR-0001\ntitle: Use cache-aside for sessions\nstatus: superseded\ndate: 2024-03-02\nsuperseded-by: ADR-0003\n---\nBody\n",
"ADR-0002-queue.md": "---\nid: ADR-0002\ntitle: Use a managed queue\nstatus: accepted\ndate: 2024-05-10\n---\nBody\n",
"ADR-0003-session-store.md": "---\nid: ADR-0003\ntitle: Move sessions to a write-through store\nstatus: accepted\ndate: 2025-09-01\nsupersedes: ADR-0001\ndepends-on: ADR-0002\n---\nBody\n",
"ADR-0004-bad.md": "---\nid: ADR-0004\ntitle: Adopt gRPC\nstatus: approved\ndate: 2025-13-40\n---\nBody\n",
"ADR-0005-missing.md": "---\nid: ADR-0005\nstatus: proposed\n---\nBody\n",
}
for name, body in docs.items():
(Path(root) / name).write_text(body, encoding="utf-8")
if __name__ == "__main__":
if len(sys.argv) > 1:
root = sys.argv[1]
errors, graph = lint(root)
for e in errors:
print("ERROR", e)
Path("adr-graph.json").write_text(json.dumps(graph, indent=2))
sys.exit(1 if errors else 0)
with tempfile.TemporaryDirectory() as tmp:
make_demo(tmp)
errors, graph = lint(tmp)
for e in errors:
print("ERROR", e)
print(json.dumps(graph))
print("exit code would be", 1 if errors else 0)
Output (running it exactly as above with no arguments builds five sample files in a temporary folder and lints them)
ERROR ADR-0004-bad.md: invalid status 'approved'
ERROR ADR-0004-bad.md: date '2025-13-40' is not YYYY-MM-DD
ERROR ADR-0005-missing.md: missing required field 'title'
ERROR ADR-0005-missing.md: missing required field 'date'
{"nodes": ["ADR-0001", "ADR-0002", "ADR-0003", "ADR-0004", "ADR-0005"], "edges": [{"from": "ADR-0003", "to": "ADR-0001", "type": "supersedes"}, {"from": "ADR-0003", "to": "ADR-0002", "type": "depends-on"}]}
exit code would be 1
In CI you would run python adr_lint.py docs/adr, which prints the same kind of errors, writes adr-graph.json and returns a failing exit code.
Walk-through of the two cross-file loops
The per-file checks cannot see other files, so after every file has been read into the adrs dictionary (id mapped to file name and metadata), two more passes run over the whole folder:
- Edge loop. For each ADR, read its
supersedesanddepends-onvalues (comma-separated ids). If an id is not a key inadrs, report "points at unknown". Otherwise record an edge for the graph. In the demo, ADR-0003 listssupersedes: ADR-0001anddepends-on: ADR-0002, both exist, so two edges are recorded. - Consistency loop. For each ADR that supersedes another, look up the target and confirm the target's own status is
superseded. Also confirm anysuperseded-byid exists. In the demo ADR-0001 is marked superseded, so ADR-0003 passes. If ADR-0001 still said accepted, the output would gainADR-0003-session-store.md: supersedes ADR-0001, but ADR-0001 is not marked superseded.
These need their own passes because a file can point at an ADR that has not been read yet; checking only after everything is loaded avoids order-dependent false errors.
Key points
- Collect all errors: a contributor fixes everything in one pass instead of one error per CI run.
- Required fields and valid status: the two checks the question asks for. Status comes from a fixed set, so a typo like
approvedfails. - Cross-file checks: unknown link targets fail, and if ADR-0003 supersedes ADR-0001 then ADR-0001 must be marked superseded. This catches half-finished supersession.
- Graph output: nodes are ADR ids and edges are
supersedesordepends-on. Feed it to a diagram tool or use it to spot an accepted ADR depending on a rejected one. - Only ADR-named files are read, so a README or template in the same folder is ignored.
Complexity
Time is O(total bytes read) (Big-O notation says how work grows with input size: here, twice as much ADR text takes about twice as long): each file is read once and each metadata line handled once. Cross-reference checks use a dictionary, so they are O(number of links). Memory is O(number of ADRs) for metadata (file text is not kept).
Edge cases
- Missing or unclosed
---block: reported, not crashed on. - Duplicate ids, id not matching the file name, wrong id format.
- Impossible dates such as month 13 are caught by
date.fromisoformat, but on Python 3.11 and later that function also accepts20250901and2025-W36-1, so the script first requires the exactYYYY-MM-DDshape with a regular expression and then letsfromisoformatreject impossible calendar dates. Both checks are needed: the regex alone would accept 2025-13-40. - A title containing a colon (for example
Use cache: aside) parses correctly, because each line is split only at the first colon; quoted values would keep their quote characters, which is one reason to move to a YAML library. - Windows line endings:
splitlineshandles them. - Limits of the mini parser: no nested values or multi-line values. If ADRs need those, switch to a real YAML library. That is a trade-off between zero dependencies and flexibility.
What else I would lint
Core (most useful first): required sections present (Context, Decision, Consequences); relative links inside the ADR resolve; the title matches the top heading; filenames follow ADR-NNNN-slug.md with no numbering gaps or duplicates; proposed ADRs older than a set number of days are flagged as stale.
Rarely needed (mention, do not lead with): cycles in the supersedes chain, valid status transitions (for example accepted cannot go back to proposed), and a review-by date in the past.
Keep new lint rules in warn mode first, and never fail a whole pull request on a rule that a human would reasonably override.
How would you show latency budgets, throughput, error rates, contracts and owners directly on service diagrams, and what conventions would you standardise so the annotations stay readable?
Sample Answer
Direct answer
I annotate the diagram with a small, fixed set of fields on a fixed place: owner as a badge on each box, contract as a link ID on each arrow, and the numbers (latency budget, throughput, error rate) as a short one-line label per arrow. Anything beyond that goes in a table under the diagram, and the numbers are generated from monitoring and contract files so they carry an "as of" date and do not rot.
The conventions I would standardise
- Node (a service box): name from the service catalog, plus a badge with the owning team, for example
pricing-api | team: pricing. - Edge (an arrow): one line, always in the same order:
contract | p95 budget | rate | error-rate target, for examplepricing-v2 | p95 150 ms | 300 rps | err < 0.1%. (p95 means the 95th percentile: 95 percent of requests are faster than this. rps means requests per second. The last field is the highest acceptable share of failed requests, 0.1 percent here. It is a per-request target, not the SRE "error budget", which is the total failure allowance over a period, derived from a reliability target.) - Colour is reserved for one meaning, health against target (green, amber, red), never for teams.
- A legend on every diagram, and an "as of" date.
- Limit: about 12 boxes and one annotation line per arrow; more goes into a table or a separate diagram.
- Contract means a link to the API specification (for example an OpenAPI document, a machine-readable description of an HTTP API's endpoints and message shapes) and its version, not prose.
Component-diagram tutorial box
pricing-api (owner: pricing team, contract: pricing-v2 OpenAPI spec, p95 budget 150 ms). Someone reading the box knows what it is, whom to page, which interface to read, and how slow it is allowed to be.
Latency budgets on the critical path
The critical path is the longest chain of calls the user waits for. Steps that run in parallel count as the slowest of them. Illustrative checkout with an end-to-end p95 target of 800 ms:
| Step | Budget (ms) |
|---|---|
| Gateway | 50 |
| checkout-service own work | 100 |
| pricing (150) in parallel with inventory (120) | 150 (the larger of the two) |
| payments | 300 |
| Total on the critical path | 600 |
50 + 100 + 150 + 300 = 600 ms, leaving 200 ms (25 percent) of headroom under 800 ms. Adding percentiles is an allocation rule, not exact statistics, and it is conservative only for calls in series. For the parallel pair the larger budget slightly understates the risk: if pricing and inventory were independent and each met its own p95, both would be within budget only about 90 percent of the time (0.95 x 0.95 = 0.9025), and reaching 95 percent for the pair would need each side at about its 97.5th percentile (0.95 ** 0.5 = 0.9747). That is one more reason to keep the 200 ms of headroom. The p95 of a chain is not the sum of each step's p95: a request is rarely unlucky in every step at once, so the true combined p95 is usually lower than the sum, but correlated slowness (for example a busy database slowing several steps together) can break that. So I treat the sum as a budget to allocate and keep headroom instead of budgeting to the last millisecond.
Instrumentation placement
Put a trace span (a timed record of one operation, from distributed tracing) at each service entry and around each outbound client call, on the critical path only. Then the diagram's numbers come from the measured span durations, and the difference between a caller's client span and the callee's server span reveals network or queueing time. For example, if checkout's client span around the call to payments lasts 340 ms but payments' own server span lasts 290 ms, the other 50 ms went to the network, connection setup or waiting in a queue before payments started work, and that 50 ms is invisible if only one side is measured.
Trade-offs and pitfalls
- Hand-typed numbers go stale; generate annotations from the monitoring configuration and contracts, or show them with a date.
- Annotating every arrow makes the diagram unreadable: annotate the critical path and the arrows that have a service-level objective, and leave the rest. (A service-level objective, SLO, is the reliability target a team commits to, such as 99.9 percent of requests succeeding.)
- Throughput and error rate are observed values, while budgets are targets. Label which is which, otherwise readers cannot tell "allowed" from "actual".
How would you represent, and keep current, the dependency map between many teams' APIs so that single points of failure are visible? What signals tell you the map has drifted from reality?
Sample Answer
Direct answer
I keep two maps and compare them. The declared map is what teams say they depend on, written in each service's descriptor file and rendered as a graph. The observed map is what traffic really does, built from distributed traces (records that follow one request across services) or from logs of the API gateway (the front door that receives external calls) and of the service mesh (infrastructure that proxies every service-to-service call and records it). Single points of failure are visible on the declared graph by fan-in (how many services depend on a node) and by whether the dependency is hard (no fallback), and drift is the difference between the two maps.
Representing it
A descriptor file is a small file kept in each service's repository that describes the service for tools (name, owner, dependencies). A sketch (illustrative field names):
name: checkout
owner: checkout-team
dependsOn:
- service: payments
contract: payments-api@v2
criticality: hard
- service: recommendations
contract: reco-api@v1
criticality: soft
- service: order-events
contract: OrderCreated@v3
criticality: async
A script reads every such file and turns each dependsOn entry into one directed edge (checkout -> payments, and so on), tagged hard, soft or async. In more detail:
- Each service declares
dependsOnentries with the callee's name, the API contract version, and a criticality flag: hard (requests fail without it), soft (degrades gracefully) or async (event-based, buffered). - A script builds a graph from all descriptors, in the catalog or as generated diagrams per domain.
- SPOF (single point of failure) flags: a node with fan-in of 2 or more, a hard dependency, and no redundancy or fallback recorded.
- Ownership is attached to each node, so a risky node always has a team behind it.
Keeping it current: declared versus observed
Illustrative example:
| Edges | |
|---|---|
| Declared | A to B, A to C, B to D, C to D, C to E |
| Observed in the last 30 days | A to B, B to D, C to D, C to E, A to E, B to F |
Comparing the two sets:
- Observed but undeclared: A to E and B to F. These are the dangerous ones, since nobody documented them and nobody is planning around them.
- Declared but never seen: A to C. Either the dependency is dead code, or it is a rarely used path such as a failover, which the team should confirm.
- Fan-in: D is called by B and C (fan-in 2), and E by C and A (fan-in 2, counting the undeclared edge). If D is a hard dependency with no fallback, it is the first SPOF to review, and the undeclared A to E edge shows E was underrated.
What to fix first. Fan-in of 2 or more only makes a node a candidate, because almost every shared service has that. Rank the candidates: first those where the dependency is hard and there is no fallback, then by how much traffic or revenue passes through, then by how weak the node is (single instance, unowned, poor reliability history). In the example, D (hard, no fallback, called by B and C) goes first; E is next only if the A to E path turns out to be hard.
Signals that the map has drifted
- Undeclared observed edges, and declared edges with zero traffic for 30 days (the set difference above).
- Services in traces or in the deploy system that have no descriptor, or descriptors whose service no longer deploys.
- Descriptors whose last-modified or last-verified date is older than a threshold such as 90 days while the service's code changed.
- A new owner or on-call team that no descriptor mentions.
- An incident where the postmortem (written review after an outage) says "we did not know X depended on Y".
- Contract version mismatch: the caller declares
v1while traffic showsv2.
Process to keep it honest
A weekly job diffs declared against observed and opens a ticket on the owning team for each mismatch. Declarations are updated in the same pull request as a new integration, and a CI (continuous integration, the automated build) check refuses names that do not resolve to a descriptor.
Trade-offs and pitfalls
- Traces are sampled (only a fraction of requests are recorded as traces to save cost), so a rare path may be missing from the observed map; use the mesh or gateway logs for completeness, and treat "not seen" as a question, not proof.
- The observed map cannot say whether an edge is hard or soft; that judgement stays with the owning team, and chaos or failure-injection tests (deliberately switching a dependency off in a controlled way to see whether callers survive) can confirm it.
- If the tool requires heavy manual entry, teams will skip it. Generate as much as possible from traces and treat declarations as the reviewed layer.
Unlock Full Question Bank
Get access to all 38 Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.