Architecture Documentation and Communication Questions
Making an architecture legible to others: architecture decision records (templates, alternatives and consequences, lifecycle, supersession, tagging and discoverability), diagramming and visualization (C4, sequence, deployment, data-flow and trust-boundary diagrams, notation conventions), and communicating designs to technical and non-technical stakeholders. Covers capturing rationale, documenting what a system does under failure, load and consistency trade-offs, recording observability, SLO and security content, running design reviews, keeping docs and diagrams current, and presenting a system clearly under time pressure. The communication skill that separates a good design from an understood one.
An auditor asks you to show, from your architecture documentation, where personal data is stored and how it flows, who can access it, and how predictions or transactions can be traced. What artifacts and diagram elements do you provide?
Sample Answer
Direct answer
Give the auditor an evidence pack that answers each question with a specific artifact: a register of personal data (where it is stored), a data-flow diagram (how it moves), an access matrix (who can reach it), audit logs with correlation ids (a unique id stamped on every log line of one request, so the lines can be joined into one story) that show what happened and when, and lineage records (a record of where a result came from: which model version and data produced it) tying a prediction or transaction back to the exact model version and data. Show the boundaries that matter (card data, and residency regions, meaning the places where rules require personal data to stay) on the diagrams, and show where personal data is redacted. Core for every audit: the register, the data-flow diagram, the access matrix and the audit log. Add the card-data boundary only if you handle card payments, and the lineage record only if models make decisions.
Evidence pack: question to artifact
| Auditor asks | Artifact | How it is kept current |
|---|---|---|
| Where is personal data stored? | Data inventory, also called a register of processing activities (a list of what personal data you handle, why, where and for how long; GDPR, the EU privacy law, Article 30 requires one from many organisations): each dataset, its fields, purpose, retention, region, owner | Reviewed on a schedule and on any new datastore |
| How does it flow? | Data-flow diagram with each store, each arrow labelled with what data and how it is protected, classification colours | Generated from or reviewed with architecture diagrams |
| Who can access it? | Access matrix: role by datastore by permission, how access is granted, review cadence, emergency access process | Exported from the identity system, reviewed quarterly |
| How can a transaction or prediction be traced? | Audit log design plus lineage record (below) | Logging is a system property, tested |
Diagram elements to include
- Data stores with classification labels (personal, payment, public).
- Trust boundaries and the residency regions:
flowchart LR
C[Customer] --> R[Edge router picks region by customer home region]
R --> A1
R --> A2
subgraph EU[EU region]
A1[App services] --> D1[(EU customer database)]
A1 --> L1[(EU audit log)]
end
subgraph US[US region]
A2[App services] --> D2[(US customer database)]
A2 --> L2[(US audit log)]
end
D1 -.->|aggregated non-personal metrics only| W[Global analytics warehouse]
D2 -.->|aggregated non-personal metrics only| W
Also mark where backups live and that they stay in the same region.
- Payment data: the cardholder data environment (the systems that store, process or transmit card data under PCI DSS, the payment card industry security standard) drawn as its own boundary. Mark where the PAN (primary account number, the full card number) is tokenised (swapped for a random stand-in value that is useless without a separate, tightly guarded lookup), and show that logs receive only the token.
- Redaction points: mark where personal identifiers (PII, personally identifiable information) are masked or dropped before reaching logs, analytics or support tools.
- Audit trail: an immutable log store (append-only, restricted write and read) recording who accessed which record, when and why, including access to card and personal data.
Traceability for predictions (ML)
Every prediction writes a lineage record so an auditor can walk backwards from one decision:
{
"prediction_id": "pred-88213",
"transaction_id": "T-1001",
"correlation_id": "req-5f3a91",
"timestamp": "2026-09-28T09:15:02Z",
"model_name": "fraud-scorer",
"model_version": 14,
"feature_set_version": 6,
"training_data_snapshot": "2026-08-01",
"input_hash": "sha256 of the feature values",
"score": 0.83,
"decision": "manual_review",
"region": "EU"
}
Read it like this: transaction T-1001 got score 0.83 from fraud-scorer version 14, which used feature set 6 (the versioned list of input fields and how each is calculated) and was trained on the 2026-08-01 snapshot (a frozen copy of the training data as of that date). The model registry (the catalogue of trained model versions) then shows who approved version 14 and its evaluation, and the audit log shows who viewed the case. Store the input hash (a fixed-length fingerprint of the inputs made with the SHA-256 algorithm: the same inputs always give the same fingerprint, so you can show what was scored without keeping the raw values) rather than raw personal values when the raw data is not needed, and keep the record for the retention period.
Worked walk-through for the auditor: pick transaction T-1001; find its dataset in the register (region, retention, owner); find the audit log entry by correlation id (req-5f3a91, which the lineage record also carries, so the two join); open the lineage record; open the registry entry for version 14; show the access matrix row for who could see the case. Five artifacts, one trace.
Pitfalls: a diagram that shows the design but not the deployed system; logging raw card numbers or emails "for debugging"; residency drawn for primary databases but forgotten for backups and caches; lineage that records the model name but not the version.
What security and authentication content must an architecture document for a public API contain, and how do you present it to a security reviewer versus a developer?
Sample Answer
Direct answer
The document must state, for a public API, who can call it, how callers prove who they are, what they are allowed to do, how data is protected in transit and at rest, and how abuse is limited and detected. A security reviewer wants threats, controls and evidence; a developer wants how to authenticate correctly and what errors to expect. I would write one source of truth and give each reader a different entry point into it.
Required content
- Authentication: the mechanism (for example OAuth 2.0, the standard for delegated access, with short-lived access tokens (an access token is a temporary string the caller sends with each request to prove it was authorised), or API keys for server-to-server use), where tokens are issued, lifetime, refresh and revocation.
- Authorization: the model (scopes are named permissions carried on the token, such as
orders:read; roles group permissions by job) and where it is enforced (gateway versus service), including object-level checks (a caller must own the record it requests: if user A asks for/orders/1002and that order belongs to user B, the API must refuse, even though A's token is valid; broken object-level authorization is the top risk in the OWASP API Security list, OWASP being an open security-community project). - Trust boundaries and data flow: what is exposed publicly, what sits behind the gateway, and where TLS ends (TLS is the encryption that protects traffic on the network; "where it ends" is the point where traffic is decrypted and must be trusted from then on).
- Data protection: encryption in transit and at rest, classification of fields, what is never logged.
- Abuse controls: rate limits, quotas, input validation, size limits, and how bad clients are blocked.
- Threat model: a threat model is a structured list of what could go wrong and who might attack, each threat paired with the control that answers it. Record the residual risk (what remains after the controls) and the named person who accepts it.
- Secrets and key management, audit logging, incident contacts.
Presenting to each reader
| Security reviewer | Developer | |
|---|---|---|
| Opens with | Threat table and trust-boundary diagram | "How to get a token" with a copyable request |
| Format | Threat, control, evidence (config link, test, scan result), residual risk | Step-by-step, examples, error codes, scope table |
| Cares about | Gaps, assumptions, compliance mapping (which control satisfies which regulation or standard requirement) | Time to a working, correct call |
Worked example (illustrative fragment)
Security reviewer fragment, threat table rows:
| Threat | Control | Evidence | Residual risk |
|---|---|---|---|
| Stolen access token replayed | 15-minute lifetime; sender-bound refresh tokens (a refresh token that only the client it was issued to can use, so a copy is useless); revocation list at the gateway (token ids refused even before they expire) | Gateway policy file, revocation test | Up to 15 minutes of misuse, accepted by the security owner |
| User A reads user B's order by changing the id in the URL | Service checks the order's owner equals the caller on every read | Automated test that requests another user's order and expects a refusal | None known; retested each release |
Developer fragment, the "how to get a token" page (client-credentials grant means a server application exchanges its own id and secret for a token, with no human user involved; a bearer token is one that works for whoever holds it, so it must be sent only over TLS):
POST /oauth/token
grant_type=client_credentials&client_id=<your-id>&client_secret=<your-secret>&scope=orders:read
-> {"access_token": "<token>", "expires_in": 900}
GET /orders/1002
Authorization: Bearer <token>
Errors: 401 means the API does not know who you are (no token, malformed or expired), so fetch a new token. 403 means it knows who you are but you may not do this (the token lacks the scope, or the order is not yours), so requesting a new token will not help. That difference is authentication (proving who you are) versus authorization (what you are allowed to do).
Pitfalls: describing the design without evidence a reviewer can check; documenting the mechanism but not the failure responses; putting secrets or real tokens in examples; letting the doc and the gateway configuration diverge. Generate the developer reference from the API's OpenAPI description (a machine-readable file listing every endpoint, parameter and response) where you can, so it cannot drift.
How do you show network and trust boundaries on architecture diagrams for a multi-tenant platform: tenant isolation, ingress and egress control, where secrets live, and how encryption in transit is applied end to end?
Sample Answer
Direct answer
Draw the platform as trust zones (areas with different levels of trust, each with its own controls) and label every arrow that crosses a boundary with what protects it: protocol, encryption, how the caller is authenticated, and where the tenant's identity comes from. Show tenant isolation as a property of each data store, ingress and egress as controlled doors, secrets as living in one management zone, and mark each place TLS (transport layer security, the encryption of data moving over a network) ends, because plaintext exists there. One diagram per concern (network, identity, data) beats one crowded picture.
Terms: a tenant is one customer organisation using the shared platform (Acme and Globex are two tenants); multi-tenant means their data and traffic share the same system but must never mix. An API gateway is the front-door service that authenticates and routes every API call. A private subnet is a network segment with no public internet address. Sidecar proxies are small helper proxies running beside each service that handle encryption and identity for its network calls. Workload identity means a running service proves who it is with short-lived credentials the platform issues to it, not a stored password. An egress proxy allowlist is a gate for outbound traffic that lets only named destinations through; default deny means anything not explicitly allowed is blocked. Inference is running a trained model to get a prediction, and the model server is the service that does it.
Diagram: zones and numbered flows
flowchart LR
U[Tenant user browser] -->|F1 TLS| E
subgraph Z1[Edge zone, public]
E[WAF and load balancer]
end
E -->|F2 TLS re-encrypted| G
subgraph Z2[App zone, private subnet]
G[API gateway verifies token, sets tenant_id]
S[Tenant services with sidecar proxies]
M[Model server for ML scoring]
end
G -->|F3 mTLS, tenant_id from verified token| S
S -->|F4 mTLS| M
subgraph Z3[Data zone, no internet route]
D[(Shared database, row-level security by tenant_id)]
B[(Object storage, one prefix per tenant)]
end
S -->|F5 TLS, tenant-scoped credentials| D
S -->|F5 TLS| B
subgraph Z4[Management zone]
K[Secrets manager and key service]
end
S -.->|F6 workload identity, TLS| K
S -->|F7 TLS via allowlist| X[Egress proxy to external APIs]
WAF is a web application firewall. mTLS is mutual TLS: both sides present certificates, so each proves who it is.
Flow table (the legend a reviewer reads)
| Flow | From to | Encryption and identity | Tenant context |
|---|---|---|---|
| F1 | Browser to edge | TLS; user authenticates with a token | Tenant chosen at login |
| F2 | Edge to gateway | TLS re-established so traffic is not plaintext inside the network | Passed through |
| F3 | Gateway to services | mTLS, service identity via certificates | Taken from the verified token, never from a client-supplied header |
| F4 | Services to model server | mTLS; inference endpoint not reachable from outside | Tenant id in the request for logging and model selection |
| F5 | Services to data | TLS; credentials scoped to the tenant | Enforced by the store |
| F6 | Services to secrets | Short-lived workload identity, TLS | n/a |
| F7 | Services to internet | TLS through an allowlisted proxy | Per-tenant egress rules where required |
Traced request (tenant Acme cannot read tenant Globex)
Illustrative values: Alice works for tenant acme. Invoice 42 belongs to tenant globex.
- F1, F2: Alice's browser sends
GET /invoices/42with her login token over TLS. - Gateway: verifies the token's signature and reads
tenant_id=acmefrom inside it. A headerX-Tenant: globexthat Alice adds is ignored. - F3, F5: the invoice service opens a database session tagged
acme(SET app.tenant_id = 'acme'), connecting as a role that does not own the table (owners bypass policies). - The database applies a row-level security policy to every query:
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON invoices
USING (tenant_id = current_setting('app.tenant_id'));
Invoice 42 has tenant_id = 'globex', which does not equal acme, so the query returns no rows and the service answers 404, even if the service code forgot its own WHERE tenant_id = .... Object storage works the same way: the credentials only allow the prefix tenants/acme/.
Tenant isolation
Label each store with its model: pooled (shared, separated by a tenant column with row-level security, meaning the database filters rows by tenant regardless of the query), bridge (separate schemas), or silo (dedicated database or account). Pooled is cheaper but a bug or missing filter can leak across tenants, so the diagram notes the control that stops it. Regulated or large tenants often get silos, and the diagram marks which.
Ingress, egress and secrets
- Ingress: one public door (edge zone). Everything behind it has no public address. Show the rate limit and authentication step at the gateway.
- Egress: default deny, with all outbound traffic through a proxy that allows named destinations. Show it explicitly, since unmapped egress is how data leaves unnoticed.
- Secrets: stored in the secrets manager and key service, fetched at runtime with workload identity, never baked into images or config files. Per-tenant encryption keys (each tenant's data is encrypted under its own key, so one key compromise or one tenant's deletion request affects only that tenant) sit in the key service (envelope encryption: a data key protected by a master key).
ML-serving variant
Give the model server its own segment with no internet route. Apply TLS on the inference call (F4), and use network rules so only the gateway-facing services can reach it. Model artifacts are pulled read-only from storage. If tenants have their own models, show the model-to-tenant mapping and the tenant check on the inference call.
What reviewers look for: an unlabelled arrow is a finding; encryption stops somewhere, so show where; tenant identity must come from something the platform verified.
Pitfalls: one giant diagram; drawing the intended design instead of the deployed one; forgetting backups, logs and analytics, which also hold tenant data.
Draw the data flow for a payment or signup flow that crosses a frontend, a gateway, several services, a database and an external provider. Which elements and data labels do you include, and where do you mark trust boundaries?
Sample Answer
Direct answer
A data flow diagram (DFD) follows data, not services: what information moves, between which elements, and where it crosses from a less trusted zone to a more trusted one. For a signup or payment flow I draw the browser, gateway, services, database and external provider; label every arrow with the data it carries and its sensitivity; and draw dashed trust-boundary boxes around zones (internet, edge, internal network, data tier, third party). A trust boundary is any line where the level of trust changes, so it is where authentication, validation and encryption must be checked.
Elements and labels to include
flowchart LR
subgraph Z1[Zone 1: open internet]
B[Browser]
end
subgraph Z2[Zone 2: edge]
GW[API gateway]
end
subgraph Z3[Zone 3: internal services]
S[Signup service]
end
subgraph Z4[Zone 4: data tier]
DB[(User database)]
end
subgraph Z5[Zone 5: third parties]
P[Payment provider]
E[Email provider]
A[Analytics]
end
B -->|"email, password (TLS)"| GW
B -->|"card details, direct"| P
P -->|"card token"| B
GW -->|"email, password (internal call)"| S
S -->|"PII: email, name"| DB
S -->|"verification link"| E
S -->|"pseudonymous events, no email or name"| A
P -.->|"signed webhook: paid"| GW
Each dashed-box idea is drawn as a zone above (a subgraph). Any arrow that leaves one zone box and enters another crosses a trust boundary: internet to edge is boundary 1, edge to services is 2, services to data tier is 3, and services or edge to third parties is 4. (Data tier means the layer that stores data, such as databases.)
Trust boundaries I mark: (1) browser to gateway (the open internet, so validate everything and use TLS, transport encryption); (2) gateway to services (authenticated internal calls); (3) services to database (data tier, least privilege, meaning each component gets only the permissions it needs, so the signup service can insert and read users but cannot drop tables); (4) services to provider and analytics (leaving our control; outbound data must be minimised, inbound webhooks verified by signature). A webhook is an HTTP call the provider makes to our system to tell us something happened; "signed" means it carries a cryptographic signature made with a secret only we and the provider share, so we can reject forged calls.
Data labels that earn their place
Name the fields and classify them: PII (personally identifiable information: email, name), secret (password hash), payment (only a token, never the raw card number, PAN). Sending card entry straight to the provider means the PAN never crosses our boundary, which keeps our PCI DSS (card-industry security standard) scope small. Also note retention (how long each store keeps the data) and who may read each store.
Worked example: signup variant, step by step
Suppose alice@example.com signs up at 09:15:00.
- 09:15:00, boundary 1. The browser sends email
alice@example.comand her password over TLS to the gateway. The gateway checks the format, rate limit and TLS. - Boundary 2. The gateway forwards the request to the signup service with an internal service token, so only trusted callers reach it.
- Boundary 3. The signup service hashes the password (a one-way scramble, so the plain password is never stored) and writes one row: id 4821, email, name, password hash, status "unverified". The database account it uses may only insert and read the users table (least privilege).
- Boundary 4, outbound. It asks the email provider to send a verification link with a one-time token valid for 24 hours. The provider receives her email address and the link, never her password.
- Boundary 4, outbound. It emits an analytics event
{event: signup, user_id: 4821}with no email or name, because analytics does not need personal fields. Note this is pseudonymous, not anonymous: user_id 4821 can be joined back to the users table by anyone holding both, so label the flow as pseudonymous personal data and apply retention to it; truly anonymous events would carry no stable user identifier. - 24 hours later. A background job finds account 4821 still "unverified" and expires it, which enforces the retention rule.
Optional extension for machine-learning systems: real-time inference variant. Same discipline, and only needed if your system serves model predictions. A request arrives with features (model inputs, such as user_id 4821 and a country), the feature store (a database of precomputed model inputs) is read, an online model (a model kept loaded in memory so predictions are fast) returns a score, and the prediction is logged. For a training pipeline: event ingest, feature store, nightly training, then serving. Label each arrow with what data it carries, mark that raw user events are more sensitive than aggregated features, and mark boundaries where training data leaves production.
What each element should link to
Each box links to its owner, schema or data dictionary (a catalogue of each field, its meaning and its classification), runbook, and data classification; each arrow to the contract (API spec or event schema).
Trade-offs and pitfalls
- Do not draw every service call; draw only flows that carry sensitive or regulated data.
- A boundary drawn without saying what is enforced there (authentication, schema validation, encryption) is decoration.
During a sales cycle a client insists on a proprietary protocol they already use. As the architect, how would you record the requirement, evaluate alternatives, capture the client's business reasoning and communicate the trade-offs?
Sample Answer
Direct answer
Treat "we must use our protocol" as a stated solution, not yet a requirement. Record the underlying need and the constraint with its source, evaluate alternatives in a short decision record (an ADR, architecture decision record: a dated note capturing one decision, its options and reasons), write down the client's reasoning in their words (separated from what you have verified), and get the trade-off accepted in writing. My default recommendation: support the protocol through an adapter at the edge that translates to our standard API, with a review trigger. A proprietary protocol is one whose specification is owned and controlled by a single vendor.
1. Record the requirement (illustrative excerpt)
REQ-014 Existing client systems must connect using ProtocolX (proprietary binary protocol over TCP)
Source: client technology lead, discovery call 1
Client's stated reason: about 300 field devices and one partner's software already speak it;
replacing them means re-certification (unverified claim)
Underlying need: keep current devices working without replacement
Type: client constraint (not yet validated)
Evidence requested: protocol specification, device list, the certification rule
Owner: solutions architect
2. Questions to turn the ask into a need
- Why this protocol: devices, a partner, a certification rule, or habit?
- Can they share the specification? Is it authenticated and encrypted?
- Who else depends on it? Is there any plan to retire it?
3. Alternatives and criteria
| Option | Time | Security risk | Long-term cost | Lock-in |
|---|---|---|---|---|
| A. Native support inside the core platform | Slow | Highest: unreviewed parser in the core | High: every team carries it | High |
| B. Adapter at the edge translating to our standard API (an anti-corruption layer, a translation layer that keeps a foreign model out of the core) | Medium | Contained: parser isolated, rate-limited | Medium: one component | Low |
| C. Client migrates to the standard, with a support window | Client-dependent | Lowest | Lowest | None |
| D. Decline | Immediate | None | None | None |
Option D is right when the protocol has no specification we can review, the client will not fund support, or the security risk cannot be contained. Then say no early and offer C.
Terms: lock-in is being tied to one component or vendor so leaving is costly. Re-certification is repeating a formal approval process after a change. The edge is the outer boundary where outside traffic first enters our systems.
4. Capture the client's reasoning
Put it in the decision record's context section, quoted and attributed, with a column "claimed" versus "verified". If the certification claim is false, the case for B weakens and C gains ground, and the record shows why.
A filled-in record for the choice (illustrative):
ADR-007: Support ProtocolX through an edge adapter
Status: Proposed Date: 2026-09-10 Deciders: solutions architect, security architect
Context: Client needs about 300 existing devices to connect. Certification
claim: unverified. Specification: requested, not yet received.
Options: A native, B adapter, C client migrates, D decline (see table)
Decision: B, subject to receiving the specification.
Consequences: one extra component to run; parser isolated from the core.
Revisit when: a second client needs ProtocolX, or the client can migrate.
Claimed vs verified: certification rule (claimed by client), device count (claimed).
5. Communicate the trade-off
- A one-page options memo for the client's business owner in their terms: time, cost, risk, and what they must supply (the specification, test devices). Illustrative opening: "We recommend Option B. It keeps your 300 devices working with no replacement. It adds about one component to run and needs your protocol specification within 30 days. If the specification is late, go-live moves by the same amount."
- A security appendix (for the client's technical team): fuzz-test the parser (feed it large volumes of malformed or random input to find crashes), terminate the protocol at the gateway (end the foreign protocol at the edge so only our standard API travels inward), wrap the traffic in an encrypted tunnel (an encrypted channel between two endpoints), and rate-limit (cap how many requests a sender may make per minute).
- Ask for a decision and a signature on the chosen option, including who pays for ongoing support.
Recommendation and what flips it
Choose B, with a review trigger: revisit if a second client needs the protocol (then native support may pay back) or if the client can migrate within an agreed window (then C, with B as the bridge). Flip to A only if the protocol is an industry norm across many prospects.
Pitfalls
- Arguing the protocol is bad. It is their asset; you are pricing the consequence.
- Agreeing in the sales meeting. An unrecorded yes becomes a support commitment nobody priced.
- Skipping the security review because the client is a strategic account.
Unlock Full Question Bank
Get access to all 10 Architecture Documentation and Communication interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.