Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
Your infrastructure is ephemeral, for example disposable Kubernetes clusters. How do you keep runbooks and recovery procedures accurate? Discuss templates, capturing real cluster state as examples, anchoring the docs to infrastructure code, and how you test the steps.
Sample Answer
Direct answer
When clusters are disposable, a hand-written runbook is wrong almost immediately, so stop treating the doc as prose and treat it as a build artifact. Generate the environment-specific parts from the infrastructure code (IaC, meaning tools like Terraform or Helm, which define infrastructure in files), keep the doc in the same repository and version as that code, show real captured output rather than invented examples, and run the procedure against a fresh cluster on a schedule. If the run fails, the doc is stale, and you find out from a red build in CI (continuous integration, the automation that runs checks on every change) instead of a bad incident.
Terms in plain words
- Kubernetes cluster: a group of machines that runs containerized applications. A namespace is a named partition inside it (here, everything belonging to
payments). A pod is one running copy of an application, and it is Ready when it has started and passes its health check. - kubectl: the command-line tool for talking to a cluster. A context is a saved "which cluster am I talking to" setting, so
use-contextswitches clusters. - Velero: a Kubernetes backup and restore tool.
velero backup getlists backups andvelero restore create --from-backup Xrecreates the saved resources. - Terraform: a tool that builds infrastructure from code, and its outputs are named values (cluster name, bucket) it prints after building. A module is a reusable bundle of that code. Helm packages Kubernetes applications the same way.
- Tribal knowledge: things everyone on a team "just knows" but nobody wrote down.
1. Templates: separate the stable steps from the values that change
The sequence of steps (check backup, restore namespace, verify pods) changes rarely. Cluster names, regions and bucket names change with every cluster. Write the steps once as a template with placeholders, and fill the values in from IaC outputs at publish time. A missing value fails the build.
import json, re
from string import Template
# Shape of `terraform output -json`: each output is {"value": ..., "type": ...}
tf_outputs = json.loads("""{
"cluster_name": {"value": "prod-eu-3", "type": "string"},
"region": {"value": "eu-west-1", "type": "string"},
"backup_bucket": {"value": "s3://acme-velero-eu-3", "type": "string"}
}""")
values = {k: v["value"] for k, v in tf_outputs.items()}
runbook = Template("""## Restore namespace `payments` on $cluster_name ($region)
1. `kubectl config use-context $cluster_name`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in $backup_bucket. Owner: $owning_team
""")
def render(t, vals):
try:
return t.substitute(vals)
except KeyError as e:
# CI fails here: the doc references a value the infrastructure code no longer outputs
return f"RENDER FAILED: template needs {e} but Terraform does not output it"
print(render(runbook, values))
print(render(runbook, {**values, "owning_team": "platform-sre"}))
It prints:
RENDER FAILED: template needs 'owning_team' but Terraform does not output it
## Restore namespace `payments` on prod-eu-3 (eu-west-1)
1. `kubectl config use-context prod-eu-3`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in s3://acme-velero-eu-3. Owner: platform-sre
The first line shows the point: the template needed an owning_team that Terraform did not provide, so the build fails loudly instead of publishing a runbook with a blank owner.
2. Capturing real cluster state as examples
Invented sample output goes stale and can hide format changes. Instead, a CI job creates a throwaway cluster (for example with kind, a tool that runs Kubernetes inside containers), runs the read-only commands the runbook references (kubectl get nodes, kubectl get pods -n payments), and writes the output to files that the docs include. Each capture is stamped with the commit and date. Scrub secrets and account IDs in the capture step, not by hand.
3. Anchoring docs to infrastructure code
- The runbook lives beside the module that builds the cluster and is released with the same version tag, so "the runbook for module v1.14" is a real thing.
- Runbooks reference IaC names (module outputs, variable names), and a CI check scans the template for every placeholder and confirms each one still exists in the outputs. A minimal version, run for real after the bucket output was renamed:
import re
tf_outputs = {"cluster_name", "region", "velero_bucket", "owning_team"} # after the rename
runbook = "Restore on $cluster_name ($region). Backups live in $backup_bucket. Owner: $owning_team"
refs = set(re.findall(r"\$(\w+)", runbook))
for name in sorted(refs - tf_outputs):
print(f"runbook references '{name}' but Terraform no longer outputs it")
print("checked", len(refs), "references")
runbook references 'backup_bucket' but Terraform no longer outputs it
checked 4 references
- A pull request that changes a module output must touch or acknowledge the runbooks that use it (CODEOWNERS, a file that assigns required reviewers by path, makes the doc owner a reviewer). One line does it:
runbooks/restore-namespace.md @acme/platform-sremakes that team an automatically requested reviewer on any PR touching the file. It becomes a hard requirement only when branch protection has "Require review from Code Owners" enabled. Because a module-output change would not touch the runbook file, also add a CODEOWNERS line for the module path that requests the runbook owners, and the PR template asks "did you update or confirm the affected runbooks?", which is the acknowledge step.
4. Testing the steps
- Executable steps: each numbered step maps to a script or
maketarget, and a scheduled pipeline builds a new cluster, seeds sample data, runs the restore steps, and asserts the result (for example that the expected pods are Ready). - Game days (planned drills where a human follows the doc cold) catch what automation cannot: unclear wording, missing permissions, steps that assume tribal knowledge.
- Record time to complete and where the tester hesitated, then fix the doc, not the tester.
Worked example
The nightly pipeline stands up a clean cluster, restores a sample namespace from the newest backup, and asserts three pods reach Ready. One night a Terraform change renames backup_bucket to velero_bucket. The reference check fails the PR that made the rename, because the runbook still uses the old name. The author updates the template in the same PR. Without the check, the mismatch would have surfaced during a real restore.
Trade-offs and pitfalls
- Full automation of every step is expensive. Automate the recovery paths that matter most and rely on game days for the rest.
- Generated docs still need human sentences explaining why and what to do when a step fails. Generate values, hand-write judgement.
- A test cluster can differ from production (size, network, data), so state the gap in the doc rather than implying the drill proves everything.
- Old versions of the runbook must remain reachable for clusters that have not been rebuilt yet.
What rules of thumb do you follow for clear technical writing aimed at engineers, product managers and lawyers at once? Show one rule applied to a sentence you would actually write.
Sample Answer
Direct answer
Write one document that works for three readers by leading with the conclusion, using one word for one thing, and making every sentence concrete: a named actor, a number, a unit. Engineers want checkable claims, product managers want the decision and its impact, and lawyers want commitments that cannot be read two ways. Plain, exact sentences serve all three.
Rules of thumb
- Put the conclusion first. The first two sentences say what happened, what is decided, or what is asked. Detail follows for those who need it.
- One term, one meaning. Define a term once, then never swap in a synonym. If "customer" means the paying company, do not later use "user" for the same thing. Lawyers read a change of word as a change of meaning.
- Name the actor and use active voice. "The service deletes the file" says who does what. "The file will be deleted" hides who.
- Use numbers, units and dates, not adjectives. "Fast", "soon" and "a short period" cannot be checked.
- One idea per sentence, and short. It survives translation and skimming.
- Separate facts, opinions and commitments. Label recommendations as recommendations. Use "must" for a requirement, "should" for a recommendation and "may" for a permission, the convention of RFC 2119 (an internet standard whose keywords are MUST, SHOULD and MAY, written in capitals in the standard itself). Avoid "can" for permission, since it can also mean ability.
- Say what you do not know. "We have not tested above 500 requests per second" is more trustworthy than silence.
- Expand acronyms on first use and keep a glossary.
One rule applied to a real sentence (rule 3 plus rule 4)
Before:
Customer data may be retained for a period following termination as deemed appropriate and will be handled in accordance with policy.
After (example figures, to be confirmed with the owner):
We delete customer data 30 days after the contract ends. We delete backups within 90 days after that.
What changed and why each reader benefits:
| Reader | Before | After |
|---|---|---|
| Engineer | Cannot build or test "a period" | A checkable deletion job with two deadlines |
| Product manager | Cannot tell customers anything | Can quote the retention promise |
| Lawyer | "May" and "as deemed appropriate" are vague, so risk is hidden | The commitment is explicit and reviewable |
Trade-offs and pitfalls
- Precision creates commitments. "We delete" is a promise, so do not publish a number until the owning team and legal have confirmed it. If the true answer is "it depends", write the dependency: "We delete data 30 days after the contract ends unless a legal hold applies."
- Do not hide behind qualifiers to stay safe ("generally", "in most cases"). Either state the exception or drop the hedge.
- Layer the document: a short summary on top for the PM and the lawyer, detail below for the engineer, rather than three separate versions that drift apart.
- Have a person from each group read the draft for one thing only: the sentence they would act on. Their confusion shows where the text is weak.
A client reports inconsistent and translated guides across regions, and misconfigurations followed in production. How do you harmonise the documentation and make it translation-ready?
Sample Answer
Direct answer
Two separate failures are hiding here. The regional guides drifted because there is no single source, and the misconfigurations happened because translation touched things that must never change (config keys, commands, values). I would fix both: make one canonical source, mark which parts are translatable and which are literal, translate through a controlled process, and add a check that catches altered code in any language. I would start with the install and configuration pages, since those caused production harm.
1. One source of truth
- Keep a single canonical language (say English). Regional guides are translations of it, plus small region overlays for facts that truly differ (endpoints, data-residency notes; data residency means rules about which country or region data must be stored in), kept as structured data rather than as copied prose.
- Each translated page records
source_commitin its front matter (the header block at the top of a page). If the source moves past that commit, CI (continuous integration, the automatic checks run on every proposed change) marks the page stale and shows a banner: "This translation is behind the English version." Readers then know when not to trust it.
2. Make the source translation-ready
- Short sentences, one term per concept, a glossary (termbase: an approved list of terms and their required translations) that translators must follow. No idioms or jokes, no sentence fragments glued together in code, no text inside screenshots (use text labels or alt text).
- Say units and formats explicitly, and let locale settings (the regional formatting rules for dates and numbers) render them.
- Never translate literals: code blocks, commands, flags, config keys, values, file paths, UI labels quoted exactly as they appear. Mark them so translation tools skip them.
3. Controlled translation
Use a translation memory (a database of past translated segments that is reused so wording stays consistent), machine translation for a first pass, and a human reviewer for high-risk pages (install, security, configuration). Tier languages (rank them into groups) by user base: fully reviewed for the largest, machine-translated with a visible label for the rest.
4. The check that would have prevented the incident
import re
FENCE = re.compile(r"```[a-z]*\n(.*?)```", re.S)
def code_blocks(markdown):
return FENCE.findall(markdown)
def compare(source, translation):
problems = []
src, tr = code_blocks(source), code_blocks(translation)
if len(src) != len(tr):
problems.append(f"block count differs: {len(src)} vs {len(tr)}")
for i, (a, b) in enumerate(zip(src, tr), start=1):
if a != b:
problems.append(f"code block {i} was altered in translation")
return problems
english = "Set the region:\n```yaml\nregion: eu-west\nretries: 3\n```\nThen restart.\n"
spanish = "Configure la region:\n```yaml\nregion: eu-west\nreintentos: 3\n```\nLuego reinicie.\n"
for problem in compare(english, spanish):
print(problem)
print("checked", len(code_blocks(english)), "block(s)")
It prints code block 1 was altered in translation and checked 1 block(s). How the checker works: FENCE finds text between triple-backtick markers, and re.S lets . match line breaks so a multi-line block is captured whole. zip pairs the Nth English block with the Nth translated block, and enumerate numbers them so the message can say which block changed.
The Spanish page translated the key retries to reintentos. The application would silently ignore an unknown key and fall back to its default, which is exactly how a "followed the guide" misconfiguration reaches production. CI now blocks the merge.
What an overlay looks like
# overlays/eu.yaml (only the facts that differ from the English source)
api_endpoint: https://api.eu.example.com
data_residency_note: Data is stored only in EU data centres.
The page text references these values instead of copying them, so a fix in English reaches every region.
Worked example
A client has US, EU and APAC guides. Week 1: diff each against English, list mismatches, and choose English as canonical. Week 2: add overlays for the three real regional differences and front matter to every page. Week 3: turn on the code-block check and stale banners. Later: add remaining languages by tier.
Trade-offs
- Full human review of every language is expensive. Tiering by risk and user base spends it where errors hurt.
- Machine translation is fast but can alter meaning; label it and never let it touch literals.
- What would change my call: with two regions and a few dozen pages, skip the translation memory and rely on the glossary and the code-block check.
What is an RFC (Request for Comments) process, and when is writing one better than an ad-hoc design discussion? What is the minimum a good RFC must contain, who should review it, and how do you settle disagreements that come out of the comments?
Sample Answer
Direct answer
An RFC (Request for Comments) process is a lightweight way for engineers to propose a significant change in writing and collect feedback before building it. The term here means an internal design document, not the public internet standards of the same name. It beats an ad-hoc design discussion when the change is costly to reverse, affects several teams, or needs a record of why. For small, reversible changes a quick chat is faster and better.
When to write one
- The change crosses team boundaries or changes a shared interface.
- It is expensive or risky to undo (data migrations, new infrastructure, public API changes).
- Several plausible designs exist and the reasoning should be preserved.
- People are in different time zones or the decision needs asynchronous input.
Skip it for a bug fix, a contained refactor, or anything you could reverse in an afternoon.
Minimum contents
- Context and problem: what is wrong today and for whom.
- Goals and non-goals: what success means and what is deliberately out of scope.
- Proposal: the design, detailed enough for someone to evaluate it.
- Alternatives considered: at least one real option and why it lost.
- Risks and impact: what could break, who is affected.
- Rollout and rollback plan.
- Open questions, plus the decision owner (the one named person who makes the final call if reviewers disagree) and the deadline for comments.
A non-goal is something the proposal could plausibly do but deliberately does not, stated so reviewers do not argue about it.
Who should review
- Owners of the systems affected and teams that consume the interface.
- One or two people with relevant expertise (security, operations, data) when the change touches those areas.
- A reviewer outside the author's team for a fresh perspective.
- Keep the list small and named. Too many reviewers means no one feels responsible.
Settling disagreements
- Ask commenters to separate blocking concerns (this will break something) from preferences.
- If a thread passes a few rounds without progress, take it to a short call and write the outcome back into the RFC.
- The named decision owner makes the call after the deadline and records the reasoning, including the dissent.
- People who lost the argument then disagree and commit (voice the objection once, then support the decision fully).
- Record the result as a decision record (also called an architecture decision record, or ADR): a short note saying what was decided, why, and what was rejected. The RFC is the discussion before the decision; the decision record is the lasting summary of the outcome. Many teams simply update the RFC's status to "accepted" and keep it as the record.
Worked example
An engineer proposes replacing a nightly cron job with a message queue for order processing. The RFC states the problem (orders wait up to a day), goals (process within minutes), non-goals (no change to the billing service), the proposal, one alternative (run cron every 5 minutes), and a rollback (switch the queue consumer off and cron resumes). Two reviewers disagree on queue technology. After a 20-minute call, the tech lead as decision owner picks the managed service already used elsewhere and records the operational-cost reason.
What the RFC looks like on the page (skeleton for the example above)
# RFC: Replace nightly order cron with a message queue
Status: In review | Decision owner: Tech lead, orders | Comments close: Fri
## Context and problem
Orders placed after 02:00 wait up to 24 hours because processing runs once nightly.
## Goals / Non-goals
Goals: process orders within minutes.
Non-goals: no change to the billing service; no new data store.
## Proposal
Publish an event per order; a consumer processes it and retries on failure.
## Alternatives considered
Run the cron every 5 minutes. Rejected: still polls the whole table and risks overlapping runs.
## Risks and rollout
Duplicate delivery is possible, so the consumer must be idempotent (safe to run twice on the same order). Roll out to 10% of orders first. Rollback: switch the consumer off and cron resumes.
## Open questions
Which managed queue service do we use?
Trade-offs and pitfalls
- RFCs can turn into bureaucracy. Keep them short and time-boxed.
- Writing after the code is done turns review into theatre. Circulate before committing to the design.
- Silence is not agreement. Ask directly for sign-off from affected owners.
- Without a decision owner, comment threads never end.
You delivered architecture documentation, yet the engineering team still makes wrong assumptions during implementation. How do you assess whether readers understood it, and what would you change in the document?
Sample Answer
Direct answer
Wrong assumptions after delivery mean the document was read but not understood the way I intended, or it never said the thing that matters. I would stop asking "is it complete?" and test comprehension directly with a few of the engineers, then change the document based on the specific misunderstandings found, not on general polish.
Step 1: Find the actual misunderstandings
- Collect the wrong assumptions already made: review PR comments, bug tickets and the design questions from the last few weeks, and write each as "engineer believed X, document intends Y".
- Teach-back sessions (15-20 minutes each, with 3 to 5 engineers who did not write the doc): ask them, without the document open, to explain the design in their own words and to answer scenario questions ("what happens to an order if the payment service is down?"). Where they diverge is where the doc fails.
- Watch someone use it: give a reader a task ("add a new event type") and see where they look and where they stop. Note the section names they never opened.
- Ask what they skipped: most readers scan headings, diagrams and the first paragraph only.
Step 2: Diagnose the pattern
| Symptom | Likely cause | Change |
|---|---|---|
| Same wrong assumption from several people | Doc is silent or ambiguous on it | State it explicitly, near the top |
| Correct info exists but people missed it | Buried in prose | Move into a diagram, a table or a callout |
| People built the wrong thing correctly | Constraints and non-goals missing | Add "what this design does not do" |
| Different readers, different readings | Undefined terms | Glossary, consistent naming |
Step 3: Change the document
Put decisions and constraints first (a one-paragraph summary a rushed reader can act on), use one C4 diagram (a standard way to draw software architecture at levels: context, containers, components, code) per question instead of a huge diagram, add explicit "guarantees and non-guarantees", add a worked example that traces one request end to end, and link to ADRs (architecture decision records, short notes on why a choice was made) so the reasons are visible.
Worked example (illustrative)
Engineers assumed the orders service publishes each event exactly once, so their consumer did not handle duplicates. The document said only "the service publishes an event when an order changes". Teach-back with four engineers: three said "once". Fix: add a callout "Delivery is at least once: the same event can arrive twice; consumers must be idempotent (safe to process twice), keyed on event_id", plus a sequence diagram showing a redelivery. Re-test with two new engineers a week later: both answer the duplicate-event question correctly.
Trade-offs and pitfalls
- Adding more text rarely helps; restructuring and adding explicit non-guarantees does.
- Do not blame the readers, and do not test with the author in the room, since the author fills gaps without noticing.
- Some wrong assumptions come from missing conversation, not missing docs. Pair the document with a short walkthrough at the start of implementation, and re-check after the change.
Unlock Full Question Bank
Get access to all Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.