Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
How do you write a deployment guide so a reader can confirm each step worked without asking the author? Give me a concrete example of the evidence you would build into one step.
Sample Answer
Direct answer
A deployment guide is a numbered procedure for putting a change into an environment. Every step should have four parts: the action, the expected observable result, a check the reader runs themselves, and what a failure looks like with the next move. The reader then confirms success from evidence on their own screen instead of trusting the author. The key design choice is to verify the outcome (what is now running), not merely that the command did not error.
Concrete example: one step with its evidence
Illustrative service orders-api on Kubernetes (a system that runs copies of a service, called pods):
STEP 4. Deploy version 1.8.2 to production.
Do:
kubectl -n orders set image deployment/orders-api api=registry.example.com/orders-api:1.8.2
Wait for it (a rollout is Kubernetes replacing the old pods with new ones, a few at a time):
kubectl -n orders rollout status deployment/orders-api --timeout=120s
Expected last line: deployment "orders-api" successfully rolled out
Prove it is the version you wanted (do not skip). The jsonpath expression below just prints the full image reference (registry, name and tag) of the first container in the deployment's pod template, that is, which version Kubernetes has been told to run:
kubectl -n orders get deployment orders-api -o jsonpath='{.spec.template.spec.containers[0].image}'
Expected output: registry.example.com/orders-api:1.8.2
Prove it serves traffic:
curl -fsSi https://orders.example.com/health
Expected: a first line of HTTP/2 200 (or HTTP/1.1 200 OK) and a body containing "status":"ok"
(-i prints the status line and headers; without it curl shows only the body. -f makes curl exit non-zero on a 4xx or 5xx response.)
If the wait ends with "error: timed out waiting for the condition" (the progress lines "Waiting for deployment ... rollout to finish" are printed during every normal rollout, so on their own they are not a failure):
run kubectl -n orders get pods and look for CrashLoopBackOff (a pod status meaning the container keeps starting, crashing and being restarted with growing delays), then go to
"Rollback" at the end of this guide. Do not continue to step 5.
Why this step works
- The second check compares the image the Deployment is now configured to run with the version in the step, which catches "the command succeeded but the wrong version was applied" (a typo or a stale tag). The rollout wait is what proves the old pods were actually replaced; to see what is running right now, read the pods' images with
kubectl -n orders get pods -o jsonpath='{.items[*].spec.containers[*].image}'. - Expected output is written out, so the reader can compare text instead of judging.
- The failure branch tells the reader where to go, so they do not improvise.
- The step is bounded (
--timeout=120s), so "still waiting" becomes a decision, not an endless hang.
Practices that generalise
- Preconditions before step 1: access, tool versions, the approved release number.
- One action per step, and outcome checks placed before the next action.
- Match on stable substrings in expected output, not banners with version numbers that change monthly.
- Make risky steps reversible: each has an undo, or an explicit note that it is irreversible and needs a second person.
- Test the guide: give it to someone who was not involved and do a dry run in staging. Each question they ask is a missing piece of evidence.
Another evidence example. For a database migration step, the check is a query against the migrations table that returns the latest applied migration name, compared with the name in the release notes.
Pitfalls
- "Verify the deployment succeeded" with no method.
- Expected output copied once and then wrong after a tool upgrade.
- Verifying only that a process is up, not that it does the job.
You are setting up review for documentation pull requests in a repository. What would you check, what would you automate, and how do you keep the review from becoming a rubber stamp?
Sample Answer
Direct answer
Split the review by who is better at each check. Machines check what is mechanical and repeatable: links, formatting, spelling, snippets that must run, metadata, secrets. People check what only a person can judge: is it true, is it complete for the reader's task, does it fit the audience. To stop human review becoming a rubber stamp (approving without really looking), make the reviewer produce evidence of what they did, keep changes small, and audit outcomes, not approvals.
What is automated versus checked by a person
| Check | Who | Blocks the merge? |
|---|---|---|
| Broken internal links, formatting, spelling and style rules | CI (automatic checks on every pull request) | Yes |
| Code snippets and commands run in a clean environment | CI | Yes |
| Required front matter (metadata block at the top of the page: owner, audience, last verified date, version) | CI script | Yes |
| Secrets and personal-data scan | CI | Yes |
| External links | CI, allowed to warn only | No (they flake: they fail sometimes because the other site is slow or down, not because your page is wrong) |
| Accuracy against the real system | Subject-matter owner (the person who knows that system best), listed in CODEOWNERS (a file mapping paths to owners; on GitHub it only requests their review, and it blocks the merge only when the branch protection rule "Require review from Code Owners" is switched on) | Yes, once that rule is on |
| Complete for the reader's task: preconditions, verify step, rollback | Second reviewer | Yes |
| Audience and clarity | Second reviewer | Advisory (a comment the author may act on, but it does not stop the merge) |
For ML artifacts, add: a model card (a short document stating a model's purpose, training data, evaluation and limits), data provenance (where the data came from and what it may be used for), and monitoring readiness (what alerts if the model degrades and who owns it). Missing any of these blocks the merge, because each covers a risk that cannot be fixed after readers rely on the page: a model with no stated limits gets misused, data with unknown origin may not be legal to use, and a model with no alert or owner degrades unnoticed. A monitoring-readiness check looks like this: "The page names the alert (for example accuracy drops below the stated floor for 3 days), the dashboard link, and the on-call owner; the reviewer opened the dashboard and saw it."
Keeping review honest
- Evidence line: for procedure docs the reviewer writes "Followed steps 1 to 6 on a clean environment; step 4 needed X" in the review. "LGTM" (looks good to me) is not accepted for pages that instruct actions.
- Required approvals by ownership, so the approver is someone who knows the system.
- Small pull requests, since a 30-page change gets skimmed.
- Outcome measures: how often a merged page needs a correction within a month, and how many reviews found anything. A reviewer who never finds anything is a signal to look, not a compliment.
- Occasional audit: a monthly sample re-checked by someone else.
Worked example
A PR updates a database-restore runbook (a step-by-step procedure for an operational task). CI passes links and formatting. The owner approves accuracy. The second reviewer actually runs the restore in a test environment and posts: "Step 3 fails because the snapshot name flag changed." That is the catch a rubber stamp would have missed.
Trade-offs and pitfalls
- Too many blocking checks slow contributors and encourage bypassing. Keep flaky checks advisory.
- Evidence lines can be copied. Audits and outcome measures are the backstop.
- What would change my call: for a tiny team, one owner approval plus CI is enough.
Your infrastructure is ephemeral, for example disposable Kubernetes clusters. How do you keep runbooks and recovery procedures accurate? Discuss templates, capturing real cluster state as examples, anchoring the docs to infrastructure code, and how you test the steps.
Sample Answer
Direct answer
When clusters are disposable, a hand-written runbook is wrong almost immediately, so stop treating the doc as prose and treat it as a build artifact. Generate the environment-specific parts from the infrastructure code (IaC, meaning tools like Terraform or Helm, which define infrastructure in files), keep the doc in the same repository and version as that code, show real captured output rather than invented examples, and run the procedure against a fresh cluster on a schedule. If the run fails, the doc is stale, and you find out from a red build in CI (continuous integration, the automation that runs checks on every change) instead of a bad incident.
Terms in plain words
- Kubernetes cluster: a group of machines that runs containerized applications. A namespace is a named partition inside it (here, everything belonging to
payments). A pod is one running copy of an application, and it is Ready when it has started and passes its health check. - kubectl: the command-line tool for talking to a cluster. A context is a saved "which cluster am I talking to" setting, so
use-contextswitches clusters. - Velero: a Kubernetes backup and restore tool.
velero backup getlists backups andvelero restore create --from-backup Xrecreates the saved resources. - Terraform: a tool that builds infrastructure from code, and its outputs are named values (cluster name, bucket) it prints after building. A module is a reusable bundle of that code. Helm packages Kubernetes applications the same way.
- Tribal knowledge: things everyone on a team "just knows" but nobody wrote down.
1. Templates: separate the stable steps from the values that change
The sequence of steps (check backup, restore namespace, verify pods) changes rarely. Cluster names, regions and bucket names change with every cluster. Write the steps once as a template with placeholders, and fill the values in from IaC outputs at publish time. A missing value fails the build.
import json, re
from string import Template
# Shape of `terraform output -json`: each output is {"value": ..., "type": ...}
tf_outputs = json.loads("""{
"cluster_name": {"value": "prod-eu-3", "type": "string"},
"region": {"value": "eu-west-1", "type": "string"},
"backup_bucket": {"value": "s3://acme-velero-eu-3", "type": "string"}
}""")
values = {k: v["value"] for k, v in tf_outputs.items()}
runbook = Template("""## Restore namespace `payments` on $cluster_name ($region)
1. `kubectl config use-context $cluster_name`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in $backup_bucket. Owner: $owning_team
""")
def render(t, vals):
try:
return t.substitute(vals)
except KeyError as e:
# CI fails here: the doc references a value the infrastructure code no longer outputs
return f"RENDER FAILED: template needs {e} but Terraform does not output it"
print(render(runbook, values))
print(render(runbook, {**values, "owning_team": "platform-sre"}))
It prints:
RENDER FAILED: template needs 'owning_team' but Terraform does not output it
## Restore namespace `payments` on prod-eu-3 (eu-west-1)
1. `kubectl config use-context prod-eu-3`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in s3://acme-velero-eu-3. Owner: platform-sre
The first line shows the point: the template needed an owning_team that Terraform did not provide, so the build fails loudly instead of publishing a runbook with a blank owner.
2. Capturing real cluster state as examples
Invented sample output goes stale and can hide format changes. Instead, a CI job creates a throwaway cluster (for example with kind, a tool that runs Kubernetes inside containers), runs the read-only commands the runbook references (kubectl get nodes, kubectl get pods -n payments), and writes the output to files that the docs include. Each capture is stamped with the commit and date. Scrub secrets and account IDs in the capture step, not by hand.
3. Anchoring docs to infrastructure code
- The runbook lives beside the module that builds the cluster and is released with the same version tag, so "the runbook for module v1.14" is a real thing.
- Runbooks reference IaC names (module outputs, variable names), and a CI check scans the template for every placeholder and confirms each one still exists in the outputs. A minimal version, run for real after the bucket output was renamed:
import re
tf_outputs = {"cluster_name", "region", "velero_bucket", "owning_team"} # after the rename
runbook = "Restore on $cluster_name ($region). Backups live in $backup_bucket. Owner: $owning_team"
refs = set(re.findall(r"\$(\w+)", runbook))
for name in sorted(refs - tf_outputs):
print(f"runbook references '{name}' but Terraform no longer outputs it")
print("checked", len(refs), "references")
runbook references 'backup_bucket' but Terraform no longer outputs it
checked 4 references
- A pull request that changes a module output must touch or acknowledge the runbooks that use it (CODEOWNERS, a file that assigns required reviewers by path, makes the doc owner a reviewer). One line does it:
runbooks/restore-namespace.md @acme/platform-sremakes that team an automatically requested reviewer on any PR touching the file. It becomes a hard requirement only when branch protection has "Require review from Code Owners" enabled. Because a module-output change would not touch the runbook file, also add a CODEOWNERS line for the module path that requests the runbook owners, and the PR template asks "did you update or confirm the affected runbooks?", which is the acknowledge step.
4. Testing the steps
- Executable steps: each numbered step maps to a script or
maketarget, and a scheduled pipeline builds a new cluster, seeds sample data, runs the restore steps, and asserts the result (for example that the expected pods are Ready). - Game days (planned drills where a human follows the doc cold) catch what automation cannot: unclear wording, missing permissions, steps that assume tribal knowledge.
- Record time to complete and where the tester hesitated, then fix the doc, not the tester.
Worked example
The nightly pipeline stands up a clean cluster, restores a sample namespace from the newest backup, and asserts three pods reach Ready. One night a Terraform change renames backup_bucket to velero_bucket. The reference check fails the PR that made the rename, because the runbook still uses the old name. The author updates the template in the same PR. Without the check, the mismatch would have surfaced during a real restore.
Trade-offs and pitfalls
- Full automation of every step is expensive. Automate the recovery paths that matter most and rely on game days for the rest.
- Generated docs still need human sentences explaining why and what to do when a step fails. Generate values, hand-write judgement.
- A test cluster can differ from production (size, network, data), so state the gap in the doc rather than implying the drill proves everything.
- Old versions of the runbook must remain reachable for clusters that have not been rebuilt yet.
How should sensitive data, such as personal data and credentials, be handled in documentation and example code? How do you stop it from leaking into a docs site?
Sample Answer
Direct answer
Assume anything typed into docs will be public and permanent, so real credentials and real personal data never go in. The approach has four layers: use obviously fake, reserved-for-examples values while writing, scan for leaks before merge, scan what gets published, and have a rehearsed response for the day something slips through. A checker in CI (the automatic checks run on every proposed change) makes this routine instead of relying on people remembering.
1. Write with safe values
- Credentials: show environment variables and placeholders such as
YOUR_API_KEY, never a working token. Where a shaped example is needed, use one the vendor documents as a placeholder, for example theAKIAIOSFODNN7EXAMPLEkey ID that AWS publishes for examples. - Personal data (also called PII, personally identifiable information: names, emails, phone numbers, addresses): use synthetic records generated by a seeded script (a fixed starting value so the output repeats). Use reserved names:
example.com(RFC 2606), IP ranges192.0.2.0/24,198.51.100.0/24,203.0.113.0/24(RFC 5737). An RFC (Request for Comments) is a numbered internet standards document. These two reserve names and ranges for examples only, so they are guaranteed never to belong to a real person or machine, which is what makes them safe to publish. The/24means the first 24 bits fix the network, leaving 256 addresses (192.0.2.0 to 192.0.2.255). Use fictional phone numbers such as555-0100. An "anonymised" copy of production data is not safe, because rows can often be re-identified (matched back to real people by combining columns such as postcode and birth date). - Hidden channels: screenshots, notebook outputs, log excerpts, curl examples with tokens in headers, and query results. These leak more often than the prose does.
2. Detect before merge
A scanner (open-source tools such as gitleaks or detect-secrets, or a small script like this one) runs as a pre-commit hook (a script git runs automatically just before a commit is recorded) and again in CI, over the source and the built site:
import re
RULES = {
"aws-access-key-id": re.compile(r"\bAKIA[0-9A-Z]{16}\b"),
"private-key-block": re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
"bearer-token": re.compile(r"Bearer\s+[A-Za-z0-9\-_.]{20,}"),
"real-looking-email": re.compile(r"[\w.+-]+@(?!example\.(?:com|org|net)\b)[\w-]+\.[\w.]+"),
}
ALLOWED = ("AKIAIOSFODNN7EXAMPLE",) # the placeholder key AWS documents
def scan(text):
findings = []
for lineno, line in enumerate(text.splitlines(), start=1):
for rule, pattern in RULES.items():
for match in pattern.finditer(line):
if match.group(0) not in ALLOWED:
findings.append((lineno, rule))
return findings
page = """\
export AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
export AWS_ACCESS_KEY_ID=AKIAZ3MPLEFAKE234567
curl -H "Authorization: Bearer abcdefghijklmnopqrstuvwx" https://api.example.com/v1/me
Contact jo.smith@corp-mail.io or support@example.com
"""
for lineno, rule in scan(page):
print(f"line {lineno}: {rule}")
It prints line 2: aws-access-key-id, line 3: bearer-token and line 4: real-looking-email. Line 1 (the documented placeholder) and support@example.com pass. How to read the rules: AKIA[0-9A-Z]{16} means the letters AKIA followed by exactly 16 capital letters or digits, the shape of an AWS access key ID, and \b marks a word edge. In the email rule, (?!example\.(?:com|org|net)\b) is a negative lookahead: it matches only when the domain is not example.com, .org or .net.
The allowlist is tiny and reviewed, so a real key cannot hide behind it.
3. Stop it reaching the docs site
- Build the site in an environment with no production credentials, and run notebooks or examples against a sandbox, so generated output cannot contain live values.
- Scan the built output and the search index, not only the Markdown.
- Require a PR field: "Contains no real customer data, and how you know."
4. If something leaks
Assume compromise. Revoke and rotate the credential first (rotate means issue a new one and disable the old one), then remove the content, purge caches and history if needed (delete copies held by the search index, CDN caches and git history), and check access logs. Deleting the page alone fixes nothing, since the secret was already copied.
Trade-offs and pitfalls
- Scanners are noisy. Keep a small reviewed allowlist, or people will disable the tool.
- Pattern scanners find keys but not a real name in a screenshot. Human review still covers images and sample datasets.
- What would change my call: for an internal-only wiki, keep the scan and the placeholder rules but relax the publishing layer.
A developer changed the code but not the architecture docs, and later an outage exposed the mismatch. Walk me through the analysis, the short-term fix, and what you would change so docs stay in step with code.
Sample Answer
Direct answer
Treat it as a blameless postmortem (a written review of an incident that asks how the system allowed the failure, not who to punish). The developer did what the process let them do: nothing in the workflow connected a code change to the document that describes it. I would find out how the mismatch caused the outage, fix the docs now, and then change the workflow so the code change itself prompts a doc update.
1. Analysis (lead the postmortem)
- Timeline: when the code changed, when the doc went stale, when the outage started, when responders consulted the doc, and how long it misled them. The measure that matters is the extra time to resolve (mean time to resolve, MTTR) caused by wrong information.
- Causal chain (the five whys): outage, because the on-call engineer (the person on duty to respond to alerts) followed the diagram, because the diagram showed the old dependency, because the change did not touch the doc, because nothing required it, because docs live outside the pull request (PR, a proposed code change reviewed before merge) flow.
- Separate three questions: was the doc wrong, was it unowned, and was it findable? Also ask whether the responders should have trusted the doc at all, since a doc without a "last verified" date invites over-trust.
- Blameless framing: name the gap in the system, and the developer helps write the fix.
2. Short-term fix (this week)
- Correct the affected docs (architecture diagram, runbook steps (the step-by-step procedure for handling an alert), dependency list) and get the service owner to sign off.
- Sweep the neighbouring docs for the same change; the same change usually broke more than one.
- Add a banner or "last verified" date to the doc, and post the corrected version where on-call actually looks.
3. Long-term: keep docs in step with code
- Docs-as-code (treat docs like source code: plain files, in version control, reviewed in PRs): keep docs in the same repository as the code, so one PR can change both and reviewers see both.
- PR template checkbox: "Does this change architecture, dependencies, config, or runbook steps? Link the doc update or write 'no'." Add CODEOWNERS (a repo file assigning required reviewers by path) so changes under
docs/orarchitecture/need the owning team's review. - Automate what can be checked: generate diagrams and dependency lists from infrastructure-as-code (the deployment config files that define what runs where) or a service catalog (a registry listing each service, its owner and its dependencies) rather than drawing them by hand; add a CI (continuous integration, the automated checks on each PR) job that fails on broken links or an API spec that no longer matches the code.
- Staleness detection: a "last verified" date with a reminder after, say, 90 days; verify runbooks in a game day (a rehearsal where the team practises an incident).
Worked example (illustrative)
A service was changed to read from a new cache cluster. The diagram still showed the database as the only dependency. During the outage, on-call restarted the database and lost 40 minutes because the cache was the real fault. Postmortem actions: (a) fix the diagram today; (b) generate the dependency diagram from deployment config; (c) add the PR checkbox; (d) owner: the service team's tech lead; due dates and a follow-up review at 30 days to check the checkbox was used.
Trade-offs and pitfalls
- Mandatory checkboxes turn into rubber stamps unless review actually enforces them; generation beats discipline where possible.
- Do not "fix" it by telling developers to be more careful. That is the answer that changes nothing.
- Preventive measures should be specific to documentation: ownership, verification dates, generation, PR linkage. Generic action items ("improve communication") fail the same follow-up test.
Unlock Full Question Bank
Get access to all Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.