Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
What is the difference between a playbook, a runbook and a README? For each, say who reads it, what it contains, how often it needs maintenance, and a situation where you would choose it over the other two.
Sample Answer
Direct answer. All three are documents, but they answer different questions. A README answers "what is this and how do I start?", a runbook answers "what exact steps do I follow for this task or alert?", and a playbook answers "how do we handle this kind of situation, and who decides what?". Naming conventions differ between companies, so I state my definitions when I write one. In everyday terms: a runbook is an operating instruction sheet, like a recipe you follow exactly; a playbook is a team's game plan, which tells people their roles and when to make a call; a README is the front-door sign of a project.
| README | Runbook | Playbook | |
|---|---|---|---|
| Reader | New developer or user opening the repository | The on-call engineer or operator under time pressure | A team or incident commander (the person coordinating a response) plus other roles |
| Contents | What the project is, install and run, how to contribute, where to find more docs | Step-by-step commands for one procedure (restart a service, fail over a database), with checks and rollback | Higher-level plan for a scenario type: roles, decision points, escalation, communication, links to runbooks |
| Maintenance | Whenever setup or usage changes (on each pull request that alters them); reviewed each release | Every time the procedure or system changes, plus a periodic test (for example, quarterly, by running it) | After every real use or exercise, and at least yearly |
| Choose it when | You want a stranger to run the project in ten minutes | An alert fires and the response is a known, repeatable procedure | The situation needs judgement and coordination, such as a security breach or a regional outage |
Worked example: a payments service.
- README: "Payments API. Requires Python 3.12. Run
make dev. Tests:make test. Docs: link." - Runbook: "Alert: queue depth above threshold. 1. Check the consumer dashboard. 2. Run
kubectl rollout restart deployment/payments-consumer. 3. Confirm the queue drains within 10 minutes. 4. If not, roll back to the previous release (command)." Each step is unambiguous and checkable. - Playbook: "Suspected payment data leak. Incident commander declares severity. Security lead assesses scope. Legal decides on notification. Communications drafts customer message. Decision point: take the service offline? criteria listed." It links to the runbooks for isolating a host.
Choosing between them. If the reader must think, it is a playbook; if the reader must follow, it is a runbook; if the reader is arriving cold, it is a README. A typical mistake is writing a runbook full of judgement calls (it fails at 3am) or a playbook full of shell commands (it goes stale quickly).
On-call engineers want highly detailed runbooks, while auditors want concise documentation. How would you design the documentation for one system so it serves both without two copies drifting apart?
Sample Answer
Direct answer
I would keep one source of truth written for the on-call engineer (the detailed runbook, a step-by-step procedure for handling an operational task or incident) and generate the auditor's view from it, instead of maintaining a second, shorter document. Auditors rarely need less truth, they need a different entry point: which control (a required safeguard, such as "we can recover a failed service") is covered, who owns it, how often it is tested, and where the evidence is. Those are metadata and links, and they can be derived from the runbook, so nothing can drift.
Design
- Docs-as-code (documentation stored as files in version control and reviewed like code). The runbook is a Markdown file in the service's repository, so every change has an author, a review and a timestamp. That history is itself audit evidence.
- Structured metadata at the top of each runbook (front matter, a small block of fields the file starts with):
id: RB-PAY-014
title: Recover a failed payments-api instance
owner: payments-oncall
last_verified: 2025-06-02
verified_by: a.rivera
drill_link: https://wiki.example.com/drills/2025-06-02
control_ids: [IR-2, BC-4]
review_every_days: 90
- A generated audit index. A small script reads the fields and produces a table, so the auditor's page is regenerated on every merge and never hand-edited. The script runs in the same CI job as the checks below, reads every runbook file's front matter in the repository, and writes the table to the audit page. Here are the two runbooks' front matter blocks (the second one differs in
id: RB-ORD-003,owner: orders-oncall,last_verified: 2025-02-10,control_ids: [IR-2]) and the whole script, run with a fixed "today" of 2025-08-15. One trap worth knowing: PyYAML'ssafe_loadalready turns2025-06-02into a Pythondate, so callingdate.fromisoformaton it raises a TypeError.
import glob
from datetime import date
import yaml # PyYAML
today = date(2025, 8, 15) # fixed so the output is reproducible
print("| Runbook | Owner | Controls | Days since verified | Status |")
print("|---|---|---|---|---|")
for path in sorted(glob.glob("*.md"), reverse=True):
front = open(path).read().split("---")[1] # text between the first two --- lines
fields = yaml.safe_load(front)
age = (today - fields["last_verified"]).days
status = "STALE" if age > int(fields["review_every_days"]) else "current"
print(f"| {fields['id']} | {fields['owner']} | {', '.join(fields['control_ids'])} | {age} | {status} |")
The script reads every runbook file in the folder and prints one table row per runbook. The output:
| Runbook | Owner | Controls | Days since verified | Status |
|---|---|---|---|---|
| RB-PAY-014 | payments-oncall | IR-2, BC-4 | 74 | current |
| RB-ORD-003 | orders-oncall | IR-2 | 186 | STALE |
Age is the date difference. RB-PAY-014: June 2 to August 15 is 74 days, under its 90-day interval, so current. RB-ORD-003: February 10 to August 15 is 186 days, over 90, so STALE.
- Summary layer by link, not copy. An audit-facing page holds two to three sentences per control plus permalinks (stable links that keep pointing at the same section even if the page is renamed or moved) to the runbook section. It never restates steps.
- Evidence is produced by doing, not by writing. Drill records (dated records of a practice run, such as deliberately failing over a service to prove the steps work), CI results and incident tickets are linked from the metadata. Auditors trust a dated drill link more than a long paragraph.
- Enforcement in CI (the automated checks that run on every change): fail the build if owner or last_verified is missing, and open a ticket when an entry passes its review interval.
Worked example
Auditor asks: "Show me how you recover a failed payments instance and how you know it works." They open the index row for RB-PAY-014, see control IR-2 (an internal control ID), owner, verification date and the drill link, and click through to the exact steps. The on-call engineer never sees the audit wrapper and the auditor never has to read a 4-page procedure unless they choose to.
Trade-offs and what would change my call
- Generation adds tooling to maintain, which is worth it once there are more than a handful of runbooks.
- Metadata can become theatre (a show of compliance: dates bumped without real verification). Guard by requiring a linked drill or run record with the date, not just a field edit.
- If an audit regime (the rules an auditor applies and how often they check) demands frozen point-in-time documents (copies that show exactly what was true on one date and never change afterwards), I would tag a snapshot of the generated view at each audit period, keeping one source of truth.
- If the wiki has no version control, I would use page properties (structured fields attached to a wiki page, which play the role of front matter) plus a scheduled export (a job that copies the pages out on a timer), but I would still refuse two hand-written copies.
Pitfalls
- Writing the "short version" by hand once and promising to update it.
- Making the runbook shorter to please auditors, which harms the reader who needs it most under pressure.
Design a one-page postmortem template that works for both engineering readers and business stakeholders. Then fill in a short example for an outage that affected ten percent of users for three hours.
Sample Answer
Direct answer. A one-page postmortem answers, in plain order: what happened and how bad, why it happened, how we responded, and what we will change. A top block written for business readers (impact, duration, cause in one sentence, status) sits above an engineering block (timeline, root cause, contributing factors, action items). It is blameless: it describes systems and decisions, not individuals.
Template
TITLE / date / severity (how bad it was, on a scale such as SEV1 total outage to SEV3 minor) / author / status (draft, reviewed, closed)
SUMMARY (business readers): impact, duration, one-sentence cause, customer action needed
IMPACT: who, how many, what they experienced, measured how (include uncertainty)
TIMELINE (UTC): detection, escalation, mitigation, resolution
ROOT CAUSE and CONTRIBUTING FACTORS: technical cause, why it was not caught earlier
WHAT WENT WELL / WHAT WENT POORLY / WHERE WE GOT LUCKY
ACTION ITEMS: owner, due date, priority, how we will verify it worked
Example: an outage that affected ten percent of users for three hours (illustrative details)
Summary. On 14 March, 10% of users could not complete checkout for 3 hours (10:05 to 13:05 UTC). Cause: a configuration change reduced the connection limit (the cap on simultaneous connections the application may open to the database, also called pool size) on one of the ten database shards (the database is split into ten slices, each holding a share of users). Fixed by reverting the change. No data was lost. Affected customers were offered a retry link.
Impact. Users are assigned to shards by ID, so one shard of ten means about 10% of users. If we had 400,000 daily active users, that is roughly 40,000 users with failed checkouts; the estimate rests on the assumption that shard membership is even, and we checked actual failed-request counts in logs to confirm the range.
Timeline (UTC).
- 09:58 config change deployed to shard 7.
- 10:05 error rate on shard 7 rises; no alert fires (the alert threshold, the error level that triggers a page, covers the global rate only).
- 10:40 support tickets spike; the on-call engineer (whoever is currently responsible for responding to problems) is paged by support, not by monitoring.
- 11:20 root cause identified; 12:50 revert approved; 13:05 recovery confirmed.
Root cause and contributing factors. The root cause is the single change without which the outage would not have happened; contributing factors are conditions that let it last or go unnoticed. The change set the pool size too low for the traffic on that shard: the new limit was below what the shard needs even off-peak, and the old value had been sized for peak. Contributing: the global error-rate alert averages ten shards, so a 10% failure looks like noise; no canary (staged rollout to a small share first); the change review did not include capacity numbers.
What went well/poorly/luck. Well: once the revert was approved at 12:50, recovery was confirmed 15 minutes later. Poorly, three delays: 35 minutes from first errors (10:05) to the first page (10:40); 40 minutes from the page to an identified root cause (11:20), which is 75 minutes after the errors began; and 90 minutes from root cause to an approved revert (11:20 to 12:50). The approval delay is the biggest single lesson, since it was the longest wait and the fix was already known. Luck: it happened off-peak; at peak the same 10% of users would have seen failures on a larger share of their requests, and the pool would have been exhausted sooner.
Action items.
| Action | Owner | Due | Verified by |
|---|---|---|---|
| Per-shard error-rate alert | SRE | 2 weeks | Fire drill (a practice run that injects fake errors on purpose to see the alert fire) |
| Canary rollout for config changes | Platform | 6 weeks | Next change goes through canary |
| Capacity check in config review template | Engineering manager | 1 week | Template updated |
| On-call may revert a config change immediately, without waiting for approval | Engineering manager | 3 weeks | Written policy, then a timed revert drill |
Variants (same skeleton, different emphasis). Data-quality incident: replace "users affected" with which tables, which dashboards, and which decisions used bad numbers, and add a backfill plan. ML outage or recurring degradation: the impact includes a metric with a range ("precision fell 3 to 5 points, estimated from a labeled sample of 400 predictions"), the timeline includes feature snapshots and drift checks before and after, and the root-cause section separates data drift (inputs changed) from a code or pipeline change. The five-section short form is: what happened, impact, why, what we did, what changes, each opened with one sample sentence like "For 3 hours, 1 in 10 users could not check out."
Your infrastructure is ephemeral, for example disposable Kubernetes clusters. How do you keep runbooks and recovery procedures accurate? Discuss templates, capturing real cluster state as examples, anchoring the docs to infrastructure code, and how you test the steps.
Sample Answer
Direct answer
When clusters are disposable, a hand-written runbook is wrong almost immediately, so stop treating the doc as prose and treat it as a build artifact. Generate the environment-specific parts from the infrastructure code (IaC, meaning tools like Terraform or Helm, which define infrastructure in files), keep the doc in the same repository and version as that code, show real captured output rather than invented examples, and run the procedure against a fresh cluster on a schedule. If the run fails, the doc is stale, and you find out from a red build in CI (continuous integration, the automation that runs checks on every change) instead of a bad incident.
Terms in plain words
- Kubernetes cluster: a group of machines that runs containerized applications. A namespace is a named partition inside it (here, everything belonging to
payments). A pod is one running copy of an application, and it is Ready when it has started and passes its health check. - kubectl: the command-line tool for talking to a cluster. A context is a saved "which cluster am I talking to" setting, so
use-contextswitches clusters. - Velero: a Kubernetes backup and restore tool.
velero backup getlists backups andvelero restore create --from-backup Xrecreates the saved resources. - Terraform: a tool that builds infrastructure from code, and its outputs are named values (cluster name, bucket) it prints after building. A module is a reusable bundle of that code. Helm packages Kubernetes applications the same way.
- Tribal knowledge: things everyone on a team "just knows" but nobody wrote down.
1. Templates: separate the stable steps from the values that change
The sequence of steps (check backup, restore namespace, verify pods) changes rarely. Cluster names, regions and bucket names change with every cluster. Write the steps once as a template with placeholders, and fill the values in from IaC outputs at publish time. A missing value fails the build.
import json, re
from string import Template
# Shape of `terraform output -json`: each output is {"value": ..., "type": ...}
tf_outputs = json.loads("""{
"cluster_name": {"value": "prod-eu-3", "type": "string"},
"region": {"value": "eu-west-1", "type": "string"},
"backup_bucket": {"value": "s3://acme-velero-eu-3", "type": "string"}
}""")
values = {k: v["value"] for k, v in tf_outputs.items()}
runbook = Template("""## Restore namespace `payments` on $cluster_name ($region)
1. `kubectl config use-context $cluster_name`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in $backup_bucket. Owner: $owning_team
""")
def render(t, vals):
try:
return t.substitute(vals)
except KeyError as e:
# CI fails here: the doc references a value the infrastructure code no longer outputs
return f"RENDER FAILED: template needs {e} but Terraform does not output it"
print(render(runbook, values))
print(render(runbook, {**values, "owning_team": "platform-sre"}))
It prints:
RENDER FAILED: template needs 'owning_team' but Terraform does not output it
## Restore namespace `payments` on prod-eu-3 (eu-west-1)
1. `kubectl config use-context prod-eu-3`
2. `velero backup get` (pick the newest with status Completed)
3. `velero restore create --from-backup <backup-name from step 2>`
Backups live in s3://acme-velero-eu-3. Owner: platform-sre
The first line shows the point: the template needed an owning_team that Terraform did not provide, so the build fails loudly instead of publishing a runbook with a blank owner.
2. Capturing real cluster state as examples
Invented sample output goes stale and can hide format changes. Instead, a CI job creates a throwaway cluster (for example with kind, a tool that runs Kubernetes inside containers), runs the read-only commands the runbook references (kubectl get nodes, kubectl get pods -n payments), and writes the output to files that the docs include. Each capture is stamped with the commit and date. Scrub secrets and account IDs in the capture step, not by hand.
3. Anchoring docs to infrastructure code
- The runbook lives beside the module that builds the cluster and is released with the same version tag, so "the runbook for module v1.14" is a real thing.
- Runbooks reference IaC names (module outputs, variable names), and a CI check scans the template for every placeholder and confirms each one still exists in the outputs. A minimal version, run for real after the bucket output was renamed:
import re
tf_outputs = {"cluster_name", "region", "velero_bucket", "owning_team"} # after the rename
runbook = "Restore on $cluster_name ($region). Backups live in $backup_bucket. Owner: $owning_team"
refs = set(re.findall(r"\$(\w+)", runbook))
for name in sorted(refs - tf_outputs):
print(f"runbook references '{name}' but Terraform no longer outputs it")
print("checked", len(refs), "references")
runbook references 'backup_bucket' but Terraform no longer outputs it
checked 4 references
- A pull request that changes a module output must touch or acknowledge the runbooks that use it (CODEOWNERS, a file that assigns required reviewers by path, makes the doc owner a reviewer). One line does it:
runbooks/restore-namespace.md @acme/platform-sremakes that team an automatically requested reviewer on any PR touching the file. It becomes a hard requirement only when branch protection has "Require review from Code Owners" enabled. Because a module-output change would not touch the runbook file, also add a CODEOWNERS line for the module path that requests the runbook owners, and the PR template asks "did you update or confirm the affected runbooks?", which is the acknowledge step.
4. Testing the steps
- Executable steps: each numbered step maps to a script or
maketarget, and a scheduled pipeline builds a new cluster, seeds sample data, runs the restore steps, and asserts the result (for example that the expected pods are Ready). - Game days (planned drills where a human follows the doc cold) catch what automation cannot: unclear wording, missing permissions, steps that assume tribal knowledge.
- Record time to complete and where the tester hesitated, then fix the doc, not the tester.
Worked example
The nightly pipeline stands up a clean cluster, restores a sample namespace from the newest backup, and asserts three pods reach Ready. One night a Terraform change renames backup_bucket to velero_bucket. The reference check fails the PR that made the rename, because the runbook still uses the old name. The author updates the template in the same PR. Without the check, the mismatch would have surfaced during a real restore.
Trade-offs and pitfalls
- Full automation of every step is expensive. Automate the recovery paths that matter most and rely on game days for the rest.
- Generated docs still need human sentences explaining why and what to do when a step fails. Generate values, hand-write judgement.
- A test cluster can differ from production (size, network, data), so state the gap in the doc rather than implying the drill proves everything.
- Old versions of the runbook must remain reachable for clusters that have not been rebuilt yet.
Tell me about a technical proposal or RFC you wrote for reliability work. How did you structure it, what convinced reviewers, and what would you change now?
Sample Answer
Direct answer
The proposal (an RFC, "request for comments": a written design put out for peer review before building) that I would describe was an illustrative composite: it fixed repeated overload of a shared job queue. I structured it around evidence from real incidents, three options compared on cost and risk, a phased reversible rollout, and success measures agreed up front. It convinced reviewers because it showed the failure reproduced, offered the cheapest fix first, and could be switched off. Today I would establish the baseline earlier and involve the most affected team sooner.
Structured elaboration (situation, task, action)
Situation. Several teams sent jobs to one shared queue. Roughly every two months, one tenant's (one team or customer sharing the platform) failing jobs retried in a tight loop, filled the queue, and delayed everyone else's work. Each time the on-call engineer (the person responsible for responding to alerts out of hours) drained the queue by hand.
Task. I owned the reliability of that queue and had to move the team from firefighting to a fix that other teams would accept, since some of them had to change their retry code.
Action. I wrote the RFC with these sections:
- Problem and evidence: three postmortems (written incident reviews) laid on one timeline, showing the same trigger each time. This made it a pattern, not an anecdote.
- Goals and non-goals: goal, one tenant can no longer delay others. Non-goal, rewriting the queue technology.
- Options: (a) per-tenant rate limits (a cap on how many jobs one tenant may submit per minute), (b) retry backoff with jitter (random delay so retries do not all fire together) plus a dead-letter queue (a holding area for jobs that keep failing), (c) migrate to a different queue product. I recommended a plus b, and rejected c as costing months for a problem we could fix in weeks.
- Rollout: a flag, shadow mode first (log what would be limited without limiting), then one tenant, then all.
- Success measures, defined before rollout: count of overload pages per quarter, and the age of the oldest waiting job. Written with a baseline and a target, for example (illustrative numbers): "baseline: 3 overload pages in the last two quarters, taken from the pager history; target: 0 over the next two; oldest waiting job under 10 minutes." Without the baseline, nobody can say later whether it worked.
- Risks and reviewers' asks: what each dependent team had to change.
Worked example: what convinced reviewers, and the result
- We reproduced it. I replayed recorded incident traffic in staging and showed the queue saturating (filling to capacity so nothing new could get through), then showed it holding with the limits on.
- Cheapest reversible option first. Shadow mode meant nobody had to trust the design; they could see its decisions before it acted.
- I pre-reviewed with the noisiest team. Their objection (a hard limit could drop legitimate bursts) became a design change: a burst allowance (letting a tenant briefly exceed its cap for a short spike, then refilling slowly).
- Result (qualitative): the overload pages stopped recurring in the following quarters, and the on-call runbook (the step-by-step guide an engineer follows when an alert fires) lost its manual-drain step (emptying the queue by hand). I would not quote a precise percentage I cannot reconstruct.
Trade-offs, pitfalls and what I would change now
- Change 1: capture the baseline before writing, not during rollout, so the "before" number is not a reconstruction.
- Change 2: bring the affected team in at the outline stage, not the draft stage. Their burst objection cost a rewrite that an earlier conversation would have avoided.
- Change 3: the first draft ran to eight pages; reviewers replied to the parts they read. A two-page body with an appendix got better comments.
- Change 4: add a date to review the outcome (say, 90 days) so the proposal is judged against its own success measures.
- A pitfall in telling this story: naming virtues ("I collaborated") instead of the specific action. Interviewers remember the timeline overlay and the shadow mode, not the adjectives.
Unlock Full Question Bank
Get access to all 9 Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.