Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
What is the difference between a playbook, a runbook and a README? For each, say who reads it, what it contains, how often it needs maintenance, and a situation where you would choose it over the other two.
Sample Answer
Direct answer. All three are documents, but they answer different questions. A README answers "what is this and how do I start?", a runbook answers "what exact steps do I follow for this task or alert?", and a playbook answers "how do we handle this kind of situation, and who decides what?". Naming conventions differ between companies, so I state my definitions when I write one. In everyday terms: a runbook is an operating instruction sheet, like a recipe you follow exactly; a playbook is a team's game plan, which tells people their roles and when to make a call; a README is the front-door sign of a project.
| README | Runbook | Playbook | |
|---|---|---|---|
| Reader | New developer or user opening the repository | The on-call engineer or operator under time pressure | A team or incident commander (the person coordinating a response) plus other roles |
| Contents | What the project is, install and run, how to contribute, where to find more docs | Step-by-step commands for one procedure (restart a service, fail over a database), with checks and rollback | Higher-level plan for a scenario type: roles, decision points, escalation, communication, links to runbooks |
| Maintenance | Whenever setup or usage changes (on each pull request that alters them); reviewed each release | Every time the procedure or system changes, plus a periodic test (for example, quarterly, by running it) | After every real use or exercise, and at least yearly |
| Choose it when | You want a stranger to run the project in ten minutes | An alert fires and the response is a known, repeatable procedure | The situation needs judgement and coordination, such as a security breach or a regional outage |
Worked example: a payments service.
- README: "Payments API. Requires Python 3.12. Run
make dev. Tests:make test. Docs: link." - Runbook: "Alert: queue depth above threshold. 1. Check the consumer dashboard. 2. Run
kubectl rollout restart deployment/payments-consumer. 3. Confirm the queue drains within 10 minutes. 4. If not, roll back to the previous release (command)." Each step is unambiguous and checkable. - Playbook: "Suspected payment data leak. Incident commander declares severity. Security lead assesses scope. Legal decides on notification. Communications drafts customer message. Decision point: take the service offline? criteria listed." It links to the runbooks for isolating a host.
Choosing between them. If the reader must think, it is a playbook; if the reader must follow, it is a runbook; if the reader is arriving cold, it is a README. A typical mistake is writing a runbook full of judgement calls (it fails at 3am) or a playbook full of shell commands (it goes stale quickly).
You are preparing a technical whitepaper for a public launch of a new hybrid-cloud architecture. How do you structure it for developers, architects and customers, and what editorial and peer-review steps do you run before publication?
Sample Answer
Direct answer. A whitepaper for three audiences works when it is layered: a self-contained front section for customers, a core architecture section for architects, and implementation detail for developers pushed into appendices and linked resources. Before publication it runs a claims audit, a technical review by the people who built the system, a security and legal check, and a copy edit, each with a named sign-off.
Terms. A hybrid-cloud architecture combines on-premises or private infrastructure with a public cloud. A whitepaper is a long-form, authoritative document that argues for an approach. An SME (subject matter expert) is the person who actually knows the area deeply.
More terms. Data residency means a legal or policy requirement that data stays in a specific country or region. A reference architecture is a proven blueprint others can copy. Trust and network boundaries are the lines on a diagram where one security zone ends and another begins. Shared responsibility spells out which security tasks the provider handles and which the customer handles. A context diagram shows the whole system as one box with the people and systems around it; a component diagram then opens the box. The benchmark method is how a performance test was set up, so others can repeat it.
Structure by audience
| Section | Reader | Content |
|---|---|---|
| Executive summary (1 page) | Customers, decision makers | The business problem, what the architecture achieves, who it suits, next step. No jargon. |
| Problem and requirements | All | Why hybrid: data residency, existing hardware, latency to on-premises systems. |
| Reference architecture | Architects | Diagrams (context, then components), the data flow, trust and network boundaries, failure modes, and design trade-offs including where the design is a poor fit. |
| Implementation guidance | Developers | Prerequisites, integration points, sample configuration, links to code and API docs (kept in a versioned repo, not pasted). |
| Security, compliance, operations | Architects, customers' security teams | Shared responsibility, encryption, identity, monitoring. |
| Appendix | Developers | Glossary, detailed configs, benchmark method. |
Each section opens with "who this is for and what you will get", so readers can skip safely.
What the one-page executive summary looks like (illustrative)
Run regulated workloads in the cloud without moving your data out of your data center. Companies in finance and healthcare often cannot move customer records off their own hardware, yet want cloud scale for peaks. This paper describes a hybrid design that keeps sensitive data on-premises and runs burst processing in the public cloud, connected by an encrypted link. It suits teams with an existing data center and a legal requirement to keep records in-country. It is a poor fit if you are starting from scratch with no on-premises systems. Next step: read the reference architecture (section 3) with your architect, or start the 30-minute walk-through in section 4.
It states the problem, the outcome, who it fits (and who it does not), and the next step in about 100 words, with no product jargon.
Editorial and peer-review steps, in order
- Audience and message brief signed by the product owner, before writing. Decide the one claim the paper defends.
- Outline review with two architects, cheap to change now.
- Technical review by SMEs who did not write it, checking every diagram against the real system.
- Claims audit. Every number, comparison and "supports X" statement gets a source or a reproducible method. Benchmarks state the setup. Anything unverifiable gets softened or cut.
- Security and legal/compliance review, including customer names, trademark use and regulatory statements.
- Developer walk-through: someone follows the implementation section on a clean environment.
- Copy edit against a style guide: terminology consistency, defined acronyms, accessible diagrams with alt text.
- Final gate: product owner and engineering lead sign off; the launch date does not override an open claims-audit finding.
- After publication: a named owner, a version number, and a review date, so it can be corrected when the product changes.
If time is short: steps 3 (SME technical review), 4 (claims audit) and 5 (security and legal check) are never cut, because errors there hurt customers or the company. Steps 2, 6 and 7 can be compressed; step 9 is cheap and should not be skipped.
Pitfalls. Marketing claims that engineering has never seen; one voice trying to serve all three audiences in every paragraph; a public paper that reveals unreleased roadmap or internal topology. If the schedule slips, I cut scope (fewer appendices), never the claims audit.
Design a one-page postmortem template that works for both engineering readers and business stakeholders. Then fill in a short example for an outage that affected ten percent of users for three hours.
Sample Answer
Direct answer. A one-page postmortem answers, in plain order: what happened and how bad, why it happened, how we responded, and what we will change. A top block written for business readers (impact, duration, cause in one sentence, status) sits above an engineering block (timeline, root cause, contributing factors, action items). It is blameless: it describes systems and decisions, not individuals.
Template
TITLE / date / severity (how bad it was, on a scale such as SEV1 total outage to SEV3 minor) / author / status (draft, reviewed, closed)
SUMMARY (business readers): impact, duration, one-sentence cause, customer action needed
IMPACT: who, how many, what they experienced, measured how (include uncertainty)
TIMELINE (UTC): detection, escalation, mitigation, resolution
ROOT CAUSE and CONTRIBUTING FACTORS: technical cause, why it was not caught earlier
WHAT WENT WELL / WHAT WENT POORLY / WHERE WE GOT LUCKY
ACTION ITEMS: owner, due date, priority, how we will verify it worked
Example: an outage that affected ten percent of users for three hours (illustrative details)
Summary. On 14 March, 10% of users could not complete checkout for 3 hours (10:05 to 13:05 UTC). Cause: a configuration change reduced the connection limit (the cap on simultaneous connections the application may open to the database, also called pool size) on one of the ten database shards (the database is split into ten slices, each holding a share of users). Fixed by reverting the change. No data was lost. Affected customers were offered a retry link.
Impact. Users are assigned to shards by ID, so one shard of ten means about 10% of users. If we had 400,000 daily active users, that is roughly 40,000 users with failed checkouts; the estimate rests on the assumption that shard membership is even, and we checked actual failed-request counts in logs to confirm the range.
Timeline (UTC).
- 09:58 config change deployed to shard 7.
- 10:05 error rate on shard 7 rises; no alert fires (the alert threshold, the error level that triggers a page, covers the global rate only).
- 10:40 support tickets spike; the on-call engineer (whoever is currently responsible for responding to problems) is paged by support, not by monitoring.
- 11:20 root cause identified; 12:50 revert approved; 13:05 recovery confirmed.
Root cause and contributing factors. The root cause is the single change without which the outage would not have happened; contributing factors are conditions that let it last or go unnoticed. The change set the pool size too low for the traffic on that shard: the new limit was below what the shard needs even off-peak, and the old value had been sized for peak. Contributing: the global error-rate alert averages ten shards, so a 10% failure looks like noise; no canary (staged rollout to a small share first); the change review did not include capacity numbers.
What went well/poorly/luck. Well: once the revert was approved at 12:50, recovery was confirmed 15 minutes later. Poorly, three delays: 35 minutes from first errors (10:05) to the first page (10:40); 40 minutes from the page to an identified root cause (11:20), which is 75 minutes after the errors began; and 90 minutes from root cause to an approved revert (11:20 to 12:50). The approval delay is the biggest single lesson, since it was the longest wait and the fix was already known. Luck: it happened off-peak; at peak the same 10% of users would have seen failures on a larger share of their requests, and the pool would have been exhausted sooner.
Action items.
| Action | Owner | Due | Verified by |
|---|---|---|---|
| Per-shard error-rate alert | SRE | 2 weeks | Fire drill (a practice run that injects fake errors on purpose to see the alert fire) |
| Canary rollout for config changes | Platform | 6 weeks | Next change goes through canary |
| Capacity check in config review template | Engineering manager | 1 week | Template updated |
| On-call may revert a config change immediately, without waiting for approval | Engineering manager | 3 weeks | Written policy, then a timed revert drill |
Variants (same skeleton, different emphasis). Data-quality incident: replace "users affected" with which tables, which dashboards, and which decisions used bad numbers, and add a backfill plan. ML outage or recurring degradation: the impact includes a metric with a range ("precision fell 3 to 5 points, estimated from a labeled sample of 400 predictions"), the timeline includes feature snapshots and drift checks before and after, and the root-cause section separates data drift (inputs changed) from a code or pipeline change. The five-section short form is: what happened, impact, why, what we did, what changes, each opened with one sample sentence like "For 3 hours, 1 in 10 users could not check out."
What makes a runbook usable by a tired on-call engineer in the middle of an outage? Walk me through the practices you follow when you write one for a data pipeline.
Sample Answer
Direct answer
A runbook is a step-by-step operating guide for one specific situation, usually opened from an alert. A tired on-call engineer (the person currently carrying the pager for the team) has almost no spare attention, so a usable runbook needs no interpretation: it starts from the alert they received, gives one action per step with the exact command, shows what a correct result looks like, and says plainly when to stop and call someone. For a data pipeline (a scheduled chain of jobs that moves and transforms data) I write each entry around a symptom such as "table not fresh by 06:00", never around the architecture.
Practices I follow for a pipeline runbook
- One entry per alert, linked from the alert itself. The page that fires should contain the URL of its entry, so nobody searches a wiki at 3 a.m.
- Symptom first, then a short decision path. Step 1 answers "is it still running, failed, or never started?" and each answer points to a numbered step. No background essays inside the steps.
- Exact commands and exact locations. Say where the command runs (a jump host, meaning a locked-down server you log into first to reach production, or a laptop with a named cluster context, the setting that says which cluster a command talks to) and paste it ready to copy, with placeholders in angle brackets.
- State the expected result after each step, so the reader knows whether to continue.
- Say whether an action is safe to repeat. Re-running a load is only calm if the job is idempotent (running it twice leaves the same data as running it once). Write "safe to repeat: yes, the load merges on order_id" or "NOT safe: creates duplicates, ask the owner first".
- Name the blast radius (how much else is affected if this breaks or your fix goes wrong): which dashboards, models or teams read this table, and who to notify.
- Escalation is a step, not an afterthought: after 30 minutes without progress, page the owning team, with the channel and rotation name (the rotation is the schedule that decides who is on call).
- Known failure table: last three real incidents as "log line seen, cause, fix", because pipelines fail in repetitive ways.
Worked example (illustrative excerpt)
ALERT: orders_daily table not updated by 06:00 UTC
Owner: data-platform-oncall Last verified: 2025-06-02 by a.rivera
Impact: finance dashboard and churn model read this table.
1. Is the run still going? (run from the jump host)
airflow dags list-runs -d orders_daily --state running
Rows returned -> wait 30 min, then go to step 4.
No rows -> go to step 2.
2. Open the DAG's runs in the orchestrator UI.
No run at all for today (it never started) -> check the DAG is not paused and the
scheduler is healthy, then go to step 4
A red (failed) run: open the failed task's log and match the last line:
"Connection refused ... source-db" -> source outage, go to step 4
"duplicate key" -> STOP, do not re-run, escalate
3. Clear (re-run) the failed task load_orders. Safe to repeat: yes (MERGE on order_id).
Expected: task turns green in about one normal run time.
4. Escalate to data-platform-owners if not green after 30 minutes.
Reading the excerpt: Airflow is a scheduler that runs a chain of jobs in order, and a DAG (directed acyclic graph) is one such chain of tasks with a fixed order and no loops. airflow dags list-runs -d orders_daily --state running lists the runs of the orders_daily DAG that are in progress right now. The orchestrator UI is Airflow's web page showing each task's status and logs. To "clear" a failed task means to tell Airflow to run it again. MERGE on order_id is a SQL statement that updates rows that already exist and inserts new ones, matched by order_id, which is why repeating it does not create duplicates.
Fields I validate before publishing (and re-validate on a schedule)
- A named owner that is a team or rotation, not a person who may leave.
- A last-verified date and who verified it, with a review interval (for example 90 days) after which the page is flagged stale.
- Tested commands: someone other than the author ran every command against staging or a copy and the output matched what the step promises.
- Links resolve, and the alert's runbook link points to this exact page.
Pitfalls
- Paragraph prose instead of numbered steps.
- "Check the logs" without saying which logs and what to look for.
- Commands that only work with context the author carries in their head.
- A runbook that has never been used under pressure. Run a drill and fix every place the reader hesitated.
Design the pull-request checks for a repository of SRE runbooks: dead links, required front-matter keys and Markdown linting. Which tools would you use, how do failures surface, and how do you let an emergency runbook fix through during an incident without gutting the checks?
Sample Answer
Direct answer
Treat the runbook repository like code: every pull request (PR, a proposed change awaiting review) runs three automated checks in continuous integration (CI, the robot that runs on each PR). A link checker (lychee) catches dead links, a small script validates required front matter (the YAML metadata block at the top of each file, such as owner and last-reviewed date), and markdownlint-cli2 enforces Markdown style. Failures show up as inline annotations on the PR and as a red required check. For emergencies, an on-call engineer adds an emergency-fix label. That label makes only the checks that are cosmetic (lint) or depend on outside websites (external links) non-blocking, keeps the check that protects accountability (front matter) blocking, and automatically opens a follow-up issue so the debt is paid rather than forgotten.
Terms in plain words
- GitHub Actions: GitHub's built-in automation. A YAML file in
.github/workflows/tells GitHub what commands to run on each PR. - Required status check / branch protection: branch protection is a repository setting on the main branch. Marking a check as "required" there means GitHub refuses to merge a PR until that check is green.
- Triage rights: a GitHub permission level. Only people with at least that level (or higher) on the repository can add or remove labels, so the label cannot be applied by a random outsider.
- Inline annotation: an error message drawn on the exact file and line in the PR's Files tab.
continue-on-error: a step setting that lets the step fail without turning the whole job red.
Tools and what each one guards
| Check | Tool | Why this one |
|---|---|---|
| Dead links | lychee (a fast link checker that reads Markdown directly, with retries and caching) | Broken links to dashboards and other runbooks are the most common way a runbook fails someone at 3 a.m. |
| Required front matter | ~30-line Python script with PyYAML (a Python library that reads YAML) | Rules are specific to your team (owner, service, last_reviewed), so a small script beats a generic tool |
| Markdown style | markdownlint-cli2 (a linter: it flags style problems such as skipped heading levels or trailing spaces) | Consistent headings and numbered lists make steps skimmable under stress |
The workflow (GitHub Actions syntax, parsed as valid YAML)
name: runbook-checks
on:
pull_request:
types: [opened, synchronize, reopened, labeled] # 'labeled' re-runs checks when the emergency label is added
# No `paths:` filter on purpose: a required check that is skipped by path filtering stays "Pending" and blocks the PR
schedule:
- cron: "0 6 * * 1" # weekly full link crawl catches link rot nobody's PR touched
permissions: { contents: read, issues: write, pull-requests: read }
jobs:
docs-checks: # the ONE job name set as a required status check
runs-on: ubuntu-latest
env:
EMERGENCY: ${{ contains(github.event.pull_request.labels.*.name, 'emergency-fix') }}
steps:
- uses: actions/checkout@v4
- name: Front matter (always blocking, because no owner means nobody is accountable)
run: pip install pyyaml==6.0.3 && python scripts/check_front_matter.py $(git ls-files 'runbooks/*.md')
- name: Markdown lint
continue-on-error: ${{ env.EMERGENCY == 'true' }}
run: npx --yes markdownlint-cli2@0.23.3 "runbooks/**/*.md"
- name: Links
continue-on-error: ${{ env.EMERGENCY == 'true' }}
uses: lycheeverse/lychee-action@v2
with:
args: --no-progress --max-retries 2 --accept 200,429 "runbooks/**/*.md"
fail: true
- name: Open follow-up issue when the emergency path was used
if: env.EMERGENCY == 'true'
env: { GH_TOKEN: "${{ github.token }}" }
run: gh issue create --title "Post-incident doc cleanup for PR ${{ github.event.pull_request.number }}" --label docs-debt --body "Lint and link checks were non-blocking on this PR. Re-run and fix within 2 business days."
Reading the workflow step by step
- The workflow starts when a PR is opened, updated, reopened or labeled, on every PR (there is deliberately no
paths:filter, see the pitfalls). A weekly schedule also runs it. permissionsgives the job read access to the code and write access to issues, so it can open the follow-up ticket, and nothing more.EMERGENCYistrueonly when the PR carries the labelemergency-fix(thecontains(...)expression asks "is that name in the PR's label list?").- Front matter always runs and always blocks: if it fails, the job is red.
- Lint and links have
continue-on-errorset to that flag: normally a failure turns the job red, but on an emergency PR the failure is still printed and the job stays green. - The last step runs only when the flag is
trueand files the cleanup issue.
The front-matter check, run for real
import os, sys, pathlib, tempfile
import yaml
REQUIRED = ["title", "owner", "service", "last_reviewed"]
def check(path, text):
"""Return a list of problems for one runbook."""
text = text.replace("\r\n", "\n") # files saved on Windows use CRLF line endings
if not text.startswith("---\n"):
return ["no front matter block"]
end = text.find("\n---", 4)
if end == -1:
return ["front matter block is never closed"]
try:
meta = yaml.safe_load(text[4:end]) or {}
except yaml.YAMLError as e:
return [f"front matter is not valid YAML: {e.__class__.__name__}"]
if not isinstance(meta, dict):
return ["front matter is not a set of key: value pairs"]
# a key that is present but empty (`owner:`) is as useless as a missing one
return [f"missing key: {k}" for k in REQUIRED if meta.get(k) in (None, "")]
def main(paths):
failures = 0
for p in paths:
for problem in check(p, pathlib.Path(p).read_text()):
print(f"::error file={p}::{problem}")
failures += 1
return 1 if failures else 0
if __name__ == "__main__":
if len(sys.argv) > 1:
sys.exit(main(sys.argv[1:]))
with tempfile.TemporaryDirectory() as d:
os.chdir(d)
good = pathlib.Path("db-failover.md")
good.write_text("---\ntitle: DB failover\nowner: team-data\nservice: orders-db\nlast_reviewed: 2026-01-05\n---\n# Steps\n")
bad = pathlib.Path("cache-flush.md")
bad.write_text("---\ntitle: Cache flush\nservice: edge-cache\n---\n# Steps\n")
empty = pathlib.Path("dns-cutover.md")
empty.write_text("---\ntitle: DNS cutover\nowner:\nservice: dns\nlast_reviewed: 2026-02-01\n---\n# Steps\n")
crlf = pathlib.Path("vpn-reset.md")
crlf.write_bytes(b"---\r\ntitle: VPN reset\r\nowner: team-net\r\nservice: vpn\r\nlast_reviewed: 2026-03-01\r\n---\r\n# Steps\r\n")
code = main([str(good), str(bad), str(empty), str(crlf)])
print("exit code:", code)
Running it prints:
::error file=cache-flush.md::missing key: owner
::error file=cache-flush.md::missing key: last_reviewed
::error file=dns-cutover.md::missing key: owner
exit code: 1
The third file has owner: with nothing after it, which YAML reads as null. A plain k not in meta test would let it pass, defeating the rule that "no owner means nobody is accountable", so the check tests for an empty value too. The fourth file (Windows line endings) passes because the script normalises \r\n first.
The ::error file=...:: format is what makes GitHub draw the message on the exact file in the PR's Files tab, so the author does not have to open CI logs.
How failures surface
- One job name (
docs-checks) is marked as a required status check in branch protection, so a red result blocks merge. - Each tool writes file-level annotations, so the fix is visible where the author is already looking.
- Link checks accept HTTP 429 (rate limited) and retry twice, because a flaky third-party site should not fail the PR.
- The weekly scheduled run crawls everything, because links rot even when nobody edits the runbook.
Emergency path without gutting the checks
- Only people with triage rights can apply the label. GitHub itself enforces this (the workflow does not need to), so it is not a self-service bypass for anyone.
- Front matter stays blocking: it is local and takes seconds, but the real reason is that an emergency runbook with no owner or review date is the one nobody will fix later, which is worse than a lint warning.
- Lint and external links become non-blocking (
continue-on-error). Lint findings are cosmetic and cannot mislead a reader, and link checks depend on third-party sites that may be down or rate-limiting during the same incident, so a red result there does not mean the runbook is wrong. Lint is not slow; the point is that a style complaint should never delay a fix. - The bypass is not a branch-protection override or a disabled workflow. The checks still run and still report, so the author sees exactly what to clean up.
- The last step files a
docs-debtissue with a two-business-day target, and a weekly report counts how many emergency labels were used. A rising count is a signal to fix the checks or the on-call process.
Worked example
At 02:10 on-call needs to change step 3 of a failover runbook. They open a PR, add emergency-fix, and a teammate approves. Front matter passes, lint fails on a trailing-space rule and is shown but non-blocking, and the merge happens in minutes. The follow-up issue lands in the team queue the same night, and the author fixes the lint error next morning.
Trade-offs and pitfalls
- Making every check non-blocking during emergencies is tempting but teaches people to always use the label. Keep the blocking set as small as the cheapest useful checks.
- Do not run link checks against private internal URLs without an allowlist or auth, or every PR fails. Exclude those patterns explicitly and verify them another way.
- Pin the action and tool versions so a tool upgrade does not turn every PR red on the same morning.
- A required check whose workflow is skipped by a
paths:filter stays "Pending" forever (GitHub documents this), so a PR touching only a README would be unmergeable. Either drop the filter, as above, or add a tiny always-running job that reports the same check name. gh issue create --label docs-debtfails if that label does not exist, which would turn the emergency path red. Create the label once up front, and remember each new push to an emergency PR re-runs the step, so de-duplicate (for example search for an open issue with that title first).- The workflow file itself is part of the PR, so an author could edit it to weaken the checks. Put
.github/workflows/under CODEOWNERS with "Require review from Code Owners" enabled. - If the label is used often, the real problem is that the checks are too slow or too strict for normal work.
Unlock Full Question Bank
Get access to all 15 Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.