Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
You are setting up review for documentation pull requests in a repository. What would you check, what would you automate, and how do you keep the review from becoming a rubber stamp?
Sample Answer
Direct answer
Split the review by who is better at each check. Machines check what is mechanical and repeatable: links, formatting, spelling, snippets that must run, metadata, secrets. People check what only a person can judge: is it true, is it complete for the reader's task, does it fit the audience. To stop human review becoming a rubber stamp (approving without really looking), make the reviewer produce evidence of what they did, keep changes small, and audit outcomes, not approvals.
What is automated versus checked by a person
| Check | Who | Blocks the merge? |
|---|---|---|
| Broken internal links, formatting, spelling and style rules | CI (automatic checks on every pull request) | Yes |
| Code snippets and commands run in a clean environment | CI | Yes |
| Required front matter (metadata block at the top of the page: owner, audience, last verified date, version) | CI script | Yes |
| Secrets and personal-data scan | CI | Yes |
| External links | CI, allowed to warn only | No (they flake: they fail sometimes because the other site is slow or down, not because your page is wrong) |
| Accuracy against the real system | Subject-matter owner (the person who knows that system best), listed in CODEOWNERS (a file mapping paths to owners; on GitHub it only requests their review, and it blocks the merge only when the branch protection rule "Require review from Code Owners" is switched on) | Yes, once that rule is on |
| Complete for the reader's task: preconditions, verify step, rollback | Second reviewer | Yes |
| Audience and clarity | Second reviewer | Advisory (a comment the author may act on, but it does not stop the merge) |
For ML artifacts, add: a model card (a short document stating a model's purpose, training data, evaluation and limits), data provenance (where the data came from and what it may be used for), and monitoring readiness (what alerts if the model degrades and who owns it). Missing any of these blocks the merge, because each covers a risk that cannot be fixed after readers rely on the page: a model with no stated limits gets misused, data with unknown origin may not be legal to use, and a model with no alert or owner degrades unnoticed. A monitoring-readiness check looks like this: "The page names the alert (for example accuracy drops below the stated floor for 3 days), the dashboard link, and the on-call owner; the reviewer opened the dashboard and saw it."
Keeping review honest
- Evidence line: for procedure docs the reviewer writes "Followed steps 1 to 6 on a clean environment; step 4 needed X" in the review. "LGTM" (looks good to me) is not accepted for pages that instruct actions.
- Required approvals by ownership, so the approver is someone who knows the system.
- Small pull requests, since a 30-page change gets skimmed.
- Outcome measures: how often a merged page needs a correction within a month, and how many reviews found anything. A reviewer who never finds anything is a signal to look, not a compliment.
- Occasional audit: a monthly sample re-checked by someone else.
Worked example
A PR updates a database-restore runbook (a step-by-step procedure for an operational task). CI passes links and formatting. The owner approves accuracy. The second reviewer actually runs the restore in a test environment and posts: "Step 3 fails because the snapshot name flag changed." That is the catch a rubber stamp would have missed.
Trade-offs and pitfalls
- Too many blocking checks slow contributors and encourage bypassing. Keep flaky checks advisory.
- Evidence lines can be copied. Audits and outcome measures are the backstop.
- What would change my call: for a tiny team, one owner approval plus CI is enough.
Documentation pull requests on your team keep stalling because nobody owns the examples and sample data. What changes to roles, templates and checks would you make?
Sample Answer
Direct answer
Documentation PRs (pull requests, the proposed changes reviewers approve before they merge) stall on examples because an example has three parts, prose, code and sample data, and only the prose has an obvious reviewer. I would fix it with three changes: give each example a named owner (the person who ships the feature owns its example), change the PR template so an example is a required field rather than an afterthought, and add automated checks so a reviewer never has to judge whether an example still works.
1. Roles: who owns what
- Feature author owns the example. Writing the example is part of the definition of done for the feature PR (the checklist that must be satisfied before work counts as finished), not a follow-up ticket. The author knows the real behaviour and the real edge cases.
- Docs owner (a rotating weekly role, not a full-time job) owns unblocking. A rota is a rotating schedule of who takes the duty each week. If a docs PR has had no review for 2 business days, the rota person reviews it or reassigns it. Stalling becomes someone's job.
- Sample-data steward per shared fixture area. A steward is a named person responsible for keeping something in good shape; a fixture is fixed sample data used by examples and tests. Shared sample data (the fake customers and orders used across many pages) lives in one directory, for example
docs/samples/, with a named steward listed in a CODEOWNERS file (a repository file that maps paths to the people or teams whose approval is required).
/docs/samples/ @org/api-team
/docs/guides/ @org/docs-rota
2. Templates: make the gap visible without inviting box-ticking
Add fields, not checkboxes, to the PR template: "Example that demonstrates this change (file path)", "Where the sample data comes from (generated by script, or hand-written and why)", and "Contains no real customer data (yes, and how you know)". A path or a sentence is harder to fake than a tick.
3. Checks: CI (continuous integration, the automatic checks run on every PR) does the judging
- Run every example file, or every fenced code block (a code sample wrapped in triple backticks in Markdown) in the guides, in CI. A broken example fails the PR before a human reads it.
- Generate sample data from a script with a fixed random seed (a fixed starting value so the output is identical every run) instead of editing JSON by hand. CI regenerates it and fails if the committed file differs.
- Validate sample JSON against the API's schema (a machine-readable description of which fields a response must have, and their types), and run a secrets and personal-data scan (a tool that searches files for things like API keys, passwords or real customer emails) over
docs/samples/.
Worked example
A "Add pagination guide" PR sits for three weeks. The reviewer's only comment is "the sample response looks old, who has current data?" Nobody answers because nobody owns it. After the changes: the author of the pagination feature had to attach the example path in the PR, the sample response is produced by make samples (a command that runs the seeded generator script and rewrites the sample files) from the seeded generator, CI validates it against the schema, and the rota person merged it on day 2. The reviewer only had to read the prose.
Trade-offs and pitfalls
- Every extra check slows contributors. Start with the two or three most-used examples and widen only when failures show up.
- Ownership by the feature author fails for cross-team examples. Give those a steward in CODEOWNERS, otherwise you recreate the orphan.
- Do not let the docs owner become the person who writes everyone's examples. That moves the bottleneck instead of removing it.
- What would change my call: on a team of three, skip the rota and just make the feature author and one named reviewer responsible. Measure success by median PR age (the middle value of the ages when sorted: if five docs PRs were open 3, 9, 14, 21 and 30 days the median is 14; after the change, 1, 2, 3, 4 and 20 days gives a median of 3, and the one slow outlier does not hide the improvement) and by how often an example breaks after merge, not by how many boxes were ticked.
Your codebase has thin documentation and new hires take weeks to become productive. Propose a prioritised documentation plan for the first three months, including how you decide which documents to fix first and how you measure the effect.
Sample Answer
Direct answer
Do not try to document everything. Spend the first weeks finding the few documents that, if they existed, would remove the most repeated questions and the most days from a new hire's ramp-up, write those first, and measure onboarding time before and after. Plan in three months: discover and fix the top items (month 1), fill the core set with owners (month 2), and make upkeep routine and measured (month 3).
How to decide what to fix first
- Ask the people who just joined (the last three or four hires): where did you get stuck, what did you ask twice, which question did you answer for the next person? Their answers are better evidence than a guess.
- Mine repeated questions in the team chat channel and support threads, and note which systems get the most "how do I..." asks.
- Score each candidate by (how many people it blocks) times (how often) times (how painful when missing) divided by (effort to write). Write the highest scores first.
Worked example (rate people, frequency and pain 1 to 5, effort in days): setup README 5 x 3 x 5 = 75, divided by 2 days = 37.5; "how we ship" guide 3 x 4 x 5 = 60, divided by 3 = 20; architecture overview 4 x 3 x 4 = 48, divided by 5 = 9.6; a rarely used legacy billing-job page 1 x 1 x 5 = 5, divided by 4 = 1.25. Order: README, ship guide, architecture, billing page last. - Default first set for a codebase: a README that gets a laptop to a running app, an architecture overview, a "how we ship" guide, then service pages.
Three-month plan (with owners and check-ins)
| Period | Deliverables | Owner | Check-in |
|---|---|---|---|
| Days 1-30 | Interview 4 recent hires; rank the gaps; ship the setup README and "first-week checklist"; baseline measurement | Doc lead (the person driving this, possibly a new hire) | Weekly 15 minutes with the engineering manager |
| Days 31-60 | Architecture overview, service pages for the top services, runbooks for the top alerts; each doc gets a named owner | Each service's owning team | Bi-weekly review; demo to a fresh hire |
| Days 61-90 | Upkeep: PR template prompt, "last verified" dates, a quarterly review; retire dead pages; second measurement | Engineering manager | End-of-plan review of the metrics |
The 30/60/90 framing (goals for days 30, 60 and 90) also suits a new hire paying down doc debt (the pile of missing or stale documentation that builds up like unpaid bills): they are the ideal tester, because they discover the gaps as they hit them, and each fix is a visible deliverable.
Optional variant: if the newcomer joins an AI project, first two weeks (short checklist)
- A README that runs the project end to end.
- A model card (a short document stating what the model is for, how it was trained and evaluated, and its limits).
- A data schema (columns, meaning, source, freshness).
- An experiment reproducibility guide (which code, data snapshot, environment and seeds reproduce the baseline result).
The payoff is fewer interruptions for senior people and a faster first useful contribution.
Public versus internal split (optional, secondary to the core plan)
Internal docs (runbooks, architecture, incident history) are for engineers and can be blunt. Public engineering content (blog posts, API docs, talks) should reflect the real platform, so derive it from the internal source of truth and review it for accuracy, rather than writing a separate, prettier version that drifts. Example: the internal runbook says "restart with svcctl reload, then check the queue depth dashboard"; the public API guide keeps the same reload behavior and error codes but drops internal hostnames and the dashboard link. Each part gets its own owner and timeline, and its own metrics.
How to measure the effect
- Ramp time (how long until a new hire is productive): days from start date to first merged PR, and to first on-call shift (their first turn being the engineer who responds to alerts). Compare hires before and after.
- Questions to mentors counted for each new hire (or the fraction of onboarding questions answered by a link).
- Docs feedback: a simple "did this help" at the bottom of each page, and a survey of each new hire at week 4.
- Treat before/after comparisons with caution: hires differ in seniority, so compare like with like and report the count of hires.
Pitfalls
Writing many pages nobody needs; giving no owners, so they rot; and reporting page counts as success.
Design an automated process that detects stale engineering documentation and prompts owners to update it. Which signals would you use, how do you map documents to owners, and how do you keep false positives and alert fatigue down?
Sample Answer
Direct answer
Build a scheduled job that gives every document a staleness score from several signals, maps each document to an owner, and sends owners a short, capped, weekly digest instead of one alert per document. No single signal is trustworthy (an old doc may still be correct, a fresh edit may have been a typo fix), so combine them, tune thresholds so precision beats recall (nudges you send should nearly always be deserved, even if that means missing some stale docs), and measure the outcome so the system earns trust. The goal is fewer, better nudges.
Terms in plain words
- False positive: the system flags a doc that is actually fine. False negative: it misses a doc that really is wrong.
- Precision: of the docs you flagged, the share that truly needed attention. Flag 10, 8 needed work: precision is 80%.
- Recall: of all docs that truly needed attention, the share you flagged. If 20 needed work and you caught 8, recall is 40%.
- "Precision beats recall" means preferring fewer, more accurate flags over catching everything.
- Embedding similarity: turning text into lists of numbers so that similar meaning gives similar numbers. It is a heavier alternative to matching names, and the worked code below does not use it.
- Normalised to 0-1: rescaled so 0 is "no problem" and 1 is "worst case". Saturates means it stops growing at 1 (a doc older than a year scores the same as one exactly a year old).
Signals and what each really tells you
| Signal | What it measures | Weakness |
|---|---|---|
| Doc last-commit date (age) | Time since anyone touched it | A stable doc is old but fine |
| Code drift | Commits to the code the doc describes since the doc last changed | Needs a doc-to-code mapping |
| Similarity to recent code diffs | Do names in the doc (backticked commands, flags, config keys, function names) appear in the removed or renamed lines of recent diffs to the same component? Embedding similarity is a heavier alternative | Regex is cheap but misses prose changes |
| Broken references | Dead links, deleted files, renamed symbols | Only catches concrete breakage |
| Access logs | Page views in the last 90 days | Popular does not mean accurate, unread does not mean fine |
| Author and owner activity | Is the owner still committing to the repo, or still employed on the team | A quiet owner may just be stable |
Age alone is the weakest signal. Drift plus broken references is the strongest, because it says something concrete changed.
Scoring, weights and thresholds (run with pinned data)
Each signal is normalised to 0-1, then combined with weights (drift 0.35, age 0.30, broken references 0.25, owner inactive 0.10). Traffic does not raise the score. It decides the action: unread stale docs go to archive review, read ones go to an owner.
from datetime import date
TODAY = date(2026, 9, 1) # pinned so the run is reproducible
# name, last_doc_edit, code_commits_to_linked_paths_since, broken_refs, views_90d, owner_active
DOCS = [
("payments-runbook", date(2025, 3, 1), 18, 3, 420, True),
("onboarding-guide", date(2026, 7, 20), 2, 0, 900, True),
("legacy-batch-job", date(2024, 1, 10), 1, 1, 0, False),
("api-gateway-setup", date(2026, 2, 1), 12, 1, 150, False),
("style-guide", date(2025, 6, 1), 0, 0, 60, True),
]
def score(last_edit, drift, broken, owner_active):
age_n = min((TODAY - last_edit).days / 365, 1) # 0..1, saturates at one year
drift_n = min(drift / 20, 1) # 20+ code changes since the doc moved = max
broken_n = min(broken / 3, 1) # dead links or renamed symbols the doc mentions
gone = 0 if owner_active else 1
return round(0.30 * age_n + 0.35 * drift_n + 0.25 * broken_n + 0.10 * gone, 2)
def action(s, views):
if s >= 0.40 and views == 0:
return "archive review"
if s >= 0.60:
return "nudge owner"
if s >= 0.40:
return "weekly digest"
return "ignore"
for name, edit, drift, broken, views, active in DOCS:
s = score(edit, drift, broken, active)
print(f"{name:18} score={s:.2f} views={views:4} -> {action(s, views)}")
It prints:
payments-runbook score=0.86 views= 420 -> nudge owner
onboarding-guide score=0.07 views= 900 -> ignore
legacy-batch-job score=0.50 views= 0 -> archive review
api-gateway-setup score=0.57 views= 150 -> weekly digest
style-guide score=0.30 views= 60 -> ignore
Tracing one score by hand (payments-runbook, run on 2026-09-01):
age = 549 days / 365 = 1.50, capped at 1 -> 0.30 x 1 = 0.300
drift = 18 commits / 20 = 0.90 -> 0.35 x 0.9 = 0.315
broken = 3 refs / 3 = 1.00 -> 0.25 x 1 = 0.250
owner active, so gone = 0 -> 0.10 x 0 = 0.000
total = 0.865
The hand total is exactly 0.865 but the program prints 0.86. That is floating-point rounding: computers store 0.865 as 0.86499999999999999..., so round(..., 2) rounds down. The action does not change (0.86 and 0.87 are both above the 0.60 line), but expect this kind of one-cent difference when you check scripts by hand.
Two thresholds are in play. 0.60 is the alerting line (a direct nudge to the owner). The 0.40 to 0.59 band is not an alert: those docs only appear in the capped weekly digest, and if nobody reads them they go to archive review instead.
Reading the result: payments-runbook is old, has 18 code changes since, three broken references, and is viewed often, so it gets a direct nudge. onboarding-guide is heavily read but fresh, so nothing happens (popularity alone never triggers a nudge). legacy-batch-job scores 0.50 but nobody reads it, so the right action is archive review, not an update. api-gateway-setup is in the middle band and its owner is inactive, so it goes to the digest and reassignment.
Mapping documents to owners (fallback chain)
ownerin the doc's front matter (the metadata block at the top of the file).- The CODEOWNERS entry (file that assigns responsible people or teams by path) for the doc.
- The owners of the code paths the doc links to.
- The last substantive author (skip bulk reformat commits).
- The team channel as a last resort, and a person is never left unassigned.
If the owner has left, reassign automatically to their team and say so in the nudge.
False positives, false negatives and alert fatigue
- A false positive is nudging an owner about a doc that is fine. A false negative is missing a doc that is wrong. Here false positives cost more, because each wasted nudge teaches people to ignore the next one, so set the alerting threshold high (0.60) and accept some misses.
- Cap the digest (for example five items per owner per week), highest score first, and never nudge the same doc twice in 30 days.
- Give every nudge two one-click answers: "Still accurate" (records a review date and resets the clock) and "Snooze 60 days". Both feed the score.
- Add a review-by date in front matter for docs where the owner knows they are stable.
- Track precision: of the nudges sent, what fraction led to an edit or a "still accurate" confirmation versus being ignored? If ignored nudges climb, raise the threshold or drop a signal.
Review workflow for candidates
Weekly job writes the candidate list to a dashboard and posts each owner's digest. Owner triages within two weeks: update, confirm, snooze, or archive. Untouched items escalate once to the team lead, not the whole channel. Archived docs get a banner and redirect rather than deletion.
Trade-offs and pitfalls
- Weights are judgement calls. Start simple, sample twenty flagged docs by hand, and adjust from what you find.
- Access logs can be gamed or skewed by bots, so use them to prioritise, never to declare a doc correct.
- Do not auto-edit or auto-archive without a human step, because a wrong automatic change destroys trust faster than a missed nudge.
Design a README-driven development process for your team to improve documentation quality. How do you enforce updates in pull requests, and how do you avoid it becoming box-ticking?
Sample Answer
Direct answer
README-driven development means writing the README (the front-page document explaining what a project does and how to use it) before the code, as a lightweight spec: if you cannot explain the feature simply, the design is not ready. To keep docs current I would combine a CI (continuous integration, automatic PR checks) rule that ties code changes to docs changes, an executable README so wrong instructions fail a build, and human review that reads the docs diff against the code diff. To avoid box-ticking, every waiver (an explicit, recorded exception to the rule) needs a real reason, and a monthly audit checks the docs the way a new joiner would use them.
The process
- Design PR first. A new feature starts as a pull request (PR, a proposed change awaiting review) that changes only the README or a docs page: what it does, the command or API call, one example, and what it does not do. Reviewers argue about behaviour here, when changes are cheap.
- Implementation PR second. The code PR must keep that documented behaviour true, or change the docs in the same PR.
- Enforcement in the PR. The CI rule is the script below: it fails any PR that touches public code without touching the docs, unless the author supplies a reason.
- Executable README. CI extracts the quickstart commands from the README and runs them in a clean container (a fresh, disposable environment with nothing left over from earlier runs, so hidden setup cannot make broken instructions look like they work). If the instructions stop working, the build breaks.
Enforcement script (tested in a scratch repo)
#!/usr/bin/env bash
# Fail a pull request that changes public code without touching docs, unless the PR says why not.
# Usage: check_docs.sh <base-ref> [skip-reason]
set -euo pipefail
base="$1"; reason="${2:-}"
changed=$(git diff --name-only "$base"...HEAD)
code=$(echo "$changed" | grep -E '^src/' || true)
docs=$(echo "$changed" | grep -E '^(README\.md|docs/)' || true)
if [ -n "$code" ] && [ -z "$docs" ]; then
if [ ${#reason} -ge 15 ]; then echo "docs skipped, reason: $reason"; exit 0; fi
echo "src/ changed but no README/docs change. Update docs or give a real reason (15+ chars)."; exit 1
fi
echo "docs check passed"
Running it against a PR that changes src/cli.py only:
$ ./check_docs.sh main
src/ changed but no README/docs change. Update docs or give a real reason (15+ chars).
$ ./check_docs.sh main "n/a"
src/ changed but no README/docs change. Update docs or give a real reason (15+ chars).
$ ./check_docs.sh main "internal refactor, no behaviour change"
docs skipped, reason: internal refactor, no behaviour change
After a commit that adds a README section, the same command prints docs check passed.
Reading the script line by line
set -euo pipefail: stop at the first error (-e), treat an unset variable as an error (-u), and fail a pipeline if any part of it fails (pipefail).reason="${2:-}": the second argument, or an empty string if none was given.git diff --name-only "$base"...HEAD: list the names of files changed on this branch since it split frombase. The three dots compare against the common ancestor, so commits made onmainafter you branched do not show up.grep -E '^src/'keeps only lines that start withsrc/(-Eenables extended regular expressions).|| trueis needed becausegrepexits with status 1 when nothing matches, which underset -ewould kill the script.${#reason} -ge 15: the reason's length is at least 15 characters.
Be honest about that last check: it only stops lazy waivers such as n/a. internal refactor, no behaviour change passes it, as the trace shows, whether or not it is true. The length rule is a speed bump; the real defence is a reviewer reading the reason and the team watching the waiver rate.
Worked example
The team adds a --retries flag to a command-line tool. PR 1 adds a "Retries" section to the README: default of 3, one example command, and a note that retries apply only to network errors. Reviewers say "network errors only is surprising, should timeouts count?" and the answer is settled before any code exists. PR 2 implements it, and CI runs the README example command.
Avoiding box-ticking
- The waiver needs a reason of substance, and the reviewer must agree with it. A repeated "n/a" gets discussed, not waved through.
- Reviewers are told to check the docs diff against the code diff, not just that a docs diff exists.
- Monthly fresh-eyes audit: pick five merged PRs, and someone who has not seen the feature follows the README cold. Every place they get stuck becomes a fix.
- Watch the waiver rate. If 60 percent of PRs use it, the rule is wrongly aimed (too broad a path filter, meaning the rule that decides which changed paths must come with docs) or being gamed. Tighten the paths to public interfaces only.
Trade-offs and pitfalls
- Path rules only detect that a docs file changed, not that it is right. That gap is why the executable README and human review matter.
- README-first can be slow for tiny changes. Limit the design-PR step to user-visible behaviour, and let internal refactors skip it with a stated reason.
- Executable docs need a clean environment. Flaky setup steps (ones that pass on some runs and fail on others) will train people to ignore red builds.
Unlock Full Question Bank
Get access to all 45 Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.