Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
A model starts returning NaN predictions after a library upgrade. Write the bug report you would file, including how you would shrink the problem to a minimal reproduction.
Sample Answer
Direct answer
A model (a program that turns input features, the columns of numbers describing each case, into predictions) is returning NaN (not-a-number, the floating-point value produced by invalid arithmetic) after a library upgrade. My report has to answer four things: what broke, what changed, the smallest way to reproduce it, and how bad it is. The method for shrinking is to freeze the inputs, then repeatedly halve the rows, keeping whichever half still fails, until one row remains, and then drop columns one at a time, keeping each drop that still fails (with only four columns, halving them is not worth the code).
Step 1: confirm and pin
- Check it is deterministic and record the old and new library versions (a comparison of the two environments' package lists shows exactly what changed).
- Freeze one input batch to a file so nothing else varies.
Step 2: shrink with a script
Below, the two predict modes stand in for the old and new library behaviour (old fills missing values with the column mean, new passes them through). In a real report you would replace them with two environments.
import numpy as np
rng = np.random.default_rng(7)
X = rng.normal(size=(512, 4))
X[301, 2] = np.nan # one row has a missing value in feature 2
w = np.array([0.5, -1.0, 0.25, 2.0])
def predict(batch, weights, fill_missing):
batch = batch.copy()
if fill_missing: # old behaviour: replace NaN with the column mean
means = np.nanmean(batch, axis=0)
idx = np.where(np.isnan(batch))
batch[idx] = means[idx[1]]
# batch @ weights = weighted sum per row; 1/(1+exp(-z)) squashes it to a 0-1 probability
return 1 / (1 + np.exp(-(batch @ weights)))
def fails(rows, cols, fill_missing=False):
return bool(np.isnan(predict(X[np.ix_(rows, cols)], w[cols], fill_missing)).any())
ALL = [0, 1, 2, 3]
print("old (fills) has NaN:", fails(list(range(512)), ALL, True))
print("new (no fill) has NaN:", fails(list(range(512)), ALL))
rows = list(range(512)) # shrink rows: keep whichever half still fails
while len(rows) > 1:
half = rows[: len(rows) // 2]
rows = half if fails(half, ALL) else rows[len(rows) // 2 :]
print("minimal failing row index:", rows)
cols = list(ALL) # shrink columns: drop one at a time, keep the drop if it still fails
for c in ALL:
trial = [k for k in cols if k != c]
if trial and fails(rows, trial):
cols = trial
print("columns still needed:", cols)
print("input for that row:", X[rows[0]].tolist()[:4])
print("prediction:", predict(X[np.ix_(rows, ALL)], w, False).tolist())
Output:
old (fills) has NaN: False
new (no fill) has NaN: True
minimal failing row index: [301]
columns still needed: [2]
input for that row: [-1.7415144478985787, -0.8905443058032078, nan, 0.8884531740936087]
prediction: [nan]
Reading it: the old behaviour gives no NaN, the new one does, halving isolates row 301, and dropping columns one at a time shows only feature 2 is needed, the one with the missing value. Halving keeps the half that still fails, which works when one thing triggers the bug. If neither half fails alone, two rows or columns interact: then remove smaller chunks instead of halves (a technique called delta debugging, rarely needed). The failing input is 1 row out of 512, so it affects 1/512 of rows (about 0.2%).
Step 3: shrink further and write it up
Cut to a standalone script that still fails without the real model or data: paste the code above with the data made inline (one row [-1.74, -0.89, NaN, 0.89], the four weights and the predict function), which is about 20 lines, and confirm it prints NaN in a fresh environment. Then check the release notes for the suspect change.
Title: Predictions are NaN for rows with a missing feature after upgrading <library> <old> -> <new>
Severity: Sev-2 (serious but not total). Only rows with a missing value are affected.
Impact: in the sample, 1 of 512 rows (0.2%); downstream ranking drops those users.
Timeline (illustrative): upgrade deployed 03:05 UTC; first NaN in logs 03:12 UTC.
Environment: old and new package lists attached (differences: <library> only).
Minimal repro: attached script, 20 lines, no model file needed.
Sample failing input: [-1.74, -0.89, NaN, 0.89]
Expected: a finite probability (old version imputes the column mean).
Actual: NaN.
Suspected cause (hypothesis, unverified): default handling of missing values changed.
Workaround: impute (fill in the missing value, here with the column mean) before calling predict.
If the trigger is a downstream data pipeline change instead of a library upgrade
Do the same halving on the pipeline's output before and after the change: keep the snapshot timestamps, the first bad output time, and one failing sample record. Report severity and the share of affected records in the same way.
Pitfalls
- Shrinking before freezing the input, so the failure moves.
- Filing a screenshot of the model output instead of the smallest script.
- Presenting the suspected cause as a fact.
Here's a short bug report submitted by a teammate:
Title: Model scores dropped
Description: Model performance dropped on validation dataset.
Steps to reproduce: run eval.py
Expected: scores stable
Observed: much lower
Logs: see attached
Identify at least five missing or ambiguous pieces of information that make this report hard to act on, and rewrite the report into a clear, actionable bug report that a remediation engineer could follow.
Sample Answer
Direct answer
The report tells an engineer that something is worse, but not what, where, compared with what, or how to see it. I would find at least five gaps (the ten below show how thorough a review can be; if you can only name a few, name these first: the dataset version, the baseline, and the exact command with its seed, because without them nobody can reproduce or measure the drop), then rewrite it so a stranger can reproduce the problem in ten minutes and knows how urgent it is.
Missing or ambiguous information
- Which model and version, and the code version (commit) that produced the scores.
- Which validation dataset and which version of it. The validation dataset is the set of labelled examples kept aside from training to score the model, and it might have changed, and a changed dataset is a very different bug from a changed model.
- Which metric, and what numbers. "Much lower" and "stable" have no values.
- The baseline (the reference result to compare against): last known good score, run and date, so the drop is measurable.
- Exact command, arguments, config and random seed (the number that fixes a program's random choices so a rerun gives the same result). "Run eval.py" cannot be repeated without them.
- Environment: machine, Python and library versions, hardware. A library upgrade can change scores.
- When it started and what changed since (code, data, dependencies, configuration).
- Reproducibility: every time, or intermittent? Is the run deterministic (same inputs always give the same output)?
- Logs "attached" with no link, filename or the relevant lines.
- Impact and urgency: does this block a release, or affect users?
Rewritten report (values illustrative, substitute your own)
Title: churn-classifier v3.2: AUC on val_2025_05 dropped 0.91 -> 0.78 since commit 4f2a1c9
(AUC = area under the ROC curve, a 0-to-1 score of how well the model ranks churners above non-churners)
Summary: Running the standard evaluation on the same validation set now gives a much
lower AUC than the last passing run. Blocks Thursday's release candidate (the build proposed for release).
Environment: commit 4f2a1c9, Python 3.11, scikit-learn <version>, CPU only, seed 42.
Data: val_2025_05 (snapshot 2025-05-27, taken before the 2025-05-28 baseline run, 12,000 rows), checksum (a short fingerprint of the file that changes if any byte changes) in the ticket. Parquet is a compressed column-oriented file format for tables.
Steps to reproduce:
1. git checkout 4f2a1c9
2. python eval.py --model models/churn-v3.2.pkl --data data/val_2025_05.parquet --seed 42
Expected: AUC about 0.91 (run 2025-05-28, commit 9be03d1, link).
Actual: AUC 0.78, reproduced 3 out of 3 runs (deterministic).
What changed between good and bad: 14 commits, plus scikit-learn upgraded in requirements.txt.
Already checked: data checksum unchanged; other models score normally.
Logs: eval-run-2025-05-31.log, lines 210-240 (warnings about feature order, meaning the input columns arrived in a different sequence than the model was trained on).
Severity: high, release blocker. Owner asked to triage: model team.
Why this rewrite is actionable
- The title states the symptom, the size of the change and the trigger.
- Steps can be run unchanged, and expected and actual results carry values and a date.
- The "already checked" line stops the fixer repeating work and narrows the search.
Pitfalls
- Adding guesses about the cause into the facts. Keep hypotheses in their own labelled line.
- Attaching a 200 MB log instead of the relevant lines.
Docs in your repository include code snippets, SQL examples and small notebooks that keep breaking silently. How do you test them in CI, sandbox them, and deal with flaky or slow examples?
Sample Answer
Direct answer
Treat every code block in the docs as a small test: extract it, run it in CI (continuous integration, checks that run on every change), and fail the pull request when it breaks. Run each snippet in a throwaway directory with a time limit and a stripped environment, use fixtures (small fixed sample inputs or fake services used only for testing) instead of real services, fix flakiness (a snippet passing on some runs and failing on others with no change to the code) at its cause instead of retrying blindly, and move genuinely slow examples to a scheduled job.
The approach in layers
| What breaks | How to test it |
|---|---|
| Python snippets in Markdown | Extract fenced blocks and run each one (script below) |
| Python examples that show output | Doctest (examples written as >>> code followed by the expected output): python -m doctest docs/api.md, or pytest --doctest-glob='*.md' |
| README quick start | A smoke test (a fast test of the basic path) that runs the README's commands or its doctests |
| JavaScript snippets | The same extractor with node as the runner |
| SQL examples | Run against a throwaway database with a small fixture; keep the expected result beside the query and compare |
| Notebooks | pytest --nbmake notebooks/ (a plugin for pytest, the Python test runner) runs every cell (set a limit with --nbmake-timeout) |
A runnable extractor (runs python and js blocks, reports failures as GitHub annotations (messages GitHub draws on the exact line of the pull request diff), skips blocks marked <!-- docs-test: skip -->):
import pathlib, re, shutil, subprocess, sys, tempfile
TICKS = "`" * 3
FENCE = re.compile(r"(<!-- docs-test: skip -->\s*\n)?" + TICKS + r"(python|js)\n(.*?)" + TICKS, re.S)
RUNNERS = {"python": "python3", "js": "node"}
TIMEOUT_SECONDS = 2
def check(root):
failed = 0
for md in sorted(pathlib.Path(root).rglob("*.md")):
text = md.read_text()
for m in FENCE.finditer(text):
line = text[:m.start(3)].count("\n") + 1 # first line of the snippet
if m.group(1):
print(f"skip {md}:{line}"); continue
exe = shutil.which(RUNNERS[m.group(2)]) # absolute path, so the stripped PATH below is fine
if not exe:
failed += 1
print(f"::error file={md},line={line}::no {RUNNERS[m.group(2)]} on this runner"); continue
suffix = ".py" if m.group(2) == "python" else ".js"
with tempfile.TemporaryDirectory() as sandbox: # scratch cwd, minimal environment
script = pathlib.Path(sandbox, "snippet" + suffix)
script.write_text(m.group(3))
try:
r = subprocess.run([exe, str(script)], cwd=sandbox, capture_output=True,
text=True, timeout=TIMEOUT_SECONDS, env={"PATH": "/usr/bin:/bin"})
ok, why = r.returncode == 0, (r.stderr.strip().splitlines() or ["exit " + str(r.returncode)])[-1]
except subprocess.TimeoutExpired:
ok, why = False, f"timed out after {TIMEOUT_SECONDS}s"
if ok: print(f"ok {md}:{line}")
else:
failed += 1
print(f"::error file={md},line={line}::snippet failed: {why}") # GitHub Actions annotation
return failed
if __name__ == "__main__":
sys.exit(1 if check(sys.argv[1]) else 0)
How the regex reads. FENCE has three parts. Group 1 is an optional skip comment followed by a newline. Then comes the opening fence: three backticks and the language, python or js (group 2), and a newline. Then group 3, (.*?), is the snippet body, taken lazily (as little as possible) up to the next three backticks; re.S lets . match newlines so the body can span many lines. So m.group(1) is the skip marker if present, m.group(2) the language, m.group(3) the code. The line number is the count of newlines before m.start(3) (where the body begins) plus one.
What the demo does and does not do. It implements the skip marker, the time limit, the stripped environment, the temporary working directory and the annotations. The policies described below (require a reason on every skip, a slow tag, one retry for a network snippet) are team rules you layer on top; the demo accepts a bare skip marker and does not enforce any of them.
Test it on this fixture, saved as docs/guide.md (it has one good snippet, one broken, one skipped, one JavaScript, one slow):
# Guide
```python
print(sum([1, 2, 3]))
```
```python
print(undefined_name)
```
<!-- docs-test: skip -->
```python
requires_a_real_api_key()
```
```js
console.log([1, 2, 3].map((x) => x * 2).join(","));
```
```python
import time; time.sleep(5)
```
Then run python3 check_docs.py docs. Output:
ok docs/guide.md:4
::error file=docs/guide.md,line=8::snippet failed: NameError: name 'undefined_name' is not defined
skip docs/guide.md:13
ok docs/guide.md:17
::error file=docs/guide.md,line=21::snippet failed: timed out after 2s
The process exits with status 1, so the job fails. The ::error file=...,line=...:: format makes GitHub show the message on the exact line in the pull request, which is how failing docs get in front of the author.
Sandboxing
- Each snippet runs in a fresh temporary directory with a minimal environment, so no secrets or files from the runner are visible.
- Every snippet has a time limit (2 seconds in the demo; use something realistic like 30 seconds in real docs).
- In CI, deny network access for the job (for example by running it in a container without a network), and give examples fixtures or a local fake server instead of real services.
- Pin the interpreter and dependency versions, or "docs broke" will just mean "the world moved".
Flaky examples: fix the cause
Flakiness in docs snippets almost always comes from the network, the clock, randomness or output ordering. Answer each directly: recorded responses or a fake server, a fixed date, a pinned random seed, sorted output or doctest's ELLIPSIS flag (an option that lets ... in the expected output match any text, so a timestamp does not break the test) for unstable parts. Do not retry blindly: a retry hides real breakage. If retries are unavoidable for one snippet you have tagged as network-dependent, allow one, and report that it needed one.
Slow examples
Put a hard timeout on each snippet, tag heavy ones slow (a label, for example a pytest marker, that lets CI select or skip them), and run them in a nightly job instead of on every pull request. The better fix is a smaller dataset in the example.
Quarantine with an expiry
Quarantine means temporarily excluding a snippet from the check, with an owner and an end date. A snippet that cannot run yet (needs a real API key, or is deliberately partial) gets the skip marker plus a reason and a ticket. Fail CI when a skip has no reason (a policy you would add: require text after the marker, which the demo regex does not yet do), and review skips periodically so "skipped" does not become "abandoned".
Trade-offs and pitfalls
- Docs break when the code changes, not only when the docs change, so run this check on every pull request that touches code, not only on docs-only ones.
- Running examples proves they execute, not that they teach well. Keep reviewing wording.
- Snippets that print a timestamp or random value make bad doctests. Rewrite the example so it is stable.
That is every published Technical Writing and Documentation question for QA Engineer so far. Browse the other topics in this category, or practice this one interactively.