Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
What should code comments explain and what should they leave out? Give a good and a bad example from a model training script.
Sample Answer
Direct answer
Comments should explain what the code cannot say for itself: why a choice was made, a constraint that is not visible, and a trap for the next editor. They should leave out what the code already says, history that version control (git) records, and anything you would not want public. The test is: "would a competent reader be surprised or misled without this line?"
What belongs in a comment
- Why: the reason for a choice that looks arbitrary.
- Constraints and hidden coupling: things that break if changed.
- Pointers: where the evidence or longer explanation lives.
- Warnings: ordering, leakage (information from validation data sneaking into training, which makes scores look better than they are), units.
What to leave out
- Restating the line, commented-out old code, jokes, and author or date notes (git blame, a command showing who last changed each line and when, answers those).
- Vague TODOs with no owner or ticket.
- Secrets and internal-only details.
Good and bad, from a model training script
# BAD: says what the code says
epochs = 20 # set epochs to 20
optimizer.zero_grad() # zero the gradients
# lr = 1e-3 # old value
# GOOD: says why
# Split by user_id, not by row. The same user's events in both train and
# validation would leak and make the validation score look better than reality.
train_users, val_users = split_users(users, seed=SEED)
# 3e-4 was the largest rate that kept loss stable in the sweep,
# see experiments/lr_sweep.md.
learning_rate = 3e-4
Terms in the snippet: a gradient is the signal that tells the model how to adjust, and zero_grad clears it before each step. Validation data is the held-out portion used to test the model. A seed is a fixed starting value so random choices repeat. The learning rate is how big a step the model takes when adjusting, and 3e-4 means 0.0003. A sweep is a set of runs trying several values of one setting, so "largest rate that kept loss stable" means the biggest tried value where training did not blow up.
The bad comments add noise, and the commented-out lr invites confusion about which value is live. The good ones save a future reader from re-deriving a decision, or from "fixing" the split back to per-row, which would silently inflate results. If the good comments are long, the reason belongs in a short doc that the comment links to.
Comments that help new contributors onboard
Comments explain local reasoning. The README carries the map a newcomer needs: how to set up, how to run a small smoke training run (a very short run that only proves the pipeline works end to end), where the data comes from, where decisions are recorded, and how to run the tests. Together they cut the questions a newcomer must ask a person.
Pitfalls
- Comments rot. A comment that contradicts the code (
# uses 5 foldsabovek = 10) is worse than none. Update it in the same change as the code. - If you need a comment to explain a confusing line, first try a clearer name or a small function.
Compliance asks you for documentation about how a model uses personal data. How does what you write, and the way you write it, differ from the notes you would give fellow data scientists?
Sample Answer
Direct answer
For compliance I write a factual, complete, plain-language record of what personal data goes in, why, where it flows, who can see it, how long it is kept, and what controls exist, and I leave the legal conclusion ("this is lawful") to them. For fellow data scientists I write shorthand about features, experiments and pitfalls. The facts overlap; the reader, vocabulary, completeness and burden of proof differ.
How the two documents differ
| Dimension | Notes for data scientists | Document for compliance |
|---|---|---|
| Reader question | "How does it work and what should I watch out for?" | "What personal data do you use, on what basis (the legal reason you are allowed to use it, such as consent or contract), and can you show control?" |
| Vocabulary | Feature names, embeddings (lists of numbers a model uses to represent an item), shorthand | Defined plain terms; a feature name is always translated ("cust_age_bucket: customer age group, derived from date of birth") |
| Scope | What matters for modelling | Everything touching personal data, including data you consider incidental |
| Tone | Informal, exploratory | Precise, stable, dated and versioned; no marketing words like "anonymous" unless demonstrated |
| Evidence | "I believe it is fine" | A link to the access-control setting, retention job or test that proves each claim |
What the compliance document contains
- Purpose of the model, in one paragraph.
- Personal data inventory: each field, whether it is personal data (any information relating to an identifiable person) or sensitive (special categories such as health or religion that carry stricter legal rules), its source, and why the model needs it.
- Data flow: collection, storage, training, inference, outputs, third parties.
- Retention and deletion: how long training data and logs live, and how a person's deletion request is handled, including the honest answer if a trained model cannot be edited and must wait for a retrain.
- Access and security controls.
- Model outputs and decisions: whether a person is affected by an automated decision (a decision made by software with no human in the loop, such as an automatic loan refusal), and whether a human can review it.
- Known risks and mitigations, and open questions marked as open.
Worked example (excerpt)
"Field: postcode. Personal data: yes, can help identify a person in combination with other fields. Source: customer signup form. Purpose: regional demand feature. Stored in the feature table; access limited to the modelling team group. Retention: raw postcode deleted after 24 months by a retention job (a scheduled job that deletes data past its keep-until date; see linked ticket). The 24 months is an example value; the real period comes from the company's retention policy. Reduced to a 3-character region code before training." A data scientist note would just say "region_code, dropped postcode, re-identification risk with city" (re-identification risk: region code combined with city could narrow a record down to very few people; this is different from target leakage, where a feature carries information that would not exist at prediction time).
Pseudonymised versus anonymised: pseudonymised data has names or IDs replaced by codes, but someone holding the key or other fields can still work out who it is, so it is still personal data. Anonymised data cannot be traced back to a person by anyone with reasonable effort, which is a much higher bar. Calling pseudonymised data anonymised overstates the protection and can mean the wrong legal rules are assumed.
Pitfalls
Overclaiming ("fully compliant", "anonymised" for merely pseudonymised data), leaving gaps silently, and writing a legal argument yourself. If I do not know something, I write "unknown, owner X will confirm by date Y".
A model card reads: 'Model v2 is better and faster. See code.' Rewrite it so an engineer, a product owner and a regulator could each rely on it, and tell me what you had to find out to do so.
Sample Answer
Direct answer
"Model v2 is better and faster. See code." fails because it has no comparison baseline, no metric, no dataset, no scope and no owner, and it sends every reader to source code that only one of them can read. I would first find out the facts below, then rewrite it as a card with a section for each reader. What I could not find out must be written down as unknown, never guessed.
What I would have to find out first
- Better than what, measured how: which metric, on which evaluation set, and is it the same set for v1 and v2?
- Better for whom: overall or for subgroups; did any segment get worse?
- Faster: latency at what percentile (for example P95, the time under which 95% of requests finish), on what hardware, batch size and input size.
- What changed: data, features, architecture, threshold?
- What the model is for, and what decisions it feeds, including whether a human reviews outputs.
- Who owns it, how it is monitored, and what triggers rollback.
- Regulatory context: does it affect people (credit, hiring, safety)? If so, what must be disclosed?
Rewrite (illustrative model: a support-ticket urgency classifier; figures derived from stated counts)
Ticket Urgency Classifier, v2 (replaces v1). Owner: Support Tooling. Date and version pinned to git tag (a permanent label on one exact code commit); linked training data snapshot.
Summary (product owner): v2 finds more of the truly urgent tickets, so agents see more of them first (the holdout figures below measure how many it finds, not how early). On the same 20,000-ticket holdout (tickets kept aside from training and used only for scoring; 1,000 truly urgent), v1 caught 700 urgent tickets (recall 70%, the share of real urgent tickets found) while flagging 1,500; v2 caught 780 (78%) while flagging 1,560, so precision (the share of flagged tickets that are truly urgent) moved from 700/1,500 = 46.7% to 780/1,560 = 50%. Non-English tickets were not in the training data, so v2 should not be used for them. Latency figures: [from the load test; not yet measured, do not publish a claim].
Engineer section: training data snapshot ID, feature list, model type, evaluation script and command, threshold, latency method (P95, hardware, batch size), how to reproduce, known dependencies, rollback: redeploy v1 artifact.
Regulator section: purpose and decision impact (prioritisation only; no customer is denied service), data sources and legal basis for using ticket text (the lawful reason for processing customers' messages, for example "legitimate interest in prioritising support, disclosed in the privacy notice"), personal data handling, human oversight (agent can override), monitoring plan, change history and approver.
Limits and unknowns: no evaluation on non-English text; performance by customer tier not measured.
The rewrite as one readable card (compressed)
"Ticket Urgency Classifier v2, owner Support Tooling. What it does: ranks incoming support tickets so likely-urgent ones are seen first; agents can override, no customer is refused service. Evidence: on a 20,000-ticket holdout, v1 found 70% of urgent tickets at 46.7% precision, v2 finds 78% at 50% precision. Limits: English only; performance by customer tier unmeasured; latency not yet measured. Rollback: redeploy the v1 artifact. Reproduce: git tag and evaluation command in the engineer section."
Why this works
- The claim "better" becomes checkable: same holdout, two stated counts, two derived percentages.
- Each of the three readers finds their section without reading the others.
- Unknowns are labelled, which protects the team more than a confident vague claim.
Trade-offs
- A card that is complete for all three readers is longer; solve with layered sections, not by dropping one audience.
- If I cannot get the v1 baseline numbers, I say so and present v2 alone rather than fabricate a comparison.
Review this documentation excerpt for a preprocessing step and identify its ambiguities: 'Resize images to 224 and normalize between 0 and 1.' Then rewrite it so another engineer can reproduce the step exactly.
Sample Answer
Direct answer
The sentence hides at least six decisions, and each one changes the numbers the model sees. "Resize to 224" does not say which side, what happens to aspect ratio, or how pixels are blended. "Normalize between 0 and 1" does not say what is divided by what. Another engineer following it literally would get a different input tensor (the array of numbers fed to the model) and a silently different model score. (The model was trained on inputs prepared one particular way; any other preparation gives it inputs it never learned from.) The rewrite must pin every choice, name the library, and give a check value.
Ambiguities
- "224": 224 by 224, or the shorter side set to 224 (keeping aspect ratio)? If the shorter side, is there a crop, and where (center, random)?
- Aspect ratio: squash the image to a square, or preserve the ratio and pad or crop?
- Interpolation (the rule for blending neighbouring pixels when the size changes): nearest, bilinear or bicubic. Libraries also differ in antialiasing (smoothing before shrinking), so the same method name can give different pixels.
- Input state: colour space and channel order (RGB versus BGR; OpenCV loads BGR by default), grayscale versus 3 channels, and integer 0-255 versus float.
- "Between 0 and 1": divide by 255, or min-max per image (subtract the image minimum, divide by its range)? These agree only when the image happens to contain both 0 and 255.
- Missing: dtype (float32), whether the dataset's mean is subtracted and the result divided by its standard deviation afterwards (a common extra rescaling that changes every value), tensor layout (height, width, channels versus channels first), and order of operations (resize before or after scaling).
Seeing the difference, not just asserting it
import numpy as np
from PIL import Image
# A 4x4 grayscale image with pixel values 0, 16, 32, ... 240
pixels = (np.arange(16, dtype=np.uint8) * 16).reshape(4, 4)
img = Image.fromarray(pixels, mode="L")
# "Resize to 2": same target size, three different interpolation methods
for name, method in [("nearest", Image.NEAREST),
("bilinear", Image.BILINEAR),
("bicubic", Image.BICUBIC)]:
out = np.asarray(img.resize((2, 2), resample=method))
print(name, out.flatten().tolist())
# "Normalize between 0 and 1": two different recipes
x = np.array([0, 51, 255], dtype=np.float64)
print("divide by 255:", (x / 255.0).round(3).tolist())
print("min-max of this image:", ((x - x.min()) / (x.max() - x.min())).round(3).tolist())
dark = np.array([20, 40, 60], dtype=np.float64)
print("divide by 255 (dark image):", (dark / 255.0).round(3).tolist())
print("min-max (dark image):", ((dark - dark.min()) / (dark.max() - dark.min())).round(3).tolist())
Output:
nearest [80, 112, 208, 240]
bilinear [57, 83, 157, 183]
bicubic [47, 77, 163, 193]
divide by 255: [0.0, 0.2, 1.0]
min-max of this image: [0.0, 0.2, 1.0]
divide by 255 (dark image): [0.078, 0.157, 0.235]
min-max (dark image): [0.0, 0.5, 1.0]
The same 4x4 image resized to 2x2 gives three different results depending on interpolation (for example 80 versus 57 versus 47 in the first cell). For normalization, on the image containing 0 and 255 both recipes agree, but on a dark image (20, 40, 60) dividing by 255 gives 0.078 to 0.235 while min-max stretches it to 0.0 to 1.0. A model trained on one recipe and served with the other sees a different input distribution (the overall spread of values it receives, which it never learned to handle).
Rewritten step (the values are an example; they must match what the model was trained with)
Input: an 8-bit RGB image (channel order R, G, B; convert grayscale or BGR sources first).
- Resize so the shorter side is 256 px, keeping aspect ratio, bilinear interpolation, using Pillow
Image.resize(pin the Pillow version in the lockfile, the file listing the exact version of every package).- Center-crop (cut a 224 x 224 window out of the middle of the resized image) to 224 x 224 px.
- Convert to float32 and divide every value by 255.0, so values lie in [0, 1]. Do not rescale per image. No mean or standard deviation subtraction.
- Output layout: height x width x channels (224, 224, 3); transpose only inside the model wrapper.
Check: the provided sample imagedocs/sample.jpgmust produce a tensor whose first row sums to the value stored intests/expected_sum.txt, and a unit test asserts this (within 1e-3).
Tiny illustration of the idea: a first row of pixels 0, 51, 255 divided by 255 gives 0.0, 0.2, 1.0, which sums to 1.2, so the file would contain1.2. For the real 224 x 224 image the file holds whatever number the reference function printed when the team first ran it; anyone whose pipeline prints a different number has diverged.
Trade-offs and pitfalls
- Writing the rule is not enough: ship the preprocessing as one shared function used by both training and serving, and test it. Prose drifts, code does not.
- The most damaging bug in this family is training and serving using different pipelines (training-serving skew: the model sees differently prepared inputs in production than it saw in training, so it scores worse there with no error message). The check value catches it at review time.
- Pinning the library version matters because implementations change between releases.
What does a good pull request description contain for a complex change? Show me a short example for a change to a data preprocessing step.
Sample Answer
Direct answer
A good PR (pull request, a proposed change that reviewers read before it is merged) description lets a reviewer who has not seen your work understand why the change exists, what it does, how you know it works, and what could go wrong, without reading the whole diff first. For ML changes it also records what is needed to reproduce the run. It stays short: a reviewer's attention is the scarce resource.
What it contains
- Why: the problem, with a link to the ticket or experiment.
- What changed: behaviour, not a line-by-line diff.
- How tested: the tests or runs, with the seed and data version.
- Risk and compatibility: who is affected, what needs retraining or migrating, how to roll back.
- Review focus and non-goals: where to look hard, and what is deliberately left alone.
Key to the example, in plain words: to impute is to fill in a missing value. The training split is the portion of data the model learns from, and validation is the portion held back to test it. Serving is the live system that uses the model to answer requests. Leakage is validation information sneaking into training and inflating scores. A seed is a fixed starting value so random steps repeat. A dataset version labels an exact snapshot of the data. A fixture is a small fixed test input. Retrained means trained again on data. The fix(...), feat and tune prefixes follow a naming convention (Conventional Commits) where a commit message starts with its type.
Example: a data preprocessing change
fix(preprocess): impute session_length with the training median, not 0
Why: Missing session_length was filled with 0, which the model reads as
"a session of zero seconds" and treats as a strong signal.
What: Missing values are now filled with the median computed on the
TRAINING split only and saved to preprocess_stats.json. Validation and
serving load that file, so no validation data influences the fill value.
Tested: New unit test on a 5-row fixture (3, 5, missing, 8, missing).
Column mean was 3.2 with zero-fill and is 5.2 with median-fill (5, from
the observed 3, 5, 8). Full run used seed 42 and dataset version v7.
Compatibility: Models trained before this change expect the old fill and
must be retrained. Rollback: revert; preprocess_stats.json is ignored.
Review focus: that the median is never computed from validation rows.
Why concise text matters for reproducing ML workflows
Months later someone asks what changed between model v12 and v13. The commit or PR is the index: its message plus the seed, dataset version and config identifies the run. A vague message forces archaeology, and a wall of text goes unread.
Three short commit-message examples
- Hyperparameter:
tune: lower learning rate from 1e-3 to 3e-4 for stable loss (experiments/lr_sweep.md). The learning rate is the size of each adjustment step in training, and 3e-4 is 0.0003. The message gives the change, the reason (loss was unstable at the larger value) and where the evidence lives. - Preprocessing fix:
fix(preprocess): impute session_length with training median, not 0 - Feature addition:
feat: add days_since_last_login, computed at prediction time to avoid leakage. The reason is in the message: computing it from data that only exists after the outcome would leak the answer to the model.
Pitfalls
- "Fixed stuff" says nothing. A pasted diff duplicates what reviewers can already see.
- Skipping the compatibility line is how a model quietly stops matching its serving code.
Unlock Full Question Bank
Get access to all 21 Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.