Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
Write a one-page quick-start for engineers adopting a new internal shared component, for instance a feature store. What must a first-time user be able to do in ten minutes, and which pitfalls belong on the page?
Sample Answer
Direct answer
The quick-start's job is one measurable outcome: a first-time user gets real value from the component in ten minutes, without asking anyone. For a feature store (a shared system that keeps precomputed model input values, called features, so training and live serving use the same definitions), that means install, authenticate, read one existing feature, and see it verified. An entity is the thing the features describe (here a customer, identified by an entity id such as customer_id = 1042). Online features are the latest values served quickly for live predictions, as opposed to offline features, which are historical values used to build training data. A namespace is a named folder that keeps sandbox features apart from production ones, and the catalogue is the searchable list of registered features and their owners. The pitfalls on the page are the ones that bite first-time users, not a full reference.
What a first-time user must do in ten minutes
| Minutes | Step | Success check |
|---|---|---|
| 0-2 | Install the client with a pinned version and set the project name | featurestore --version prints a version |
| 2-4 | Authenticate with your own team credentials (never a shared production key) | A test call returns "authenticated as <you>" |
| 4-6 | Read features for one known entity id (for example customer_id = 1042) | You see the feature values and a timestamp |
| 6-9 | Define one new feature in a sandbox namespace and fetch it | The new value appears in the read |
| 9-10 | Find your feature in the catalogue or UI and see who owns it | You can name its owner and freshness |
The steps use an illustrative client, since this is an internal component:
from featurestore import Client # illustrative internal client
client = Client(project="sandbox")
row = client.get_online_features(
features=["customer.orders_30d", "customer.avg_basket"],
entity={"customer_id": 1042},
)
print(row) # expect two values plus the time each was computed
The page states the expected result after each step, so success is visible without guessing. In the code: Client(project="sandbox") connects to the sandbox namespace, get_online_features asks for two named features for one customer, and the print should show one value per feature plus the time each was computed.
Pitfalls that belong on the page
Pitfalls 1 and 2 are the ones that fail silently and cost the most; 3 to 6 are housekeeping the page can state in one line each.
- Point-in-time correctness. When building training data, fetch each feature value as it was at the label's timestamp. Using today's value leaks the future into training, so the model looks better offline than it will in production. Example (illustrative): customer 1042 churned (stopped buying) on 2025-03-15, and the label timestamp is that date. Their
orders_30dwas 5 on 2025-03-15, but by today (2025-06-01) it is 0 because they left. If training uses today's 0, the model learns "zero orders means churn", which looks perfect offline, yet in production the model only ever sees the value 5 before the customer leaves. The correct training value is the one as of 2025-03-15. - Training-serving skew. Reuse the registered definition. Recomputing the feature separately in your own code is how the two sides drift apart.
- Freshness. A feature has a refresh interval and a time-to-live (how long a value stays valid). Say which features are hourly and which daily, so a user does not expect real-time.
- Entity keys. Use the documented key name and type (an id stored as text vs integer silently returns nothing).
- Sandbox vs production. Where to experiment and what needs a review before promotion.
- Cost and quotas. Large backfills (recomputing history) consume shared capacity: ask the owner first.
Worked example
A new engineer reads the page at 10:00. By 10:04 they are authenticated. At 10:06 they see orders_30d = 7 for customer 1042 with a computed-at time. They ask, "is this current?" The page's freshness note says daily at 03:00, so a 10:00 read is 7 hours old, which they can decide is acceptable for their model. They never opened the full reference.
Trade-offs and pitfalls of the page itself
- The page must stay one screen long. Link to the reference, the on-call channel, and the design doc rather than embedding them.
- Test the ten-minute claim by watching someone new do it, then cut every step where they hesitated.
- Put the sharpest pitfall (point-in-time leakage) near the top of the pitfalls list, since it is silent and costly.
- Examples must use the sandbox, so copy-paste never touches production.
How would you tell whether your team's documentation is working? Which signals would you collect, how would you collect them, and which popular measure would you distrust?
Sample Answer
Direct answer
I would judge documentation by what happens after someone reads it: do they find what they need, act on it, and stop asking people? So I collect a few behavioural signals (search that fails, questions that repeat, ramp time for new hires), one or two opinion signals (a page rating and a short survey), and I would distrust raw page views, which go up when docs are good and also when they are confusing.
Signals and how to collect them
| Signal | What it tells you | How to collect |
|---|---|---|
| Search abandonment (searches with no click) divided by all searches | People cannot find the answer | Docs site search logs, weekly |
| "Avoidable ticket" share: support tickets or team-chat questions whose answer was already in the docs, divided by all tickets | Docs exist but are not found or not trusted | Tag tickets at closing; sample if volume is high |
| Time to first successful task for new hires (days to first merged PR, first deploy) | End-to-end usefulness | Onboarding tracker |
| Page feedback ("helpful: yes/no" plus a comment box) | Quality on that page, qualitatively | Widget on each page |
| Freshness: share of tier-1 pages (pages for the most critical services) verified in the last 90 days | Trustworthiness | Metadata "last verified" |
| Time to resolve incidents (MTTR, mean time to resolve) where a runbook (step-by-step procedure for an alert) existed versus not | Link to a real outcome | Incident tool, by tag |
MTTR by runbook tag, with made-up numbers: incidents with a runbook took 45, 30 and 60 minutes (mean 135/3 = 45); incidents without one took 90, 120 and 75 (mean 285/3 = 95). The 50-minute gap is a hint, not proof, since the incident types may differ.
Worked example (illustrative data, run it)
# Tiny, made-up docs-effectiveness data (illustrative). Each row is one week.
weeks = [
# (page_views, searches, searches_with_no_click, tickets, tickets_where_answer_was_in_docs)
(900, 200, 90, 40, 14),
(1500, 210, 60, 33, 9), # after rewriting the top pages
]
for i, (views, searches, no_click, tickets, in_docs) in enumerate(weeks, start=1):
print(f"week {i}: views={views}",
f"| search abandonment={no_click / searches:.0%}",
f"| avoidable tickets={in_docs / tickets:.0%} ({in_docs}/{tickets})")
# Onboarding signal: days from start date to first merged pull request
before = sorted([21, 24, 19, 30, 26])
after = sorted([14, 16, 12, 20, 15])
median = lambda xs: xs[len(xs) // 2]
print("median days to first merged PR:", median(before), "->", median(after))
Output:
week 1: views=900 | search abandonment=45% | avoidable tickets=35% (14/40)
week 2: views=1500 | search abandonment=29% | avoidable tickets=27% (9/33)
median days to first merged PR: 24 -> 15
Reading it: after rewriting the top pages, search abandonment fell from 45% (90 of 200 searches) to 29% (60 of 210), avoidable tickets from 35% (14 of 40) to 27% (9 of 33), and the median time to first merged PR from 24 to 15 days. Note that page views rose too; on their own, views could not have told us whether that was good.
Which popular measure I distrust: page views
Views measure traffic, not usefulness. A confusing page produces repeat visits and more views. A great page that answers in one glance also produces views but no follow-up. Other gameable ones are "number of pages written" and "words". Guard by pairing every activity measure with an outcome measure (views with search abandonment; docs written with avoidable tickets).
Quantitative versus qualitative KPIs (key performance indicators, the few numbers you track to judge success) and the response when one drops
Quantitative: the rates above. Qualitative: free-text feedback, and watching a new hire try a task. When a KPI drops, look at the pages contributing most, read the comments, run a 20-minute test with someone who has not seen the page, then fix and re-measure a few weeks later. Also link the improvement to an outcome such as incident time to resolve: compare incidents where the runbook was used with those where it was not.
Pitfalls
Small samples (say which counts are behind a percentage), changing two things at once, and treating a correlation as proof. Say the numbers back a claim about direction, not exact cause.
You are writing public-facing release notes for an ML feature that changes what users see. What do you check before publishing, and how do they differ from internal notes?
Sample Answer
Direct answer
Public release notes for an ML feature must be checked for truth, tone, and exposure. Truth: they describe what users will actually see and the limits of it. Tone: they use the user's language, not the team's. Exposure: nothing internal (model names, datasets, unreleased metrics, vendor details) leaks out. Internal notes can be technical and candid, while public notes promise only what you can support.
What I check before publishing
- Matches reality: the feature flag (a switch in the code that turns a feature on or off without redeploying), rollout percentage (the share of users who currently have it), dates and versions in the note equal what is actually live.
- User-visible change, in user terms: what looks different, for whom, and when. If ordering, suggestions or scores can change for the same input, say so.
- Claims are supportable: no "smarter" or "more accurate" unless it was evaluated and you can say on what. Prefer describing behaviour ("suggestions now prefer recent items").
- Limits and failure modes: results can be wrong or vary; say what users should double-check.
- Data and privacy: what data the feature uses, whether customer data trains it, and how to opt out. Privacy or legal review signs off.
- Control and support: where to turn it off, and support has the help article before launch.
- Confidentiality sweep: strip internal model names, experiment IDs, dataset details, unreleased numbers, and anything under a partner agreement.
Public versus internal, for the same change
| Public note | Internal note | |
|---|---|---|
| Audience | Customers | Engineers, support, product |
| Language | Behaviour and benefit | Model, flag, offline evaluation link (results from testing on saved historical data) |
| Limits | Plain caveats | Known regressions (things that got worse), failure modes (the ways it can go wrong) |
| Rollback | Where to switch it off | Flag name, owner, on-call (the engineer currently responsible for responding to problems) |
Redaction example (exposure check)
Before: "Model rank-v3 (experiment EXP-4471) trained on the clicks dataset, with embeddings from a third-party vendor." After: "Search now puts items you opened recently nearer the top." The model name, experiment ID (an internal label for one test run), dataset and vendor all stay internal.
Worked example
Public: "Search now puts items you opened recently nearer the top, so the same query may show a different order than before. Turn this off under Settings, Personalised ranking. Rolling out to all accounts over the next two weeks."
Internal: "A new ranking model replaces the previous one behind the feature flag. Offline evaluation and the A/B result (a comparison of the old and new versions on live users) are linked. Known weaker on rare queries. Rollback: turn the flag off; owner is on-call."
Pitfalls
- Overclaiming ("AI-powered accuracy boost") creates support tickets and trust damage when the feature is wrong once.
- Publishing before rollout finishes makes the note false for part of your users.
In your own words, what makes a piece of technical documentation good? Walk me through how you would judge a document you have just inherited from a colleague.
Sample Answer
Direct answer
Good technical documentation lets a specific reader finish a specific task correctly without asking the author. So I judge it on four things: is it aimed at a clear reader and task, is it true today, can the reader find what they need fast, and can they verify it by doing what it says. Everything else (polish, length, diagrams) is secondary.
How I would judge an inherited document, in order
- Name the reader and the job. Who was this written for (a new hire, an on-call engineer, an analyst) and what should they be able to do afterwards? If I cannot tell, that is the first defect. Different deliverables serve different readers: a README (the front-page file of a code repository) gets someone running the thing, a slide deck persuades a decision-maker in one sitting, and an API reference is scanned by a developer looking up one parameter. A good document fits its reader, so judging "good" without naming the reader is meaningless.
- Check it is current. Compare the last-edited date to the last meaningful change in the system. Look for a named owner. Spot-check five concrete claims (a command, a config name, a diagram box) against reality.
- Run it. Follow the steps literally in a clean environment. Every place I have to guess or ask someone is a bug in the document.
- Test findability. Can I answer three realistic questions in under two minutes using only headings, the table of contents and search? Are there prerequisites up front and a clear "if this fails, do this"?
- Decide: keep, fix, or retire. Fix the top few defects, delete what is wrong or duplicated elsewhere, and add an owner and a review date. A stale document that looks authoritative is worse than none.
Worked example
I inherit a README for a nightly ingestion service. It has a 40-line "Overview" of the team's history but "Setup" says only "install dependencies and run main". Step 1: the reader is a new engineer who needs to run it locally. Step 2: the document names an environment variable that no longer exists in the code. Step 3: following it literally fails at the second command. Step 4: the failure-handling section is missing. Verdict: fix. I move setup to the top, correct the variable, add a "known failure: expired credentials, how to renew" entry, cut the history to two lines, and record myself as owner with a quarterly review.
Trade-offs and pitfalls
- Length is not quality: a short accurate page beats a long stale one.
- Do not judge only prose style. A beautifully written document with a wrong command fails its reader.
- Do not rewrite everything on inheritance. Fix what blocks the main task first, then improve iteratively.
You have just shipped a new internal microservice. What goes in its README and in what order, and how would the README change if this were a public library instead?
Sample Answer
Direct answer
A README is the front page of a repository, so I order it by the questions a newcomer asks in sequence: what is this, is it the right thing, how do I run it, how do I use it, who owns it. For an internal microservice that order ends with operating and owning it. For a public library it starts with installing and using it, and adds versioning and contribution rules, because the readers are strangers who cannot walk over and ask.
Internal microservice README, in order
- Name, one-line purpose, status and owner: "billing-events: turns payment webhooks into internal events. Owner: payments team, #payments-help."
- What it does and does not do: two short lists. Stops people using it for the wrong job.
- Quick start: clone, install, run locally, and how to see a request work, in copy-paste commands.
- Configuration: every environment variable, default, and whether it is a secret.
- API and contracts: link to the API spec or schema, plus one request and response example.
- Run the tests: the one command.
- Deploy and operate: where it runs, dashboards, alerts, the runbook link (a step-by-step guide for handling each alert or failure), and its SLA (service-level agreement, the availability and response promise to its users).
- Dependencies and architecture: upstream services (the ones that send it data) and downstream services (the ones that consume its output), plus one small diagram.
- Where the rest lives and how to contribute: design docs, ADRs (architecture decision records, short notes on why a choice was made), change process.
Rule of thumb: the first screen plus quick start must get a new engineer to a working local run without reading further.
# billing-events
Turns payment webhooks into internal `payment.settled` events.
**Owner:** payments team | **Status:** production | **Runbook:** link
## Quick start
make install && make run
curl localhost:8080/health
## Configuration
| Variable | Default | Secret |
|---|---|---|
| `WEBHOOK_SIGNING_KEY` | none | yes |
What changes for a public library
| Aspect | Internal service | Public library |
|---|---|---|
| Opening | Purpose and owner | Purpose plus a badge row (small status images: build passing, latest version, license type) |
| First action | Run locally | Install (pip install / npm install) then a minimal usage example |
| Middle | Config, deploy, runbook | API reference link, more examples, supported versions |
| Operating info | Dashboards, SLA, on-call | Removed. Not the users' concern |
| Added | Versioning policy (semantic versioning: major.minor.patch, where major means breaking), changelog (list of changes per release), migration notes (step-by-step upgrade instructions between versions), license, security reporting address, contribution guide |
The internal reader can ask a colleague; the public reader can only read, so the library README must be complete at first contact.
Worked example: adapting the README for an ML repo and a multi-team service
- An ML repository for analysts and new hires keeps the same order but adapts the middle: what the training code does, how to run training on a small sample, how to run the containerised inference (the model packaged in a container so it runs identically anywhere), and what the CI (continuous integration, automated checks on each change) verifies.
- A service used by several teams adds an onboarding block: ownership, the SLA, review cadence for changes, and how to request a new capability.
- Documenting a new module or public API for maintainers and new hires: usage examples for the common cases, a diagram of how the parts connect, the schema or API contract, migration notes for anyone upgrading, and one line saying where longer docs live so the README stays short.
A concrete excerpt for the ML repository (illustrative):
# churn-model
Predicts which subscribers may cancel in the next 30 days. Owner: data-science team.
## Try it in 10 minutes
make sample-data # downloads a 5,000-row sample
make train-small # trains in a few minutes on a laptop
make predict ROW=17 # prints one prediction and its score
## What CI checks
Lint, unit tests, and one training run on the sample. It does not check full-data accuracy.
The multi-team onboarding block is a short list: "Owner: payments team. SLA: 99.9% monthly availability. Changes reviewed by two owners, weekly. To request a capability, open a ticket with the label payments-request."
Trade-offs and pitfalls
- A README that grows into a manual is skipped. Link out for depth.
- Commands that were never run are the top cause of distrust. Copy them from a working session.
- Do not put secrets or environment specifics that go stale in it; link to the source of truth.
Unlock Full Question Bank
Get access to all 19 Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.