Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
You lead a team where documentation is poor and engineers resist writing it. Propose a multi-quarter plan to change the documentation culture. What do you do first, how do you handle cross-team governance, and how do you show progress each quarter?
Sample Answer
Direct answer
Do not start with a mandate or a big tool migration. Spend the first quarter finding out why engineers do not write docs (usually: no time, no clear owner, docs nobody reads, and no credit), fix the cheapest of those causes, and prove value on one high-pain area. Then add ownership and light governance in quarters two and three, and report progress with a small set of outcome metrics, not page counts.
Quarter 1: diagnose, then win one thing
- Interview 8-10 engineers across teams and read the last few months of onboarding questions and repeated Slack questions. Group them: "could not find it", "found it but wrong", "does not exist".
- Pick ONE painful area (typically on-call runbooks or service onboarding) and make it good, with the team's leads. A visible win before any policy is what buys credibility.
- Lower the cost of writing: a short template, docs stored next to code, and writing time budgeted in sprint planning as real capacity (meaning the team plans fewer feature points so writing is not unpaid overtime). A filled-in template for a service runbook is only four lines:
Purpose: Restart and drain the invoice worker when the queue backs up
Owner: payments team (#payments-oncall)
How to operate: 1) check queue depth on the dashboard 2) run the drain command 3) confirm depth falls
Known failure modes: drain hangs if the database is read-only; escalate to the database team
- Record a baseline: median time for a new hire to make a first merged change, and the count of repeat questions in the help channel. For example, the baseline might be 21 days and 30 repeat questions a month (illustrative), so later quarters have something to be compared against.
Quarter 2: ownership and governance across teams
- Every doc has an owning team (a field in the doc header plus a CODEOWNERS file, a repository file naming who must review changes in a path). No owner means it is archived after a notice period.
- A lightweight cross-team docs guild (a group with one rep per team, meeting monthly) owns the template, the style guide and the definition of "done" (the team's checklist of what must be true before a piece of work counts as finished). Teams keep autonomy on content; the guild owns only standards. This is a hybrid between central control and free-for-all.
- Add CI checks (automated checks on each pull request): broken links, missing owner, doc past its review date. Warn first, block later.
- Make docs part of the "definition of done" (the checklist a piece of work must satisfy before it counts as finished) for new services and for changes to public interfaces, and reflect it in review, not in a separate approval board.
Quarter 3: reinforce and scale
- Recognition: docs work counts in promotion and performance packets (the written evidence engineers submit when they are reviewed for a raise or promotion), and a short "doc of the month" is shared. Incentives beat exhortation.
- Onboarding curriculum: a 30-day reading path made of the docs that exist, where the new hire files an issue for each gap, and fixing it is their first contribution.
- Retire or merge stale docs so trust in what remains rises.
How to show progress each quarter (worked example)
A leading indicator is an early signal of behavior you control (services with owners); an outcome indicator is the result you actually want (faster onboarding), which moves later.
| Quarter | Leading indicator | Outcome indicator |
|---|---|---|
| 1 | Baseline recorded (for example 21 days to first accepted pull request, 30 repeat questions a month); pain area chosen | Repeat questions in that area (count) |
| 2 | Share of services with an owner, for example 40 of 50 = 80% | Docs past review date, falling |
| 3 | Share of new services shipped with a runbook | New-hire time to first merged change vs the Q1 baseline (for example 21 days down to 14, a one-third drop; illustrative) |
Report the trend against the baseline, and never celebrate volume (pages written), because it rewards padding.
ML organization variant (reproducibility as the goal)
In a machine learning (ML) team, the doc that matters most is whatever lets someone else reproduce a result. Replace the generic template with: a model card (a short standard doc stating what a model is for, its training data, evaluation results and known limits), the exact dataset version, code commit, random seed (the fixed number that makes random steps in training repeat identically) and environment, and the metrics table. A minimal model card reads: "Purpose: rank support tickets by urgency. Training data: tickets Jan to Jun, dataset v4. Evaluation: 0.82 accuracy on held-out tickets (illustrative). Limits: not tested on non-English tickets." CI checks verify the card exists and its links to data and code resolve before a model can be promoted. Ownership is per model, and the curriculum includes re-running one past experiment from docs alone; failure to reproduce it is a documentation bug to fix. Quality over time is measured by the share of promoted models whose results a second person reproduced from the docs.
Pitfalls
- Mandating with no time budget produces low-quality docs and resentment.
- Central review boards become bottlenecks; keep standards central and content federated.
- Measuring page counts, or reading only engagement, hides whether docs are trusted.
- What would change the plan: if the survey shows the problem is findability, not absence, invest in search and structure before asking anyone to write more.
A long chat thread ended in a decision, but the outcome and the follow-up actions are buried in it. How do you turn it into a durable decision note the team can find later, what does it always include, and where does it link back to the discussion?
Sample Answer
Direct answer
Write a short decision note (also called a decision record) that leads with the outcome, states why, lists the actions with owners and dates, and links back to the thread. Store it in a searchable, permanent place (a decisions folder or wiki page tied to the project), and post the link in the original thread so the chat points to the note.
Terms
- Decision record (ADR-style note): a short doc capturing one decision. ADR means architecture decision record, and the same shape works for product or team decisions.
- Owner: the one named person who ensures an action happens.
- Durable: findable and still understandable a year later.
What every decision note includes
- Title as a statement: "Use weekly releases for the mobile app", not "Release discussion".
- Status and date: proposed (under discussion), accepted (agreed and in force), or superseded (replaced by a newer decision; keep the old note, mark it superseded and link to its replacement).
- Decision: one or two sentences.
- Context: the problem and constraints at the time.
- Options considered and why each was rejected.
- Consequences and trade-offs accepted.
- Actions: each with owner and due date.
- Decision makers and who was consulted.
- Links: to the thread, related tickets and any superseded note.
- Review trigger: what would cause us to revisit.
Process
- Find the decision. Scroll to where the group converged, and check the last few messages for dissent.
- Draft the note from the thread within a day, while memory is fresh.
- Ask the decision maker to confirm it in the thread, which turns a summary into a shared record.
- Publish it and reply in the thread: "Decision recorded here: link."
- Copy actions into the ticket tracker with links back to the note.
Short note versus full ADR
The shape is the same; the difference is depth. The short note above (about 200 words) fits reversible or low-cost decisions: a tool choice for one team, a release cadence. Write a full ADR when the decision is hard or costly to reverse, affects several teams or systems, or needs sign-off: it adds a fuller options analysis with pros and cons for each option, cost and risk estimates, an explicit list of who was consulted and who approved, and it goes through a review (usually a pull request) before it is accepted. A rule of thumb: if undoing the decision would take more than a few weeks of work, or another team will build on it, use the full form.
Where it links back
Chat threads are hard to search and may expire, so the note quotes the essential reasoning itself and carries a permalink (a permanent direct link to the thread). If the chat tool deletes history, the note still stands. Export or screenshot the thread into an attachment when the retention policy is short.
Worked example (illustrative)
Thread: 60 messages about whether the data team should migrate a dashboard tool. Note:
- Title: "Keep the existing dashboard tool through Q4."
- Decision: postpone migration.
- Context: two analysts on leave, migration needs about six weeks of effort, and the current tool meets the compliance need.
- Options: migrate now, migrate in Q1, keep. Rejected "now" for capacity.
- Actions: Sam to add a migration estimate to the Q1 plan by the 15th.
- Review trigger: if license renewal price rises materially.
- Thread link.
Someone who joins in six months reads 200 words instead of 60 messages.
Trade-offs and pitfalls
- Summaries written by one person can misrepresent, so get confirmation.
- Recording what was chosen but not why leaves the same debate to recur.
- Keep it short. A long note that nobody reads is another buried thread.
Engineers at your company rarely write documentation unless someone forces them to. Design an incentive program that encourages meaningful contributions. How would you judge quality, stop people from gaming it, and pilot it before a wide rollout?
Sample Answer
Direct answer
Pay for outcomes that readers experience (docs that get used, kept current and that reduce questions), not for volume. I would build a small program with three parts: visible recognition and career credit rather than cash, a quality signal that comes from readers and from freshness, and guardrails against gaming. Then pilot it with two or three teams for a quarter before scaling.
Terms
- Gaming: producing what the metric rewards without the value it was meant to stand for (for example, splitting one page into ten).
- Docs-as-code: writing docs as files in the code repository, reviewed like code.
- Stale doc: one whose owner has not confirmed it is correct within an agreed period.
- Pilot: a limited trial with a comparison baseline.
- README: the front-page file of a code repository that says what it is and how to run it.
- Swag: small branded goodies such as t-shirts or stickers.
- Freshness: whether the doc has been confirmed correct recently (for example within 90 days).
- Extrinsic rewards crowd out pride: paying or prizing people for something they would otherwise do because they care can make them care less, so the reward replaces the pride instead of adding to it.
Why people do not write docs today (diagnose first)
Usually it is not laziness: writing is invisible in reviews, there is no time budgeted, docs rot so writing feels wasted, and the payoff goes to someone else. An incentive on top of a broken environment fails, so the program pairs incentives with removing friction (templates, docs in the repo, docs time in sprint planning).
Incentives (ranked by what I would try first)
- Career credit: documentation impact counts in promotion and performance write-ups, with managers trained to ask about it.
- Recognition: a monthly "most useful doc" chosen by readers, plus named owners on high-traffic pages.
- Protected time: a small documentation allowance per sprint.
- Small rewards (swag, conference budget) as a garnish. Avoid cash per page.
Judging quality without counting words
| Signal | What it tells you | How it is gamed |
|---|---|---|
| Reader "was this helpful" votes and comments | Usefulness | Friends vote for friends; cap votes per reader per author and weight by distinct teams |
| Usage (views, links from code or tickets) | Whether people find it | Self-viewing; exclude the author's team |
| Freshness (owner confirmed within 90 days, code changed with the doc) | Whether it is current | Rubber-stamp "reviewed" clicks; audit a random sample by asking a peer to follow the doc |
| Reduction in repeat questions on that topic | Real impact | Slow to measure; use as a trend, not a score |
| A monthly sample of ten docs is scored by a peer using a short rubric (accurate, complete for the task, clear next step; 0 to 2 points each, so 0 to 6) and this sample outranks the automated numbers. |
Worked check on one page: last confirmed 2026-07-01, so on 2026-09-28 it is 89 days old and passes the 90-day rule. A page confirmed 2026-06-20 is 100 days old and is flagged stale. The first page has 14 helpful votes from 5 different teams and a peer score of 5 of 6, so it is a strong result. A page with 14 votes all from the author's own team and a peer score of 2 of 6 is not, however popular it looks.
Anti-gaming guardrails
- Never reward raw count or length.
- Credit revisions and deletions of duplicates as much as new pages.
- Rate limits on credited edits per person and per page, and a minimum age before the same page earns credit again.
- Publish the rules so gaming is visible, and review outliers by hand.
Pilot design
Pick two teams that volunteer plus one comparison team without the program. Baselines before start: share of services with a current README, and median time for a new hire to make a first change (measure, don't guess). Run for one quarter, then interview participants and read the sampled docs. Success means participants' freshness and reader ratings rise versus the comparison team without a jump in trivial pages. Illustrative numbers: at baseline 6 of 20 services (30 percent) have a current README and a new hire's median time to a first change is 12 days. After a quarter, the pilot teams are at 14 of 20 (70 percent) and 7 days while the comparison team went from 30 to 35 percent. The gap, not the raw rise, is the signal. Adjust rules, then extend in waves.
Worked example (illustrative)
A team splits one 3-page setup guide into nine tiny pages to raise its count. Under the program the credit came from reader-rated usefulness and freshness, so nine thin pages earn less, the peer sample flags fragmentation, and the reviewer merges them, crediting the merge.
Trade-offs and pitfalls
- Extrinsic rewards can crowd out pride in the work, so keep rewards modest and tied to helping others.
- Heavy scoring becomes a job of its own. Keep to three or four signals.
- What would change my call: in a very small team, direct manager expectations beat a formal program.
Teams complain that searching the documentation returns irrelevant results. How would you diagnose the problem: what data would you collect, what are the most likely root causes, and how would you improve relevance over time?
Sample Answer
Direct answer
Treat "irrelevant results" as a measurable problem. First collect data on what people search for and what happens next, then sort failures into a few root causes (nothing relevant exists, it exists but is not found, it ranks too low, or it is stale or duplicated), fix the cheapest and most frequent cause first, and keep improving with a small recurring review of failed searches.
Terms
- Rank: a result's position in the list (rank 1 is the top result).
- Index: the searchable copy of your content that the search tool actually looks through; a page missing from it can never be found.
- Synonym list / alias metadata: a table telling search that "k8s" means "Kubernetes", or a field on a page listing other names people use for it.
- Boost: a setting that pushes matching pages up (or down) in the ranking, for example boosting recently updated pages.
Data to collect
- Search logs: query text, number of results, which result was clicked and at what rank, time on page after click, and whether the user searched again immediately (a reformulation).
- Zero-result rate: share of queries returning nothing.
- Abandonment: share of searches that showed results but ended with no click. Zero-result searches also end with no click, but count them under the zero-result rate so the two are not mixed.
- Content facts: last-updated date, owner, page views, duplicates per topic.
- Qualitative: ask ten engineers for the last thing they could not find, and watch them search.
Worked example (small, made-up log of 20 queries)
| Result of the search | Count | Share |
|---|---|---|
| Clicked a result at rank 1 to 3 | 9 | 45% |
| Clicked a result at rank 4 or lower | 3 | 15% |
| Zero results | 4 | 20% |
| Results shown but no click | 4 | 20% |
Rates: zero-result rate is 4 of 20 = 20%; abandonment (results shown, no click) is 4 of 20 = 20%. Together, 8 of 20 searches (40%) ended without a click, which is why the two are tracked separately.
Now trace every failure to a cause. The 4 zero-result queries: "k8s runbook" and "oncall rota" (the pages say "Kubernetes" and "on-call schedule"), a vocabulary mismatch that synonyms fix, plus "vpn setup for contractors" and "ssl cert renewal steps", which match nothing because no such page exists (missing content). The 4 no-click searches: two show a stale 2021 page first (stale content), and two show pages titled "Notes 3" and "Infra misc" that do not say what they cover (weak titles). The 3 low-rank clicks are ranking problems: the right page exists but sits at rank 4 or lower.
Tally: vocabulary 2, missing content 2, stale content 2, weak titles 2, ranking 3. That accounts for all 11 failed searches (9 succeeded). Eight of the 11 are content or vocabulary problems that are cheap to fix without touching ranking, so fix those first.
Most likely root causes and fixes
The first four rows are the core causes; the last two are checks to make once those are clean.
| Root cause | Sign in the data | Fix |
|---|---|---|
| Vocabulary mismatch (abbreviations, team jargon) | Zero results; users retry with other words | Synonym list, alias metadata |
| Poor ranking (old page above the good one) | Clicks at low rank; reformulations | Boost fresh or high-traffic or owned pages; demote stale ones |
| Duplicated or stale content | Several pages per topic; old dates | Merge, archive, add banners and review dates |
| Missing content | Zero or no-click on real needs | Write it; feed the top failed queries to owners |
| Weak titles and structure | Right page exists but low rank | Titles that say the question answered; headings |
| Index gaps (some spaces not indexed, permissions hiding pages) | Known page never appears | Fix the index and permission handling |
Improving relevance over time
- Build a test set of the 30 to 50 most common queries with the page that should win, and run it after every change to search settings; the measure is the share where the right page is in the top three.
- Review the top failed queries weekly for the first month, then monthly, and assign each to an owner.
- Add ranking signals gradually (freshness, popularity, ownership), one at a time, re-running the test set to see which helped.
- Keyword ranking formulas (such as BM25, a standard scoring method that favors pages containing rare query words) are often enough; add semantic (meaning-based) search only if failures persist after cleanup.
Pitfalls
- Tuning ranking before cleaning content wastes effort: no ranking can surface a page that is missing or wrong.
- Using click-through alone misleads: a click on a bad page counts as a success unless you also check whether the user searched again.
- What would change my order of work: if the zero-result rate dominates, fix vocabulary and gaps before touching ranking.
What documentation standards would you set for an engineering team, and how would you get people to follow them without turning it into bureaucracy? Cover how the standards are kept up to date as the work changes.
Sample Answer
Direct answer
Set a small, tiered minimum standard (a few artifacts every service, pipeline or model must have), keep the docs in the same repository as the code and change them in the same pull request (PR), and enforce the standard with automation and defaults rather than meetings and sign-offs. Freshness comes from ownership plus a review date on every doc, not from goodwill. The test of "not bureaucracy" is that following the standard is the easiest path and the checks are run by tooling.
Structured elaboration
1. The standard itself: small and tiered
- Tier 1 (anything on-call, customer-facing or shared): a README (the front-page file of a repo: what it is, who owns it, how to run it, where to ask), a runbook (step-by-step instructions for responding to an alert or failure), and a decision record (a short note recording what was chosen, the options rejected and why) for hard-to-reverse choices.
- Tier 2 (internal tools, experiments): README only.
- ML and data assets: a model card (short document stating what a model is for, its training data, known limits and owner) or a dataset datasheet, plus a data dictionary (what each column means, its unit and source). For analytics projects add required metadata: owner, refresh schedule, source tables and last-reviewed date.
- One template per artifact, pre-filled, with the required fields marked. Templates are what stop inconsistent deliverables and the rework that follows.
- The ML and data items are optional add-ons for teams that ship models or datasets, not part of the core README, runbook and decision record set.
A minimal template skeleton, so the required fields are concrete (the CI check simply fails if a heading is missing):
README.md RUNBOOK.md decision record
# <service> # Alert: <alert name> # Decision: <statement>
Owner: <team> Severity: <P1-P3> Status: proposed | accepted | superseded
Last reviewed: <d> Symptoms: <what you see> Context / Options / Consequences
## What it does Diagnosis: <steps> Owner and review date
## Run locally Fix or rollback: <steps>
## Ask for help Escalate to: <who>
2. Governance as one framework (owners, templates, review cadence, tooling, incentives)
- Templates: the pre-filled skeleton above is the only format accepted for each artifact, stored in one template repository, so the CI check has a fixed set of required headings to test.
- Owners: every doc names a team owner (a CODEOWNERS file, which routes PR review to the named owners, plus a service catalog entry, meaning a central list of every service with its owner, docs link and on-call contact). An unowned doc is treated as deletable.
- Review cadence: each doc carries a
last_revieweddate. Tier 1 is reviewed at least every 6 months, and also whenever the system changes shape. - Tooling: docs-as-code (markdown in the repo, reviewed like code), a CI (continuous integration) job that checks required headings, broken links and stale dates, and one searchable site built from all repos.
- Incentives: doc work counts in performance and promotion criteria, authors are credited by name, and "a new hire followed the README and it worked" is celebrated. Punishment alone breeds checkbox docs.
3. Keeping standards current as work changes
- A PR template asks "docs updated or not needed, and why". Changes to alerts, APIs or schemas require touching the linked doc (a CI rule can map code paths to doc paths).
- A bot opens a ticket on the owner when
last_reviewedpasses its limit. Two missed cycles moves the doc toarchivedrather than leaving a confident but wrong page. - The standard itself is versioned and reviewed twice a year by a small docs working group (a few volunteers from different teams who own the standard), and any rule that nobody can justify with an incident or onboarding pain is removed.
4. Measuring compliance and handling non-compliance
- Coverage = Tier 1 assets with all required docs / all Tier 1 assets. Freshness = docs reviewed inside their window / docs required.
- Escalation ladder: bot nudge, ticket on the owner, visible on the team's monthly ops review, and only for Tier 1 a launch-readiness blocker, meaning the service cannot go live until its Tier 1 docs exist. Nothing else is gated.
Worked example
One team owns 15 Tier 1 services. The dashboard shows 12 have a README and runbook, and 9 of those are inside the 180-day review window.
coverage = 12 / 15 = 80%
freshness = 9 / 15 = 60% (9 of the 15 required sets are fresh)
The team lead sees two numbers and six named services to fix (three with no docs, three with docs past their review window), not a lecture. The runbook and decision record carry the same Last reviewed field as the README, so the CI stale-date check and the 6-month cadence apply to every Tier 1 artifact, not only the README. Note freshness is computed against the required 15, so a missing doc cannot hide behind a fresh one.
Rollout at scale (500 engineers, ten teams, two regions), phased over 12 months
- Months 1-2: pilot with two willing teams, build templates and the CI check, fix what the pilot hates.
- Months 3-5: add four teams, stand up the single search site and the service catalog, run a 2-hour training per engineer (500 x 2 hours = 1,000 engineer-hours, which is 25 engineer-weeks at 40 hours, a small fraction of the organization's yearly capacity).
- Months 6-9: remaining four teams; migrate only Tier 1 legacy docs (archive the rest, do not migrate everything).
- Months 10-12: switch on the launch-readiness gate for Tier 1, publish the first compliance report, review the standard.
- Rough cost, in people rather than invented dollars: a docs platform owner and a tooling engineer (about 2 full-time), plus one steward per team (the team's named docs champion who helps peers and chases reviews) at about 10% time (ten teams is about 1 full-time equivalent). Two regions: one site, async PR review so no time zone waits on another, and one regional steward each.
Trade-offs and pitfalls
- Too many required docs produces empty template filling. Cut the list until each item has a story of an incident or onboarding delay it prevents.
- Central mandate without team stewards makes the standard feel imposed. Stewards give teams a peer to ask.
- Gating everything on compliance encourages gaming (docs written to satisfy a linter). Gate only Tier 1 and spot-check quality by having a new joiner try the README.
- What would change my approach: for a 15-person team I would skip the catalog and the phases and adopt README, runbook and a PR checkbox in one week.
Unlock Full Question Bank
Get access to all 12 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.