Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
Your knowledge base is stale and teams are repeating incidents because playbooks are obsolete. Propose a remediation program: what you clean up first, how stale content is detected automatically, how humans review, what stops recurrence, and a realistic timeline and staffing.
Sample Answer
Direct answer
Run a time-boxed remediation program aimed at the runbooks that pages and incidents actually touch. Clean up tier-1 procedures first, detect staleness automatically with checks against real infrastructure and incident history, require a human to walk each procedure before it is marked verified, and stop recurrence by making a working runbook a condition for shipping an alert and for closing a postmortem. Archive everything nobody claims instead of reviewing it.
Terms. A playbook or runbook is a step-by-step procedure for handling an operational event. A postmortem is a written review of an incident; blameless means it looks for system causes, not someone to blame. A tier-1 service is one whose failure pages people or hurts revenue; tier 2 is important but a failure degrades it rather than stopping it (for example an internal reporting tool). A paging alert is an alert that wakes someone up. A staging environment is a safe copy of production for testing. The inventory is the current list of live hosts and services, taken from your configuration database. A decommissioned host is one that has been shut down and removed.
1. What to clean up first
Rank runbooks by consequence, not by age:
- Runbooks linked from paging alerts on tier-1 services.
- Runbooks named in the last 12 months of postmortems as missing, wrong or ignored (those are the ones that already repeated an incident).
- Runbooks for tier-2 services.
- Everything else: put an archive notice on it and delete or archive after 30 days if no owner claims it.
2. Automatic staleness detection
- Nightly job: links, hostnames, service names and commands referenced in the runbook checked against the current inventory. A reference to a decommissioned host is a hard flag.
- Alert-to-runbook check: every paging alert's runbook URL must resolve and must have a verified date inside its interval.
- Incident hook: responders press "runbook was wrong or missing" during or after an incident, which creates a ticket automatically.
- What the nightly check produces, for one runbook:
db-failover step 4 references host db-eu-2 (decommissioned 2026-06-30): HARD FLAG. Ticket opened for payments-team, due in 5 business days. - Age since last verified, used as a weak signal only (age alone is a poor predictor; broken references and incident flags are stronger).
3. Human review
The owner walks the runbook step by step in a staging environment, or as a tabletop drill (a talked-through rehearsal), with a second person who did not write it. They fix wrong steps, record the date and both names, and only then is the page marked verified. Reading it and clicking "still accurate" does not count.
4. Preventing recurrence
- CI (continuous integration) lint: an automated check that runs on every proposed change and rejects it when a rule is broken, here: an alert definition cannot merge without a runbook link. Illustrative output of a small custom lint script (no standard tool prints exactly this; you define the wording) when it fails:
FAIL alerts/checkout.yaml: "CheckoutLatencyHigh" has severity=page but no runbook_url
FAIL alerts/search.yaml: runbook_url https://wiki.example.com/runbooks/search-old returns 404
- Postmortem action items include "update the runbook" and the postmortem cannot close until the change is merged.
- Runbook lives next to the service code and its owner appears in the repo's CODEOWNERS file (a file that maps paths to owning teams), so changes get reviewed.
- Quarterly drill of the tier-1 runbooks (a game day, a rehearsal where you deliberately trigger a failure).
5. Realistic timeline and staffing
Illustrative sizing: 400 runbooks, of which 40 are tier 1 and 100 are tier 2.
| Phase | Weeks | Work | Effort |
|---|---|---|---|
| 0. Inventory and triage | 1-2 | Export all runbooks, tag tier and owner | Program lead, part-time |
| 1. Tier 1 | 3-8 | Walk and fix 40 runbooks, 1.5 hours each: 60 hours | Owners of each service: 60 hours over 6 weeks is about 10 hours per week in total, so with roughly 10 owning teams that is 1 to 2 hours per week each |
| 2. Detection and CI guard | 4-10 | Build nightly reference check, alert lint | One engineer, about two weeks |
| 3. Tier 2 | 9-16 | 100 runbooks at 1 hour: 100 hours | Owners |
| 4. Long tail | 12-16 | Archive unclaimed after 30-day notice | Program lead |
Staffing: one part-time program lead, one engineer for tooling, and a small time budget from each owning team. The human review effort totals about 160 hours (60 for tier 1 plus 100 for tier 2), which is roughly 4 person-weeks spread across teams, not one team's burden.
Success measures. No paging alert without a verified runbook; postmortems citing "runbook wrong" decline quarter over quarter; nightly reference-check failures trend to zero.
Pitfalls
- Reviewing all 400 equally wastes the program's budget on the long tail.
- A one-time cleanup decays; without the CI and postmortem gates you are back here in a year.
- What would change my call: if the team is tiny, skip the tooling and run a monthly tier-1 drill instead.
An internal wiki and search platform saw heavy use for a month after launch and then usage fell sharply. How would you diagnose why, and what prioritized plan would you set to bring adoption back, including quick fixes, governance and how you would measure success?
Sample Answer
Direct answer
A month of heavy use followed by a sharp fall is usually a novelty spike hitting a product that does not fit people's daily work: the content was thin or stale, search did not find things, or nothing sent people back. So diagnose from data first (search behaviour, who left, what they wanted), then interview a handful of lapsed users, and only then choose fixes. Plan in four steps: quick fixes within two weeks, habit hooks within six weeks, governance within a quarter, and a review of the platform choice last, with metrics defined before you start so success is not just "more page views".
Terms
- Novelty spike: a burst of use because something is new, not because it is useful.
- Launch cohort retention: of the people who used it in the launch month, the share still using it now.
- Zero-result search: a search that returned nothing.
- Habit hook: a place in people's normal daily tools that sends them to the wiki without them having to remember it.
- Chat link previews: when a wiki link pasted in chat expands to show the page title and first lines.
- Bounce: a visit where the person leaves after one page without doing anything else.
- Synonyms (in search): telling the search engine that different words mean the same thing.
Step 1: Diagnose (week 1)
| Question | Evidence |
|---|---|
| Who left: everyone, or one team? | Weekly active readers by team, launch cohort retention |
| Did people find what they searched for? | Zero-result searches, searches with no click |
| Was content missing or wrong? | Top queries versus existing pages, page last-verified dates |
| Is there a habit hook? | Are there links to the wiki from the tools people use all day (chat, alerts, code repos)? |
| Do people trust it? | Five to eight short interviews with people who stopped: "the last time you needed X, where did you go?" |
Metric definitions (numerator, denominator, window)
- Weekly active readers = distinct engineers who opened a page or searched in a 7-day window, divided by engineers eligible to use it.
- Search success rate = searches followed by a click on a result within the same session, divided by all searches.
- Zero-result rate = searches returning nothing, divided by all searches.
- Freshness = pages viewed in the window that were verified within their review interval, divided by pages viewed.
Illustrative worked example. Suppose 200 engineers. Launch month: 150 weekly active (75 percent). Now: 60 (30 percent). Of 1,000 searches last week, 320 returned nothing (32 percent) and 400 led to a click (40 percent). Sort the 1,000 searches into three groups: 320 returned nothing (32 percent), 280 returned results but nobody clicked (28 percent), and 400 led to a click (40 percent). So 600 searches, six in ten, ended without a click, and only 40 percent succeeded. Roughly half of the failures (320 of 600) mean missing content or a vocabulary mismatch, and the other 280 mean results were shown but looked irrelevant or stale. Both are search problems, so search is the main leak and search fixes come before governance.
Step 2: Prioritized plan
- Quick fixes (weeks 1-2): create or redirect pages for the top zero-result queries, add synonyms (for example "on-call" and "pager"), fix or archive the 20 most-viewed stale pages, and add chat link previews so wiki links are useful where people already talk.
- Habit hooks (weeks 2-6): link runbooks from alerts, link onboarding checklists, and make "put it in the wiki" the default answer in team channels.
- Governance (weeks 4-12): each space has an owner, each page a review-by date, stale pages get a visible "unverified" banner then archive, and there is a simple contribution guide with a page template.
- Review the platform choice last. Replacing the tool is expensive and rarely the root cause; only reopen it if search relevance cannot be fixed.
Measuring success
Target weekly active readers recovering to a stated level (say back toward 60 percent: that is a judgment call, not a law. Launch was 75 percent, and part of that was curiosity, so recovering about half of the gap between today's 30 percent and the launch peak of 75 percent (a 45-point gap, so about 22 points back), roughly 52 percent, is the floor and 60 percent is a stretch. Agree the number with the sponsor up front) sustained for eight weeks, search success rising, zero-result rate falling, and the freshness ratio rising. Also track contribution (pages edited by non-authors).
Pitfalls
- Raw page views are easy to game and inflated by bots or bounces, so pair them with search success and freshness.
- A mandate ("everyone must use the wiki") gets compliance, not trust.
- What would change my call: if interviews show the content lives in a different tool people already prefer, integrate with that instead of competing.
Teams complain that searching the documentation returns irrelevant results. How would you diagnose the problem: what data would you collect, what are the most likely root causes, and how would you improve relevance over time?
Sample Answer
Direct answer
Treat "irrelevant results" as a measurable problem. First collect data on what people search for and what happens next, then sort failures into a few root causes (nothing relevant exists, it exists but is not found, it ranks too low, or it is stale or duplicated), fix the cheapest and most frequent cause first, and keep improving with a small recurring review of failed searches.
Terms
- Rank: a result's position in the list (rank 1 is the top result).
- Index: the searchable copy of your content that the search tool actually looks through; a page missing from it can never be found.
- Synonym list / alias metadata: a table telling search that "k8s" means "Kubernetes", or a field on a page listing other names people use for it.
- Boost: a setting that pushes matching pages up (or down) in the ranking, for example boosting recently updated pages.
Data to collect
- Search logs: query text, number of results, which result was clicked and at what rank, time on page after click, and whether the user searched again immediately (a reformulation).
- Zero-result rate: share of queries returning nothing.
- Abandonment: share of searches that showed results but ended with no click. Zero-result searches also end with no click, but count them under the zero-result rate so the two are not mixed.
- Content facts: last-updated date, owner, page views, duplicates per topic.
- Qualitative: ask ten engineers for the last thing they could not find, and watch them search.
Worked example (small, made-up log of 20 queries)
| Result of the search | Count | Share |
|---|---|---|
| Clicked a result at rank 1 to 3 | 9 | 45% |
| Clicked a result at rank 4 or lower | 3 | 15% |
| Zero results | 4 | 20% |
| Results shown but no click | 4 | 20% |
Rates: zero-result rate is 4 of 20 = 20%; abandonment (results shown, no click) is 4 of 20 = 20%. Together, 8 of 20 searches (40%) ended without a click, which is why the two are tracked separately.
Now trace every failure to a cause. The 4 zero-result queries: "k8s runbook" and "oncall rota" (the pages say "Kubernetes" and "on-call schedule"), a vocabulary mismatch that synonyms fix, plus "vpn setup for contractors" and "ssl cert renewal steps", which match nothing because no such page exists (missing content). The 4 no-click searches: two show a stale 2021 page first (stale content), and two show pages titled "Notes 3" and "Infra misc" that do not say what they cover (weak titles). The 3 low-rank clicks are ranking problems: the right page exists but sits at rank 4 or lower.
Tally: vocabulary 2, missing content 2, stale content 2, weak titles 2, ranking 3. That accounts for all 11 failed searches (9 succeeded). Eight of the 11 are content or vocabulary problems that are cheap to fix without touching ranking, so fix those first.
Most likely root causes and fixes
The first four rows are the core causes; the last two are checks to make once those are clean.
| Root cause | Sign in the data | Fix |
|---|---|---|
| Vocabulary mismatch (abbreviations, team jargon) | Zero results; users retry with other words | Synonym list, alias metadata |
| Poor ranking (old page above the good one) | Clicks at low rank; reformulations | Boost fresh or high-traffic or owned pages; demote stale ones |
| Duplicated or stale content | Several pages per topic; old dates | Merge, archive, add banners and review dates |
| Missing content | Zero or no-click on real needs | Write it; feed the top failed queries to owners |
| Weak titles and structure | Right page exists but low rank | Titles that say the question answered; headings |
| Index gaps (some spaces not indexed, permissions hiding pages) | Known page never appears | Fix the index and permission handling |
Improving relevance over time
- Build a test set of the 30 to 50 most common queries with the page that should win, and run it after every change to search settings; the measure is the share where the right page is in the top three.
- Review the top failed queries weekly for the first month, then monthly, and assign each to an owner.
- Add ranking signals gradually (freshness, popularity, ownership), one at a time, re-running the test set to see which helped.
- Keyword ranking formulas (such as BM25, a standard scoring method that favors pages containing rare query words) are often enough; add semantic (meaning-based) search only if failures persist after cleanup.
Pitfalls
- Tuning ranking before cleaning content wastes effort: no ranking can surface a page that is missing or wrong.
- Using click-through alone misleads: a click on a bad page counts as a success unless you also check whether the user searched again.
- What would change my order of work: if the zero-result rate dominates, fix vocabulary and gaps before touching ranking.
What documentation standards would you set for an engineering team, and how would you get people to follow them without turning it into bureaucracy? Cover how the standards are kept up to date as the work changes.
Sample Answer
Direct answer
Set a small, tiered minimum standard (a few artifacts every service, pipeline or model must have), keep the docs in the same repository as the code and change them in the same pull request (PR), and enforce the standard with automation and defaults rather than meetings and sign-offs. Freshness comes from ownership plus a review date on every doc, not from goodwill. The test of "not bureaucracy" is that following the standard is the easiest path and the checks are run by tooling.
Structured elaboration
1. The standard itself: small and tiered
- Tier 1 (anything on-call, customer-facing or shared): a README (the front-page file of a repo: what it is, who owns it, how to run it, where to ask), a runbook (step-by-step instructions for responding to an alert or failure), and a decision record (a short note recording what was chosen, the options rejected and why) for hard-to-reverse choices.
- Tier 2 (internal tools, experiments): README only.
- ML and data assets: a model card (short document stating what a model is for, its training data, known limits and owner) or a dataset datasheet, plus a data dictionary (what each column means, its unit and source). For analytics projects add required metadata: owner, refresh schedule, source tables and last-reviewed date.
- One template per artifact, pre-filled, with the required fields marked. Templates are what stop inconsistent deliverables and the rework that follows.
- The ML and data items are optional add-ons for teams that ship models or datasets, not part of the core README, runbook and decision record set.
A minimal template skeleton, so the required fields are concrete (the CI check simply fails if a heading is missing):
README.md RUNBOOK.md decision record
# <service> # Alert: <alert name> # Decision: <statement>
Owner: <team> Severity: <P1-P3> Status: proposed | accepted | superseded
Last reviewed: <d> Symptoms: <what you see> Context / Options / Consequences
## What it does Diagnosis: <steps> Owner and review date
## Run locally Fix or rollback: <steps>
## Ask for help Escalate to: <who>
2. Governance as one framework (owners, templates, review cadence, tooling, incentives)
- Templates: the pre-filled skeleton above is the only format accepted for each artifact, stored in one template repository, so the CI check has a fixed set of required headings to test.
- Owners: every doc names a team owner (a CODEOWNERS file, which routes PR review to the named owners, plus a service catalog entry, meaning a central list of every service with its owner, docs link and on-call contact). An unowned doc is treated as deletable.
- Review cadence: each doc carries a
last_revieweddate. Tier 1 is reviewed at least every 6 months, and also whenever the system changes shape. - Tooling: docs-as-code (markdown in the repo, reviewed like code), a CI (continuous integration) job that checks required headings, broken links and stale dates, and one searchable site built from all repos.
- Incentives: doc work counts in performance and promotion criteria, authors are credited by name, and "a new hire followed the README and it worked" is celebrated. Punishment alone breeds checkbox docs.
3. Keeping standards current as work changes
- A PR template asks "docs updated or not needed, and why". Changes to alerts, APIs or schemas require touching the linked doc (a CI rule can map code paths to doc paths).
- A bot opens a ticket on the owner when
last_reviewedpasses its limit. Two missed cycles moves the doc toarchivedrather than leaving a confident but wrong page. - The standard itself is versioned and reviewed twice a year by a small docs working group (a few volunteers from different teams who own the standard), and any rule that nobody can justify with an incident or onboarding pain is removed.
4. Measuring compliance and handling non-compliance
- Coverage = Tier 1 assets with all required docs / all Tier 1 assets. Freshness = docs reviewed inside their window / docs required.
- Escalation ladder: bot nudge, ticket on the owner, visible on the team's monthly ops review, and only for Tier 1 a launch-readiness blocker, meaning the service cannot go live until its Tier 1 docs exist. Nothing else is gated.
Worked example
One team owns 15 Tier 1 services. The dashboard shows 12 have a README and runbook, and 9 of those are inside the 180-day review window.
coverage = 12 / 15 = 80%
freshness = 9 / 15 = 60% (9 of the 15 required sets are fresh)
The team lead sees two numbers and six named services to fix (three with no docs, three with docs past their review window), not a lecture. The runbook and decision record carry the same Last reviewed field as the README, so the CI stale-date check and the 6-month cadence apply to every Tier 1 artifact, not only the README. Note freshness is computed against the required 15, so a missing doc cannot hide behind a fresh one.
Rollout at scale (500 engineers, ten teams, two regions), phased over 12 months
- Months 1-2: pilot with two willing teams, build templates and the CI check, fix what the pilot hates.
- Months 3-5: add four teams, stand up the single search site and the service catalog, run a 2-hour training per engineer (500 x 2 hours = 1,000 engineer-hours, which is 25 engineer-weeks at 40 hours, a small fraction of the organization's yearly capacity).
- Months 6-9: remaining four teams; migrate only Tier 1 legacy docs (archive the rest, do not migrate everything).
- Months 10-12: switch on the launch-readiness gate for Tier 1, publish the first compliance report, review the standard.
- Rough cost, in people rather than invented dollars: a docs platform owner and a tooling engineer (about 2 full-time), plus one steward per team (the team's named docs champion who helps peers and chases reviews) at about 10% time (ten teams is about 1 full-time equivalent). Two regions: one site, async PR review so no time zone waits on another, and one regional steward each.
Trade-offs and pitfalls
- Too many required docs produces empty template filling. Cut the list until each item has a story of an incident or onboarding delay it prevents.
- Central mandate without team stewards makes the standard feel imposed. Stewards give teams a peer to ask.
- Gating everything on compliance encourages gaming (docs written to satisfy a linter). Gate only Tier 1 and spot-check quality by having a new joiner try the README.
- What would change my approach: for a 15-person team I would skip the catalog and the phases and adopt README, runbook and a PR checkbox in one week.
Documentation for services keeps becoming undiscoverable when code is refactored or renamed. Design a way to link documentation to the code it describes so that the links survive change: what metadata you record, where checks run, and how broken links are detected and repaired.
Sample Answer
Direct answer
Link docs to code through stable identifiers (a service name and the paths or symbols it owns), not through line numbers or pasted URLs. Store those identifiers as metadata in each doc, run a check in CI (continuous integration, the automated build that runs on every pull request) that fails when a link no longer resolves, and run a scheduled sweep for links that rot without a PR touching them. When a break is found, the tooling proposes the repair (for example the renamed path) so a human only approves.
1. What metadata you record
Put it in the doc's front matter (a small YAML header at the top of a Markdown file), so it travels with the doc in the same repository ("docs-as-code": docs live in version control and go through review like code).
---
service: billing
owner_team: payments
covers:
- src/billing/
last_verified_commit: abc1234
review_by: 2026-12-01
---
covers: paths (or symbol names for API docs, where a symbol is a named code element such as a function or class) the doc describes. This is the forward link, doc to code.owner_teamplus a CODEOWNERS file (a repository file mapping paths to the people who must review changes there) gives the reverse link and someone to notify.last_verified_commitandreview_by: when a human last confirmed the doc matched the code. A commit SHA (the unique hash Git gives every commit, such asabc1234) records exactly which version of the code was checked.- The
service: billingvalue is an identifier from a service catalog (an internal registry listing each service, its owning team and where its code lives). If the code moves from one repository to another, only the catalog entry (billing -> repo: platform-monorepo, path: services/billing/) is updated, and every doc that saysservice: billingstill resolves. Docs that hard-code a repository URL would all break. - Links from doc prose to code use a stable anchor (a symbol or section name). A permalink pinned to a commit SHA never breaks but shows old code, so use it only as evidence ("this was true at abc1234"), and use the live path for navigation.
2. Where checks run
- On every pull request: compute the files changed; find docs whose
coversoverlap; require either a doc edit or an explicit "docs not needed" label. Also run the link resolver below on the whole docs folder. - Nightly sweep: catches external links, docs past
review_by, and services whose owner team no longer exists. - Reverse index: a generated page listing each service and its docs, so engineers finding code can find docs.
3. Detecting and repairing broken links (worked example)
A PR renames src/billing/ to src/invoicing/. The check below reads each doc's covers, confirms each prefix still matches a tracked file, and uses git diff -M (which detects renames) to suggest the new location. Setup: pip install pyyaml; a repo with src/billing/charge.py, src/billing/refund.py and the doc above (saved as docs/billing-runbook.md) committed on branch main; then create a branch, run git mv src/billing src/invoicing and commit.
import subprocess, sys, pathlib, yaml
def front_matter(path):
text = path.read_text()
if not text.startswith("---\n"):
return None
return yaml.safe_load(text.split("---\n", 2)[1])
def git(*args):
return subprocess.run(["git", *args], capture_output=True, text=True, check=True).stdout
base = sys.argv[1] # e.g. origin/main
tracked = git("ls-files").splitlines()
renames = {}
for line in git("diff", "--name-status", "-M", base, "HEAD").splitlines():
parts = line.split("\t")
if parts[0].startswith("R"):
renames[parts[1]] = parts[2]
problems = 0
for doc in pathlib.Path("docs").rglob("*.md"):
meta = front_matter(doc)
if not meta:
print(f"{doc}: MISSING front matter"); problems += 1; continue
for prefix in meta.get("covers", []):
if any(f.startswith(prefix) for f in tracked):
continue
hint = next((new for old, new in renames.items() if old.startswith(prefix)), None)
msg = f"{doc}: covers '{prefix}' matches no file"
if hint:
msg += f" (renamed to {hint.rsplit('/', 1)[0]}/)"
print(msg); problems += 1
print(f"{problems} problem(s)")
sys.exit(1 if problems else 0)
Run as python3 check_doc_links.py main (the argument is the branch to compare against). Output:
docs/billing-runbook.md: covers 'src/billing/' matches no file (renamed to src/invoicing/)
1 problem(s)
How it works, block by block: front_matter reads the YAML header of a doc. renames is a dictionary built from git diff -M (the -M flag makes Git detect that a file was moved rather than deleted and re-added), mapping old path to new path, for example src/billing/charge.py to src/invoicing/charge.py. For each doc and each covers prefix, the loop asks whether any tracked file still starts with that prefix (startswith). If none does, the link is broken, and the loop looks for a renamed file whose old path started with that prefix to suggest the repair.
The CI job exits non-zero, so the rename PR cannot merge until the doc's covers is updated. A bot can open that one-line fix automatically, because the rename target is already known.
4. Trade-offs and pitfalls
- Strict blocking on every code change creates label-spam ("docs not needed" clicked reflexively). Start warn-only for a quarter, then block only for services tagged tier-1 (the most business-critical services, for example those that can page someone at night).
- Path prefixes are coarse: a rename inside the folder is not detected. Symbol-level links are more precise but need per-language tooling, so reserve them for public API docs.
- Metadata rots too.
review_bywith a nightly report of expired docs is the guard. - Docs living outside the repo (a wiki) cannot get a PR check. Either move the ones that matter into the repo, or accept only the nightly sweep for them.
Unlock Full Question Bank
Get access to all 20 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.