Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
How would you measure whether an SRE team's documented procedures are going stale or being forgotten? Define the inputs and how you would combine them, when a score should trigger action, and how you would avoid alert fatigue.
Sample Answer
Direct answer
Treat "stale or forgotten" as several independent signals, score each 0 to 1, combine them with fixed weights into a 0 to 100 staleness score per runbook, and set the action threshold by how critical the service is. Add a hard override: if a real incident said the procedure was wrong, act regardless of the score. Avoid alert fatigue by sending weekly digests by default and opening a ticket only past the threshold, capped per owner.
Terms. A runbook is a step-by-step procedure for an operational task. An SRE (site reliability engineer) team runs production and relies on runbooks during incidents. A tier ranks a service's criticality (tier 1 is paging or revenue-critical). A paging alert is an alert that wakes someone up. Drilled means rehearsed on purpose before a real emergency. A dry run executes the steps against a safe copy with no real effect. The inventory is the current list of live hosts and services (from your configuration database). A digest is one summary message instead of many alerts. To snooze is to silence a reminder until a set date. An archive candidate is a runbook that looks safe to retire. To deduplicate is to keep one item where there would be several.
Inputs (each scaled 0 to 1, where 1 is worst)
| Signal | Definition | Weight |
|---|---|---|
| Age | Days since last review divided by the tier's review interval, capped at 2, then halved | 0.25 |
| Unexercised | Same calculation using days since the runbook was last used or drilled | 0.15 |
| Proved wrong | Incidents in 90 days where responders flagged it wrong, divided by incidents that used it | 0.30 |
| Dead references | Links, hosts or commands that no longer exist, divided by total references (found by a nightly check against the inventory) | 0.20 |
| Owner gone | 1 if the owner left or the team was reorganized, else 0 | 0.10 |
Why these shapes and weights:
- Capped at 2, then halved. Age is measured against the review interval. A runbook exactly at its interval scores 1 before halving, so 0.5 after. Capping at 2 means anything more than twice overdue is simply maximally stale (1.0 after halving) rather than letting a 10-times-overdue page dominate the score.
- Proved wrong weighs most (0.30) because it is direct evidence from a real incident. Dead references (0.20) are hard, checkable facts. Age and unexercised (0.25, 0.15) are only proxies. Owner gone (0.10) is a flag more than a measurement.
Weights are a judgement. Calibrate them by backtesting, meaning you replay the past: compute scores as they stood last quarter, then check them against what actually went wrong afterwards. Illustrative example: 4 runbooks were flagged wrong in this quarter's incidents. If 3 of those 4 were among last quarter's top-scoring runbooks, the score caught 75%. If only 1 of 4 was, raise the weight on dead references or add a signal, and re-check.
When to act
Tier 1 opens a ticket at 45 or above, tier 2 at 60, tier 3 at 80. Tier 1 opens earliest because a stale runbook there costs the most during an outage, and its review interval is shortest (90 days, versus 180 and 365). These numbers are starting values you tune from the backtest, not laws. Below the threshold the runbook appears only in the weekly digest. Override: one or more "proved wrong" flags opens a ticket whatever the score.
Runnable computation (pinned data, prints the result)
from datetime import date
TODAY = date(2026, 9, 28)
WEIGHTS = {"age": 0.25, "unexercised": 0.15, "wrong": 0.30, "dead_refs": 0.20, "owner": 0.10}
# tier -> (score that opens a ticket, review interval in days)
TIERS = {1: (45, 90), 2: (60, 180), 3: (80, 365)}
# name, tier, last_reviewed, last_exercised (used or drilled), uses_in_90d, flagged_wrong_in_90d,
# dead_refs, total_refs, owner_active
RUNBOOKS = [
("db-failover", 1, date(2026, 2, 10), date(2026, 3, 1), 1, 1, 3, 10, True),
("cache-flush", 2, date(2026, 8, 20), date(2026, 9, 10), 2, 0, 0, 6, True),
("legacy-batch-job", 3, date(2025, 6, 1), date(2025, 6, 1), 0, 0, 4, 5, False),
("cert-rotation", 1, date(2026, 8, 30), date(2026, 9, 1), 0, 0, 1, 12, True),
("queue-backlog", 2, date(2026, 3, 15), date(2026, 8, 2), 3, 0, 1, 8, True),
]
def score(tier, reviewed, exercised, uses, wrong, dead, total, owner_active):
interval = TIERS[tier][1]
parts = {
"age": min((TODAY - reviewed).days / interval, 2) / 2,
"unexercised": min((TODAY - exercised).days / interval, 2) / 2,
"wrong": (wrong / uses) if uses else 0.0,
"dead_refs": dead / total,
"owner": 0.0 if owner_active else 1.0,
}
return 100 * sum(WEIGHTS[k] * v for k, v in parts.items()), parts
def action(tier, s, wrong):
if wrong >= 1:
return "OPEN TICKET (incident said it was wrong)"
return "open ticket" if s >= TIERS[tier][0] else "weekly digest only"
print(f"{'runbook':<17}{'tier':>4}{'score':>7} action")
for name, tier, rv, ex, uses, wrong, dead, total, owner in RUNBOOKS:
s, _ = score(tier, rv, ex, uses, wrong, dead, total, owner)
print(f"{name:<17}{tier:>4}{s:>7.1f} {action(tier, s, wrong)}")
Output:
runbook tier score action
db-failover 1 76.0 OPEN TICKET (incident said it was wrong)
cache-flush 2 3.5 weekly digest only
legacy-batch-job 3 52.5 weekly digest only
cert-rotation 1 7.9 weekly digest only
queue-backlog 2 18.6 weekly digest only
Check one by hand. db-failover (tier 1, interval 90 days): reviewed 230 days ago, so age is capped at 1.0. Last exercised 211 days ago, also capped at 1.0. It was used once and flagged wrong once, so wrong is 1.0. Dead references are 3 of 10, so 0.3. Owner is active, so 0.
100×(0.25×1+0.15×1+0.30×1+0.20×0.3+0.10×0)=76.0
Now a low row. cache-flush (tier 2, interval 180 days): reviewed 39 days ago, so age is 39/180 = 0.217, halved to 0.108. Last exercised 18 days ago, so 18/180 = 0.1, halved to 0.05. It was used twice with no wrong flags, has no dead references and an active owner, so those are 0.
100×(0.25×0.108+0.15×0.05)=3.5
Recent review plus recent use gives a near-zero score, so it stays out of the ticket queue.
Also note legacy-batch-job scores 52.5, which is below tier 3's threshold of 80, so it stays in the digest. Its dead references and missing owner make it an archive candidate, not something to nag anyone about.
Avoiding alert fatigue
- Default channel is one weekly digest per team, sorted by score.
- Tickets, not pages: a stale runbook is never a 3 a.m. alert.
- Cap open stale-runbook tickets per owner (for example three) so a backlog does not flood them.
- Deduplicate: one ticket per runbook, updated, not re-created.
- Snooze with a reason and an expiry date, so silence is a decision.
How it gets gamed, and the guard
- Rubber-stamp reviews: someone bumps the date without reading. Guard: a review counts only if it includes a checked step (a dry run or a drill signed by someone other than the author).
- Deleting dead links to zero the reference score: guard by tracking the reference count as well.
- Archiving a runbook to remove it from the list: guard by requiring that a paging alert never points to an archived page.
Pitfall. A single number hides the reason. Show the top contributing signal next to the score so the owner knows what to fix.
Design a knowledge management system for a global engineering organization of about two thousand people. It has to support versioned, code-linked documentation, fast search, access controls, doc checks in CI, and usage, gap and staleness analytics. Describe the components, the governance, and where you would deliberately keep it simple.
Sample Answer
Direct answer
For about 2,000 engineers I would build a small set of well-understood components around docs-as-code (docs in the repos), a central index for search, identity-driven access, CI checks, and an analytics layer, with lightweight governance built on owners and review dates. Keep simple: one authoring format, one search engine, no bespoke editor, and no machine-learning features until search logs prove a need.
Terms
- Docs-as-code: docs stored as Markdown in repos and reviewed like code.
- CI (continuous integration): automated checks on each change.
- ACL / RBAC: access control lists / role-based access control (permissions granted by role or group).
- SSO: single sign-on, one login for all tools.
- Staleness: a doc likely out of date because time passed or its code changed.
- Content gap: a question people search for that has no good page.
- Front matter: a small metadata header at the top of a Markdown file (owner, type, sensitivity, last_reviewed).
- Documentation guild: a small volunteer group with one representative per major area that maintains shared standards. It advises and sets templates, unlike a central team that writes or approves everything.
- Federating ownership: each team owns its own docs instead of one central team owning all of them.
- Directory-to-group mapping: the company's employee directory (who reports to which team) is used to create the user groups that permissions follow, so a reorg updates access automatically.
- Boost: nudging certain results higher in search ranking. Drift: a doc slowly becoming wrong as the code changes.
Sizing (illustrative assumptions, not measurements)
Suppose 2,000 engineers in about 250 teams, each team owning roughly 20 pages: about 5,000 pages. With a review every 6 months on average, that is about 830 reviews a month, or 3 to 4 per team per month, which is why review dates must be automated rather than chased by hand. If each engineer searches about 3 times a week, that is roughly 6,000 searches a week, a small load that one managed search engine handles comfortably. This is the reasoning behind keeping the platform simple.
Components
| Component | Choice and reason |
|---|---|
| Authoring | Markdown in each repo with front matter (owner, type, sensitivity, last_reviewed). Versioned by git, so docs branch and tag with code |
| Code linking | Front matter or a mapping file links a page to code paths and services, so CI can spot drift |
| Build and CI checks | Link check, required fields, formatting, and "code changed, doc not touched" warning; blocking only for a short list of critical rules: (1) an owner field is present, (2) no broken internal links, (3) no credential-like strings (keys, passwords). Everything else, including "code changed, doc not touched", starts as a warning |
| Publishing | One static site build per merge, plus a portal page and redirects |
| Search | One managed search engine indexing all sources with synonyms and boosted freshness, filtering by user groups at query time |
| Access control | SSO groups from the identity system; sensitivity classes map to groups; page-level ACL for the small sensitive set, defaults open |
| Analytics | Search logs, page views, zero-result queries, and stale-page counts, in a dashboard |
Governance
- Ownership: every page has an owner team; reorganizations transfer ownership via a directory-to-group mapping.
- Review cadence: reviewed within its window (for example quarterly for runbooks, longer for concepts).
- Standards: a short style guide and templates, maintained by a small documentation guild with representatives from major areas rather than a central gatekeeping team.
- Escalation: unowned or repeatedly stale critical pages go to the engineering director for that area.
Analytics that drive action
- Usage: the top viewed and searched pages get quality investment first.
- Gaps: queries with zero results or immediate re-search become a backlog for owners.
- Staleness: pages past review date, grouped by team, in a monthly leadership report.
- Onboarding: time for a new hire to complete first tasks, and which pages they visit.
Global scale extras
- Multi-language: keep English as the source of truth, translate the top pages (by usage) rather than everything, and label translations that lag the source.
- Personalized ranking: boost by team, region and role from directory data. Simple boosts beat learned models here.
- Page-level access across business units: group-based ACLs with an audit log; keep the restricted set small so defaults stay open.
- Optional machine learning summarization: defer until search analytics show people struggle with long pages, run it as a pilot on non-sensitive content, and require it to cite its sources. The justification test is a measurable improvement in time to answer against the baseline, and a cost within budget.
Where I deliberately keep it simple
- One authoring format instead of supporting many.
- Off-the-shelf search and hosting instead of building.
- Permissions by group, not per-person exceptions.
- Automated freshness signals instead of manual audits.
- No custom ranking algorithm at launch.
Worked example (illustrative)
A payments engineer changes a service's timeout config. CI notes that the linked runbook was not touched, and the owner adds a one-line update in the same pull request. Meanwhile analytics shows "how to rotate keys" returns no results, and the security guild writes the missing page. Both are cheap loops driven by data.
Trade-offs and pitfalls
- Strict CI gates breed workarounds; use warnings first, and block only for the critical rules.
- Analytics can mislead: high views may mean confusion, so pair with feedback.
- Federating ownership scales better than a central team but needs the guild for consistency.
- What would change my call: heavy regulation would raise access and audit requirements, and a smaller organization could skip the guild and analytics dashboards.
Your organization must choose between building an internal knowledge platform and buying a hosted documentation product. Walk through how you would make and defend that decision: which criteria matter, how you weight them, and what would change your answer.
Sample Answer
Direct answer
Decide from criteria tied to the actual problem (people cannot find or trust knowledge), not from a preference for building. My default recommendation for most organizations is to buy or adopt a hosted product for the reading and editing experience, and keep the parts that must live with the code (reference docs, runbooks tied to services) in the repository. Build only if a specific requirement rules every product out, because a knowledge platform is a long-lived product that needs an owner, on-call, and roadmap forever.
Terms used below
- Docs-as-code: writing docs as Markdown files in the same Git repository as the code, reviewed by pull request like code.
- Data residency: a rule about which country or region your data may be stored in. Data sovereignty is the stricter legal version: the data must stay under one country's laws.
- Group sync: automatically copying your company's user groups (for example "SRE team") into the tool so page permissions follow them.
- Audit log: a tamper-resistant record of who viewed or changed what. Retention: how long content and history are kept before deletion.
- Per-seat pricing: paying per user per month, so cost grows with headcount. Exit (lock-in): how hard it is to leave, mainly whether you can export everything cleanly.
- Self-hosted open-source: free software you install and run yourself, so you own the operations work.
Criteria and weights (for a mid-sized engineering organization)
| Criterion | Weight | What to test |
|---|---|---|
| Permissions and access control | 20% | Per-space and per-page rights, single sign-on (SSO, one company login), group sync, guest access |
| Search quality | 20% | Run 20 real past questions; does the right page rank in the top three? |
| Compliance and audit | 15% | Audit log, retention, data residency, export of everything |
| Code linking and docs-as-code fit | 15% | Can docs live in Git, be reviewed by pull request, and link to code? |
| Total cost over 3 years | 15% | Licences plus engineer time; include build's hidden run cost |
| Migration cost and exit | 10% | Import tooling, link preservation, ability to leave |
| Extensibility and integrations | 5% | API, chat and ticket integrations |
Weights depend on context; the point is to write them down before demos, so the vendor with the best presentation does not set them.
Worked example: SRE organization, three options
| Wiki-style tool | Repo-based docs (Markdown in Git, published as a site) | Newer all-in-one workspace | |
|---|---|---|---|
| Permissions | Strong, per page | Only as good as repo access; coarse | Good, still maturing |
| Code linking | Manual links, drift | Best: same PR as code | Integrations, partial |
| Search | Good within the tool | Needs a search layer added | Good, often AI-assisted |
| Compliance | Mature audit and retention | Git history is audit; retention is DIY | Varies by vendor, verify |
| Migration cost | Low if already wiki | Medium: convert and restructure | Medium; import quality varies |
Score each option 1 to 5 per criterion (5 is best), multiply by the weight, and add up. The options: a hosted wiki product, repo-based docs, and a newer all-in-one workspace (a hosted tool combining notes, docs and small databases). Illustrative scores for this SRE org:
| Criterion (weight) | Wiki | Repo docs | All-in-one |
|---|---|---|---|
| Permissions (0.20) | 5 (1.00) | 2 (0.40) | 3 (0.60) |
| Search (0.20) | 4 (0.80) | 2 (0.40) | 4 (0.80) |
| Compliance and audit (0.15) | 5 (0.75) | 3 (0.45) | 3 (0.45) |
| Code linking (0.15) | 2 (0.30) | 5 (0.75) | 3 (0.45) |
| 3-year cost (0.15) | 3 (0.45) | 4 (0.60) | 3 (0.45) |
| Migration and exit (0.10) | 4 (0.40) | 3 (0.30) | 3 (0.30) |
| Integrations (0.05) | 4 (0.20) | 3 (0.15) | 4 (0.20) |
| Total (max 5.0) | 3.90 | 3.05 | 3.25 |
For repo docs, code linking is 5 so 0.15 x 5 = 0.75, while search is 2 so 0.20 x 2 = 0.40. Across all content the wiki wins (3.90). But the weights assume one blended pile of content. Runbooks and service docs are a different class: they change with code and need review, so re-weight for them (code linking 0.35, search 0.15, permissions 0.10, compliance 0.10, cost 0.15, migration 0.10, integrations 0.05, which again sum to 1.00) and recompute: repo docs 3.60, wiki 3.35, all-in-one 3.20. Repo docs win that class.
That is why the recommendation is a hybrid: repo-based docs for runbooks and service docs, plus the wiki for broad non-code material (policies, onboarding, meeting notes). The matrix picks the best tool per content class, and a hybrid is defensible when two classes rank differently. Then sanity check the ranking against your gut feeling, and if they disagree a criterion is probably missing.
Build versus buy
- Build makes sense only with a hard constraint: data may not leave your environment, a deep integration no product offers, or scale where per-seat pricing is prohibitive.
- Buy wins when time to value matters and the team cannot staff a product. Include the 3-year staffing cost of build: at least one engineer's ongoing attention.
What would change my answer
- A strict data-sovereignty rule pushes to self-hosted open-source or build.
- A pilot showing search relevance below a threshold (say fewer than half of the 20 test questions land in the top three) disqualifies that product regardless of price.
- If exit is hard (no clean export), demote that product unless the discount is large.
Defending the decision
Run a two-week pilot with real content and real users, publish the scored matrix and the assumptions, and record an ADR (architecture decision record, a short document capturing a decision, options and reasons) so the decision can be revisited when assumptions change.
Your knowledge base is stale and teams are repeating incidents because playbooks are obsolete. Propose a remediation program: what you clean up first, how stale content is detected automatically, how humans review, what stops recurrence, and a realistic timeline and staffing.
Sample Answer
Direct answer
Run a time-boxed remediation program aimed at the runbooks that pages and incidents actually touch. Clean up tier-1 procedures first, detect staleness automatically with checks against real infrastructure and incident history, require a human to walk each procedure before it is marked verified, and stop recurrence by making a working runbook a condition for shipping an alert and for closing a postmortem. Archive everything nobody claims instead of reviewing it.
Terms. A playbook or runbook is a step-by-step procedure for handling an operational event. A postmortem is a written review of an incident; blameless means it looks for system causes, not someone to blame. A tier-1 service is one whose failure pages people or hurts revenue; tier 2 is important but a failure degrades it rather than stopping it (for example an internal reporting tool). A paging alert is an alert that wakes someone up. A staging environment is a safe copy of production for testing. The inventory is the current list of live hosts and services, taken from your configuration database. A decommissioned host is one that has been shut down and removed.
1. What to clean up first
Rank runbooks by consequence, not by age:
- Runbooks linked from paging alerts on tier-1 services.
- Runbooks named in the last 12 months of postmortems as missing, wrong or ignored (those are the ones that already repeated an incident).
- Runbooks for tier-2 services.
- Everything else: put an archive notice on it and delete or archive after 30 days if no owner claims it.
2. Automatic staleness detection
- Nightly job: links, hostnames, service names and commands referenced in the runbook checked against the current inventory. A reference to a decommissioned host is a hard flag.
- Alert-to-runbook check: every paging alert's runbook URL must resolve and must have a verified date inside its interval.
- Incident hook: responders press "runbook was wrong or missing" during or after an incident, which creates a ticket automatically.
- What the nightly check produces, for one runbook:
db-failover step 4 references host db-eu-2 (decommissioned 2026-06-30): HARD FLAG. Ticket opened for payments-team, due in 5 business days. - Age since last verified, used as a weak signal only (age alone is a poor predictor; broken references and incident flags are stronger).
3. Human review
The owner walks the runbook step by step in a staging environment, or as a tabletop drill (a talked-through rehearsal), with a second person who did not write it. They fix wrong steps, record the date and both names, and only then is the page marked verified. Reading it and clicking "still accurate" does not count.
4. Preventing recurrence
- CI (continuous integration) lint: an automated check that runs on every proposed change and rejects it when a rule is broken, here: an alert definition cannot merge without a runbook link. Illustrative output of a small custom lint script (no standard tool prints exactly this; you define the wording) when it fails:
FAIL alerts/checkout.yaml: "CheckoutLatencyHigh" has severity=page but no runbook_url
FAIL alerts/search.yaml: runbook_url https://wiki.example.com/runbooks/search-old returns 404
- Postmortem action items include "update the runbook" and the postmortem cannot close until the change is merged.
- Runbook lives next to the service code and its owner appears in the repo's CODEOWNERS file (a file that maps paths to owning teams), so changes get reviewed.
- Quarterly drill of the tier-1 runbooks (a game day, a rehearsal where you deliberately trigger a failure).
5. Realistic timeline and staffing
Illustrative sizing: 400 runbooks, of which 40 are tier 1 and 100 are tier 2.
| Phase | Weeks | Work | Effort |
|---|---|---|---|
| 0. Inventory and triage | 1-2 | Export all runbooks, tag tier and owner | Program lead, part-time |
| 1. Tier 1 | 3-8 | Walk and fix 40 runbooks, 1.5 hours each: 60 hours | Owners of each service: 60 hours over 6 weeks is about 10 hours per week in total, so with roughly 10 owning teams that is 1 to 2 hours per week each |
| 2. Detection and CI guard | 4-10 | Build nightly reference check, alert lint | One engineer, about two weeks |
| 3. Tier 2 | 9-16 | 100 runbooks at 1 hour: 100 hours | Owners |
| 4. Long tail | 12-16 | Archive unclaimed after 30-day notice | Program lead |
Staffing: one part-time program lead, one engineer for tooling, and a small time budget from each owning team. The human review effort totals about 160 hours (60 for tier 1 plus 100 for tier 2), which is roughly 4 person-weeks spread across teams, not one team's burden.
Success measures. No paging alert without a verified runbook; postmortems citing "runbook wrong" decline quarter over quarter; nightly reference-check failures trend to zero.
Pitfalls
- Reviewing all 400 equally wastes the program's budget on the long tail.
- A one-time cleanup decays; without the CI and postmortem gates you are back here in a year.
- What would change my call: if the team is tiny, skip the tooling and run a monthly tier-1 drill instead.
You are moving scattered documentation into one centralized knowledge base. Lay out your migration plan from first inventory to cutover: how you decide what to keep, merge or retire, how content gets owners and stays findable, what happens to old material, and which measures tell you the migration worked.
Sample Answer
Direct answer
Treat the migration as content curation first, tooling second: inventory everything, decide keep, merge or retire using a simple scoring rule, put owners and metadata on every kept page BEFORE moving it, migrate in waves ordered by operational criticality, leave redirects behind, and judge success by whether people find answers faster, not by how many pages moved.
Terms
- Inventory: a spreadsheet listing every doc, where it lives, who wrote it and when it last changed.
- Redirect: an old link that automatically forwards to the new location.
- Cutover: the moment the new system becomes the official source.
- Runbook: step-by-step instructions for operating or fixing a system.
Phase 1: Inventory and triage (weeks 1-2)
Crawl the wikis, shared drives and repos, and add the informal sources: chat channels and personal notes. Record for each item: location, owner, last edit, views if known, and which system it describes.
Decide keep, merge or retire with a short rule:
| Decision | Rule |
|---|---|
| Keep | Still accurate, describes a live system, and is used or operationally needed |
| Merge | Overlaps another page on the same topic; combine into one canonical page |
| Retire | Describes decommissioned systems, unedited for a long time, and nobody uses it |
| Ask a domain owner to confirm every retire in bulk, so nothing critical vanishes silently. |
Worked example (illustrative arithmetic)
Inventory finds 400 items. Triage results: 120 keep, 90 merge (they collapse into 30 pages), 190 retire. The new knowledge base receives 120 + 30 = 150 pages, 37.5 percent of the original. Every item still has a recorded fate in the inventory, so the numbers reconcile: 120 + 90 + 190 = 400.
Phase 2: Structure, ownership, findability
- Target structure: a few top-level areas (for example Services, Runbooks, Processes, Onboarding), not a mirror of the org chart.
- Tags: a small controlled list (team, system, lifecycle, audience); no free-form tags at launch.
- Every page has an owner team and a last-reviewed date; pages without an owner are not migrated.
- Search: index titles, headings and tags, boost pages by verified freshness, and test with a list of the 20 most common real questions.
Phase 3: Migrate in waves, by operational criticality
For knowledge that lives only in chat channels and personal notes, rank by "what hurts if it is missing at 3am": on-call runbooks and recovery steps first, then setup and onboarding, then background context. Capture chat knowledge by turning a decision or fix into a proper page (with the thread linked), not by dumping the export. Set a measurable onboarding-speed goal, for example that a new engineer completes the first supervised on-call shift using only the knowledge base, and measure the median days before and after.
Phase 4: Old material and cutover
Freeze the old wikis to read-only with a banner pointing to the new location, keep redirects for a defined period, archive (do not delete) retired items for legal or historical reasons, then switch off after traffic to old links falls.
Success measures
- Search: share of searches with a click and no immediate re-search, and zero-result queries falling.
- Coverage: percent of critical systems with an owned, current runbook.
- Freshness: percent of pages reviewed within their window.
- Onboarding speed (above) and fewer "where is the doc for X" questions in chat.
Trade-offs and pitfalls
- Migrating everything "just in case" copies the rot into the new home.
- No owners means the new system decays as the old one did.
- A big-bang cutover is riskier than waves, but waves prolong two sources of truth, so set a firm end date.
- What would change my call: a small team with 50 pages should skip scoring and simply re-read them.
Unlock Full Question Bank
Get access to all 11 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.