Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
You join a team where documentation on key systems is missing, inconsistent or owned by nobody. What do you do in your first 30, 60 and 90 days to fix it and make it stay fixed?
Sample Answer
Direct answer
I would spend 30 days listening and ranking by risk, 60 days building a minimal ownership model and fixing the highest-risk gaps together with the people who know the systems, and 90 days making it self-sustaining by wiring documentation into the normal workflow (change tickets, reviews, on-call). I would not try to document everything: the goal is that the systems that page people or lose money are documented first, and that new gaps stop appearing.
Structured elaboration
Days 1-30: understand and prioritize
- Inventory the key systems and, for each, who is on call, who really understands it and where any docs live. Read the last few months of incidents and change tickets to see where missing docs hurt.
- Interview the team: "what did you wish existed the last time you were paged?" That is the demand signal.
- Rank the systems by risk (below) and pick the top few. Ship one quick win in the first weeks: for example a runbook for the noisiest alert. It earns credibility.
Days 31-60: ownership and first fixes
- Give every key system a named owning team, recorded in a service catalog (a searchable list of every system with its owner, purpose and links to its docs) or a CODEOWNERS file (the file that says who must review changes to a path). Unowned means the team decides: own it, or retire it.
- Fix the top 5 by pairing (working on the same document together, side by side): I interview the engineer who knows the system, write it with them, and they review it. It is cheaper than asking experts to write alone, and it spreads the knowledge.
- Set a minimal standard: one template, a
last_revieweddate on each page and one searchable home. - Network and other changes documented only in tickets: define a change ticket (the record that requests and describes a change to a live system) standard with required fields: what changes, why, risk, rollback plan (the exact steps to undo the change if it goes wrong), how to verify, and a link to the config diff (the side-by-side view of the configuration before and after). A filled-in example: "Change: add route 10.4.0.0/16 via gateway B. Why: new office. Risk: medium, could drop traffic to the payments subnet. Rollback: remove the route with the saved command, about 5 minutes. Verify: ping the payments subnet from the office. Diff: link to the merged config change." Backlog the poorly documented past changes by risk: those touching core routing (the network paths all traffic depends on), firewalls or shared services first, and let low-risk changes stay as they are. Win buy-in by showing that a good ticket shortens post-incident investigation, and by having changes peer-reviewed before they happen, which also catches mistakes.
Days 61-90: make it stay fixed
- Definition of done (the team's checklist of what must be true before work counts as finished): a change is not done until the affected doc is updated (a line in the pull request or change ticket checklist).
- Reviews and cadence: the doc owner gets an automatic reminder on the review date; two missed cycles marks the page archived.
- Test: a new joiner or an engineer from another team follows a runbook. Where they fail becomes the next fix.
- Report progress with two numbers (share of key systems with an owner and a runbook, and share reviewed in the last 6 months), plus one story of an incident that went better.
Worked example
Ranking with a simple score, risk = business criticality (1 to 3) x documentation gap (1 to 3):
| System | Criticality | Gap | Score |
|---|---|---|---|
| Payments gateway | 3 | 3 | 9 |
| Core network routing | 3 | 2 | 6 |
| Internal wiki | 1 | 3 | 3 |
| Batch reports | 2 | 1 | 2 |
Payments and core routing go into the first 60-day push. The wiki waits. If incident history shows batch reports paged people twice last month, I would raise its score: the missing runbook clearly hurts, so I would move the gap from 1 to 3, making the score 2 x 3 = 6, tied with core routing and inside the 60-day push. The formula guides, evidence overrides.
Trade-offs and pitfalls
- Documenting everything at once yields stale pages nobody trusts. Prioritize by risk and archive the rest.
- A big-bang mandate breeds resentment. Start with the pain the team already feels.
- What would change my call: if the team is in constant firefighting, I would narrow the first 30 days to the three noisiest alerts and only then do the wider inventory.
Your knowledge base is stale and teams are repeating incidents because playbooks are obsolete. Propose a remediation program: what you clean up first, how stale content is detected automatically, how humans review, what stops recurrence, and a realistic timeline and staffing.
Sample Answer
Direct answer
Run a time-boxed remediation program aimed at the runbooks that pages and incidents actually touch. Clean up tier-1 procedures first, detect staleness automatically with checks against real infrastructure and incident history, require a human to walk each procedure before it is marked verified, and stop recurrence by making a working runbook a condition for shipping an alert and for closing a postmortem. Archive everything nobody claims instead of reviewing it.
Terms. A playbook or runbook is a step-by-step procedure for handling an operational event. A postmortem is a written review of an incident; blameless means it looks for system causes, not someone to blame. A tier-1 service is one whose failure pages people or hurts revenue; tier 2 is important but a failure degrades it rather than stopping it (for example an internal reporting tool). A paging alert is an alert that wakes someone up. A staging environment is a safe copy of production for testing. The inventory is the current list of live hosts and services, taken from your configuration database. A decommissioned host is one that has been shut down and removed.
1. What to clean up first
Rank runbooks by consequence, not by age:
- Runbooks linked from paging alerts on tier-1 services.
- Runbooks named in the last 12 months of postmortems as missing, wrong or ignored (those are the ones that already repeated an incident).
- Runbooks for tier-2 services.
- Everything else: put an archive notice on it and delete or archive after 30 days if no owner claims it.
2. Automatic staleness detection
- Nightly job: links, hostnames, service names and commands referenced in the runbook checked against the current inventory. A reference to a decommissioned host is a hard flag.
- Alert-to-runbook check: every paging alert's runbook URL must resolve and must have a verified date inside its interval.
- Incident hook: responders press "runbook was wrong or missing" during or after an incident, which creates a ticket automatically.
- What the nightly check produces, for one runbook:
db-failover step 4 references host db-eu-2 (decommissioned 2026-06-30): HARD FLAG. Ticket opened for payments-team, due in 5 business days. - Age since last verified, used as a weak signal only (age alone is a poor predictor; broken references and incident flags are stronger).
3. Human review
The owner walks the runbook step by step in a staging environment, or as a tabletop drill (a talked-through rehearsal), with a second person who did not write it. They fix wrong steps, record the date and both names, and only then is the page marked verified. Reading it and clicking "still accurate" does not count.
4. Preventing recurrence
- CI (continuous integration) lint: an automated check that runs on every proposed change and rejects it when a rule is broken, here: an alert definition cannot merge without a runbook link. Illustrative output of a small custom lint script (no standard tool prints exactly this; you define the wording) when it fails:
FAIL alerts/checkout.yaml: "CheckoutLatencyHigh" has severity=page but no runbook_url
FAIL alerts/search.yaml: runbook_url https://wiki.example.com/runbooks/search-old returns 404
- Postmortem action items include "update the runbook" and the postmortem cannot close until the change is merged.
- Runbook lives next to the service code and its owner appears in the repo's CODEOWNERS file (a file that maps paths to owning teams), so changes get reviewed.
- Quarterly drill of the tier-1 runbooks (a game day, a rehearsal where you deliberately trigger a failure).
5. Realistic timeline and staffing
Illustrative sizing: 400 runbooks, of which 40 are tier 1 and 100 are tier 2.
| Phase | Weeks | Work | Effort |
|---|---|---|---|
| 0. Inventory and triage | 1-2 | Export all runbooks, tag tier and owner | Program lead, part-time |
| 1. Tier 1 | 3-8 | Walk and fix 40 runbooks, 1.5 hours each: 60 hours | Owners of each service: 60 hours over 6 weeks is about 10 hours per week in total, so with roughly 10 owning teams that is 1 to 2 hours per week each |
| 2. Detection and CI guard | 4-10 | Build nightly reference check, alert lint | One engineer, about two weeks |
| 3. Tier 2 | 9-16 | 100 runbooks at 1 hour: 100 hours | Owners |
| 4. Long tail | 12-16 | Archive unclaimed after 30-day notice | Program lead |
Staffing: one part-time program lead, one engineer for tooling, and a small time budget from each owning team. The human review effort totals about 160 hours (60 for tier 1 plus 100 for tier 2), which is roughly 4 person-weeks spread across teams, not one team's burden.
Success measures. No paging alert without a verified runbook; postmortems citing "runbook wrong" decline quarter over quarter; nightly reference-check failures trend to zero.
Pitfalls
- Reviewing all 400 equally wastes the program's budget on the long tail.
- A one-time cleanup decays; without the CI and postmortem gates you are back here in a year.
- What would change my call: if the team is tiny, skip the tooling and run a monthly tier-1 drill instead.
Documentation for services keeps becoming undiscoverable when code is refactored or renamed. Design a way to link documentation to the code it describes so that the links survive change: what metadata you record, where checks run, and how broken links are detected and repaired.
Sample Answer
Direct answer
Link docs to code through stable identifiers (a service name and the paths or symbols it owns), not through line numbers or pasted URLs. Store those identifiers as metadata in each doc, run a check in CI (continuous integration, the automated build that runs on every pull request) that fails when a link no longer resolves, and run a scheduled sweep for links that rot without a PR touching them. When a break is found, the tooling proposes the repair (for example the renamed path) so a human only approves.
1. What metadata you record
Put it in the doc's front matter (a small YAML header at the top of a Markdown file), so it travels with the doc in the same repository ("docs-as-code": docs live in version control and go through review like code).
---
service: billing
owner_team: payments
covers:
- src/billing/
last_verified_commit: abc1234
review_by: 2026-12-01
---
covers: paths (or symbol names for API docs, where a symbol is a named code element such as a function or class) the doc describes. This is the forward link, doc to code.owner_teamplus a CODEOWNERS file (a repository file mapping paths to the people who must review changes there) gives the reverse link and someone to notify.last_verified_commitandreview_by: when a human last confirmed the doc matched the code. A commit SHA (the unique hash Git gives every commit, such asabc1234) records exactly which version of the code was checked.- The
service: billingvalue is an identifier from a service catalog (an internal registry listing each service, its owning team and where its code lives). If the code moves from one repository to another, only the catalog entry (billing -> repo: platform-monorepo, path: services/billing/) is updated, and every doc that saysservice: billingstill resolves. Docs that hard-code a repository URL would all break. - Links from doc prose to code use a stable anchor (a symbol or section name). A permalink pinned to a commit SHA never breaks but shows old code, so use it only as evidence ("this was true at abc1234"), and use the live path for navigation.
2. Where checks run
- On every pull request: compute the files changed; find docs whose
coversoverlap; require either a doc edit or an explicit "docs not needed" label. Also run the link resolver below on the whole docs folder. - Nightly sweep: catches external links, docs past
review_by, and services whose owner team no longer exists. - Reverse index: a generated page listing each service and its docs, so engineers finding code can find docs.
3. Detecting and repairing broken links (worked example)
A PR renames src/billing/ to src/invoicing/. The check below reads each doc's covers, confirms each prefix still matches a tracked file, and uses git diff -M (which detects renames) to suggest the new location. Setup: pip install pyyaml; a repo with src/billing/charge.py, src/billing/refund.py and the doc above (saved as docs/billing-runbook.md) committed on branch main; then create a branch, run git mv src/billing src/invoicing and commit.
import subprocess, sys, pathlib, yaml
def front_matter(path):
text = path.read_text()
if not text.startswith("---\n"):
return None
return yaml.safe_load(text.split("---\n", 2)[1])
def git(*args):
return subprocess.run(["git", *args], capture_output=True, text=True, check=True).stdout
base = sys.argv[1] # e.g. origin/main
tracked = git("ls-files").splitlines()
renames = {}
for line in git("diff", "--name-status", "-M", base, "HEAD").splitlines():
parts = line.split("\t")
if parts[0].startswith("R"):
renames[parts[1]] = parts[2]
problems = 0
for doc in pathlib.Path("docs").rglob("*.md"):
meta = front_matter(doc)
if not meta:
print(f"{doc}: MISSING front matter"); problems += 1; continue
for prefix in meta.get("covers", []):
if any(f.startswith(prefix) for f in tracked):
continue
hint = next((new for old, new in renames.items() if old.startswith(prefix)), None)
msg = f"{doc}: covers '{prefix}' matches no file"
if hint:
msg += f" (renamed to {hint.rsplit('/', 1)[0]}/)"
print(msg); problems += 1
print(f"{problems} problem(s)")
sys.exit(1 if problems else 0)
Run as python3 check_doc_links.py main (the argument is the branch to compare against). Output:
docs/billing-runbook.md: covers 'src/billing/' matches no file (renamed to src/invoicing/)
1 problem(s)
How it works, block by block: front_matter reads the YAML header of a doc. renames is a dictionary built from git diff -M (the -M flag makes Git detect that a file was moved rather than deleted and re-added), mapping old path to new path, for example src/billing/charge.py to src/invoicing/charge.py. For each doc and each covers prefix, the loop asks whether any tracked file still starts with that prefix (startswith). If none does, the link is broken, and the loop looks for a renamed file whose old path started with that prefix to suggest the repair.
The CI job exits non-zero, so the rename PR cannot merge until the doc's covers is updated. A bot can open that one-line fix automatically, because the rename target is already known.
4. Trade-offs and pitfalls
- Strict blocking on every code change creates label-spam ("docs not needed" clicked reflexively). Start warn-only for a quarter, then block only for services tagged tier-1 (the most business-critical services, for example those that can page someone at night).
- Path prefixes are coarse: a rename inside the folder is not detected. Symbol-level links are more precise but need per-language tooling, so reserve them for public API docs.
- Metadata rots too.
review_bywith a nightly report of expired docs is the guard. - Docs living outside the repo (a wiki) cannot get a PR check. Either move the ones that matter into the repo, or accept only the nightly sweep for them.
How would you measure whether an SRE team's documented procedures are going stale or being forgotten? Define the inputs and how you would combine them, when a score should trigger action, and how you would avoid alert fatigue.
Sample Answer
Direct answer
Treat "stale or forgotten" as several independent signals, score each 0 to 1, combine them with fixed weights into a 0 to 100 staleness score per runbook, and set the action threshold by how critical the service is. Add a hard override: if a real incident said the procedure was wrong, act regardless of the score. Avoid alert fatigue by sending weekly digests by default and opening a ticket only past the threshold, capped per owner.
Terms. A runbook is a step-by-step procedure for an operational task. An SRE (site reliability engineer) team runs production and relies on runbooks during incidents. A tier ranks a service's criticality (tier 1 is paging or revenue-critical). A paging alert is an alert that wakes someone up. Drilled means rehearsed on purpose before a real emergency. A dry run executes the steps against a safe copy with no real effect. The inventory is the current list of live hosts and services (from your configuration database). A digest is one summary message instead of many alerts. To snooze is to silence a reminder until a set date. An archive candidate is a runbook that looks safe to retire. To deduplicate is to keep one item where there would be several.
Inputs (each scaled 0 to 1, where 1 is worst)
| Signal | Definition | Weight |
|---|---|---|
| Age | Days since last review divided by the tier's review interval, capped at 2, then halved | 0.25 |
| Unexercised | Same calculation using days since the runbook was last used or drilled | 0.15 |
| Proved wrong | Incidents in 90 days where responders flagged it wrong, divided by incidents that used it | 0.30 |
| Dead references | Links, hosts or commands that no longer exist, divided by total references (found by a nightly check against the inventory) | 0.20 |
| Owner gone | 1 if the owner left or the team was reorganized, else 0 | 0.10 |
Why these shapes and weights:
- Capped at 2, then halved. Age is measured against the review interval. A runbook exactly at its interval scores 1 before halving, so 0.5 after. Capping at 2 means anything more than twice overdue is simply maximally stale (1.0 after halving) rather than letting a 10-times-overdue page dominate the score.
- Proved wrong weighs most (0.30) because it is direct evidence from a real incident. Dead references (0.20) are hard, checkable facts. Age and unexercised (0.25, 0.15) are only proxies. Owner gone (0.10) is a flag more than a measurement.
Weights are a judgement. Calibrate them by backtesting, meaning you replay the past: compute scores as they stood last quarter, then check them against what actually went wrong afterwards. Illustrative example: 4 runbooks were flagged wrong in this quarter's incidents. If 3 of those 4 were among last quarter's top-scoring runbooks, the score caught 75%. If only 1 of 4 was, raise the weight on dead references or add a signal, and re-check.
When to act
Tier 1 opens a ticket at 45 or above, tier 2 at 60, tier 3 at 80. Tier 1 opens earliest because a stale runbook there costs the most during an outage, and its review interval is shortest (90 days, versus 180 and 365). These numbers are starting values you tune from the backtest, not laws. Below the threshold the runbook appears only in the weekly digest. Override: one or more "proved wrong" flags opens a ticket whatever the score.
Runnable computation (pinned data, prints the result)
from datetime import date
TODAY = date(2026, 9, 28)
WEIGHTS = {"age": 0.25, "unexercised": 0.15, "wrong": 0.30, "dead_refs": 0.20, "owner": 0.10}
# tier -> (score that opens a ticket, review interval in days)
TIERS = {1: (45, 90), 2: (60, 180), 3: (80, 365)}
# name, tier, last_reviewed, last_exercised (used or drilled), uses_in_90d, flagged_wrong_in_90d,
# dead_refs, total_refs, owner_active
RUNBOOKS = [
("db-failover", 1, date(2026, 2, 10), date(2026, 3, 1), 1, 1, 3, 10, True),
("cache-flush", 2, date(2026, 8, 20), date(2026, 9, 10), 2, 0, 0, 6, True),
("legacy-batch-job", 3, date(2025, 6, 1), date(2025, 6, 1), 0, 0, 4, 5, False),
("cert-rotation", 1, date(2026, 8, 30), date(2026, 9, 1), 0, 0, 1, 12, True),
("queue-backlog", 2, date(2026, 3, 15), date(2026, 8, 2), 3, 0, 1, 8, True),
]
def score(tier, reviewed, exercised, uses, wrong, dead, total, owner_active):
interval = TIERS[tier][1]
parts = {
"age": min((TODAY - reviewed).days / interval, 2) / 2,
"unexercised": min((TODAY - exercised).days / interval, 2) / 2,
"wrong": (wrong / uses) if uses else 0.0,
"dead_refs": dead / total,
"owner": 0.0 if owner_active else 1.0,
}
return 100 * sum(WEIGHTS[k] * v for k, v in parts.items()), parts
def action(tier, s, wrong):
if wrong >= 1:
return "OPEN TICKET (incident said it was wrong)"
return "open ticket" if s >= TIERS[tier][0] else "weekly digest only"
print(f"{'runbook':<17}{'tier':>4}{'score':>7} action")
for name, tier, rv, ex, uses, wrong, dead, total, owner in RUNBOOKS:
s, _ = score(tier, rv, ex, uses, wrong, dead, total, owner)
print(f"{name:<17}{tier:>4}{s:>7.1f} {action(tier, s, wrong)}")
Output:
runbook tier score action
db-failover 1 76.0 OPEN TICKET (incident said it was wrong)
cache-flush 2 3.5 weekly digest only
legacy-batch-job 3 52.5 weekly digest only
cert-rotation 1 7.9 weekly digest only
queue-backlog 2 18.6 weekly digest only
Check one by hand. db-failover (tier 1, interval 90 days): reviewed 230 days ago, so age is capped at 1.0. Last exercised 211 days ago, also capped at 1.0. It was used once and flagged wrong once, so wrong is 1.0. Dead references are 3 of 10, so 0.3. Owner is active, so 0.
100×(0.25×1+0.15×1+0.30×1+0.20×0.3+0.10×0)=76.0
Now a low row. cache-flush (tier 2, interval 180 days): reviewed 39 days ago, so age is 39/180 = 0.217, halved to 0.108. Last exercised 18 days ago, so 18/180 = 0.1, halved to 0.05. It was used twice with no wrong flags, has no dead references and an active owner, so those are 0.
100×(0.25×0.108+0.15×0.05)=3.5
Recent review plus recent use gives a near-zero score, so it stays out of the ticket queue.
Also note legacy-batch-job scores 52.5, which is below tier 3's threshold of 80, so it stays in the digest. Its dead references and missing owner make it an archive candidate, not something to nag anyone about.
Avoiding alert fatigue
- Default channel is one weekly digest per team, sorted by score.
- Tickets, not pages: a stale runbook is never a 3 a.m. alert.
- Cap open stale-runbook tickets per owner (for example three) so a backlog does not flood them.
- Deduplicate: one ticket per runbook, updated, not re-created.
- Snooze with a reason and an expiry date, so silence is a decision.
How it gets gamed, and the guard
- Rubber-stamp reviews: someone bumps the date without reading. Guard: a review counts only if it includes a checked step (a dry run or a drill signed by someone other than the author).
- Deleting dead links to zero the reference score: guard by tracking the reference count as well.
- Archiving a runbook to remove it from the list: guard by requiring that a paging alert never points to an archived page.
Pitfall. A single number hides the reason. Show the top contributing signal next to the score so the owner knows what to fix.
That is every published Documentation and Knowledge Management question for Cloud Engineer so far. Browse the other topics in this category, or practice this one interactively.