Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
How would you measure whether an SRE team's documented procedures are going stale or being forgotten? Define the inputs and how you would combine them, when a score should trigger action, and how you would avoid alert fatigue.
Sample Answer
Direct answer
Treat "stale or forgotten" as several independent signals, score each 0 to 1, combine them with fixed weights into a 0 to 100 staleness score per runbook, and set the action threshold by how critical the service is. Add a hard override: if a real incident said the procedure was wrong, act regardless of the score. Avoid alert fatigue by sending weekly digests by default and opening a ticket only past the threshold, capped per owner.
Terms. A runbook is a step-by-step procedure for an operational task. An SRE (site reliability engineer) team runs production and relies on runbooks during incidents. A tier ranks a service's criticality (tier 1 is paging or revenue-critical). A paging alert is an alert that wakes someone up. Drilled means rehearsed on purpose before a real emergency. A dry run executes the steps against a safe copy with no real effect. The inventory is the current list of live hosts and services (from your configuration database). A digest is one summary message instead of many alerts. To snooze is to silence a reminder until a set date. An archive candidate is a runbook that looks safe to retire. To deduplicate is to keep one item where there would be several.
Inputs (each scaled 0 to 1, where 1 is worst)
| Signal | Definition | Weight |
|---|---|---|
| Age | Days since last review divided by the tier's review interval, capped at 2, then halved | 0.25 |
| Unexercised | Same calculation using days since the runbook was last used or drilled | 0.15 |
| Proved wrong | Incidents in 90 days where responders flagged it wrong, divided by incidents that used it | 0.30 |
| Dead references | Links, hosts or commands that no longer exist, divided by total references (found by a nightly check against the inventory) | 0.20 |
| Owner gone | 1 if the owner left or the team was reorganized, else 0 | 0.10 |
Why these shapes and weights:
- Capped at 2, then halved. Age is measured against the review interval. A runbook exactly at its interval scores 1 before halving, so 0.5 after. Capping at 2 means anything more than twice overdue is simply maximally stale (1.0 after halving) rather than letting a 10-times-overdue page dominate the score.
- Proved wrong weighs most (0.30) because it is direct evidence from a real incident. Dead references (0.20) are hard, checkable facts. Age and unexercised (0.25, 0.15) are only proxies. Owner gone (0.10) is a flag more than a measurement.
Weights are a judgement. Calibrate them by backtesting, meaning you replay the past: compute scores as they stood last quarter, then check them against what actually went wrong afterwards. Illustrative example: 4 runbooks were flagged wrong in this quarter's incidents. If 3 of those 4 were among last quarter's top-scoring runbooks, the score caught 75%. If only 1 of 4 was, raise the weight on dead references or add a signal, and re-check.
When to act
Tier 1 opens a ticket at 45 or above, tier 2 at 60, tier 3 at 80. Tier 1 opens earliest because a stale runbook there costs the most during an outage, and its review interval is shortest (90 days, versus 180 and 365). These numbers are starting values you tune from the backtest, not laws. Below the threshold the runbook appears only in the weekly digest. Override: one or more "proved wrong" flags opens a ticket whatever the score.
Runnable computation (pinned data, prints the result)
from datetime import date
TODAY = date(2026, 9, 28)
WEIGHTS = {"age": 0.25, "unexercised": 0.15, "wrong": 0.30, "dead_refs": 0.20, "owner": 0.10}
# tier -> (score that opens a ticket, review interval in days)
TIERS = {1: (45, 90), 2: (60, 180), 3: (80, 365)}
# name, tier, last_reviewed, last_exercised (used or drilled), uses_in_90d, flagged_wrong_in_90d,
# dead_refs, total_refs, owner_active
RUNBOOKS = [
("db-failover", 1, date(2026, 2, 10), date(2026, 3, 1), 1, 1, 3, 10, True),
("cache-flush", 2, date(2026, 8, 20), date(2026, 9, 10), 2, 0, 0, 6, True),
("legacy-batch-job", 3, date(2025, 6, 1), date(2025, 6, 1), 0, 0, 4, 5, False),
("cert-rotation", 1, date(2026, 8, 30), date(2026, 9, 1), 0, 0, 1, 12, True),
("queue-backlog", 2, date(2026, 3, 15), date(2026, 8, 2), 3, 0, 1, 8, True),
]
def score(tier, reviewed, exercised, uses, wrong, dead, total, owner_active):
interval = TIERS[tier][1]
parts = {
"age": min((TODAY - reviewed).days / interval, 2) / 2,
"unexercised": min((TODAY - exercised).days / interval, 2) / 2,
"wrong": (wrong / uses) if uses else 0.0,
"dead_refs": dead / total,
"owner": 0.0 if owner_active else 1.0,
}
return 100 * sum(WEIGHTS[k] * v for k, v in parts.items()), parts
def action(tier, s, wrong):
if wrong >= 1:
return "OPEN TICKET (incident said it was wrong)"
return "open ticket" if s >= TIERS[tier][0] else "weekly digest only"
print(f"{'runbook':<17}{'tier':>4}{'score':>7} action")
for name, tier, rv, ex, uses, wrong, dead, total, owner in RUNBOOKS:
s, _ = score(tier, rv, ex, uses, wrong, dead, total, owner)
print(f"{name:<17}{tier:>4}{s:>7.1f} {action(tier, s, wrong)}")
Output:
runbook tier score action
db-failover 1 76.0 OPEN TICKET (incident said it was wrong)
cache-flush 2 3.5 weekly digest only
legacy-batch-job 3 52.5 weekly digest only
cert-rotation 1 7.9 weekly digest only
queue-backlog 2 18.6 weekly digest only
Check one by hand. db-failover (tier 1, interval 90 days): reviewed 230 days ago, so age is capped at 1.0. Last exercised 211 days ago, also capped at 1.0. It was used once and flagged wrong once, so wrong is 1.0. Dead references are 3 of 10, so 0.3. Owner is active, so 0.
100×(0.25×1+0.15×1+0.30×1+0.20×0.3+0.10×0)=76.0
Now a low row. cache-flush (tier 2, interval 180 days): reviewed 39 days ago, so age is 39/180 = 0.217, halved to 0.108. Last exercised 18 days ago, so 18/180 = 0.1, halved to 0.05. It was used twice with no wrong flags, has no dead references and an active owner, so those are 0.
100×(0.25×0.108+0.15×0.05)=3.5
Recent review plus recent use gives a near-zero score, so it stays out of the ticket queue.
Also note legacy-batch-job scores 52.5, which is below tier 3's threshold of 80, so it stays in the digest. Its dead references and missing owner make it an archive candidate, not something to nag anyone about.
Avoiding alert fatigue
- Default channel is one weekly digest per team, sorted by score.
- Tickets, not pages: a stale runbook is never a 3 a.m. alert.
- Cap open stale-runbook tickets per owner (for example three) so a backlog does not flood them.
- Deduplicate: one ticket per runbook, updated, not re-created.
- Snooze with a reason and an expiry date, so silence is a decision.
How it gets gamed, and the guard
- Rubber-stamp reviews: someone bumps the date without reading. Guard: a review counts only if it includes a checked step (a dry run or a drill signed by someone other than the author).
- Deleting dead links to zero the reference score: guard by tracking the reference count as well.
- Archiving a runbook to remove it from the list: guard by requiring that a paging alert never points to an archived page.
Pitfall. A single number hides the reason. Show the top contributing signal next to the score so the owner knows what to fix.
Several product teams each own documentation for their own services, and compliance reviewers now need to see who is accountable for what. Design an ownership model: roles, edit rights, approvals, how quickly docs must follow code changes, and what audit evidence you keep.
Sample Answer
Direct answer
Give every doc one accountable owning team, encode that in the repository (not in someone's memory), use pull-request review as the approval and audit mechanism, set update deadlines by how risky the doc is, and keep the evidence automatically (version history, review records, and an ownership report). Compliance reviewers then answer "who is accountable for what" by pulling a report, not by asking around.
Roles and edit rights
| Role | Responsibility | Rights |
|---|---|---|
| Owner (a named team plus a named lead) | Accountable for accuracy and review dates | Approves changes; can archive |
| Contributor | Any engineer | Proposes edits by pull request |
| Reviewer | Owner team member, plus a domain expert for sensitive docs | Approves merge |
| Compliance reviewer | Reads, does not edit | Read access to all, plus the audit report |
| Docs platform admin | Tooling and permissions | No content authority |
Ownership lives in a doc header field and a CODEOWNERS file (a repository file naming who must review changes in a path), so branch protection (a repository setting that refuses to merge a change until the required approvals, including the owner's, are in) can require owner approval automatically. A tier is just a risk label your policy defines; three tiers is a common, workable number.
A concrete header and CODEOWNERS line:
# top of docs/payments/refund-runbook.md
owner: payments-team lead: J. Doe tier: 1
last_reviewed: 2026-07-15 next_review: 2027-01-15
# CODEOWNERS
/docs/payments/ @acme/payments-team
Approvals and update deadlines
- Tier 1 (security procedures, incident runbooks, data-handling): owner approval plus one independent reviewer; docs must be updated in the same pull request as the code change when it touches described behavior, and reviewed at least every six months.
- Tier 2 (service docs): owner approval; update within, say, five business days of a behavior change.
- Tier 3 (notes, drafts): no approval; auto-archived after inactivity.
The exact numbers are policy choices your compliance team sets; the structure is what matters: shorter deadlines for higher risk.
Audit evidence you keep
- Version history with author, reviewer and time (Git provides this).
- A generated ownership register (a table built automatically from the doc headers, never typed by hand): doc, owner team, tier, last reviewed date, next review date.
- Review records: pull-request approvals, and a signed attestation for periodic tier-1 reviews (a dated statement by the owner, "I re-read this and it is accurate", kept as a record).
- Exceptions: a log of overdue docs and who accepted the risk, with an expiry. One entry looks like:
ledger-reconciliation-guide | overdue since 2026-08-01 | risk accepted by: Director of Engineering | reason: rewrite tied to the payments migration | expires: 2026-10-31. When the expiry passes, the doc is either fixed or the acceptance is renewed on the record.
Worked example
A compliance reviewer asks: "Who is accountable for the payment-refund runbook, and when was it last verified?" The register shows: owner payments team (lead named), tier 1, last reviewed 2026-07-15 by two people, next review due 2027-01-15, and the linked pull request shows the approvals. No meeting was needed.
Lifecycle (creation to archive)
Draft, then in review, then published, then due for review (a timer), then either re-verified, updated, or archived. Archived docs stay in version control and carry a banner pointing to the replacement; nothing is deleted while retention rules (how long the company or a regulator requires records to be kept) apply.
ML organization variant: governance versus agility
For models, add a model owner and require a model card (a short standard document describing purpose, training data, evaluation and limits) before release. To protect agility, apply strict controls only to models in production or that affect customers; research notebooks and experiments stay lightweight (tier 3). Update deadlines follow releases: a model card must be current at each promotion (moving a model from testing into real production use). The trade-off: strict governance slows experimentation, so scope it by risk rather than applying it everywhere.
Pitfalls
Approvals that become rubber stamps, ownership assigned to a person who leaves (assign to teams and review quarterly), and evidence collected by hand only at audit time (automate it).
Your company just merged with another engineering organization and now has conflicting decision records, duplicate runbooks and inconsistent documentation formats. How do you unify them: what happens to conflicting decisions, how do you preserve history, and how do you tell engineers where the authoritative version lives?
Sample Answer
Direct answer
Do not pick a winner org and delete the other. Run a time-boxed reconciliation: inventory both sets, sort every conflict into a small number of types, let a named decider resolve true conflicts through a new decision record that supersedes (never erases) the old ones, archive the old material read-only with redirects, and make one index the only place anyone is told to look. Start with the documents that hurt during an incident or a deploy, not the ones that are merely ugly.
Terms used below
- ADR (architecture decision record): a short dated document recording one decision, the options considered, and why one won. It has a status such as Proposed, Accepted or Superseded.
- Runbook: step-by-step instructions for an operational task, such as failing over a database.
- Authoritative version: the one copy a reader is entitled to trust and follow.
- Terraform: a tool that describes cloud infrastructure (servers, networks, databases) in code files. Ansible: a tool that configures existing servers from scripts.
- Jenkins and GitHub Actions: two competing services that automatically build and test code when it changes.
- OAuth: a standard way for one service to be granted limited access to another without sharing passwords, either on a user's behalf or, for service-to-service authentication as in the table below, as a short-lived access token issued to the calling service itself.
- Tier-1 service: a service whose outage pages people or stops revenue (tier 2 matters but degrades gracefully; tier 3 is internal or low impact).
- Staging: a safe copy of the production environment used to test changes.
- History-preserving import: copying a repository so every past change and author is kept, not just the latest files.
- Redirect: an old web address automatically forwarding to the new page. Stub: a tiny placeholder page that only points to the new location.
- Sunset date: the date after which something stops being supported. Carve-out: a deliberate exception to a new rule. Claim window: a fixed period in which owners may say "this page is still needed" before it is archived.
1. Inventory and triage (weeks 1-2)
List every ADR, runbook and standards page from both orgs with owner, service, last-verified date and format. Then classify each conflict:
| Type | Example | Resolution |
|---|---|---|
| Compatible | Both orgs chose OAuth for service auth | Keep one, link the other as "also decided" |
| Different scope | Org A: Terraform for cloud, Org B: Ansible for on-prem hosts | Keep both, record the boundary in one ADR |
| True conflict | Org A: Jenkins for CI (continuous integration, automated build and test), Org B: GitHub Actions | One decider picks, new ADR supersedes |
| Obsolete | Either org's decision about a system being retired | Mark Superseded or Archived, no debate |
In practice you expect many conflicts to fall into the first, second or fourth type (this is an expectation, not a measured figure; your own inventory will show the real split). Spend the debating effort only on true conflicts.
2. What happens to conflicting decisions
- Each true conflict gets a reconciliation ADR with one named decider (the architect or engineering lead who owns the affected system, not a committee vote). It lists both original ADRs, the criteria (cost, migration effort, who runs it, security posture), the choice and a migration deadline.
- The losing ADR is not deleted. Its status becomes Superseded by M-0004 and it keeps its text. Reason: engineers will ask "why did the other org do it that way?" for years, and the reasoning is often still valid for a corner case.
- If a decision cannot be settled quickly, record it as Both accepted, sunset date TBD with an owner. An explicit temporary state beats silent divergence.
- Rule for the interim: whichever ADR governs new work is the reconciliation ADR; existing systems follow the old one until migrated.
3. Preserving history
- Import each org's docs repo into a read-only archive keeping full version history (for git repos, a history-preserving import rather than copy-paste).
- Keep original numbers and add an org prefix:
A-0007andB-0007cannot collide, and old links in code comments and tickets still resolve. New reconciliation ADRs written after the merger use the prefixM-(M for merged company), for exampleM-0004. - Old wiki URLs get a redirect to the new page, or, if the page was archived, to the archive copy with a banner.
4. Runbook consolidation
Pair runbooks by service. For each pair, keep the one that was most recently exercised or used in a real incident, then merge in the steps the other had that the kept runbook lacks, in a single new template. Where two runbooks disagree on a command, test it in staging before choosing. Tier-1 services (paging, revenue path) first.
5. Telling engineers where the truth lives
- One documentation portal or repo with a front-page index: "Authoritative docs for service X".
- Every page shows a status banner (Authoritative, Superseded, Archived) and a last-verified date.
- Every old page becomes a stub: "Moved. Do not follow this copy." with a link.
- Pager alerts and CODEOWNERS files (a repo file that maps paths to owning teams) point to the new runbook URLs. This matters more than any announcement, because engineers follow the link in the alert.
- Announce once, with the top ten most-searched pages, then rely on redirects.
Worked example
Org A's A-0012 says "Postgres on managed cloud". Org B's B-0031 says "self-hosted MySQL on VMs". The merged company has two production databases for one product. The database architect decides: Postgres managed for all new services, MySQL kept for two legacy services until a migration ticket closes. Result: new ADR M-0004 (Accepted), B-0031 marked Superseded by M-0004 (legacy carve-out until migration), and the two db-failover runbooks merged into one Postgres runbook plus a short MySQL appendix that carries a removal date.
Trade-offs and pitfalls
- Reformatting everything to one template up front is a trap: it burns months on pages nobody reads. Migrate tier-1 now and everything else "on touch" or by archiving after a claim window.
- Letting both orgs keep their own portal "for now" is how the duplicates return.
- A committee that seeks consensus stalls; one decider with a written rationale is faster and leaves a record.
- What would change my call: if one org's systems are being retired within months, archive that org's docs instead of reconciling them.
You are asked to organize an internal knowledge base for a data platform. How would you structure the top-level categories, how would a dataset doc, a runbook and an architecture doc each be classified, and why does the taxonomy matter for people finding things?
Sample Answer
Direct answer
Organize the knowledge base around how people look for things, not around who wrote them. Use a few top-level categories, classify each document by a type (what kind of document it is), and let secondary labels such as team or system cut across. A taxonomy matters because people can only use what they can find, and a stable structure lets search, ownership and freshness rules work.
Terms
- Taxonomy: an agreed way to group and name content.
- Dataset doc: describes a table or file (columns, source, refresh, owner).
- Runbook: step-by-step instructions to operate or fix something, usually during an incident.
- Architecture doc: explains how a system is built and why.
Top-level categories for a data platform
- Getting started: onboarding, access requests, tooling setup.
- Data catalog: dataset docs, grouped by business domain (orders, payments, customers).
- Pipelines and systems: architecture docs and design decisions per system.
- Operations: runbooks, on-call guides, incident reviews.
- Standards and policies: naming conventions, quality rules, data handling.
How each document is classified
| Document | Category | Type | Key labels |
|---|---|---|---|
Dataset doc for orders_daily | Data catalog, Orders domain | dataset | owner team, refresh cadence, sensitivity |
| Runbook: "Nightly load is stuck" | Operations | runbook | system (orders pipeline), severity, on-call team |
| Architecture doc for the ingestion service | Pipelines and systems | architecture | system, status (current or proposed) |
The same page can be reached by category, type or system, but has one home so there is one canonical copy.
Why it matters (worked example, illustrative)
At 2 a.m. an engineer sees the nightly orders load fail. With a good taxonomy she opens Operations, filters by system "orders pipeline", and finds the runbook in two clicks. With no taxonomy she searches "load failed", gets forty results including old design notes and an outdated wiki page, and wastes time or follows stale steps. The taxonomy also tells people where to put new content, which prevents duplicates.
Principles
- Name categories by reader tasks (get started, fix, look up) rather than org charts, because teams reorganize and paths break.
- Keep depth to about two or three levels.
- Use one type list and reuse it everywhere so search filters are predictable.
Trade-offs and pitfalls
- Deep hierarchies hide content; flat structures with rich filters scale better.
- A misc bucket grows without limit, so review it monthly.
- Testing with real users (ask three people to find a given doc) beats debating categories.
In a large engineering organization, compare a centralized documentation model with a federated one. How do ownership, findability, governance overhead, update speed and standards enforcement differ, when is each appropriate, and what hybrid would you propose?
Sample Answer
Direct answer
A centralized model puts one team in charge of writing, structure and tooling. A federated model lets each engineering team own its docs under shared standards. In a large organization neither extreme works alone, so I would recommend a hybrid: central ownership of the platform, standards and navigation; federated ownership of content, with each doc having an accountable owning team.
Comparison
| Dimension | Centralized | Federated |
|---|---|---|
| Ownership | Clear (doc team), but far from the code | Clear per team, but uneven across teams |
| Findability | High: one structure, one search | Low unless indexed: many styles, locations |
| Governance overhead | High: intake queue, review | Lower per team, but coordination cost across teams |
| Update speed | Slow: writers wait on engineers | Fast: authors are the people making the change |
| Standards enforcement | Easy: one gatekeeper | Hard: needs automation and peer pressure |
| Accuracy | Prone to lag and errors | Better where teams care; poor where they do not |
When each fits
- Centralized: a small organization, highly regulated or customer-facing docs where consistency and legal review matter, or a young platform where standards do not exist yet.
- Federated: many teams with different stacks, fast-moving internal engineering docs, and a culture where teams already own services end to end.
The hybrid I would propose (worked example)
Imagine 40 teams:
- A small platform team (2 to 4 people) owns the publishing tool, the templates, the style guide, the global search and the navigation. It does not write team content.
- Each team owns its docs in its own repository, marked with an owner in the doc header; a CODEOWNERS file (a repository file naming who must review changes in a path) makes this enforceable.
- Automated checks in CI (automated checks that run on each change) enforce standards: required sections, broken links, owner present, review date. This replaces a human gatekeeper.
- Tier-1 material (the small set of docs where an error is costly, such as public API, incident runbooks, security policies) gets extra editorial review from a small docs guild (one representative per team).
- All content is indexed into one search, and the platform team publishes a scorecard per team (share of docs with owner, share past review date), which makes uneven quality visible without policing. An illustrative scorecard:
Team Docs with owner Past review date
Payments 38 of 40 (95%) 3 of 40 (8%)
Search 20 of 25 (80%) 9 of 25 (36%)
Mobile 9 of 30 (30%) 18 of 30 (60%)
Mobile is the team to help first, not to shame.
Pitfalls
- Federated with no platform team ends in 40 different structures and no findability.
- Centralized with too few writers turns into a backlog; engineers route around it.
- A scorecard that punishes low scores encourages gaming; pair it with support for teams that are behind.
What would change my recommendation
If most docs are customer-facing and legally reviewed, lean central. If the organization is under roughly ten teams, a shared repo with light standards may be enough and the platform team can be part-time.
Unlock Full Question Bank
Get access to all 21 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.