Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
Your organization must choose between building an internal knowledge platform and buying a hosted documentation product. Walk through how you would make and defend that decision: which criteria matter, how you weight them, and what would change your answer.
Sample Answer
Direct answer
Decide from criteria tied to the actual problem (people cannot find or trust knowledge), not from a preference for building. My default recommendation for most organizations is to buy or adopt a hosted product for the reading and editing experience, and keep the parts that must live with the code (reference docs, runbooks tied to services) in the repository. Build only if a specific requirement rules every product out, because a knowledge platform is a long-lived product that needs an owner, on-call, and roadmap forever.
Terms used below
- Docs-as-code: writing docs as Markdown files in the same Git repository as the code, reviewed by pull request like code.
- Data residency: a rule about which country or region your data may be stored in. Data sovereignty is the stricter legal version: the data must stay under one country's laws.
- Group sync: automatically copying your company's user groups (for example "SRE team") into the tool so page permissions follow them.
- Audit log: a tamper-resistant record of who viewed or changed what. Retention: how long content and history are kept before deletion.
- Per-seat pricing: paying per user per month, so cost grows with headcount. Exit (lock-in): how hard it is to leave, mainly whether you can export everything cleanly.
- Self-hosted open-source: free software you install and run yourself, so you own the operations work.
Criteria and weights (for a mid-sized engineering organization)
| Criterion | Weight | What to test |
|---|---|---|
| Permissions and access control | 20% | Per-space and per-page rights, single sign-on (SSO, one company login), group sync, guest access |
| Search quality | 20% | Run 20 real past questions; does the right page rank in the top three? |
| Compliance and audit | 15% | Audit log, retention, data residency, export of everything |
| Code linking and docs-as-code fit | 15% | Can docs live in Git, be reviewed by pull request, and link to code? |
| Total cost over 3 years | 15% | Licences plus engineer time; include build's hidden run cost |
| Migration cost and exit | 10% | Import tooling, link preservation, ability to leave |
| Extensibility and integrations | 5% | API, chat and ticket integrations |
Weights depend on context; the point is to write them down before demos, so the vendor with the best presentation does not set them.
Worked example: SRE organization, three options
| Wiki-style tool | Repo-based docs (Markdown in Git, published as a site) | Newer all-in-one workspace | |
|---|---|---|---|
| Permissions | Strong, per page | Only as good as repo access; coarse | Good, still maturing |
| Code linking | Manual links, drift | Best: same PR as code | Integrations, partial |
| Search | Good within the tool | Needs a search layer added | Good, often AI-assisted |
| Compliance | Mature audit and retention | Git history is audit; retention is DIY | Varies by vendor, verify |
| Migration cost | Low if already wiki | Medium: convert and restructure | Medium; import quality varies |
Score each option 1 to 5 per criterion (5 is best), multiply by the weight, and add up. The options: a hosted wiki product, repo-based docs, and a newer all-in-one workspace (a hosted tool combining notes, docs and small databases). Illustrative scores for this SRE org:
| Criterion (weight) | Wiki | Repo docs | All-in-one |
|---|---|---|---|
| Permissions (0.20) | 5 (1.00) | 2 (0.40) | 3 (0.60) |
| Search (0.20) | 4 (0.80) | 2 (0.40) | 4 (0.80) |
| Compliance and audit (0.15) | 5 (0.75) | 3 (0.45) | 3 (0.45) |
| Code linking (0.15) | 2 (0.30) | 5 (0.75) | 3 (0.45) |
| 3-year cost (0.15) | 3 (0.45) | 4 (0.60) | 3 (0.45) |
| Migration and exit (0.10) | 4 (0.40) | 3 (0.30) | 3 (0.30) |
| Integrations (0.05) | 4 (0.20) | 3 (0.15) | 4 (0.20) |
| Total (max 5.0) | 3.90 | 3.05 | 3.25 |
For repo docs, code linking is 5 so 0.15 x 5 = 0.75, while search is 2 so 0.20 x 2 = 0.40. Across all content the wiki wins (3.90). But the weights assume one blended pile of content. Runbooks and service docs are a different class: they change with code and need review, so re-weight for them (code linking 0.35, search 0.15, permissions 0.10, compliance 0.10, cost 0.15, migration 0.10, integrations 0.05, which again sum to 1.00) and recompute: repo docs 3.60, wiki 3.35, all-in-one 3.20. Repo docs win that class.
That is why the recommendation is a hybrid: repo-based docs for runbooks and service docs, plus the wiki for broad non-code material (policies, onboarding, meeting notes). The matrix picks the best tool per content class, and a hybrid is defensible when two classes rank differently. Then sanity check the ranking against your gut feeling, and if they disagree a criterion is probably missing.
Build versus buy
- Build makes sense only with a hard constraint: data may not leave your environment, a deep integration no product offers, or scale where per-seat pricing is prohibitive.
- Buy wins when time to value matters and the team cannot staff a product. Include the 3-year staffing cost of build: at least one engineer's ongoing attention.
What would change my answer
- A strict data-sovereignty rule pushes to self-hosted open-source or build.
- A pilot showing search relevance below a threshold (say fewer than half of the 20 test questions land in the top three) disqualifies that product regardless of price.
- If exit is hard (no clean export), demote that product unless the discount is large.
Defending the decision
Run a two-week pilot with real content and real users, publish the scored matrix and the assumptions, and record an ADR (architecture decision record, a short document capturing a decision, options and reasons) so the decision can be revisited when assumptions change.
In a large engineering organization, compare a centralized documentation model with a federated one. How do ownership, findability, governance overhead, update speed and standards enforcement differ, when is each appropriate, and what hybrid would you propose?
Sample Answer
Direct answer
A centralized model puts one team in charge of writing, structure and tooling. A federated model lets each engineering team own its docs under shared standards. In a large organization neither extreme works alone, so I would recommend a hybrid: central ownership of the platform, standards and navigation; federated ownership of content, with each doc having an accountable owning team.
Comparison
| Dimension | Centralized | Federated |
|---|---|---|
| Ownership | Clear (doc team), but far from the code | Clear per team, but uneven across teams |
| Findability | High: one structure, one search | Low unless indexed: many styles, locations |
| Governance overhead | High: intake queue, review | Lower per team, but coordination cost across teams |
| Update speed | Slow: writers wait on engineers | Fast: authors are the people making the change |
| Standards enforcement | Easy: one gatekeeper | Hard: needs automation and peer pressure |
| Accuracy | Prone to lag and errors | Better where teams care; poor where they do not |
When each fits
- Centralized: a small organization, highly regulated or customer-facing docs where consistency and legal review matter, or a young platform where standards do not exist yet.
- Federated: many teams with different stacks, fast-moving internal engineering docs, and a culture where teams already own services end to end.
The hybrid I would propose (worked example)
Imagine 40 teams:
- A small platform team (2 to 4 people) owns the publishing tool, the templates, the style guide, the global search and the navigation. It does not write team content.
- Each team owns its docs in its own repository, marked with an owner in the doc header; a CODEOWNERS file (a repository file naming who must review changes in a path) makes this enforceable.
- Automated checks in CI (automated checks that run on each change) enforce standards: required sections, broken links, owner present, review date. This replaces a human gatekeeper.
- Tier-1 material (the small set of docs where an error is costly, such as public API, incident runbooks, security policies) gets extra editorial review from a small docs guild (one representative per team).
- All content is indexed into one search, and the platform team publishes a scorecard per team (share of docs with owner, share past review date), which makes uneven quality visible without policing. An illustrative scorecard:
Team Docs with owner Past review date
Payments 38 of 40 (95%) 3 of 40 (8%)
Search 20 of 25 (80%) 9 of 25 (36%)
Mobile 9 of 30 (30%) 18 of 30 (60%)
Mobile is the team to help first, not to shame.
Pitfalls
- Federated with no platform team ends in 40 different structures and no findability.
- Centralized with too few writers turns into a backlog; engineers route around it.
- A scorecard that punishes low scores encourages gaming; pair it with support for teams that are behind.
What would change my recommendation
If most docs are customer-facing and legally reviewed, lean central. If the organization is under roughly ten teams, a shared repo with light standards may be enough and the platform team can be part-time.
You need to stand up a searchable internal knowledge base for engineering documentation that lives near the code and needs fine-grained permissions. Sketch a simple architecture, explain your search choices, the authoring and review workflow, the ownership model, and how you keep content current.
Sample Answer
Direct answer
Keep documentation as files in the code repositories (docs-as-code), build them into one site on every merge, and index the built content in a search service that filters results by the searcher's permissions. Ownership lives in the repo (a CODEOWNERS file), review is the normal pull request, and freshness is enforced by automated checks and review dates. That is enough for a modest scale; do not build a custom platform.
Terms
- Docs-as-code: writing docs in plain text (Markdown) inside the repository and reviewing them like code.
- CODEOWNERS: a repository file naming who must approve changes to given paths.
- CI (continuous integration): automated checks that run on every change.
- ACL (access control list): a rule on who may view an item.
- Audit log: a record of who accessed or changed what.
- Front matter: a small metadata header at the top of a Markdown file (owner, type, sensitivity, last_reviewed) that tools can read.
- Facets: clickable filters in search results, such as by team or page type. Typo tolerance: still finding "kubernetes" when someone types "kuberentes".
- Link unfurl: the preview card a chat tool shows when someone pastes a link. Lineage: a record of which data sources feed which datasets or reports.
- Model card: a short standard document describing what a machine learning model is for, its training data and its limits. Metric spec: a written definition of a business metric.
Architecture (simple)
Repos (Markdown + front matter) -> CI build and checks -> Static docs site
|
v
Search index (with access groups per page)
^
User -> single sign-on -> search API filters hits by user's groups
Search choices
| Option | Fits when | Watch out for |
|---|---|---|
| Database full-text search (for example PostgreSQL, a common open-source database with built-in text search) | Under a few hundred thousand pages, small team | Ranking tuning is manual |
| Dedicated search engine (for example OpenSearch, an open-source search server, or a hosted service) | Many repos, typo tolerance, synonyms, facets | Extra operations cost |
| Hosted enterprise search | Need connectors to chat, tickets, wiki | Cost and permission syncing |
| My default: start with the database or hosted engine; move up only when search quality complaints justify it. Whichever is chosen, store each page's allowed groups in the index and apply them as a filter at query time so titles and snippets of restricted pages never appear. Filtering after retrieval on the client is the classic leak: the server has already sent the restricted titles and snippets to the browser, so anyone can read them in the network response even if the page hides them. An index entry that supports the safe approach looks like this: |
{ "title": "Checkout runbook", "body": "...", "allowed_groups": ["payments-oncall"] }
query: text = "checkout runbook" AND allowed_groups contains any of (caller's groups)
The filter is part of the query, so a non-member gets zero hits back, not hits that are hidden afterwards.
Authoring and review workflow
- Author edits Markdown in the repo, using a template (runbook, architecture decision record or ADR, how-to guide) with front matter: owner, type, sensitivity, last_reviewed.
- Pull request triggers CI: link checker, formatting, required front matter, spelling, and a check that code-linked pages changed when the code did (or a note explains why not).
- CODEOWNERS approval merges, and the site rebuilds.
Ownership model
Every page has an owner team via CODEOWNERS or front matter. A weekly report lists pages past their review date, sent to the owner team. Orphans (owner team deleted) escalate to the engineering manager.
Keeping content current
- Review-by dates and a staleness banner after the date.
- CI flags pages whose linked code has changed since last review.
- Usage data: heavily used but stale pages get priority.
- Deleting is a feature: archive deprecated pages with redirects.
Common variations (optional extras, rarely the first question in an interview; same skeleton, different fields)
- Machine learning (ML) or data organization: add model cards and runbooks as typed pages, restrict sensitive ones by group, and keep an audit log of reads and edits. Auto-generated metadata (schemas, lineage) merges with human-written pages, and CI keeps the two consistent.
- A wiki serving about 100 teams: add versioned templates for runbooks and ADRs (a template for architecture decision records), a tag list with a steward, ranking by freshness and usage, and Slack and GitHub integrations (link unfurls, a bot posting stale-page reminders).
- BI artifacts: metric specs and dashboard templates as typed records with owner, definition and access, searchable alongside the rest.
I would keep all of these at component level, adding a variant only when a team needs it.
Worked example (illustrative)
An SRE edits the checkout service runbook in the repo. CI fails because the front matter (the metadata header) lacks an owner, so she adds one. The reviewer from the owning team approves, the site rebuilds, and the index updates with group "payments-oncall". A contractor outside that group searches "checkout runbook" and sees no hit. Three months later the owner gets the stale-page report.
Trade-offs and pitfalls
- Docs-as-code excludes non-engineers unless there is an editing path (a web editor over the repo), so plan one.
- Permission syncing between repos and the index is the hardest part. Test it with a deny case.
- Consistency versus speed: generated metadata is fast but shallow, human docs deeper but stale. Label which is which.
- What would change my call: heavy non-technical contributors favor a hosted wiki with an export to the repo.
Your knowledge base is stale and teams are repeating incidents because playbooks are obsolete. Propose a remediation program: what you clean up first, how stale content is detected automatically, how humans review, what stops recurrence, and a realistic timeline and staffing.
Sample Answer
Direct answer
Run a time-boxed remediation program aimed at the runbooks that pages and incidents actually touch. Clean up tier-1 procedures first, detect staleness automatically with checks against real infrastructure and incident history, require a human to walk each procedure before it is marked verified, and stop recurrence by making a working runbook a condition for shipping an alert and for closing a postmortem. Archive everything nobody claims instead of reviewing it.
Terms. A playbook or runbook is a step-by-step procedure for handling an operational event. A postmortem is a written review of an incident; blameless means it looks for system causes, not someone to blame. A tier-1 service is one whose failure pages people or hurts revenue; tier 2 is important but a failure degrades it rather than stopping it (for example an internal reporting tool). A paging alert is an alert that wakes someone up. A staging environment is a safe copy of production for testing. The inventory is the current list of live hosts and services, taken from your configuration database. A decommissioned host is one that has been shut down and removed.
1. What to clean up first
Rank runbooks by consequence, not by age:
- Runbooks linked from paging alerts on tier-1 services.
- Runbooks named in the last 12 months of postmortems as missing, wrong or ignored (those are the ones that already repeated an incident).
- Runbooks for tier-2 services.
- Everything else: put an archive notice on it and delete or archive after 30 days if no owner claims it.
2. Automatic staleness detection
- Nightly job: links, hostnames, service names and commands referenced in the runbook checked against the current inventory. A reference to a decommissioned host is a hard flag.
- Alert-to-runbook check: every paging alert's runbook URL must resolve and must have a verified date inside its interval.
- Incident hook: responders press "runbook was wrong or missing" during or after an incident, which creates a ticket automatically.
- What the nightly check produces, for one runbook:
db-failover step 4 references host db-eu-2 (decommissioned 2026-06-30): HARD FLAG. Ticket opened for payments-team, due in 5 business days. - Age since last verified, used as a weak signal only (age alone is a poor predictor; broken references and incident flags are stronger).
3. Human review
The owner walks the runbook step by step in a staging environment, or as a tabletop drill (a talked-through rehearsal), with a second person who did not write it. They fix wrong steps, record the date and both names, and only then is the page marked verified. Reading it and clicking "still accurate" does not count.
4. Preventing recurrence
- CI (continuous integration) lint: an automated check that runs on every proposed change and rejects it when a rule is broken, here: an alert definition cannot merge without a runbook link. Illustrative output of a small custom lint script (no standard tool prints exactly this; you define the wording) when it fails:
FAIL alerts/checkout.yaml: "CheckoutLatencyHigh" has severity=page but no runbook_url
FAIL alerts/search.yaml: runbook_url https://wiki.example.com/runbooks/search-old returns 404
- Postmortem action items include "update the runbook" and the postmortem cannot close until the change is merged.
- Runbook lives next to the service code and its owner appears in the repo's CODEOWNERS file (a file that maps paths to owning teams), so changes get reviewed.
- Quarterly drill of the tier-1 runbooks (a game day, a rehearsal where you deliberately trigger a failure).
5. Realistic timeline and staffing
Illustrative sizing: 400 runbooks, of which 40 are tier 1 and 100 are tier 2.
| Phase | Weeks | Work | Effort |
|---|---|---|---|
| 0. Inventory and triage | 1-2 | Export all runbooks, tag tier and owner | Program lead, part-time |
| 1. Tier 1 | 3-8 | Walk and fix 40 runbooks, 1.5 hours each: 60 hours | Owners of each service: 60 hours over 6 weeks is about 10 hours per week in total, so with roughly 10 owning teams that is 1 to 2 hours per week each |
| 2. Detection and CI guard | 4-10 | Build nightly reference check, alert lint | One engineer, about two weeks |
| 3. Tier 2 | 9-16 | 100 runbooks at 1 hour: 100 hours | Owners |
| 4. Long tail | 12-16 | Archive unclaimed after 30-day notice | Program lead |
Staffing: one part-time program lead, one engineer for tooling, and a small time budget from each owning team. The human review effort totals about 160 hours (60 for tier 1 plus 100 for tier 2), which is roughly 4 person-weeks spread across teams, not one team's burden.
Success measures. No paging alert without a verified runbook; postmortems citing "runbook wrong" decline quarter over quarter; nightly reference-check failures trend to zero.
Pitfalls
- Reviewing all 400 equally wastes the program's budget on the long tail.
- A one-time cleanup decays; without the CI and postmortem gates you are back here in a year.
- What would change my call: if the team is tiny, skip the tooling and run a monthly tier-1 drill instead.
You are moving scattered documentation into one centralized knowledge base. Lay out your migration plan from first inventory to cutover: how you decide what to keep, merge or retire, how content gets owners and stays findable, what happens to old material, and which measures tell you the migration worked.
Sample Answer
Direct answer
Treat the migration as content curation first, tooling second: inventory everything, decide keep, merge or retire using a simple scoring rule, put owners and metadata on every kept page BEFORE moving it, migrate in waves ordered by operational criticality, leave redirects behind, and judge success by whether people find answers faster, not by how many pages moved.
Terms
- Inventory: a spreadsheet listing every doc, where it lives, who wrote it and when it last changed.
- Redirect: an old link that automatically forwards to the new location.
- Cutover: the moment the new system becomes the official source.
- Runbook: step-by-step instructions for operating or fixing a system.
Phase 1: Inventory and triage (weeks 1-2)
Crawl the wikis, shared drives and repos, and add the informal sources: chat channels and personal notes. Record for each item: location, owner, last edit, views if known, and which system it describes.
Decide keep, merge or retire with a short rule:
| Decision | Rule |
|---|---|
| Keep | Still accurate, describes a live system, and is used or operationally needed |
| Merge | Overlaps another page on the same topic; combine into one canonical page |
| Retire | Describes decommissioned systems, unedited for a long time, and nobody uses it |
| Ask a domain owner to confirm every retire in bulk, so nothing critical vanishes silently. |
Worked example (illustrative arithmetic)
Inventory finds 400 items. Triage results: 120 keep, 90 merge (they collapse into 30 pages), 190 retire. The new knowledge base receives 120 + 30 = 150 pages, 37.5 percent of the original. Every item still has a recorded fate in the inventory, so the numbers reconcile: 120 + 90 + 190 = 400.
Phase 2: Structure, ownership, findability
- Target structure: a few top-level areas (for example Services, Runbooks, Processes, Onboarding), not a mirror of the org chart.
- Tags: a small controlled list (team, system, lifecycle, audience); no free-form tags at launch.
- Every page has an owner team and a last-reviewed date; pages without an owner are not migrated.
- Search: index titles, headings and tags, boost pages by verified freshness, and test with a list of the 20 most common real questions.
Phase 3: Migrate in waves, by operational criticality
For knowledge that lives only in chat channels and personal notes, rank by "what hurts if it is missing at 3am": on-call runbooks and recovery steps first, then setup and onboarding, then background context. Capture chat knowledge by turning a decision or fix into a proper page (with the thread linked), not by dumping the export. Set a measurable onboarding-speed goal, for example that a new engineer completes the first supervised on-call shift using only the knowledge base, and measure the median days before and after.
Phase 4: Old material and cutover
Freeze the old wikis to read-only with a banner pointing to the new location, keep redirects for a defined period, archive (do not delete) retired items for legal or historical reasons, then switch off after traffic to old links falls.
Success measures
- Search: share of searches with a click and no immediate re-search, and zero-result queries falling.
- Coverage: percent of critical systems with an owned, current runbook.
- Freshness: percent of pages reviewed within their window.
- Onboarding speed (above) and fewer "where is the doc for X" questions in chat.
Trade-offs and pitfalls
- Migrating everything "just in case" copies the rot into the new home.
- No owners means the new system decays as the old one did.
- A big-bang cutover is riskier than waves, but waves prolong two sources of truth, so set a firm end date.
- What would change my call: a small team with 50 pages should skip scoring and simply re-read them.
Unlock Full Question Bank
Get access to all 12 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.