Documentation and Knowledge Management Questions
Governing how teams capture and keep knowledge usable: documentation standards, ownership models (centralized, federated or dedicated team) and review cadences, incentives and culture change that get engineers to document, decision records and assumption logs, post-project reviews that preserve lessons, and knowledge-base strategy so organizational knowledge stays findable and attributed. Includes building or buying a knowledge platform, migration and consolidation plans, taxonomy, tagging and search quality, linking docs to code, capturing tacit expertise into the knowledge base, detecting and remediating stale docs across a documentation estate, and versioning and change policies for shared reports and metric definitions. Covers the governance and lifecycle of a body of documentation, not the craft of writing any single document.
Design a knowledge management system for a global engineering organization of about two thousand people. It has to support versioned, code-linked documentation, fast search, access controls, doc checks in CI, and usage, gap and staleness analytics. Describe the components, the governance, and where you would deliberately keep it simple.
Sample Answer
Direct answer
For about 2,000 engineers I would build a small set of well-understood components around docs-as-code (docs in the repos), a central index for search, identity-driven access, CI checks, and an analytics layer, with lightweight governance built on owners and review dates. Keep simple: one authoring format, one search engine, no bespoke editor, and no machine-learning features until search logs prove a need.
Terms
- Docs-as-code: docs stored as Markdown in repos and reviewed like code.
- CI (continuous integration): automated checks on each change.
- ACL / RBAC: access control lists / role-based access control (permissions granted by role or group).
- SSO: single sign-on, one login for all tools.
- Staleness: a doc likely out of date because time passed or its code changed.
- Content gap: a question people search for that has no good page.
- Front matter: a small metadata header at the top of a Markdown file (owner, type, sensitivity, last_reviewed).
- Documentation guild: a small volunteer group with one representative per major area that maintains shared standards. It advises and sets templates, unlike a central team that writes or approves everything.
- Federating ownership: each team owns its own docs instead of one central team owning all of them.
- Directory-to-group mapping: the company's employee directory (who reports to which team) is used to create the user groups that permissions follow, so a reorg updates access automatically.
- Boost: nudging certain results higher in search ranking. Drift: a doc slowly becoming wrong as the code changes.
Sizing (illustrative assumptions, not measurements)
Suppose 2,000 engineers in about 250 teams, each team owning roughly 20 pages: about 5,000 pages. With a review every 6 months on average, that is about 830 reviews a month, or 3 to 4 per team per month, which is why review dates must be automated rather than chased by hand. If each engineer searches about 3 times a week, that is roughly 6,000 searches a week, a small load that one managed search engine handles comfortably. This is the reasoning behind keeping the platform simple.
Components
| Component | Choice and reason |
|---|---|
| Authoring | Markdown in each repo with front matter (owner, type, sensitivity, last_reviewed). Versioned by git, so docs branch and tag with code |
| Code linking | Front matter or a mapping file links a page to code paths and services, so CI can spot drift |
| Build and CI checks | Link check, required fields, formatting, and "code changed, doc not touched" warning; blocking only for a short list of critical rules: (1) an owner field is present, (2) no broken internal links, (3) no credential-like strings (keys, passwords). Everything else, including "code changed, doc not touched", starts as a warning |
| Publishing | One static site build per merge, plus a portal page and redirects |
| Search | One managed search engine indexing all sources with synonyms and boosted freshness, filtering by user groups at query time |
| Access control | SSO groups from the identity system; sensitivity classes map to groups; page-level ACL for the small sensitive set, defaults open |
| Analytics | Search logs, page views, zero-result queries, and stale-page counts, in a dashboard |
Governance
- Ownership: every page has an owner team; reorganizations transfer ownership via a directory-to-group mapping.
- Review cadence: reviewed within its window (for example quarterly for runbooks, longer for concepts).
- Standards: a short style guide and templates, maintained by a small documentation guild with representatives from major areas rather than a central gatekeeping team.
- Escalation: unowned or repeatedly stale critical pages go to the engineering director for that area.
Analytics that drive action
- Usage: the top viewed and searched pages get quality investment first.
- Gaps: queries with zero results or immediate re-search become a backlog for owners.
- Staleness: pages past review date, grouped by team, in a monthly leadership report.
- Onboarding: time for a new hire to complete first tasks, and which pages they visit.
Global scale extras
- Multi-language: keep English as the source of truth, translate the top pages (by usage) rather than everything, and label translations that lag the source.
- Personalized ranking: boost by team, region and role from directory data. Simple boosts beat learned models here.
- Page-level access across business units: group-based ACLs with an audit log; keep the restricted set small so defaults stay open.
- Optional machine learning summarization: defer until search analytics show people struggle with long pages, run it as a pilot on non-sensitive content, and require it to cite its sources. The justification test is a measurable improvement in time to answer against the baseline, and a cost within budget.
Where I deliberately keep it simple
- One authoring format instead of supporting many.
- Off-the-shelf search and hosting instead of building.
- Permissions by group, not per-person exceptions.
- Automated freshness signals instead of manual audits.
- No custom ranking algorithm at launch.
Worked example (illustrative)
A payments engineer changes a service's timeout config. CI notes that the linked runbook was not touched, and the owner adds a one-line update in the same pull request. Meanwhile analytics shows "how to rotate keys" returns no results, and the security guild writes the missing page. Both are cheap loops driven by data.
Trade-offs and pitfalls
- Strict CI gates breed workarounds; use warnings first, and block only for the critical rules.
- Analytics can mislead: high views may mean confusion, so pair with feedback.
- Federating ownership scales better than a central team but needs the guild for consistency.
- What would change my call: heavy regulation would raise access and audit requirements, and a smaller organization could skip the guild and analytics dashboards.
What documentation standards would you set for an engineering team, and how would you get people to follow them without turning it into bureaucracy? Cover how the standards are kept up to date as the work changes.
Sample Answer
Direct answer
Set a small, tiered minimum standard (a few artifacts every service, pipeline or model must have), keep the docs in the same repository as the code and change them in the same pull request (PR), and enforce the standard with automation and defaults rather than meetings and sign-offs. Freshness comes from ownership plus a review date on every doc, not from goodwill. The test of "not bureaucracy" is that following the standard is the easiest path and the checks are run by tooling.
Structured elaboration
1. The standard itself: small and tiered
- Tier 1 (anything on-call, customer-facing or shared): a README (the front-page file of a repo: what it is, who owns it, how to run it, where to ask), a runbook (step-by-step instructions for responding to an alert or failure), and a decision record (a short note recording what was chosen, the options rejected and why) for hard-to-reverse choices.
- Tier 2 (internal tools, experiments): README only.
- ML and data assets: a model card (short document stating what a model is for, its training data, known limits and owner) or a dataset datasheet, plus a data dictionary (what each column means, its unit and source). For analytics projects add required metadata: owner, refresh schedule, source tables and last-reviewed date.
- One template per artifact, pre-filled, with the required fields marked. Templates are what stop inconsistent deliverables and the rework that follows.
- The ML and data items are optional add-ons for teams that ship models or datasets, not part of the core README, runbook and decision record set.
A minimal template skeleton, so the required fields are concrete (the CI check simply fails if a heading is missing):
README.md RUNBOOK.md decision record
# <service> # Alert: <alert name> # Decision: <statement>
Owner: <team> Severity: <P1-P3> Status: proposed | accepted | superseded
Last reviewed: <d> Symptoms: <what you see> Context / Options / Consequences
## What it does Diagnosis: <steps> Owner and review date
## Run locally Fix or rollback: <steps>
## Ask for help Escalate to: <who>
2. Governance as one framework (owners, templates, review cadence, tooling, incentives)
- Templates: the pre-filled skeleton above is the only format accepted for each artifact, stored in one template repository, so the CI check has a fixed set of required headings to test.
- Owners: every doc names a team owner (a CODEOWNERS file, which routes PR review to the named owners, plus a service catalog entry, meaning a central list of every service with its owner, docs link and on-call contact). An unowned doc is treated as deletable.
- Review cadence: each doc carries a
last_revieweddate. Tier 1 is reviewed at least every 6 months, and also whenever the system changes shape. - Tooling: docs-as-code (markdown in the repo, reviewed like code), a CI (continuous integration) job that checks required headings, broken links and stale dates, and one searchable site built from all repos.
- Incentives: doc work counts in performance and promotion criteria, authors are credited by name, and "a new hire followed the README and it worked" is celebrated. Punishment alone breeds checkbox docs.
3. Keeping standards current as work changes
- A PR template asks "docs updated or not needed, and why". Changes to alerts, APIs or schemas require touching the linked doc (a CI rule can map code paths to doc paths).
- A bot opens a ticket on the owner when
last_reviewedpasses its limit. Two missed cycles moves the doc toarchivedrather than leaving a confident but wrong page. - The standard itself is versioned and reviewed twice a year by a small docs working group (a few volunteers from different teams who own the standard), and any rule that nobody can justify with an incident or onboarding pain is removed.
4. Measuring compliance and handling non-compliance
- Coverage = Tier 1 assets with all required docs / all Tier 1 assets. Freshness = docs reviewed inside their window / docs required.
- Escalation ladder: bot nudge, ticket on the owner, visible on the team's monthly ops review, and only for Tier 1 a launch-readiness blocker, meaning the service cannot go live until its Tier 1 docs exist. Nothing else is gated.
Worked example
One team owns 15 Tier 1 services. The dashboard shows 12 have a README and runbook, and 9 of those are inside the 180-day review window.
coverage = 12 / 15 = 80%
freshness = 9 / 15 = 60% (9 of the 15 required sets are fresh)
The team lead sees two numbers and six named services to fix (three with no docs, three with docs past their review window), not a lecture. The runbook and decision record carry the same Last reviewed field as the README, so the CI stale-date check and the 6-month cadence apply to every Tier 1 artifact, not only the README. Note freshness is computed against the required 15, so a missing doc cannot hide behind a fresh one.
Rollout at scale (500 engineers, ten teams, two regions), phased over 12 months
- Months 1-2: pilot with two willing teams, build templates and the CI check, fix what the pilot hates.
- Months 3-5: add four teams, stand up the single search site and the service catalog, run a 2-hour training per engineer (500 x 2 hours = 1,000 engineer-hours, which is 25 engineer-weeks at 40 hours, a small fraction of the organization's yearly capacity).
- Months 6-9: remaining four teams; migrate only Tier 1 legacy docs (archive the rest, do not migrate everything).
- Months 10-12: switch on the launch-readiness gate for Tier 1, publish the first compliance report, review the standard.
- Rough cost, in people rather than invented dollars: a docs platform owner and a tooling engineer (about 2 full-time), plus one steward per team (the team's named docs champion who helps peers and chases reviews) at about 10% time (ten teams is about 1 full-time equivalent). Two regions: one site, async PR review so no time zone waits on another, and one regional steward each.
Trade-offs and pitfalls
- Too many required docs produces empty template filling. Cut the list until each item has a story of an incident or onboarding delay it prevents.
- Central mandate without team stewards makes the standard feel imposed. Stewards give teams a peer to ask.
- Gating everything on compliance encourages gaming (docs written to satisfy a linter). Gate only Tier 1 and spot-check quality by having a new joiner try the README.
- What would change my approach: for a 15-person team I would skip the catalog and the phases and adopt README, runbook and a PR checkbox in one week.
A cross-functional initiative just wrapped and everyone agrees the lessons are valuable, but past experience says they will be filed and forgotten. How would you capture them in a way that future teams actually find and use, and how would you check they changed anything?
Sample Answer
Direct answer
Lessons are forgotten because they are captured as a report and nobody has a reason to open it. I would capture a few specific, actionable lessons, attach each to the moment in a future project where it applies, convert the best ones into changes in templates, checklists or automation (so people use them without reading anything), and then check three things: were they found, were they applied, and did the same problem recur.
Structured elaboration
1. Capture: fewer, sharper, owned
- Hold the review within two weeks of wrap-up while memory is fresh, with people from every function that took part. Keep it blameless (focus on how the process allowed the problem, not who caused it).
- Distil to 5-8 lessons. Each lesson uses one card: context (what was going on), what happened, what to do next time in a concrete verb form, and the trigger ("when planning a vendor integration...").
- Each lesson has a named owner who is responsible for turning it into a change. A lesson without an owner is a wish.
2. Make them findable at the point of use
- Tag by the decision point or project phase the lesson affects (kickoff, vendor selection, launch readiness), not just by project name. People search by their current problem, not by the last project's title.
- Store in one place that the kickoff process itself points to: the project kickoff template includes a mandatory step, "review lessons tagged for this phase and list which apply".
- Push, do not just pull: a short summary goes to the owners of the next planned initiatives.
3. Convert lessons into changes
A lesson that only lives in a document depends on memory. Prefer, in order: an automated check or default, a checklist or template line, a named role or gate, and only last a document.
4. Check they changed anything
- Found: were lesson cards opened or cited at the next kickoffs (link clicks, or the kickoff template's "lessons reviewed" field filled in)?
- Applied: for each converted lesson, did the next initiative actually do the new step (audit two projects at 90 days)?
- Outcome: did the original problem recur? Track recurrence per lesson (yes, no, partly). Be honest that with a couple of projects this is evidence, not proof.
Worked example
A billing-migration initiative finished. One lesson: legal review of the new vendor contract began after the contract was drafted, which stalled launch for several weeks.
Lesson L3
Context: New payments vendor, contract drafted by procurement.
Happened: Legal saw the data-processing terms only after vendor selection, forcing renegotiation.
Next time: Start legal review at vendor shortlisting, before any pricing talks.
Trigger: Any project that adds a third party who will handle customer data.
Owner: Program manager for platform initiatives
Change: Add "Legal review started (date)" as a line in the vendor-selection checklist.
Check: At 90 days, did the next two vendor projects show that date before contract draft?
Suppose the audit finds one of the next two projects started legal review early and one did not. That is 1 of 2 applied: the owner asks why the second missed it (checklist not known to that team), and fixes the distribution, not the people.
Trade-offs and pitfalls
- A long retrospective document is comprehensive and unread. Prefer the card format.
- Generic lessons ("communicate earlier") cannot be applied. Force a trigger and a verb.
- Blame-shaped lessons make people hide problems next time.
- Recommend investing in checklist and template changes over a searchable archive, because the archive still depends on someone thinking to search. What would change my call: for a rare, one-off initiative that will never repeat, a short written note to the sponsor is enough.
Design a knowledge-base strategy for an infrastructure organization so that runbooks, architecture docs and postmortems stay findable, current and credited to their authors. What governance, ownership and incentives would you put in place?
Sample Answer
Direct answer
Treat the knowledge base as a product with owners, a lifecycle and a search experience, not a wiki people dump into. Give every runbook, architecture doc and postmortem one accountable owner (a team, not a person), a lifecycle status and a review-by date, put them behind one search entry point with a small controlled taxonomy, and give authors visible credit and career recognition. Then measure findability, freshness and use, and fix what the numbers show.
Structured elaboration
1. Content types and what "current" means for each
| Type | What it is | Owner | Freshness rule |
|---|---|---|---|
| Runbook | Step-by-step instructions for an alert or operation | Team that gets paged for the service | Verified by actually running it (a game day, meaning a planned exercise where an engineer follows the runbook against a safe, simulated failure, or a real incident) at least every 6 months |
| Architecture doc | How a system is built and why | Service or platform team | Reviewed at every major change and yearly |
| Postmortem | Blameless written review of an incident (blameless: it examines causes and process, not who to punish) | Incident lead, action items owned by named teams | Not edited after the fact except to add outcomes. Its action items are tracked to closure |
2. Findability
- One search box across all types, with filters on service, team, type and status. Titles follow a convention ("Runbook: payments-api: high 5xx alert", where 5xx means the service is returning server-error HTTP responses) so a person paged at 3 a.m. finds it by the alert name.
- Every alert links straight to its runbook, and every runbook links to the architecture doc and past postmortems for that service. Links beat search when you are in a hurry.
- A small controlled taxonomy (a fixed, curated list of tags that only the platform group changes; service names come from the service catalog, the authoritative list of services and owners). Free tagging sprawls ("db", "database" and "DB-prod" all end up meaning the same thing).
- Track searches with no results and searches followed by a second search. Those are the gaps.
3. Governance
- Ownership comes from the service catalog. When a team is reorganized, transferring catalog entries transfers the docs, so nothing becomes an orphan.
- Lifecycle: draft, current, needs-review (auto-set when the review-by date passes), archived (hidden from default search, kept for history, with a banner pointing to the replacement).
- Edit model: anyone can propose, owners approve, and history is kept (docs-as-code, meaning docs written as Markdown files in Git and reviewed like code, or wiki page history) so every version is recoverable.
- Access control: read-open by default for engineers. Restricted spaces for security-sensitive procedures. Never store credentials in docs at all: link to the secrets manager (the dedicated tool that stores passwords and API keys with access control) instead.
- Versioning: architecture docs mention the system version or date they describe. Superseded docs point forward to their replacement rather than being deleted.
4. Incentives and credit
- Bylines and a "maintained by" field on every page, for example this header at the top of a runbook:
Maintained by: payments-platform team (contact: #payments-oncall)
Authors: A. Rivera, J. Chen
Status: current | Last verified: 2026-08-14 | Review by: 2027-02-14
Replaced by: (none)
Archived pages instead show a banner such as "Archived: superseded by <link>".
- Authors appear in the monthly ops review when their runbook resolved an incident.
- Documentation contributions count in performance and promotion packets, and one recognition slot per quarter.
- Guard against gaming: do not reward raw page counts. Reward pages that other people used (link clicks from alerts, incident references) and pages verified as still correct.
5. Extending the same rules to a global data engineering organization
Combine the data catalog (the searchable inventory of datasets with owners and definitions), onboarding docs, runbooks and design patterns under the same rules. Each dataset entry names an owner and links its runbook. Access control follows the data's sensitivity classification (a label such as public, internal or restricted that decides who may see it), so a catalog entry for a restricted dataset is visible but its sample data is not. Design patterns are versioned like code so teams can say "we use pattern v2".
Worked example
An infrastructure organization has 120 runbooks. After adding review-by dates, the dashboard reports 84 verified within the last 180 days.
verified share = 84 / 120 = 70%
The target for services with an on-call rotation is 90%, so 24 more runbooks need verification. The team owning the most unverified ones gets a slot in the next game day (a planned exercise where an engineer follows the runbook on a safe failure). This gives a concrete quarterly goal and a name against each gap, instead of "docs are stale".
Trade-offs and pitfalls
- Central docs team versus owner teams: a central team gets consistency but becomes a bottleneck. Recommend owner teams writing, with a small central group owning templates, search and the metrics.
- A single wiki versus docs-in-repo: keep runbooks and architecture docs in repos next to the code where drift is visible in review, and index them into the one search site.
- Postmortems are worthless if their action items rot. Track them as tickets and report the open count monthly.
- Credit metrics get gamed. Pair usage metrics with periodic spot verification.
- What would change my call: a small organization should skip the catalog integration and start with owners, a review date field and the alert-to-runbook links.
You are introducing decision records for product and analytics decisions made under uncertainty. Draft the template you would give the team, then explain how you would get people to actually fill it in and revisit it later rather than let it rot.
Sample Answer
Direct answer
For decisions made under uncertainty, the record must capture more than the choice: it needs the assumptions, the evidence and its limits, how we will know if we were right (metrics and a date), the fallback, and who disagreed and why. I would give the team a one-page template like the one below, embed it in the places where work already happens (sprint planning, handoffs, roadmap reviews), and keep it alive with review dates that create tickets automatically.
Structured elaboration
The template
Two terms first. A pull request (PR) is a proposed change to a shared file or code that teammates comment on before it is accepted; on a non-engineering team, read it as the comment thread on a proposed change. Validate by means the date by which we check the success metric.
Decision: <one-sentence decision> ID / Date / Status (proposed | accepted | reversed | superseded)
Owner (decides): <name> Consulted: <names> Informed: <names>
Context: <what problem, what deadline, what we know>
Options considered: <option, why not chosen> (at least two)
Assumptions: A1 <statement> confidence H/M/L how we will test it
Evidence: <data links, with a note on quality and sample size>
Success metrics: <metric, threshold, when we check> ("validate by")
Risks and fallback: <what would make us reverse, and the plan B>
Dissent log: <who disagreed, their argument, link to the pull request or thread, how it was resolved>
Review date: <date> Outcome (filled later): <what actually happened>
Fields that matter most under uncertainty
- Assumptions with confidence: the decision is only as good as the beliefs under it. Marking low-confidence ones tells reviewers what to test first.
- Metrics to validate with a date and threshold: without a threshold nobody can say we were wrong.
- Stakeholders and fallbacks: who is affected and what we do if it fails, decided before we are emotionally committed.
- Dissent log (conflict log): recording disagreement is not a blame device. It preserves the reasoning of people who lost the argument, links to their pull request comments or thread, and lets a later reviewer see if the dissent was proven right.
Getting people to fill it in
- Proportionality: only for decisions above a threshold (irreversible, costly, or affecting other teams). A 15-minute fill-in, pre-populated from the ticket.
- Put it in the workflow: sprint planning asks "does this story depend on an open decision record?", and handoffs link the record so the next team inherits the assumptions.
- Owner writes it, others comment: ownership is one name.
- Visible use: open the record in the retro when the review date comes up. If people see records being used, they write them.
Keeping it from rotting
- The review date creates a calendar ticket for the owner. On review they set outcome: confirmed, partly, reversed.
- Retention: keep records permanently (small), mark them superseded rather than deleting, and each quarter review the aggregated outcomes to learn which kinds of assumptions tend to fail.
Worked example
Decision: cap the free plan at 3 projects.
- Assumption A1 (confidence: medium): users who reach a second project are more likely to upgrade. Test: compare upgrade rates for users with 1 vs 2+ projects in the event logs.
- Metric: at least 20% of new free users create a second project within 30 days (illustrative threshold). Check date: 45 days after launch.
- Fallback: the second-project metric is only a necessary precondition for the cap to bind (a cap of 3 only affects users who reach a third project), so on its own it cannot show the cap is binding. If fewer than 20% of new free users reach a second project, a cap of 3 is not what limits them, so raising it to 5 would change nothing and the cap should be dropped or replaced by a different upgrade lever. Raise the cap to 5 only if the dissent proves right, that is, if churn among students rises past the agreed threshold. Add a second metric so A1 is actually tested: upgrade rate of users with 2+ projects versus 1 project after 45 days. Also track the share of new free users who reach the 3-project limit, since that is the group the cap actually affects.
- Dissent: support lead argued the cap would raise churn among students. Logged with a link to their comment. At the review, the outcome field records what churn actually did, so the dissent is judged fairly.
Filled in for this example:
Decision: Cap the free plan at 3 projects. ID: D-014 / 2026-03-10 / accepted
Owner (decides): Head of Product Consulted: support lead, growth analyst Informed: sales, engineering
Context: Free users cost hosting money; no data on what drives upgrades.
Options considered: No cap (rejected: cost); cap at 1 (rejected: blocks evaluation).
Assumptions: A1 users who reach a second project upgrade more; confidence M; test: upgrade rate 1 vs 2+ projects
Evidence: Event log query, 3 months, about 4,000 users (small for paid conversion)
Success metrics: >= 20% of new free users create a second project within 30 days; upgrade rate 2+ vs 1 project at 45 days; validate by 2026-04-24
Risks and fallback: If <20% reach a 2nd project the cap is not binding: drop it. If student churn rises past threshold: raise cap to 5
Dissent log: Support lead: cap raises student churn. Link: comment thread on the proposal. Resolved: accepted with churn tracked.
Review date: 2026-04-24 Outcome: (filled at review)
Trade-offs and pitfalls
- A template that is too long dies. If people skip fields, cut them.
- A record with no threshold looks rigorous but cannot be wrong. Insist on a number and a date.
- Do not use the dissent log to relitigate. It records reasoning, not a running argument.
Unlock Full Question Bank
Get access to all 32 Documentation and Knowledge Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.