Technical Writing and Documentation Questions
The craft of producing durable, reference-quality written artifacts and keeping them accurate: READMEs and quick-start guides, design docs, RFCs and technical proposals, runbooks and deployment guides, model cards, datasheets and data dictionaries, bug reports and reproducible examples, postmortem write-ups, handoff documents, pull request descriptions, code comments, release notes, experiment reports, and knowledge-base articles. Covers structure and information design, writing for a specific audience and for future readers (including plain language and accessibility), templates and style standards, docs-as-code workflows with CI checks, testing of examples and snippets, documentation review and quality checks, versioning and freshness checks on the documents you own, keeping sensitive data out of docs, and measuring whether documentation works. Architecture decision records, API reference docs, PRDs and PR/FAQs, and live presentations are covered elsewhere.
Design a README-driven development process for your team to improve documentation quality. How do you enforce updates in pull requests, and how do you avoid it becoming box-ticking?
Sample Answer
Direct answer
README-driven development means writing the README (the front-page document explaining what a project does and how to use it) before the code, as a lightweight spec: if you cannot explain the feature simply, the design is not ready. To keep docs current I would combine a CI (continuous integration, automatic PR checks) rule that ties code changes to docs changes, an executable README so wrong instructions fail a build, and human review that reads the docs diff against the code diff. To avoid box-ticking, every waiver (an explicit, recorded exception to the rule) needs a real reason, and a monthly audit checks the docs the way a new joiner would use them.
The process
- Design PR first. A new feature starts as a pull request (PR, a proposed change awaiting review) that changes only the README or a docs page: what it does, the command or API call, one example, and what it does not do. Reviewers argue about behaviour here, when changes are cheap.
- Implementation PR second. The code PR must keep that documented behaviour true, or change the docs in the same PR.
- Enforcement in the PR. The CI rule is the script below: it fails any PR that touches public code without touching the docs, unless the author supplies a reason.
- Executable README. CI extracts the quickstart commands from the README and runs them in a clean container (a fresh, disposable environment with nothing left over from earlier runs, so hidden setup cannot make broken instructions look like they work). If the instructions stop working, the build breaks.
Enforcement script (tested in a scratch repo)
#!/usr/bin/env bash
# Fail a pull request that changes public code without touching docs, unless the PR says why not.
# Usage: check_docs.sh <base-ref> [skip-reason]
set -euo pipefail
base="$1"; reason="${2:-}"
changed=$(git diff --name-only "$base"...HEAD)
code=$(echo "$changed" | grep -E '^src/' || true)
docs=$(echo "$changed" | grep -E '^(README\.md|docs/)' || true)
if [ -n "$code" ] && [ -z "$docs" ]; then
if [ ${#reason} -ge 15 ]; then echo "docs skipped, reason: $reason"; exit 0; fi
echo "src/ changed but no README/docs change. Update docs or give a real reason (15+ chars)."; exit 1
fi
echo "docs check passed"
Running it against a PR that changes src/cli.py only:
$ ./check_docs.sh main
src/ changed but no README/docs change. Update docs or give a real reason (15+ chars).
$ ./check_docs.sh main "n/a"
src/ changed but no README/docs change. Update docs or give a real reason (15+ chars).
$ ./check_docs.sh main "internal refactor, no behaviour change"
docs skipped, reason: internal refactor, no behaviour change
After a commit that adds a README section, the same command prints docs check passed.
Reading the script line by line
set -euo pipefail: stop at the first error (-e), treat an unset variable as an error (-u), and fail a pipeline if any part of it fails (pipefail).reason="${2:-}": the second argument, or an empty string if none was given.git diff --name-only "$base"...HEAD: list the names of files changed on this branch since it split frombase. The three dots compare against the common ancestor, so commits made onmainafter you branched do not show up.grep -E '^src/'keeps only lines that start withsrc/(-Eenables extended regular expressions).|| trueis needed becausegrepexits with status 1 when nothing matches, which underset -ewould kill the script.${#reason} -ge 15: the reason's length is at least 15 characters.
Be honest about that last check: it only stops lazy waivers such as n/a. internal refactor, no behaviour change passes it, as the trace shows, whether or not it is true. The length rule is a speed bump; the real defence is a reviewer reading the reason and the team watching the waiver rate.
Worked example
The team adds a --retries flag to a command-line tool. PR 1 adds a "Retries" section to the README: default of 3, one example command, and a note that retries apply only to network errors. Reviewers say "network errors only is surprising, should timeouts count?" and the answer is settled before any code exists. PR 2 implements it, and CI runs the README example command.
Avoiding box-ticking
- The waiver needs a reason of substance, and the reviewer must agree with it. A repeated "n/a" gets discussed, not waved through.
- Reviewers are told to check the docs diff against the code diff, not just that a docs diff exists.
- Monthly fresh-eyes audit: pick five merged PRs, and someone who has not seen the feature follows the README cold. Every place they get stuck becomes a fix.
- Watch the waiver rate. If 60 percent of PRs use it, the rule is wrongly aimed (too broad a path filter, meaning the rule that decides which changed paths must come with docs) or being gamed. Tighten the paths to public interfaces only.
Trade-offs and pitfalls
- Path rules only detect that a docs file changed, not that it is right. That gap is why the executable README and human review matter.
- README-first can be slow for tiny changes. Limit the design-PR step to user-visible behaviour, and let internal refactors skip it with a stated reason.
- Executable docs need a clean environment. Flaky setup steps (ones that pass on some runs and fail on others) will train people to ignore red builds.
Docs in your repository include code snippets, SQL examples and small notebooks that keep breaking silently. How do you test them in CI, sandbox them, and deal with flaky or slow examples?
Sample Answer
Direct answer
Treat every code block in the docs as a small test: extract it, run it in CI (continuous integration, checks that run on every change), and fail the pull request when it breaks. Run each snippet in a throwaway directory with a time limit and a stripped environment, use fixtures (small fixed sample inputs or fake services used only for testing) instead of real services, fix flakiness (a snippet passing on some runs and failing on others with no change to the code) at its cause instead of retrying blindly, and move genuinely slow examples to a scheduled job.
The approach in layers
| What breaks | How to test it |
|---|---|
| Python snippets in Markdown | Extract fenced blocks and run each one (script below) |
| Python examples that show output | Doctest (examples written as >>> code followed by the expected output): python -m doctest docs/api.md, or pytest --doctest-glob='*.md' |
| README quick start | A smoke test (a fast test of the basic path) that runs the README's commands or its doctests |
| JavaScript snippets | The same extractor with node as the runner |
| SQL examples | Run against a throwaway database with a small fixture; keep the expected result beside the query and compare |
| Notebooks | pytest --nbmake notebooks/ (a plugin for pytest, the Python test runner) runs every cell (set a limit with --nbmake-timeout) |
A runnable extractor (runs python and js blocks, reports failures as GitHub annotations (messages GitHub draws on the exact line of the pull request diff), skips blocks marked <!-- docs-test: skip -->):
import pathlib, re, shutil, subprocess, sys, tempfile
TICKS = "`" * 3
FENCE = re.compile(r"(<!-- docs-test: skip -->\s*\n)?" + TICKS + r"(python|js)\n(.*?)" + TICKS, re.S)
RUNNERS = {"python": "python3", "js": "node"}
TIMEOUT_SECONDS = 2
def check(root):
failed = 0
for md in sorted(pathlib.Path(root).rglob("*.md")):
text = md.read_text()
for m in FENCE.finditer(text):
line = text[:m.start(3)].count("\n") + 1 # first line of the snippet
if m.group(1):
print(f"skip {md}:{line}"); continue
exe = shutil.which(RUNNERS[m.group(2)]) # absolute path, so the stripped PATH below is fine
if not exe:
failed += 1
print(f"::error file={md},line={line}::no {RUNNERS[m.group(2)]} on this runner"); continue
suffix = ".py" if m.group(2) == "python" else ".js"
with tempfile.TemporaryDirectory() as sandbox: # scratch cwd, minimal environment
script = pathlib.Path(sandbox, "snippet" + suffix)
script.write_text(m.group(3))
try:
r = subprocess.run([exe, str(script)], cwd=sandbox, capture_output=True,
text=True, timeout=TIMEOUT_SECONDS, env={"PATH": "/usr/bin:/bin"})
ok, why = r.returncode == 0, (r.stderr.strip().splitlines() or ["exit " + str(r.returncode)])[-1]
except subprocess.TimeoutExpired:
ok, why = False, f"timed out after {TIMEOUT_SECONDS}s"
if ok: print(f"ok {md}:{line}")
else:
failed += 1
print(f"::error file={md},line={line}::snippet failed: {why}") # GitHub Actions annotation
return failed
if __name__ == "__main__":
sys.exit(1 if check(sys.argv[1]) else 0)
How the regex reads. FENCE has three parts. Group 1 is an optional skip comment followed by a newline. Then comes the opening fence: three backticks and the language, python or js (group 2), and a newline. Then group 3, (.*?), is the snippet body, taken lazily (as little as possible) up to the next three backticks; re.S lets . match newlines so the body can span many lines. So m.group(1) is the skip marker if present, m.group(2) the language, m.group(3) the code. The line number is the count of newlines before m.start(3) (where the body begins) plus one.
What the demo does and does not do. It implements the skip marker, the time limit, the stripped environment, the temporary working directory and the annotations. The policies described below (require a reason on every skip, a slow tag, one retry for a network snippet) are team rules you layer on top; the demo accepts a bare skip marker and does not enforce any of them.
Test it on this fixture, saved as docs/guide.md (it has one good snippet, one broken, one skipped, one JavaScript, one slow):
# Guide
```python
print(sum([1, 2, 3]))
```
```python
print(undefined_name)
```
<!-- docs-test: skip -->
```python
requires_a_real_api_key()
```
```js
console.log([1, 2, 3].map((x) => x * 2).join(","));
```
```python
import time; time.sleep(5)
```
Then run python3 check_docs.py docs. Output:
ok docs/guide.md:4
::error file=docs/guide.md,line=8::snippet failed: NameError: name 'undefined_name' is not defined
skip docs/guide.md:13
ok docs/guide.md:17
::error file=docs/guide.md,line=21::snippet failed: timed out after 2s
The process exits with status 1, so the job fails. The ::error file=...,line=...:: format makes GitHub show the message on the exact line in the pull request, which is how failing docs get in front of the author.
Sandboxing
- Each snippet runs in a fresh temporary directory with a minimal environment, so no secrets or files from the runner are visible.
- Every snippet has a time limit (2 seconds in the demo; use something realistic like 30 seconds in real docs).
- In CI, deny network access for the job (for example by running it in a container without a network), and give examples fixtures or a local fake server instead of real services.
- Pin the interpreter and dependency versions, or "docs broke" will just mean "the world moved".
Flaky examples: fix the cause
Flakiness in docs snippets almost always comes from the network, the clock, randomness or output ordering. Answer each directly: recorded responses or a fake server, a fixed date, a pinned random seed, sorted output or doctest's ELLIPSIS flag (an option that lets ... in the expected output match any text, so a timestamp does not break the test) for unstable parts. Do not retry blindly: a retry hides real breakage. If retries are unavoidable for one snippet you have tagged as network-dependent, allow one, and report that it needed one.
Slow examples
Put a hard timeout on each snippet, tag heavy ones slow (a label, for example a pytest marker, that lets CI select or skip them), and run them in a nightly job instead of on every pull request. The better fix is a smaller dataset in the example.
Quarantine with an expiry
Quarantine means temporarily excluding a snippet from the check, with an owner and an end date. A snippet that cannot run yet (needs a real API key, or is deliberately partial) gets the skip marker plus a reason and a ticket. Fail CI when a skip has no reason (a policy you would add: require text after the marker, which the demo regex does not yet do), and review skips periodically so "skipped" does not become "abandoned".
Trade-offs and pitfalls
- Docs break when the code changes, not only when the docs change, so run this check on every pull request that touches code, not only on docs-only ones.
- Running examples proves they execute, not that they teach well. Keep reviewing wording.
- Snippets that print a timestamp or random value make bad doctests. Rewrite the example so it is stable.
Write a one-page quick-start for engineers adopting a new internal shared component, for instance a feature store. What must a first-time user be able to do in ten minutes, and which pitfalls belong on the page?
Sample Answer
Direct answer
The quick-start's job is one measurable outcome: a first-time user gets real value from the component in ten minutes, without asking anyone. For a feature store (a shared system that keeps precomputed model input values, called features, so training and live serving use the same definitions), that means install, authenticate, read one existing feature, and see it verified. An entity is the thing the features describe (here a customer, identified by an entity id such as customer_id = 1042). Online features are the latest values served quickly for live predictions, as opposed to offline features, which are historical values used to build training data. A namespace is a named folder that keeps sandbox features apart from production ones, and the catalogue is the searchable list of registered features and their owners. The pitfalls on the page are the ones that bite first-time users, not a full reference.
What a first-time user must do in ten minutes
| Minutes | Step | Success check |
|---|---|---|
| 0-2 | Install the client with a pinned version and set the project name | featurestore --version prints a version |
| 2-4 | Authenticate with your own team credentials (never a shared production key) | A test call returns "authenticated as <you>" |
| 4-6 | Read features for one known entity id (for example customer_id = 1042) | You see the feature values and a timestamp |
| 6-9 | Define one new feature in a sandbox namespace and fetch it | The new value appears in the read |
| 9-10 | Find your feature in the catalogue or UI and see who owns it | You can name its owner and freshness |
The steps use an illustrative client, since this is an internal component:
from featurestore import Client # illustrative internal client
client = Client(project="sandbox")
row = client.get_online_features(
features=["customer.orders_30d", "customer.avg_basket"],
entity={"customer_id": 1042},
)
print(row) # expect two values plus the time each was computed
The page states the expected result after each step, so success is visible without guessing. In the code: Client(project="sandbox") connects to the sandbox namespace, get_online_features asks for two named features for one customer, and the print should show one value per feature plus the time each was computed.
Pitfalls that belong on the page
Pitfalls 1 and 2 are the ones that fail silently and cost the most; 3 to 6 are housekeeping the page can state in one line each.
- Point-in-time correctness. When building training data, fetch each feature value as it was at the label's timestamp. Using today's value leaks the future into training, so the model looks better offline than it will in production. Example (illustrative): customer 1042 churned (stopped buying) on 2025-03-15, and the label timestamp is that date. Their
orders_30dwas 5 on 2025-03-15, but by today (2025-06-01) it is 0 because they left. If training uses today's 0, the model learns "zero orders means churn", which looks perfect offline, yet in production the model only ever sees the value 5 before the customer leaves. The correct training value is the one as of 2025-03-15. - Training-serving skew. Reuse the registered definition. Recomputing the feature separately in your own code is how the two sides drift apart.
- Freshness. A feature has a refresh interval and a time-to-live (how long a value stays valid). Say which features are hourly and which daily, so a user does not expect real-time.
- Entity keys. Use the documented key name and type (an id stored as text vs integer silently returns nothing).
- Sandbox vs production. Where to experiment and what needs a review before promotion.
- Cost and quotas. Large backfills (recomputing history) consume shared capacity: ask the owner first.
Worked example
A new engineer reads the page at 10:00. By 10:04 they are authenticated. At 10:06 they see orders_30d = 7 for customer 1042 with a computed-at time. They ask, "is this current?" The page's freshness note says daily at 03:00, so a 10:00 read is 7 hours old, which they can decide is acceptable for their model. They never opened the full reference.
Trade-offs and pitfalls of the page itself
- The page must stay one screen long. Link to the reference, the on-call channel, and the design doc rather than embedding them.
- Test the ten-minute claim by watching someone new do it, then cut every step where they hesitated.
- Put the sharpest pitfall (point-in-time leakage) near the top of the pitfalls list, since it is silent and costly.
- Examples must use the sandbox, so copy-paste never touches production.
Have you worked with a docs-as-code workflow? Explain what it is, how a team would adopt it, and the main benefit and main cost you have seen.
Sample Answer
Direct answer
Docs-as-code means treating documentation the way a team treats source code: plain-text files (usually Markdown) stored in Git next to the code, changed through pull requests (proposed changes that a teammate reviews before they merge), checked by automated tests in CI (continuous integration), and built into a website on every merge. If you have not used it, say so plainly and describe how you would pilot it; the shape of the answer below works either way.
What it looks like
payments-service/
src/retry.py
docs/
index.md
runbook-timeouts.md # a runbook: step-by-step guide an on-call engineer follows when something breaks
mkdocs.yml # site config for a static site generator
A change to the retry behaviour is one pull request containing the code in src/retry.py and the edit to docs/runbook-timeouts.md. The reviewer sees both together, and the merge publishes the site.
How a team adopts it
- Pilot with one repository and one team, not the whole company.
- Pick a static site generator (a tool that turns Markdown files into a website, for example MkDocs or Docusaurus) and publish from the main branch.
- Add cheap automatic checks first: Markdown lint (a tool that flags formatting mistakes such as a skipped heading level), a link checker (a tool that follows every link and reports ones that lead nowhere), and a build that fails on errors. A failed link check prints something like this (illustrative):
docs/index.md:12: broken link 'runbook-timeout.md' (file not found; did you mean runbook-timeouts.md?)
- Change the habit: add a line to the pull request template (the pre-filled text every new pull request starts with), such as "Docs updated, or not needed because ...". A filled-in line reads "Docs updated: docs/runbook-timeouts.md" or "Docs not needed because: internal refactor, no behaviour change". Also name a docs owner (the person responsible for the accuracy of a folder, often recorded in a CODEOWNERS file) for each folder.
- Move only the pages people use, starting with the runbooks and the quick start (the short getting-started page a new user follows first). Leave the archive where it is.
- Lower the barrier for non-Git users: an "edit this page" link on every page, so a product manager or support engineer can fix a typo in the browser.
Main benefit
Docs change in the same pull request as the behaviour they describe, so they are reviewed, versioned and released together. A reader can also look at the docs as they were at any past release. That is what keeps them from going stale silently.
Main cost
Contributors who do not live in Git face friction. Support staff, product managers and analysts either stop editing or file requests that engineers turn into pull requests, which recreates the bottleneck. The mitigations are the in-browser edit link, a light-touch review rule for docs-only changes, and a tool that shows a live preview. A second, smaller cost is upkeep of the pipeline itself: someone owns the build and the link checker, and broken builds must be fixed quickly or people will start ignoring them.
Pitfalls
- Moving everything and fixing nothing: a migrated pile of stale pages is still a pile of stale pages. Delete before you migrate.
- Making the checks noisy: a flaky link check (one that fails on some runs and passes on others for reasons unrelated to your change, such as a third-party site being briefly down) trains people to click past failures. Block on internal links and build errors, and treat external links as advisory.
- No owner: docs in the repository still rot if nobody is responsible for a folder.
Design the pull-request checks for a repository of SRE runbooks: dead links, required front-matter keys and Markdown linting. Which tools would you use, how do failures surface, and how do you let an emergency runbook fix through during an incident without gutting the checks?
Sample Answer
Direct answer
Treat the runbook repository like code: every pull request (PR, a proposed change awaiting review) runs three automated checks in continuous integration (CI, the robot that runs on each PR). A link checker (lychee) catches dead links, a small script validates required front matter (the YAML metadata block at the top of each file, such as owner and last-reviewed date), and markdownlint-cli2 enforces Markdown style. Failures show up as inline annotations on the PR and as a red required check. For emergencies, an on-call engineer adds an emergency-fix label. That label makes only the checks that are cosmetic (lint) or depend on outside websites (external links) non-blocking, keeps the check that protects accountability (front matter) blocking, and automatically opens a follow-up issue so the debt is paid rather than forgotten.
Terms in plain words
- GitHub Actions: GitHub's built-in automation. A YAML file in
.github/workflows/tells GitHub what commands to run on each PR. - Required status check / branch protection: branch protection is a repository setting on the main branch. Marking a check as "required" there means GitHub refuses to merge a PR until that check is green.
- Triage rights: a GitHub permission level. Only people with at least that level (or higher) on the repository can add or remove labels, so the label cannot be applied by a random outsider.
- Inline annotation: an error message drawn on the exact file and line in the PR's Files tab.
continue-on-error: a step setting that lets the step fail without turning the whole job red.
Tools and what each one guards
| Check | Tool | Why this one |
|---|---|---|
| Dead links | lychee (a fast link checker that reads Markdown directly, with retries and caching) | Broken links to dashboards and other runbooks are the most common way a runbook fails someone at 3 a.m. |
| Required front matter | ~30-line Python script with PyYAML (a Python library that reads YAML) | Rules are specific to your team (owner, service, last_reviewed), so a small script beats a generic tool |
| Markdown style | markdownlint-cli2 (a linter: it flags style problems such as skipped heading levels or trailing spaces) | Consistent headings and numbered lists make steps skimmable under stress |
The workflow (GitHub Actions syntax, parsed as valid YAML)
name: runbook-checks
on:
pull_request:
types: [opened, synchronize, reopened, labeled] # 'labeled' re-runs checks when the emergency label is added
# No `paths:` filter on purpose: a required check that is skipped by path filtering stays "Pending" and blocks the PR
schedule:
- cron: "0 6 * * 1" # weekly full link crawl catches link rot nobody's PR touched
permissions: { contents: read, issues: write, pull-requests: read }
jobs:
docs-checks: # the ONE job name set as a required status check
runs-on: ubuntu-latest
env:
EMERGENCY: ${{ contains(github.event.pull_request.labels.*.name, 'emergency-fix') }}
steps:
- uses: actions/checkout@v4
- name: Front matter (always blocking, because no owner means nobody is accountable)
run: pip install pyyaml==6.0.3 && python scripts/check_front_matter.py $(git ls-files 'runbooks/*.md')
- name: Markdown lint
continue-on-error: ${{ env.EMERGENCY == 'true' }}
run: npx --yes markdownlint-cli2@0.23.3 "runbooks/**/*.md"
- name: Links
continue-on-error: ${{ env.EMERGENCY == 'true' }}
uses: lycheeverse/lychee-action@v2
with:
args: --no-progress --max-retries 2 --accept 200,429 "runbooks/**/*.md"
fail: true
- name: Open follow-up issue when the emergency path was used
if: env.EMERGENCY == 'true'
env: { GH_TOKEN: "${{ github.token }}" }
run: gh issue create --title "Post-incident doc cleanup for PR ${{ github.event.pull_request.number }}" --label docs-debt --body "Lint and link checks were non-blocking on this PR. Re-run and fix within 2 business days."
Reading the workflow step by step
- The workflow starts when a PR is opened, updated, reopened or labeled, on every PR (there is deliberately no
paths:filter, see the pitfalls). A weekly schedule also runs it. permissionsgives the job read access to the code and write access to issues, so it can open the follow-up ticket, and nothing more.EMERGENCYistrueonly when the PR carries the labelemergency-fix(thecontains(...)expression asks "is that name in the PR's label list?").- Front matter always runs and always blocks: if it fails, the job is red.
- Lint and links have
continue-on-errorset to that flag: normally a failure turns the job red, but on an emergency PR the failure is still printed and the job stays green. - The last step runs only when the flag is
trueand files the cleanup issue.
The front-matter check, run for real
import os, sys, pathlib, tempfile
import yaml
REQUIRED = ["title", "owner", "service", "last_reviewed"]
def check(path, text):
"""Return a list of problems for one runbook."""
text = text.replace("\r\n", "\n") # files saved on Windows use CRLF line endings
if not text.startswith("---\n"):
return ["no front matter block"]
end = text.find("\n---", 4)
if end == -1:
return ["front matter block is never closed"]
try:
meta = yaml.safe_load(text[4:end]) or {}
except yaml.YAMLError as e:
return [f"front matter is not valid YAML: {e.__class__.__name__}"]
if not isinstance(meta, dict):
return ["front matter is not a set of key: value pairs"]
# a key that is present but empty (`owner:`) is as useless as a missing one
return [f"missing key: {k}" for k in REQUIRED if meta.get(k) in (None, "")]
def main(paths):
failures = 0
for p in paths:
for problem in check(p, pathlib.Path(p).read_text()):
print(f"::error file={p}::{problem}")
failures += 1
return 1 if failures else 0
if __name__ == "__main__":
if len(sys.argv) > 1:
sys.exit(main(sys.argv[1:]))
with tempfile.TemporaryDirectory() as d:
os.chdir(d)
good = pathlib.Path("db-failover.md")
good.write_text("---\ntitle: DB failover\nowner: team-data\nservice: orders-db\nlast_reviewed: 2026-01-05\n---\n# Steps\n")
bad = pathlib.Path("cache-flush.md")
bad.write_text("---\ntitle: Cache flush\nservice: edge-cache\n---\n# Steps\n")
empty = pathlib.Path("dns-cutover.md")
empty.write_text("---\ntitle: DNS cutover\nowner:\nservice: dns\nlast_reviewed: 2026-02-01\n---\n# Steps\n")
crlf = pathlib.Path("vpn-reset.md")
crlf.write_bytes(b"---\r\ntitle: VPN reset\r\nowner: team-net\r\nservice: vpn\r\nlast_reviewed: 2026-03-01\r\n---\r\n# Steps\r\n")
code = main([str(good), str(bad), str(empty), str(crlf)])
print("exit code:", code)
Running it prints:
::error file=cache-flush.md::missing key: owner
::error file=cache-flush.md::missing key: last_reviewed
::error file=dns-cutover.md::missing key: owner
exit code: 1
The third file has owner: with nothing after it, which YAML reads as null. A plain k not in meta test would let it pass, defeating the rule that "no owner means nobody is accountable", so the check tests for an empty value too. The fourth file (Windows line endings) passes because the script normalises \r\n first.
The ::error file=...:: format is what makes GitHub draw the message on the exact file in the PR's Files tab, so the author does not have to open CI logs.
How failures surface
- One job name (
docs-checks) is marked as a required status check in branch protection, so a red result blocks merge. - Each tool writes file-level annotations, so the fix is visible where the author is already looking.
- Link checks accept HTTP 429 (rate limited) and retry twice, because a flaky third-party site should not fail the PR.
- The weekly scheduled run crawls everything, because links rot even when nobody edits the runbook.
Emergency path without gutting the checks
- Only people with triage rights can apply the label. GitHub itself enforces this (the workflow does not need to), so it is not a self-service bypass for anyone.
- Front matter stays blocking: it is local and takes seconds, but the real reason is that an emergency runbook with no owner or review date is the one nobody will fix later, which is worse than a lint warning.
- Lint and external links become non-blocking (
continue-on-error). Lint findings are cosmetic and cannot mislead a reader, and link checks depend on third-party sites that may be down or rate-limiting during the same incident, so a red result there does not mean the runbook is wrong. Lint is not slow; the point is that a style complaint should never delay a fix. - The bypass is not a branch-protection override or a disabled workflow. The checks still run and still report, so the author sees exactly what to clean up.
- The last step files a
docs-debtissue with a two-business-day target, and a weekly report counts how many emergency labels were used. A rising count is a signal to fix the checks or the on-call process.
Worked example
At 02:10 on-call needs to change step 3 of a failover runbook. They open a PR, add emergency-fix, and a teammate approves. Front matter passes, lint fails on a trailing-space rule and is shown but non-blocking, and the merge happens in minutes. The follow-up issue lands in the team queue the same night, and the author fixes the lint error next morning.
Trade-offs and pitfalls
- Making every check non-blocking during emergencies is tempting but teaches people to always use the label. Keep the blocking set as small as the cheapest useful checks.
- Do not run link checks against private internal URLs without an allowlist or auth, or every PR fails. Exclude those patterns explicitly and verify them another way.
- Pin the action and tool versions so a tool upgrade does not turn every PR red on the same morning.
- A required check whose workflow is skipped by a
paths:filter stays "Pending" forever (GitHub documents this), so a PR touching only a README would be unmergeable. Either drop the filter, as above, or add a tiny always-running job that reports the same check name. gh issue create --label docs-debtfails if that label does not exist, which would turn the emergency path red. Create the label once up front, and remember each new push to an emergency PR re-runs the step, so de-duplicate (for example search for an open issue with that title first).- The workflow file itself is part of the PR, so an author could edit it to weaken the checks. Put
.github/workflows/under CODEOWNERS with "Require review from Code Owners" enabled. - If the label is used often, the real problem is that the checks are too slow or too strict for normal work.
Unlock Full Question Bank
Get access to all Technical Writing and Documentation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.