API Documentation and Developer Experience Questions
Documentation and developer experience for API consumers: reference docs for endpoints and fields, OpenAPI-driven and interactive documentation, quickstarts and runnable code samples, doc comments that feed generated reference, changelogs, deprecation notices and migration guides, error messages and error tables that explain themselves, developer portals, sandboxes and onboarding flows. Covers developer experience (DX) as a product concern: time to first successful call, DX and docs-effectiveness metrics and dashboards, diagnosing onboarding drop-off, developer feedback loops, experiments and developer research, documentation standards and CI checks across teams, prioritising DX investment, and reducing support load through better docs and self-service. Designing the API itself (versioning policy, auth, rate limits, webhooks, gateways, SDK engineering) is a separate subject.
Write the code samples for a quickstart that make an authenticated GET to a users endpoint in three languages, parse the list response, and handle an authorization failure. What makes a sample good for a copy-paste reader, and how would you stop these samples going stale?
Sample Answer
Direct answer
A good quickstart sample is a complete, minimal program that a stranger can paste, set one environment variable for, and run unchanged. It makes the authenticated request, reads the list out of the response, and handles the one failure every newcomer hits first (a rejected key) with a message that says what to do. To stop samples going stale, they must live as real files that CI (continuous integration, the automated build run on every change) executes against a sandbox, and the docs page must include them from those files instead of holding pasted copies. A contract test compares real API responses to the OpenAPI description.
Approach
Core to the answer: the three samples and the reasons they are good. Optional: the local stand-in server, which only proves the samples run.
The endpoint returns a JSON body shaped like {"data": [ {"id": ..., "email": ...} ], "next_cursor": null}. OpenAPI is the standard machine-readable description of an API. A sandbox is a safe test copy of the API with fake data. The API key travels as a bearer token (a credential sent in the Authorization header). Status 401 means the credentials are missing or invalid, 403 means the key is valid but not allowed to do this; both get the same friendly handling in a first sample. The base URL is read from an environment variable with a production default, so the identical code can be run against a test server.
Sample 1: curl (shell, needs jq, a command-line JSON reader)
#!/usr/bin/env bash
# Needs: curl, jq. Set API_KEY first. BASE_URL defaults to the production API.
# ${BASE_URL:-default} means: use BASE_URL if set, otherwise the default after ':-'.
BASE_URL="${BASE_URL:-https://api.example.com/v1}"
status=$(curl -sS -o response.json -w '%{http_code}' \
-H "Authorization: Bearer $API_KEY" \
"$BASE_URL/users?limit=2")
if [ "$status" = "401" ] || [ "$status" = "403" ]; then
echo "Authorization failed (HTTP $status). Check that API_KEY is set and current." >&2
jq -r '.error.message' response.json >&2
exit 1
fi
jq -r '.data[] | "\(.id) \(.email)"' response.json
What the curl flags do: -sS hides the progress meter but still prints errors; -o response.json saves the body to a file; -w '%{http_code}' prints just the status code, which the script captures in status; -H adds a header.
Sample 2: Python (with the requests library, saved as quick.py)
import os
import sys
import requests
BASE_URL = os.environ.get("BASE_URL", "https://api.example.com/v1")
API_KEY = os.environ.get("API_KEY", "")
resp = requests.get(
f"{BASE_URL}/users",
params={"limit": 2},
headers={"Authorization": f"Bearer {API_KEY}"},
timeout=10, # seconds to wait before giving up
)
if resp.status_code in (401, 403):
print(f"Authorization failed (HTTP {resp.status_code}). Check that API_KEY is set and current.",
file=sys.stderr)
print(resp.json()["error"]["message"], file=sys.stderr)
sys.exit(1)
resp.raise_for_status()
for user in resp.json()["data"]:
print(user["id"], user["email"])
Sample 3: JavaScript (Node 18 or later, built-in fetch, saved as quick.mjs)
const BASE_URL = process.env.BASE_URL ?? "https://api.example.com/v1";
const API_KEY = process.env.API_KEY ?? "";
const resp = await fetch(`${BASE_URL}/users?limit=2`, {
headers: { Authorization: `Bearer ${API_KEY}` },
signal: AbortSignal.timeout(10_000),
});
if (resp.status === 401 || resp.status === 403) {
const body = await resp.json();
console.error(`Authorization failed (HTTP ${resp.status}). Check that API_KEY is set and current.`);
console.error(body.error.message);
process.exit(1);
}
if (!resp.ok) throw new Error(`Unexpected HTTP ${resp.status}`);
const { data } = await resp.json();
for (const user of data) console.log(user.id, user.email);
Proving they run: a local stand-in server (save as mock_api.py)
from http.server import BaseHTTPRequestHandler, HTTPServer
import json
USERS = [{"id": "usr_1", "email": "ada@example.com"},
{"id": "usr_2", "email": "grace@example.com"}]
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
ok = self.headers.get("Authorization") == "Bearer test_key_123"
code, body = (200, {"data": USERS, "next_cursor": None}) if ok else (
401, {"error": {"code": "invalid_api_key", "message": "API key is missing or invalid."}})
payload = json.dumps(body).encode()
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def log_message(self, *args):
pass
HTTPServer(("127.0.0.1", 8765), Handler).serve_forever()
Run it and the Python sample, once with a good key and once with a bad one:
python3 mock_api.py &
export BASE_URL=http://127.0.0.1:8765/v1
API_KEY=test_key_123 python3 quick.py
API_KEY=wrong python3 quick.py; echo "exit code: $?"
Output:
usr_1 ada@example.com
usr_2 grace@example.com
Authorization failed (HTTP 401). Check that API_KEY is set and current.
API key is missing or invalid.
exit code: 1
The curl and Node versions print the same lines against the same server.
What makes a sample good for a copy-paste reader
- Runs unchanged: the only setup is
API_KEY. No placeholders inside the code that must be edited. - Key from the environment, never inline: a pasted key ends up in screenshots, shell history and git.
- Shows the response shape being used: the loop reads
data, so the reader learns the list lives under that field. - Fails helpfully: prints the cause and a next step, exits non-zero (a non-zero exit code is how a program tells scripts it failed) so scripts notice; error messages go to stderr (the error output stream, separate from normal output).
- Has a timeout: without one, a stuck network hangs the newcomer's terminal with no explanation.
- Same behaviour in all three languages, so switching language never changes the story.
Keeping samples from going stale
- Store each sample as a file in the repository and include it into the docs page at build time (docs-as-code: docs built from source files by the same build process as software).
- CI runs every sample against the sandbox on each pull request (a proposed code change awaiting review) and nightly (once a night on a schedule). A non-zero exit fails the build, which is why the samples exit properly.
- A contract test compares real API responses to the OpenAPI description, so a renamed field breaks the build before it breaks a reader.
- Pin the language versions the docs claim (for example "Node 18 or later") and test the oldest one.
- Name an owner: the team that changes the endpoint fixes the samples in the same pull request.
Complexity and edge cases
Each sample makes one request and one pass over the returned page: O(n) time and memory (work and storage grow in proportion to n, the number of users in the page), with n capped by limit. Edge cases a quickstart should acknowledge, even if it does not solve them all: an empty data list (loop prints nothing, so add a line saying so in the guide), next_cursor not null (a later guide shows how to page on), a 429 rate-limit response (wait and retry), a non-JSON error body from a proxy (an intermediate server returning its own plain-text error page instead of the API's JSON; the Python and Node samples would raise here, acceptable in a quickstart but named in the troubleshooting page), and a missing API_KEY (sent as an empty token, which produces the same 401 message).
Trade-offs
Three languages triple the maintenance, which is exactly why the CI run matters. Generated samples from the OpenAPI description are cheap but read as generic; hand-written quickstarts teach better, generated ones cover the long tail of endpoints.
You are asked to write the doc comments for a new SDK method so they serve both engineers reading the code and the automated reference docs generated from them. What practices do you apply? Show a short example for a tax-calculation method.
Sample Answer
Direct answer
Write the comment for two readers at once: the engineer who sees it in the editor tooltip, and the documentation generator that turns it into a reference page. Use the language's standard doc-comment format (a docstring in Python, Javadoc in Java, TSDoc in TypeScript: comment conventions that documentation tools parse), lead with a one-line summary, document every parameter with units and limits, and include an example that is tested so it cannot go stale. Only the Python form is needed for the example below; the Java and TypeScript names are just the same idea in other languages, so do not memorise the tools. (An SDK, software development kit, is the language-specific library that wraps an API.) This is "docs-as-code": the documentation lives next to the code and is reviewed and built with it.
Practices to apply
- Summary line first, in the imperative (a command form: "Calculate the sales tax...", not "Calculates" or "This function calculates"). Generators use it as the list description.
- Say what, and the contract: units, allowed ranges, rounding, side effects (here: no network call, no mutation).
- Document each parameter and return value, including the type and what the number means (currency units, not just "a number").
- List the errors raised, and under which conditions.
- Include a runnable example. In Python, examples written as
>>>lines can be executed by the standarddoctestmodule, so the docs fail the build when the behaviour changes. - Do not repeat the function name or narrate the implementation; comments explain what a caller needs.
- Mention versioning notes when behaviour changes.
Example: a tax-calculation method (Python)
from decimal import Decimal, ROUND_HALF_UP
RATES = {"US-CA": Decimal("0.0725"), "US-OR": Decimal("0")}
def calculate_tax(amount, jurisdiction, *, tax_exempt=False):
"""Calculate the sales tax owed on an order subtotal.
Looks up the rate for ``jurisdiction`` and rounds the result to whole
cents (half up). Makes no network call and never changes the order.
Args:
amount (Decimal): Taxable subtotal in the order currency, for example
``Decimal("100.00")``. Must not be negative. Shipping is only
taxed if the caller includes it in this amount.
jurisdiction (str): Region code such as ``"US-CA"``.
tax_exempt (bool): If True, returns zero without a rate lookup.
Defaults to False.
Returns:
Decimal: The tax amount, rounded to 2 decimal places.
Raises:
ValueError: If ``amount`` is negative or ``jurisdiction`` has no rate.
Example:
>>> calculate_tax(Decimal("100.00"), "US-CA")
Decimal('7.25')
>>> calculate_tax(Decimal("19.99"), "US-CA")
Decimal('1.45')
"""
if amount < 0:
raise ValueError("amount must not be negative")
if tax_exempt:
return Decimal("0.00")
if jurisdiction not in RATES:
raise ValueError(f"no tax rate for jurisdiction {jurisdiction!r}")
return (amount * RATES[jurisdiction]).quantize(Decimal("0.01"), ROUND_HALF_UP)
if __name__ == "__main__":
import doctest
print(doctest.testmod())
Running this file prints TestResults(failed=0, attempted=2): both examples in the docstring were executed and matched.
Two code details in plain words: the bare * in the signature makes everything after it keyword-only, so callers must write tax_exempt=True and cannot pass a bare True by position (which would be unreadable at the call site). quantize(Decimal("0.01"), ROUND_HALF_UP) rounds to two decimal places, and ROUND_HALF_UP means an exact half rounds up (1.005 becomes 1.01), the rule most tax authorities expect.
Why this serves both readers
The docstring's Args/Returns/Raises sections are parsed by a generator such as Sphinx (a Python documentation generator, with its autodoc extension that reads docstrings; mkdocstrings is a similar alternative) into a reference page, while the same text appears in the editor when an engineer hovers over the function. Worked check by hand: 19.99 x 0.0725 = 1.449275, which rounds half up to 1.45, matching the second example.
Pitfalls
- Example output copied by hand rather than executed goes stale silently.
- Money as floating point causes rounding surprises; the example uses
Decimal, and the docs say so. - Undocumented rounding rules are the most common source of "your tax is one cent off" support tickets.
Write a one-paragraph summary of a POST endpoint that creates a customer order, pitched at a product manager who will never call it. What do you include and what do you deliberately leave out?
Sample Answer
Direct answer
Write it in terms of business outcome, rules and consequences, not HTTP mechanics. A product manager needs to know what the endpoint lets the business do, what must be true for it to work, what happens after it succeeds, and what can go wrong for customers.
The paragraph
"Create Order is the call our storefront and partner systems make to place a customer's order. It needs a customer, at least one item that is in stock, and a delivery address; if any of these is missing, no order is created and the caller is told which one. When it succeeds, the order is saved in a 'pending' state, stock for those items is reserved, and the order appears in the customer's history and in the daily sales report. Payment is collected separately, so a pending order can still be abandoned if payment fails. It is safe for a caller to retry after a timeout without creating a duplicate order. Orders above our fraud-review threshold are held for manual approval, which can delay confirmation."
What I included and why
- Purpose in business words: who calls it and what it enables.
- Preconditions that change what customers experience (needs stock, needs address).
- What happens next: reserved stock, pending status, reporting. A PM plans features on these downstream effects.
- Consequences of failure: abandoned pending orders, held orders.
- Anything that changes an outcome: retry safety means no duplicate orders, so a customer is never asked to pay twice for one purchase by a retry (payment itself is collected by a separate call, which needs its own retry protection), which is a customer-facing promise.
What I left out on purpose
- Header names, authentication token format, field types and JSON structure.
- Exact status codes, the database tables touched, internal service names.
- Pagination, versioning details and other things the PM will never act on.
If the same summary were for an engineer, those omitted items become the core; the facts stay the same, the audience changes which ones matter.
Pitfalls
Copying the reference description ("POST /orders creates an order resource") gives the PM nothing to decide with. Also avoid promising business behaviour the code does not implement: check each claim (stock reservation, fraud threshold) with the owning team before it goes in the paragraph.
External developers need a sandbox to build against before they have production access. What should the sandbox give them, how do you keep it faithful to production, and how do you document where it differs?
Sample Answer
Direct answer
A sandbox is a safe copy of the API where outside developers can make real calls with fake money and fake data, before the provider trusts them with production. It should give them self-serve access, realistic and resettable data, a way to trigger every error and edge case on demand, and the same request and response shapes as production. You keep it faithful by making production and sandbox share the same contract and the same tests, and you document the remaining differences in one table that developers see before they hit them.
What the sandbox should give developers
- Instant, self-serve credentials. A sandbox key created in the portal in a minute, with no sales call. If getting a key takes days, developers stall before writing any code.
- A separate base URL (the address every request starts with) and separate keys (for example
api.sandbox.example.com). A test key must never work against production, and a production key must never work in the sandbox, so nobody moves real money by mistake. - Seeded, realistic data. A few sample customers, orders and accounts already present, so list and search calls return something meaningful on the first try.
- Triggers for every outcome. Developers must be able to test failure without waiting for one to happen. The usual technique is "magic values": a documented input that forces a result, such as amount
4003always returns a card-declined error (the number itself is arbitrary, chosen by the provider and listed in a "test values" table in the docs so nobody has to guess it), or a headerX-Sandbox-Scenario: rate_limitedforces a 429 (Too Many Requests). - Working asynchronous behaviour. If production sends webhooks (HTTP callbacks the provider makes to the developer's server), the sandbox must send them too, with a "resend this event" button. Otherwise the most bug-prone part of an integration goes untested.
- Reset and inspect tools. A "wipe my sandbox data" button and a log of recent requests, so a developer can tell whether the bug is theirs.
Keeping it faithful to production
- One contract, two environments. Both are described by the same OpenAPI file (OpenAPI is the standard machine-readable format for describing an HTTP API). Publish the spec once, not a sandbox copy that can drift.
- Run the same contract tests against both. A contract test is an automated check that a service still returns the responses its published description promises. After every deploy, an automated suite sends the same requests to sandbox and production-like builds and checks status codes and response fields. A failure blocks the release.
- Same code paths where it is safe. Validation, authentication, pagination and error formats should run the real code. Only the parts with real-world side effects (moving money, sending email, calling a bank) are replaced with stubs (fake stand-ins that return canned results). Pagination means returning long lists in pages.
- Keep the release cadence aligned. The sandbox is never behind production: a new API version reaches the sandbox first, or on the same day, so partners can test ahead of the change. Two environments on different versions is the commonest source of silent drift.
- Watch for drift. Track "worked in sandbox, failed in production" support tickets. Each one is a fidelity bug and should produce either a fix or a new row in the differences table.
Documenting where it differs
Put one page called "Sandbox vs production" in the docs, link it from the quickstart (the short guide to a first working call), and repeat the relevant row next to the affected endpoint. Illustrative content:
| Area | Sandbox | Production | What the developer should do |
|---|---|---|---|
| Rate limits | 100 requests/minute | Higher, per plan | Do not load-test in the sandbox |
| Webhooks | Delivered within seconds | Delivered with retries over hours | Handle duplicates and out-of-order events |
| Settlement (money actually moving between banks) | Instant | Takes days | Never assume instant in your UI |
| Data | Wiped on request, may be reset on upgrade | Permanent | Do not store sandbox IDs |
Also add a header on every sandbox response (for example X-Environment: sandbox) so logs make the environment obvious.
Worked example
A payments partner tests refunds. In sandbox, POST /refunds for a payment of amount 1000 (10.00 in the smallest currency unit, cents) returns 201 at once with "status": "refunded", and the webhook arrives in 2 seconds, so their code marks the order refunded immediately. In production the refund is pending for days and the webhook comes later. To let them rehearse that, the sandbox scenario is requested like this: POST /refunds with header X-Sandbox-Scenario: refund_pending, which returns 201 with "status": "pending" and sends the refund.completed webhook only when the developer clicks "advance time" in the portal. A state transition is a status change like pending to refunded. The differences table row "Settlement: instant vs takes days" and a sandbox scenario that returns status: "pending" on request would have caught this before launch. Faking the delay is cheap and prevents the most common go-live surprise.
Trade-offs and pitfalls
- Perfect fidelity is not the goal. A sandbox that reproduces every production rule is expensive and slow. Pick fidelity for shapes, validation, errors and state transitions, and accept documented differences for speed, scale and external systems.
- An undocumented difference is worse than a large one. Developers forgive "the sandbox is instant" if the docs say so.
- Stale data and shared state cause flaky partner tests, so isolate each account's data.
- Do not use the sandbox to hold production-like personal data. Synthetic data only.
How would you measure time to first successful call for a public API, for both SDK and direct HTTP users? Define what counts as success, what you would instrument from first docs view onward, and how you would report it without an average hiding the slow tail.
Sample Answer
Direct answer
An SDK (software development kit) is a client library in a language like Python that wraps the raw HTTP calls; "direct HTTP" means calling the API yourself with curl or an HTTP library. Time to first successful call (TTFC, also called time to first "hello world") is the elapsed time from a developer's first docs view to their first successful, authenticated request to a real endpoint. Define success precisely, capture events that survive across sessions and devices, compute it per developer, and report a distribution (percentiles and completion rates, split by SDK versus direct HTTP), never a single average.
1. What counts as success
- Success = the first HTTP 2xx response from a non-trivial endpoint, made with a key that developer created. A health check or ping does not count, or the number flatters you.
- For a model API, use the stricter time to first meaningful result: the first response that contains a real prediction.
- Start of the clock = first docs view (as asked), stored as one anonymous identifier. Optionally also report signup-to-first-call, which isolates the product's own share.
- Developers who never succeed are censored (their true time is unknown because they have not finished yet). Count them in the denominator, never drop them.
2. What to instrument
| Event | Source | Note |
|---|---|---|
docs_viewed (first) | Docs site | Sets an anonymous ID |
signup_completed | Auth service | Links anonymous ID to account ID |
key_created | Console | Links account to key ID |
sdk_installed | Optional install ping | Package download counts alone cannot be tied to a person |
first_request (any status) | API gateway logs (the server-side record of every request) | Shows failures before the win |
first_2xx | API gateway logs | Server-side, so it cannot be blocked by an ad blocker |
Two design points. Cross-session identity: the anonymous ID is aliased (linked, so the two IDs count as one person) to the account ID at signup, then to the key ID, so a developer who returns three days later is one journey. SDK versus HTTP: the SDK sends an identifying User-Agent (a header every HTTP client sends naming the software making the call) with name and version; a raw client shows as curl or a language HTTP library. Segment on that at the gateway. Server-side events plus consent-aware client events keep it privacy-safe.
3. Report the distribution, not the mean
Worked example. Times in minutes from first docs view to first 2xx for 16 developers (illustrative, pinned so you can reproduce it):
import math
# minutes from first docs view to first HTTP 2xx; None = never succeeded in 7 days
sdk = [3, 4, 4, 5, 6, 7, 9, 12]
http = [6, 9, 14, 22, 35, 95, 1500, None]
def pct(vals, p):
vals = sorted(vals)
k = math.ceil(p / 100 * len(vals)) # nearest-rank percentile
return vals[k - 1]
def report(name, xs):
done = [x for x in xs if x is not None]
print(f"{name}: n={len(xs)} succeeded={len(done)} "
f"within10m={sum(x <= 10 for x in done)/len(xs):.0%} "
f"within60m={sum(x <= 60 for x in done)/len(xs):.0%} "
f"median={pct(done,50)} p90={pct(done,90)} mean={sum(done)/len(done):.1f}")
report("SDK ", sdk)
report("HTTP", http)
report("ALL ", sdk + http)
SDK : n=8 succeeded=8 within10m=88% within60m=100% median=5 p90=12 mean=6.2
HTTP: n=8 succeeded=7 within10m=25% within60m=62% median=22 p90=1500 mean=240.1
ALL : n=16 succeeded=15 within10m=56% within60m=81% median=9 p90=95 mean=115.4
How to read it: pct sorts the times and returns the value at the nearest rank (with 7 values, p90 is the ceil(0.9 x 7) = 7th, the largest); report counts everyone in the denominator (so the never-succeeded developer lowers the within-10m and within-60m rates) and prints each statistic. A percentile is the value below which that share of developers fall.
The HTTP mean (240 minutes) is dragged up by one 1500-minute outlier and says nothing useful; the median (22) plus p90 (the time by which 90% of successful developers finished) and the within-10-minute rate show the real picture: SDK users are fast, direct HTTP users are the problem. Note the HTTP p90 is 1500, the same outlier: with only 7 finishers, nearest-rank p90 is just the maximum, so p90 needs a real sample (tens or hundreds of developers) to be informative; here the median (22) and the completion rates carry the message. Report by weekly signup cohort (developers grouped by the week they signed up) so changes are visible.
4. Gaming and guards
- Health-check or test-account calls (including your own employees): exclude internal accounts and require a non-trivial endpoint.
- Multiple keys or accounts resetting the clock: anchor to the earliest docs view for the person, not the key.
- Bots and scrapers: require a completed signup.
- Shipping a "hello" endpoint to game the number: pair the metric with a 7-day activation rate (share of signups making a call) and with later retention (share still calling weeks later).
This is the first funnel stage; the funnel also gives install rate and drop-off per step, and the metric pairs with end-to-end integration tracking (docs view to production traffic) using the same identity chain.
Unlock Full Question Bank
Get access to all 36 API Documentation and Developer Experience interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.