API Documentation and Developer Experience Questions
Documentation and developer experience for API consumers: reference docs for endpoints and fields, OpenAPI-driven and interactive documentation, quickstarts and runnable code samples, doc comments that feed generated reference, changelogs, deprecation notices and migration guides, error messages and error tables that explain themselves, developer portals, sandboxes and onboarding flows. Covers developer experience (DX) as a product concern: time to first successful call, DX and docs-effectiveness metrics and dashboards, diagnosing onboarding drop-off, developer feedback loops, experiments and developer research, documentation standards and CI checks across teams, prioritising DX investment, and reducing support load through better docs and self-service. Designing the API itself (versioning policy, auth, rate limits, webhooks, gateways, SDK engineering) is a separate subject.
Multiple teams publish API docs and the quality is uneven. Propose a documentation standard and the automated checks you would run in CI across repos, and explain how you would get engineers to accept the review gate.
Sample Answer
Direct answer
I would publish one short documentation standard (a minimum contract for every endpoint plus a style guide), turn it into automated checks that run in CI (continuous integration, the automated checks on every pull request) from one shared ruleset, roll it out in stages so it never blocks people on day one, and win acceptance by making the gate cheap, fair and tied to real pain. A gate that engineers experience as bureaucracy gets bypassed; a gate that saves them support tickets gets defended.
The standard (kept to one page)
- Every operation has a summary, description, documented parameters with types, at least one request example, and all error responses.
- Shared conventions: one error format, one pagination style, naming rules, auth described once.
- Docs live with the code (docs-as-code: plain text files in the repository, reviewed like code) and each service names an owner.
- Changes to public behaviour require a changelog entry.
Automated checks (same job in every repo)
Terms: OpenAPI is the standard machine-readable file that describes an API (its paths, parameters and responses), so a tool can check it; a spec is that file; an operationId is the unique name each operation carries, which SDK generators and links rely on. Start with the first two rows (spec lint, then breaking-change diff): they are cheap, need no running service, and catch most gaps. Contract tests, snippet tests, Vale and preview builds are worth adding later, once the first two are trusted.
| Check | What it catches | Example tool |
|---|---|---|
| Spec lint | Missing descriptions, operationId, examples, error responses | Spectral with a shared ruleset |
| Breaking-change diff | Removed field, renamed operation | oasdiff against main |
| Contract test (runs the real service and checks its responses match the spec) | Running service deviates from spec | Schemathesis on a test deployment |
| Prose and link check | Style, dead links | Vale, a link checker |
| Snippet test | Code samples that no longer run | Run samples against a sandbox |
| Preview build | Docs render, nav intact | Docs site build per pull request |
extends: ["spectral:oas"]
rules:
operation-description: error
operation-operationId: error
operation-summary-required:
description: Every operation needs a summary
severity: error
given: "$.paths[*][get,post,put,patch,delete]"
then:
field: summary
function: truthy
What the ruleset says, line by line: extends: ["spectral:oas"] starts from Spectral's built-in OpenAPI rules. operation-description: error and operation-operationId: error turn two of those built-in rules into failures. The custom rule operation-summary-required says: given selects every GET, POST, PUT, PATCH or DELETE operation under paths; then says its summary field must exist and be non-empty (function: truthy); severity: error makes a miss fail the build (versus warn, which only prints). To write another rule, change given (where to look) and then (what must hold).
Publish the ruleset in one central repository and version it, so teams upgrade deliberately.
Getting engineers to accept the gate
- Start warn-only (checks report problems but do not fail the build) for several weeks and publish a scoreboard (a simple table of findings per team) so nobody is surprised.
- Ratchet, do not cliff. (A ratchet only lets the bar go up, never back down; a cliff is switching every rule on at once.) Enforce as errors only on new and changed endpoints; existing debt sits on a baseline list (a file of known existing violations that are tolerated for now) with expiry dates, and shrinks over time.
- Make it easy: a template spec, editor and pre-commit checks (checks that run on the author's machine before a commit is recorded) so failures appear before the pull request, and messages that say how to fix.
- Keep it fast and quiet. A slow or noisy check gets ignored. Track the false-positive rate (how often a warning is about something that is actually fine) and fix or remove rules that cry wolf.
- Give a documented escape hatch (a sanctioned way to override a rule) (an override needing a reason and a second reviewer) and count overrides; the count is a health signal.
- Tie it to the pain. Show support tickets and integration delays traced to undocumented behaviour. Recruit one champion (a volunteer who advocates the standard) per team who helps others fix findings.
- Own the standard jointly with engineers: rule changes proposed through pull requests to the central repo.
Worked example. A team adds DELETE /invoices/{id} with no summary, description or operationId. Run against the ruleset above (executed with Spectral), the change produces three findings: operation-description, operation-operationId and operation-summary-required, plus a built-in warning that the operation has no tags. In week one (warn-only) the pull request shows them as warnings with links to fix guides. After enforcement the same change fails on those three, the author adds a summary, a description and an operationId, and it merges the same day. Note that this ruleset does not check for error responses such as a 404: the standard requires them, so add a custom rule for that or leave it to review. The old undocumented endpoints stay on the baseline until their owning team's quarterly cleanup.
Trade-offs
- Strict rules improve consistency but increase friction: start with a few high-value rules rather than fifty.
- Checking presence is easy; checking accuracy is not, so keep contract tests and human review for correctness.
- Central ownership scales the standard but can feel imposed unless teams help shape it.
You want a working feedback loop between developers using your API and the team that owns the docs and product. What signals would you collect, how do you turn them into a prioritised backlog, and how do you show developers that feedback was acted on?
Sample Answer
Direct answer
(Triage means sorting incoming items by type and urgency and deciding who handles each.) Collect feedback from three kinds of sources (asked-for, unprompted, and behavioural), route it into one tagged backlog owned by a named person, prioritise by how many developers it affects and how badly it blocks them, and close the loop publicly with a changelog and direct replies so developers see that reporting is worth their time.
Signals to collect
| Kind | Sources | What it tells you |
|---|---|---|
| Asked-for | "Was this page helpful?" thumbs on each docs page (with an optional comment), a short survey after first success, a periodic satisfaction survey (for example CSAT, customer satisfaction score) | What developers say is unclear |
| Unprompted | Support tickets, community forum and chat threads, GitHub issues on the SDKs (language-specific client libraries), sales and solutions engineers' call notes | What developers ask when they are stuck |
| Behavioural | Docs search terms with no results, pages with high exit rates (many readers leave from that page instead of continuing), error codes by frequency, drop-off in the onboarding funnel (the sequence signup, first key, first successful call, and how many developers are lost at each step) | What developers do, not what they say |
Behavioural data covers the silent majority: most stuck developers never write in.
Turning it into a prioritised backlog
- One intake, one owner: every signal becomes an item in a single tracker, tagged by area (auth, errors, pagination, SDK) and type (docs gap, bug, feature). A named owner triages weekly.
- Merge duplicates and count: each item carries the number of distinct developers affected and the source.
- Score: rank by developers affected times severity (blocked entirely, worked around, cosmetic) divided by effort. Use simple numeric weights: severity 3 = blocked entirely, 2 = worked around, 1 = cosmetic; effort 1 = under a day, 2 = a few days, 3 = weeks. Boost items that hit the first-call path, and give an enterprise-account request a multiplier (for example x2, meaning its revenue counts double) agreed with product, not decided by whoever shouts.
- Route: docs fixes go to the docs owner, bugs to engineering, feature requests into the product roadmap process.
Worked example (illustrative)
Docs search logs show "webhook signature" (the check that proves a webhook really came from us) returned no result 45 times in a month, three forum threads ask the same thing, and support has 9 tickets on it. De-duplicating by developer gives about 40 distinct developers. Blocking a common flow, so severity 3; one page to write, so effort 1.
webhook signature page: 40 x 3 / 1 = 120
cosmetic request from one enterprise customer: 5 developers x 1 / 2 = 2.5, x2 enterprise multiplier = 5
The docs page outranks the loud request by a wide margin. The enterprise item would only overtake it if its multiplier was enormous, which is a business decision to make openly.
Showing developers that feedback was acted on
- A public changelog entry that says "you asked, we did", linking the original issue where that is allowed.
- Reply directly to whoever reported it when the item ships, even a one-line message.
- A visible public roadmap or "top requested" list with status (planned, in progress, shipped, not planned), including an honest reason for "not planned".
- A quarterly "what we fixed from your feedback" note.
Pitfalls
- Prioritising by loudest voice: the single vocal developer, or the biggest account, always wins without a count.
- Collecting feedback and never replying, which teaches developers not to bother.
- Treating only ticket text as feedback and missing the search logs, where the silent majority speaks.
- Unowned backlog: without a weekly triage owner the list grows and stales within months.
Developer retention on your public API is flat. How would you set up an ongoing programme of experiments and developer research on onboarding and docs, and how do you decide what to test next?
Sample Answer
Direct answer
Set up a repeating loop: diagnose where developers leave, gather qualitative evidence about why, turn that into a ranked backlog of hypotheses, test the top one, and feed the result back in. Developer products have small user counts, so research does the finding and experiments do the confirming. Choose the next test by expected impact on the biggest leak in the funnel, not by whichever idea is loudest.
1. Define and diagnose flat retention
Define retention concretely: for example, the share of developers who made a successful call in the first week and still make calls in week 5. Build the developer funnel (the ordered steps a developer passes through): signup, first successful call, first real integration (production key), sustained use. Split it by cohort (the group of developers who signed up in the same week) and by segment (hobbyist versus company, language). Flat retention usually hides a specific leak, and the biggest drop shows where to look.
Illustrative week of 500 signups: 350 make a first successful call (70% pass), 140 reach a production key (40% of those 350 pass, so 60% are lost here), and 100 are still calling in week 5 (about 71% pass). The biggest leak is between first call and real integration, so research and the first experiments target that step, not the signup page.
2. Research feeds the hypotheses
- Interview developers who stopped after week 1 and those who kept going.
- Run usability sessions on onboarding and docs (watch someone integrate while thinking aloud).
- Mine support tickets and docs-search queries with no results.
- Add a one-question survey when someone goes inactive.
3. Deciding what to test next
Score each hypothesis: Reach (developers affected per month), Impact (expected effect on the leak, 1 small to 3 large), Confidence (how strong the research evidence is, 0 to 1), divided by Effort (engineer-weeks). Multiply the first three and divide by effort. Candidate backlog (illustrative numbers):
| Hypothesis | Reach | Impact | Confidence | Effort | Score |
|---|---|---|---|---|---|
| Tutorial for the second use case | 300 | 2 | 0.8 (interviews) | 3 | 300 x 2 x 0.8 / 3 = 160 |
| Error messages with fixes | 400 | 2 | 0.8 (tickets) | 1 | 400 x 2 x 0.8 / 1 = 640 |
| Interactive widget | 150 | 1 (unknown) | 0.3 | 6 | 150 x 1 x 0.3 / 6 = 7.5 |
Errors with fixes scores 640, four times the tutorial, so it goes first: many developers hit it, the evidence is strong and it costs one week. The widget scores 7.5 and waits for more evidence. The scores rank ideas, they do not replace judgment.
4. The power problem and how to handle it
Low traffic means small effects cannot be detected. Power is the chance a test spots a real effect (80% here), and significance (alpha 0.05) caps how often it calls noise a win. Suppose 500 new developers a week and baseline 30% retention. In the code, z_alpha and z_power are the standard-normal cut-offs for those two settings (about 1.96 and 0.84), variance sums p(1-p) for the two arms, and the answer is (z_alpha + z_power)^2 x variance / (difference)^2:
from math import ceil
from statistics import NormalDist
def n_per_arm(p1, p2, alpha=0.05, power=0.80):
z = NormalDist().inv_cdf
z_alpha, z_power = z(1 - alpha / 2), z(power)
variance = p1 * (1 - p1) + p2 * (1 - p2)
return ceil((z_alpha + z_power) ** 2 * variance / (p2 - p1) ** 2)
n = n_per_arm(0.30, 0.35)
print("retention 30% -> 35%:", n, "per arm,", 2 * n, "developers in total")
print("weeks at 500 new developers a week:", round(2 * n / 500, 1))
Output:
retention 30% -> 35%: 1374 per arm, 2748 developers in total
weeks at 500 new developers a week: 5.5
By hand: variance = 0.30 x 0.70 + 0.35 x 0.65 = 0.4375; (1.96 + 0.84)^2 is about 7.85; 7.85 x 0.4375 / 0.05^2 is about 1374 per arm. Two arms is 2748 developers, and 2748 / 500 a week is 5.5 weeks. So detecting a 5-point lift needs about 5.5 weeks of all sign-ups, and a 1-point tweak would take far longer. Consequences: test bigger changes, prefer leading indicators (early signals that move faster than retention, such as first-call rate and time to first call), and use qualitative studies where an experiment cannot reach significance. An underpowered test (too few developers to detect the effect you care about) mostly produces noise.
5. Operating the programme
- A fixed cadence (two-week cycles), one owner, and an experiment log with hypothesis, metric, sample size, decision rule, result and decision.
- Guardrail metrics on every test (numbers that must not get worse while you chase the main one): support tickets, error rate.
- Monthly review of the funnel, the research themes and the backlog ranking.
Pitfalls
Running many small underpowered tests and calling noise wins, slicing results into many segments until something looks significant (with 20 slices, one will look like a win by chance, a false discovery), and optimising activation while ignoring that the flat number is week-5 retention. Retention is often driven by product fit and reliability, which docs experiments cannot fix; say so and hand those findings to the product team.
You are designing the developer portal homepage for a B2B payments API aimed at experienced integration engineers. What would you put above the fold, how do developers get from landing to a working call, and how would you judge the page a success?
Sample Answer
Direct answer
"Above the fold" means the part of the page visible before scrolling. Experienced integration engineers want to see the API work, not read marketing. Above the fold I would put a one-sentence value statement, a primary button to get sandbox keys instantly, a secondary link to the API reference, and a copy-paste first request with the response beside it. The measure of success is how quickly and how often a new visitor reaches a successful sandbox call.
Above the fold
- One line saying what the API does ("Accept and reconcile B2B payments with one API").
- Primary call to action (CTA, the main button): "Get sandbox keys". No sales call and no approval queue for the sandbox.
- Secondary CTA: "API reference".
- A live first request: a curl command for creating a test payment with the visitor's key pre-filled once signed in, language tabs, and the JSON response shown next to it.
- Trust essentials in one row: the status page (live and historical uptime), the changelog (a dated list of what changed in the API, so they can see it is maintained), the SDK list (the client libraries, software development kits, available per language) and the security and compliance page. Integration engineers check these before committing.
Rough sketch of the layout:
[ Logo ] [ API reference ] [ Sign in ]
Accept and reconcile B2B payments with one API
[ Get sandbox keys ] (API reference)
curl https://api.example.com/v1/payments { "id": "pay_9137",
-H "Authorization: Bearer sk_test_..." "status": "pending" }
-d amount=2500 -d currency=USD [ cURL | Python | Node ]
Status | Changelog | SDKs | Security
Leave out hero videos (large autoplay marketing clips), customer logos as the main element, and long feature tours.
Path from landing to a working call
- Land, read the one line and the sample request.
- Click "Get sandbox keys", sign up with work email, keys appear immediately.
- Copy the request with the key filled in, run it, see a 201 response.
- Next-step cards: "Accept your first real payment flow", webhooks, error handling, go-live checklist.
- Production access is a separate, clearly described step (business verification) so nobody is surprised.
How to judge success
Define time to first call (TTFC): minutes from account creation to the first successful (2xx) sandbox request. Track:
- Activation rate: share of new sign-ups making a first successful call within 15 minutes.
- Median TTFC for those who do.
- Sandbox-to-production conversion weeks later, which tells you whether easy sandbox success turned into real integrations.
- Docs search queries that return nothing (missing content).
Worked example: 200 sign-ups in a week, 130 with a first 2xx inside 15 minutes: 130 / 200 = 65% activation. If the next design lifts it to 75%, that is a real gain: with 200 sign-ups in each week, the gap of 10 points is about 2.2 times the normal week-to-week wobble (standard error, roughly 4.6 points), so it is likely real, though you would confirm with another week. A lift from 65% to 67% would be only about 0.4 of the wobble, which is noise. A page-view increase would tell you nothing.
Pitfalls
Judging by page views or sign-up counts (vanity metrics, numbers that look good but do not show integration) rewards curiosity, not integration. A pre-filled sandbox key is only safe because it is a test key; never pre-fill live credentials.
Here is an API error: HTTP 400 with the body {"error": "invalid request"}. Critique it from a developer's point of view, rewrite it so someone can fix the problem without contacting support, and explain how you would keep your error messages from leaking sensitive detail.
Sample Answer
Direct answer
The error {"error": "invalid request"} with a 400 tells the developer that they are wrong but not what, where or how to fix it. A good error names the field, says what was expected, gives a stable machine-readable code, carries a request ID for support, and links to help. The safe way to do this is to build errors from a controlled catalogue of messages you wrote, never from raw exception text (the internal error message your code throws, which can reveal how your system is built).
Critique
- Tautology. "Invalid request" restates the 400. No new information.
- No location. Which field or header? Maybe the JSON did not even parse.
- No cause or fix. What was wrong and what is valid?
- No machine-readable code, so client code cannot branch on the failure without string-matching the message.
- No request ID, so support cannot find the log line, which forces a ticket.
- No documentation link.
Rewrite
{
"error": {
"code": "invalid_parameter",
"message": "amount_cents must be an integer of at least 1. Received the string \"25.00\".",
"param": "amount_cents",
"request_id": "req_8f3a2c",
"doc_url": "https://docs.example.com/errors/invalid_parameter"
}
}
The response also carries Content-Type: application/json (or application/problem+json, the media type of RFC 9457 "Problem Details for HTTP APIs", if you follow that standard; a media type is the label in the Content-Type header that says what format the body is in, and an RFC is a numbered internet standards document). For several problems at once, return an array of {param, code, message} items so the developer fixes everything in one pass rather than one round trip per mistake.
Keeping errors from leaking sensitive detail
- Catalogue, not exceptions. Map each failure class to a code and a message a person wrote. A catalogue is just a lookup table kept in your code, for example:
invalid_parameter -> "{param} must be {rule}."
not_found -> "No such resource."
rate_limited -> "Too many requests. Retry after {n} seconds."
The code picks the row and fills in only safe values such as the field name. Anything not in the table becomes a generic internal_error plus the request ID. Never pass an exception's text to the client: it can contain SQL, file paths, hostnames or stack traces (the listing of internal code locations printed when a program crashes).
2. Log the detail server-side under the request_id. The developer gets the ID, support gets the internals.
3. Do not echo secrets. If the invalid value could be a token or password, name the field but do not repeat the value. Truncate long echoed input.
4. Avoid enumeration (letting an attacker learn what exists by trying guesses and reading the differences in your answers). Do not say "no account with that email" versus "wrong password" on login; return the same message. Return 404 for another tenant's resource rather than 403, so existence is not confirmed. A tenant is one customer account in a system shared by many. If customer A requests /invoices/1042 and it belongs to customer B, a 403 Forbidden tells A that invoice 1042 exists (just not for them), so A can probe IDs and map B's data. A 404 Not Found is the same answer A would get for an ID that never existed, so nothing leaks.
5. Keep validation messages about the caller's own input, never about internal structure ("column customers.tax_id is null").
6. Test it. Send malformed and hostile inputs in CI and assert responses never contain a stack trace, file path or internal hostname. CI (continuous integration) is the automated test run on every code change.
Pitfalls
- Messages helpful to attackers ("SQL syntax error near...").
- Changing messages breaks clients that parse them: that is why the code, not the message, is the stable part.
Unlock Full Question Bank
Get access to all 31 API Documentation and Developer Experience interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.