RESTful API Design Questions
Designing resource-oriented HTTP APIs following REST constraints: resource modeling, URI structure, correct use of HTTP methods, statelessness, and HATEOAS trade-offs. Covers naming conventions, collection vs. singleton resources, filtering/sorting/pagination, and choosing appropriate status codes. The default paradigm most interview questions in this category probe.
Design a REST API for a long-running bulk job (for example a bulk data export) that must support submit, status polling, pause, resume, and cancel, plus resuming correctly after a failure. Define the job's state machine and which transitions are valid from which states, ensure operations stay idempotent under retries, and describe what the client-visible progress and error model looks like.
Sample Answer
Direct answer. Model the job as an explicit state machine (queued, running, paused, cancelled, failed, completed), expose one action endpoint per valid transition rather than a single generic status update, and make every transition idempotent under retry the same way a single mutating endpoint would be.
The state machine. States: queued -> running -> (paused <-> running) -> completed, with cancelled and failed reachable from queued, running, or paused, but not from completed. Each transition is its own endpoint: POST /jobs/{id}/pause, POST /jobs/{id}/resume, POST /jobs/{id}/cancel, plus GET /jobs/{id} for status and progress. A transition request that does not correspond to a legal edge from the job's CURRENT state (say, resuming a job that already completed) returns 409 Conflict, naming both the attempted transition and the actual current state, the same discipline as the order-lifecycle state-machine design.
Idempotency for each transition. Each action endpoint accepts an Idempotency-Key the same way a POST /orders create endpoint would: retrying POST /jobs/{id}/pause with the same key after a dropped connection replays the stored result (confirming the job is now paused) rather than erroring or double-processing the pause. This matters more here than for a single mutating endpoint precisely because a long-running job's client is far more likely to experience a network interruption mid-operation, simply because the whole point of the job is that it takes a long time.
Resuming correctly after a failure. On resume (whether client-initiated after a deliberate pause, or automatic after a worker crash mid-job), the job must pick up from its last durably-recorded checkpoint, not restart from the beginning; this requires the job's own internal progress to be persisted incrementally (a processed-so-far marker written to durable storage as the job runs), not held only in the worker process's memory, so a crash loses at most the work since the last checkpoint, not the whole job.
Client-visible progress and error model. GET /jobs/{id} returns the current state, a progress indicator (items processed out of an estimated or exact total, when knowable), and, on a failed state, a structured error explaining what went wrong and whether the failure is one the client can address (bad input data) versus one that is purely operational (a transient infrastructure failure the client should simply retry the whole job for). Progress should be a real, monotonically-increasing signal derived from checkpoints, not simply "polling started N seconds ago", so a client (or a monitoring dashboard) can distinguish a genuinely stuck job from one that is legitimately still working through a large dataset.
Trade-offs and pitfalls. The most common mistake is implementing pause as "stop processing but keep the in-progress state only in the worker's memory," which looks correct in every test where the same worker process later resumes it, and silently loses all progress the first time an actual resume happens against a different worker instance after the original one was recycled.
For each of GET, POST, PUT, PATCH, DELETE, HEAD, and OPTIONS, state whether it is safe, whether it is idempotent, and whether it is cacheable, then give a short example endpoint where using the wrong method caused a real bug (for instance, a client retry duplicating a purchase, or a caching layer serving a stale response for a method it should not have cached).
Sample Answer
Direct answer. GET, HEAD, and OPTIONS are safe (they must not change server state) and idempotent (calling them N times has the same effect as calling them once). PUT and DELETE are idempotent but not safe. POST is neither safe nor idempotent by default. PATCH is technically unspecified but should usually be treated as non-idempotent unless you deliberately design it to be. Cacheability tracks safety closely but is not identical to it: GET and HEAD are cacheable by default, OPTIONS is safe but has no real caching convention, and none of PUT, DELETE, POST, or PATCH are cacheable by default.
The full classification.
| Method | Safe | Idempotent | Cacheable | Typical use |
|---|---|---|---|---|
| GET | yes | yes | yes, by default | read a resource, no side effects |
| HEAD | yes | yes | yes, by default | GET without a body, for existence/metadata checks |
| OPTIONS | yes | yes | no | discover allowed methods, CORS preflight |
| PUT | no | yes | no | replace a resource entirely at a known URI |
| DELETE | no | yes | no | remove a resource; deleting twice leaves the same end state (gone) |
| POST | no | no | no by default | create a new resource, or trigger a non-idempotent action |
| PATCH | no | usually not | no | partially update a resource |
Why cacheability does not just follow safety. GET and HEAD are cacheable by default because a cache can reuse their response without risking a stale side effect, the same property that makes them safe in the first place. OPTIONS is also safe, but nothing about discovering allowed methods or a CORS preflight benefits from caching the way a resource representation does, so there is no real caching convention for it in practice. POST responses CAN technically be cached per the HTTP spec if the response carries explicit freshness information (Cache-Control or Expires), but this is rarely implemented, so treat POST as effectively not cacheable. PUT, DELETE, and PATCH have no meaningful default caching semantics either: caching the result of a mutation makes little sense when the whole point of the call was to change state.
Why this matters beyond vocabulary. Idempotency is the property that makes retries safe. If a client's network call to a PUT times out and it retries, the end state is identical whether the first request actually landed or not, because PUT is defined as "the resource now looks like this", not "apply this delta". A POST retried the same way can create two resources, because POST means "do this action again", and doing a creation action twice creates two things. Safety is what a cache, a browser prefetcher, or a crawler relies on: none of them should ever issue a POST speculatively, because a safe method is one where the caller assumes no side effect happened.
A real bug from getting this wrong. A checkout flow implemented "add item to cart" as a GET request (because it was convenient to trigger from a link). A corporate web-security scanner crawled every link on the page, including that one, adding dozens of items to real users' carts, because the scanner (correctly, per the HTTP contract) assumed GET was safe to call without consequence. The fix was not to block the scanner, it was to make cart mutation a POST, which is exactly what the safety property exists to protect against.
Trade-offs and pitfalls. The subtlest mistake is assuming PATCH is idempotent by default. A PATCH body of {"counter": "increment"} is not idempotent (retrying it increments twice); a PATCH body of {"counter": 5} (set to an absolute value) is. The method name alone does not tell a client which one they are getting, so this needs to be documented per endpoint, not assumed from the HTTP verb.
Write the JSON body a server should return for a failed POST /orders call, using the Problem Details for HTTP APIs format. Include type, title, status, detail, and instance from the standard, plus a machine-readable error_code and a correlation_id you add on top of it, and briefly say what each field is for.
Sample Answer
Direct answer. The body below follows RFC 7807's shape (type/title/status/detail/instance) and adds two fields the standard leaves out but any real API needs: a stable machine-readable code and a correlation id.
{
"type": "https://api.example.com/errors/validation-failed",
"title": "Request validation failed",
"status": 400,
"detail": "The 'quantity' field must be greater than 0, and 'email' is not a valid email address.",
"instance": "/orders",
"error_code": "ORDER_VALIDATION_FAILED",
"correlation_id": "req_8f2a1c9d",
"remediation": "Fix the fields listed in 'detail' and resubmit the request."
}
What each field is for.
type: a stable URI identifying this PROBLEM TYPE (not this occurrence); resolving it (or just reading it, most implementations do not require it to be a live, dereferenceable page) tells a developer or a client library "this is the kind of error where request validation failed," reusable across every endpoint that can fail this way.title: a short, generic summary of that problem type, identical every time this specific type occurs; changes to wording here are a documentation update, not a per-request detail.status: 400, repeated in the body because some client code and logging pipelines only see the JSON body, not the HTTP status line, especially once a response has passed through a proxy that logs bodies but not headers.detail: THIS occurrence's specifics, which fields failed and why; this is what actually varies request to request, unliketitle.instance: the specific resource path this occurrence happened against, useful when the same problemtypecan occur on multiple endpoints and you want to know exactly which call triggered it.error_code: the machine-readable string a client's error-handling code should actually branch on (if error_code == "ORDER_VALIDATION_FAILED"), sincetype/titleare meant to be more human-facing and are not guaranteed to be a tight enum a client can safely switch over.correlation_id: lets support or the client hand this exact value back to you, and you look up the full server-side log entry for the request instantly, instead of searching by approximate timestamp.
Trade-offs and pitfalls. Putting per-request specifics (like the exact invalid value a user typed) into title instead of detail breaks the standard's own intended use of title as a stable, cacheable label for the error TYPE, and makes it useless for grouping or aggregating errors by type later.
Compare four ways to expose a long-running operation to a client: a synchronous call with a long timeout, an asynchronous job endpoint the client polls, a webhook callback on completion, and a push mechanism like Server-Sent Events or WebSockets. For each, describe the API contract for starting the operation and getting the result, and the trade-off in scalability, reliability, and how much complexity it pushes onto the client.
Sample Answer
Direct answer. A synchronous call with a long timeout is the simplest contract but scales worst and is the least reliable; asynchronous polling adds one extra round trip per check but is simple, universally supported, and tolerant of client disconnects; a webhook callback removes polling entirely but requires the client to run a reachable, publicly addressable endpoint; and a push mechanism (SSE or WebSockets) gives the lowest latency notification but costs a held-open connection per client and, unlike the other three, loses events outright across a dropped connection unless you deliberately design around it.
Synchronous, long-timeout call. Contract: the client makes one request and the connection stays open until the operation finishes. Simplest to implement and to consume, but it ties up a connection (and, usually, a worker thread or process) on both ends for the full duration, does not survive a client disconnect or a load-balancer's own idle-connection timeout, and gives the client no way to check progress or cancel while waiting. Reasonable only for operations that reliably finish in a few seconds.
Asynchronous polling. Contract: POST starts the job and returns 202 Accepted with a Location header pointing at a status resource; the client GETs that resource repeatedly until it reports a terminal state. Trade-off: an extra round trip per check, and the client has to decide a polling interval (too frequent wastes both sides' resources, too infrequent adds latency to when the client learns of completion), but it needs no special client-side networking capability (any client that can make a plain GET can poll) and survives a client disconnecting and reconnecting later, since the job's state lives independently on the server.
Webhook callback. Contract: the client registers a callback URL at job-submission time; the server POSTs the result to that URL once the job completes, with the usual webhook discipline (signing the payload, retrying on delivery failure, the client acknowledging receipt). Trade-off: removes polling entirely and notifies the client the instant the job finishes, but requires the client to operate a publicly reachable HTTP endpoint capable of receiving the callback reliably, which is a real operational burden a purely client-side application (a mobile app, a browser tab) usually cannot meet at all.
Push (SSE or WebSockets). Contract: the client opens one connection and receives job-status events pushed over it as they happen, no polling and no callback endpoint needed on the client's side. Trade-off: lowest latency notification of the four options, but the server has to hold one open connection per subscribed client for as long as they care about updates, which is real, ongoing resource cost per client (unlike polling, whose cost is bounded and predictable, or webhooks, which cost nothing while nothing is happening) and this cost scales linearly with the number of simultaneously-watching clients, not with how many jobs are actually running. Reliability is the real weak point of this option specifically: if the connection drops mid-job (a mobile network hiccup, a laptop sleeping), any event pushed while disconnected is simply lost, unlike polling (the next poll just re-reads current state) or a webhook (the server retries delivery). SSE mitigates this with a built-in Last-Event-ID mechanism so a reconnecting client tells the server where it left off and missed events can be replayed; a raw WebSocket has no equivalent built in and needs the same idea implemented by hand. In practice, push is usually paired with a fallback GET on the status endpoint after a reconnect, so a client that missed an event still converges on the true state instead of silently believing stale information.
Choosing. Reach for polling as the default (works everywhere, no special client capability required); reach for a webhook when the caller is itself a server-side integration that can reliably host a callback endpoint; reach for push only when true low-latency notification to many simultaneously-connected clients is a real product requirement (a live collaborative dashboard), since it is the option with the highest ongoing server cost per client and the one most in need of an explicit reconnect-and-reconcile plan.
Your organization wants a single standard shape for API error responses instead of every team inventing its own. Explain what the Problem Details for HTTP APIs standard (RFC 7807) specifies as required versus optional fields, and what you would add beyond the standard (for example a machine-readable error code, a correlation id, and a retryable flag) to make it genuinely useful for SDK authors and partner integrators.
Sample Answer
Direct answer. RFC 7807 (Problem Details for HTTP APIs) standardizes a small set of fields so error responses across different APIs, and different teams within one company, all have the same recognizable shape instead of every team inventing its own.
Required-in-spirit vs. optional fields. The standard defines five members, none of which are strictly mandatory by the RFC itself, but which only earn the "Problem Details" name when used together: type (a URI identifying the problem type, defaulting to "about:blank" if you have not documented one), title (a short, human-readable summary that should be the SAME for every occurrence of this problem type, not per-instance), status (the HTTP status code, repeated in the body for convenience since some clients only see the body), detail (a human-readable explanation specific to THIS occurrence, unlike the generic title), and instance (a URI identifying this specific occurrence, useful for correlation). The spec is explicitly EXTENSIBLE: you are expected to add your own fields on top for anything domain-specific.
What you would add beyond the standard. A machine-readable error_code (the standard's type/title are meant to be somewhat human-facing and are not guaranteed unique or stable enough for client code to branch on reliably), a correlation_id (the standard has no built-in concept of a trace or request id), and a retryable boolean (the standard says nothing about whether a client should retry). None of this conflicts with the spec, since RFC 7807 is deliberately a small, extensible core rather than a complete error contract.
Why bother with a named standard instead of just inventing your own shape. Two real benefits: existing HTTP client libraries and API gateways increasingly recognize the application/problem+json content type and can surface it specially (rather than treating every error body as an opaque, one-off shape), and new engineers or partner integrators who have seen RFC 7807 elsewhere immediately recognize the shape of your errors instead of needing to learn a company-specific convention from scratch.
Trade-offs and pitfalls. The most common mistake is adopting the standard's field NAMES but not its actual DISCIPLINE: repeating the same generic title across genuinely different problem types (so it stops being useful for grouping/aggregation), or putting instance-specific detail into title instead of detail, defeating the distinction the spec draws between the two.
Unlock Full Question Bank
Get access to all 33 RESTful API Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.