RESTful API Design Questions
Designing resource-oriented HTTP APIs following REST constraints: resource modeling, URI structure, correct use of HTTP methods, statelessness, and HATEOAS trade-offs. Covers naming conventions, collection vs. singleton resources, filtering/sorting/pagination, and choosing appropriate status codes. The default paradigm most interview questions in this category probe.
Implement an idempotent POST /orders endpoint (Python, Flask or FastAPI) that reads an Idempotency-Key header and uses it to prevent duplicate order creation when a client retries. Show the database schema you would use to store keys against their result, the transaction boundaries, and what the server returns when a second request arrives with the same key while the first one is still being processed.
Sample Answer
Direct answer. Read the Idempotency-Key header; if a request with that key has already completed, replay the stored response verbatim instead of re-running the creation logic; if one is currently in flight, reject the duplicate with a distinct status instead of racing it.
Implementation (Python, Flask). This uses an in-memory store for a self-contained, runnable example; in production the same design sits on a Postgres table with a unique index on the idempotency key, so the have-I-seen-this-key check and the reserve-it step are atomic at the database level, not just inside a Python lock.
import threading
import uuid
from flask import Flask, request, jsonify
app = Flask(__name__)
# Stand-ins for a Postgres "idempotency_keys" table and an "orders" table.
idempotency_keys = {}
orders = {}
lock = threading.Lock()
@app.post("/orders")
def create_order():
key = request.headers.get("Idempotency-Key")
if not key:
return jsonify({"error": "Idempotency-Key header is required"}), 400
body = request.get_json(force=True)
with lock:
existing = idempotency_keys.get(key)
if existing is not None:
if existing["status"] == "in_progress":
return jsonify({"error": "already being processed"}), 409
return jsonify(existing["response_body"]), existing["response_status"]
idempotency_keys[key] = {"status": "in_progress", "response_body": None, "response_status": None}
order_id = str(uuid.uuid4())
orders[order_id] = {"id": order_id, "item": body.get("item"), "qty": body.get("qty")}
response_body = {"order_id": order_id, "item": body.get("item"), "qty": body.get("qty")}
response_status = 201
with lock:
idempotency_keys[key] = {"status": "done", "response_body": response_body, "response_status": response_status}
return jsonify(response_body), response_status
if __name__ == "__main__":
client = app.test_client()
r1 = client.post("/orders", json={"item": "widget", "qty": 3}, headers={"Idempotency-Key": "key-abc"})
print("first request ->", r1.status_code, r1.get_json())
r2 = client.post("/orders", json={"item": "widget", "qty": 3}, headers={"Idempotency-Key": "key-abc"})
print("retry request ->", r2.status_code, r2.get_json())
print("same order_id returned on retry:", r1.get_json()["order_id"] == r2.get_json()["order_id"])
print("total orders actually created:", len(orders))
Output (actually run):
first request -> 201 {'item': 'widget', 'order_id': '72664f4a-d7da-4618-8e70-38c7bff94a38', 'qty': 3}
retry request -> 201 {'item': 'widget', 'order_id': '72664f4a-d7da-4618-8e70-38c7bff94a38', 'qty': 3}
same order_id returned on retry: True
total orders actually created: 1
Key points. The reservation of the key (marking it in_progress) happens before the real work starts, inside the same lock as the lookup, closing the race where two near-simultaneous retries both pass the have-I-seen-this check. The stored result is the exact response (body and status), replayed verbatim, not re-derived.
Complexity. Each request does O(1) dictionary work plus the actual creation work; the lock is held only for the cheap lookup-and-reserve step, not for the full creation logic, so concurrent requests with different keys are not serialized against each other in the real (database-backed) version, only requests racing on the same key are.
Edge cases. A second request with the same key but a different body (a client bug, not a legitimate retry) is not checked in this minimal example; a production version should also compare a hash of the request body against what was stored for that key, and reject the request with a 422 if they do not match, since replaying the original response to what looks like a different logical request would silently hide the client bug.
Implement PATCH /users/{id} (Python, Flask) that accepts a JSON Merge Patch body, supports optimistic concurrency via an If-Match header carrying an ETag, returns 412 Precondition Failed on a mismatch, validates the incoming fields, and returns 200 with the updated resource on success. You may assume helper functions get_user, save_user, and compute_etag already exist.
Sample Answer
Direct answer. Read the resource's current ETag, require it on If-Match, reject a mismatch with 412 before applying anything, and only then apply the JSON Merge Patch semantics (null removes a key, any other value sets it, an absent key is untouched).
Implementation (Python, Flask).
import hashlib
from flask import Flask, request, jsonify
app = Flask(__name__)
users = {"u1": {"id": "u1", "email": "alice@example.com", "name": "Alice", "_version": 1}}
def compute_etag(user):
basis = f"{user['id']}:{user['_version']}"
return 'W/"' + hashlib.sha256(basis.encode()).hexdigest()[:16] + '"'
def get_user(user_id):
return users.get(user_id)
def save_user(user_id, updated_fields):
user = users[user_id]
user.update(updated_fields)
user["_version"] += 1
return user
@app.patch("/users/<user_id>")
def patch_user(user_id):
user = get_user(user_id)
if user is None:
return jsonify({"error": "not found"}), 404
if_match = request.headers.get("If-Match")
if not if_match:
return jsonify({"error": "If-Match header is required"}), 428
current_etag = compute_etag(user)
if if_match != current_etag:
return jsonify({"error": "modified since last read", "current_etag": current_etag}), 412
patch = request.get_json(force=True)
if "email" in patch and patch["email"] is not None and "@" not in str(patch["email"]):
return jsonify({"error": "invalid email"}), 400
updated_fields = {}
for key, value in patch.items():
if value is None:
user.pop(key, None)
else:
updated_fields[key] = value
updated = save_user(user_id, updated_fields)
resp = jsonify({k: v for k, v in updated.items() if not k.startswith("_")})
resp.headers["ETag"] = compute_etag(updated)
return resp, 200
if __name__ == "__main__":
client = app.test_client()
etag_v1 = compute_etag(users["u1"])
r1 = client.patch("/users/u1", json={"name": "Alice Smith"}, headers={"If-Match": etag_v1})
print("valid patch ->", r1.status_code, r1.get_json())
r2 = client.patch("/users/u1", json={"email": "x@y.com"}, headers={"If-Match": etag_v1})
print("stale patch ->", r2.status_code, r2.get_json())
fresh_etag = r1.headers["ETag"]
r3 = client.patch("/users/u1", json={"email": "x@y.com"}, headers={"If-Match": fresh_etag})
print("retry with fresh etag ->", r3.status_code, r3.get_json())
fresh_etag2 = r3.headers["ETag"]
r4 = client.patch("/users/u1", json={"name": None}, headers={"If-Match": fresh_etag2})
print("patch name to null ->", r4.status_code, r4.get_json())
r5 = client.patch("/users/u1", json={"email": "not-an-email"}, headers={"If-Match": r4.headers["ETag"]})
print("invalid email ->", r5.status_code, r5.get_json())
r6 = client.patch("/users/u1", json={"email": "y@z.com"})
print("missing If-Match ->", r6.status_code, r6.get_json())
Output (actually run):
valid patch -> 200 {'email': 'alice@example.com', 'id': 'u1', 'name': 'Alice Smith'}
stale patch -> 412 {'current_etag': 'W/"1938f82137da5fc0"', 'error': 'modified since last read'}
retry with fresh etag -> 200 {'email': 'x@y.com', 'id': 'u1', 'name': 'Alice Smith'}
patch name to null -> 200 {'email': 'x@y.com', 'id': 'u1'}
invalid email -> 400 {'error': 'invalid email'}
missing If-Match -> 428 {'error': 'If-Match header is required'}
Every claim below is one of these six runs, not a separate, unshown assertion: retrying with the fresh ETag from the successful response succeeds (200); patching a field to null removes it from the returned object entirely; an invalid email is rejected with 400 without mutating anything; and a request with no If-Match header at all is rejected with 428, not silently applied.
Key points. The If-Match check happens BEFORE the merge-patch logic runs at all, so a stale write never even reaches the mutation code. The ETag is derived from a cheap version counter, not a hash of the whole object, so checking it costs a lookup, not a re-serialization.
Complexity. O(1) per field in the patch body for the merge step itself; the expensive part in a real system is the database round trip to fetch the current version and, on success, to persist the update, both of which this in-memory example elides for clarity.
Edge cases. A patch with no If-Match header is rejected with 428 Precondition Required (not silently applied without a concurrency check); an invalid email is rejected with 400 before any field is mutated, so a partially-applied invalid patch never happens. Both are exercised directly above, not just asserted.
Design a shared idempotency-key service that several backend services can call to make their own POST endpoints safely retryable, rather than every team building its own key store. Cover the storage schema, the TTL policy, what happens on a key collision from two different request bodies, the consistency guarantee you offer callers, and how the service scales to thousands of checks per second without becoming a single point of failure.
Sample Answer
Direct answer. A shared idempotency-key service separates whether-this-request-is-a-duplicate from the business logic of any one team's endpoint, so every service calls one small, well-tested component instead of each team reinventing (and inevitably getting subtly wrong) the same check-then-reserve race condition.
Storage schema. A single table: idempotency_key (unique, primary key), caller_service, request_hash (a hash of the request body, to detect a key reused for a genuinely different request), status (in_progress or done), response_body, response_status, created_at, expires_at (the TTL boundary, derived from created_at plus the caller's configured TTL). The primary key on the idempotency key itself is what makes check-and-reserve atomic: an insert that violates the primary key constraint is the signal that someone else already holds this key, no separate lock needed.
TTL policy. Default the TTL to 24 hours after created_at, the same default a single-service idempotency store would use, since that comfortably covers realistic retry storms (a client retrying for a few minutes, or a batch job retrying hours later after being paged) without keeping every key forever. Because this is a SHARED service used by callers with very different retry patterns, make the TTL configurable per caller_service within a bounded range (say, 1 hour to 7 days) rather than one fixed value for everyone: a billing job replaying a batch over a weekend needs a longer window than a low-latency checkout path. Expire keys through the store's own native TTL mechanism where one exists (a Redis EXPIRE on the key) or a partition-local sweep deleting rows past their expires_at, rather than one centralized cron job scanning the whole table, since that job would itself become a scaling bottleneck as the table grows.
Collision from two different request bodies. If a key arrives with a request_hash that does not match what is stored, this is not a legitimate retry, it is a bug (a client reusing one key across two logically different requests) or an attack (a caller guessing another caller's key to hijack its cached response). Reject with 422 rather than either silently returning the wrong cached response or silently overwriting it, both of which hide the underlying problem.
Worked example. Caller A submits Idempotency-Key: order-778 with body {"amount": 4200, "currency": "USD", "recipient": "acct_9182"}; the service computes request_hash a5507621 (a short hash of the body's bytes) and stores it against the key alongside the 201 response. If caller A's connection drops and it retries with the SAME key and the SAME body, the request_hash recomputes to the identical a5507621, matches what is stored, and the service replays the original response verbatim, no new work done. Now say caller B (or a bug in caller A) reuses order-778 with a DIFFERENT body, {"amount": 5300, "currency": "USD", "recipient": "acct_9182"}; that body hashes to 16ef6b3b, which does not match the stored a5507621, so the service rejects it with 422 instead of either replaying caller A's stale result or silently overwriting it.
Consistency guarantee. Offer callers a simple, honest contract: once a key's status is done, every subsequent lookup for that key returns the exact same response, with strong (not eventual) consistency, for the configured TTL. This requires the underlying store to support a real atomic compare-and-set (the check "is this key already claimed" and the act of claiming it happen as one uninterruptible step, so two callers racing to claim the same key cannot both succeed) or unique-constraint-based insert, not an eventually-consistent cache (a store where different copies of the data can briefly disagree with each other before catching up, rather than always agreeing immediately), because the entire point of the service is to prevent a race, and an eventually-consistent store can itself let two concurrent reserve-this-key attempts both briefly believe they won.
Scaling to thousands of checks per second. Partition the table (or use a partitioned key-value store) by a hash of the idempotency key itself (partitioning means splitting one big table into several smaller, independent pieces by key, so no single machine has to hold or lock against every key in the system), so any one key's reservation only ever contends with itself, never with an unrelated key from a different caller. Concretely, with 16 partitions: hashing the key order-778 and taking it mod 16 lands it on partition 1, so only partition 1 ever needs to agree on whether THIS key is claimed; the other 15 partitions are untouched by this reservation and keep serving their own callers' keys in parallel, which is what lets total throughput scale roughly with partition count instead of being capped by one table's write rate. Each partition can then be served by an independent database shard or a Redis instance with SET-if-not-exists semantics (Redis's "only write this key if it does not already exist" command: the check and the write happen as one uninterruptible step, the same atomic-reservation guarantee as a unique-constraint insert, just against an in-memory store instead of a relational table), at much higher throughput than a single relational table. Read replicas (copies of the database that only answer read queries, never writes) can serve the already-done lookups (the overwhelming majority of traffic once a system is warmed up: most callers are asking about a key that already resolved, not racing to create one), while writes for genuinely new keys go to the primary (the one copy of the data that accepts writes and propagates the change out to its replicas). This needs one explicit guardrail to avoid quietly breaking the strong-consistency guarantee above: a read that lands on a replica before that replica has caught up with a very recent write could return a stale not-found or in-progress result for a key that in fact just completed on the primary. Route lookups for a key still inside its replication window (the first few hundred milliseconds after it was written) to the primary, or to a replica confirmed to have caught up, and only serve the read-replica fast path once a key is old enough that replication lag can no longer be in question. In practice this only affects the sliver of traffic asking about a key in the same instant it was created, not the bulk of already-resolved lookups the replicas exist to serve, so it does not undermine the point of adding replicas in the first place.
Avoiding a single point of failure. The service itself needs the same idempotent-retry discipline it offers others: if a caller's request to the idempotency service times out, their retry to the idempotency service needs its own strategy (typically: the caller retries at-least-once against the idempotency service using the network call's own natural retry semantics, since the service's operations are themselves designed to be safe to call more than once for the same key).
Trade-offs and pitfalls. The most common design mistake is skipping the request-hash comparison and trusting the key alone; without it, a client-side bug that reuses a key looks exactly like a correctly-working idempotency system, right up until it silently returns stale data for a genuinely new request.
Design a JSON error response schema that both your internal teams and external clients will consume: what fields would you include (for example a machine-readable code, a human message, field-level validation detail, and a correlation id for tracing), what belongs in the client response versus only in your logs, and how would a client tell a retryable error from one it should not retry?
Sample Answer
Direct answer. A good error schema separates three concerns that a single "message" string conflates: a stable, machine-readable code a client can branch on programmatically, a human-readable message for logs and debugging, and enough structured detail (which field, what was wrong with it) for a UI to show something more useful than a generic failure.
A concrete shape.
{
"error": {
"code": "VALIDATION_ERROR",
"message": "Request failed validation.",
"retryable": false,
"request_id": "req_8f2a1c",
"details": [
{ "field": "email", "issue": "must be a valid email address" },
{ "field": "quantity", "issue": "must be greater than 0" }
]
}
}
code: a stable string a client's error-handling logic can switch on; unlike the HTTP status code alone, it can distinguish "insufficient funds" from "card declined" even though both might return the same 402.message: for humans (logs, a developer reading a support ticket), never the primary thing client CODE should branch on, since it is free to change wording without that being a breaking change.request_id: a correlation id the client can hand back to support, letting you find the exact server-side log line for this request instantly instead of searching by timestamp and endpoint.details: field-level validation information for a form UI to highlight the specific inputs that were wrong.
Client response vs. logs. The client response should NEVER include a stack trace, an internal service name, a raw database error message, or any other detail that reveals your system's internals; that information belongs only in your server-side logs, correlated by the same request_id, so an engineer investigating a support ticket can look up the full internal detail without ever exposing it to the caller.
Retryable vs. not. A retryable boolean (or deriving retryability from the error code via a documented mapping) tells the client whether blindly retrying the exact same request is safe and potentially successful (a transient 503) versus pointless or actively harmful (a 400 validation error, which will fail identically on every retry until the request itself changes). Without this signal, clients either retry everything (wasting calls on errors that can never succeed) or retry nothing (giving up on transient failures a simple retry would have fixed).
Trade-offs and pitfalls. The most common mistake is putting the human message where client code actually parses it, so a later, purely cosmetic wording change ("Invalid email" to "Please provide a valid email") silently breaks any client that was doing string-matching against the message instead of the code.
What conventions do you use for naming REST resources and endpoints: plural versus singular nouns, when to nest a resource under its parent (for example /users/{userId}/orders) versus keep it top-level, how to represent an action that is not plain CRUD without falling back to an RPC-style verb in the URL, and how you keep the number of endpoints from exploding as relationships between resources grow.
Sample Answer
Direct answer. Use plural nouns for collections (/users, not /user), nest a resource under its parent only when it genuinely cannot exist independently of that parent (/users/{userId}/orders, since an order without a user makes no sense in this domain), and never encode an action as a verb in the path (no /getUser or /createOrder); if an operation is not naturally CRUD-shaped, model it as a sub-resource or an explicit action endpoint under the resource it acts on, not a bare verb.
Plural vs. singular. Plural nouns for every collection endpoint, consistently, even for a resource that will usually only ever have one instance per parent (a user's single /profile is a defensible, deliberate exception, since it genuinely is not a collection); consistency here matters more than any individual argument for singular naming, since an API mixing /users and /order for no principled reason forces every client developer to memorize which convention applies where.
When to nest, and when not to. Nest when the child resource's identity is meaningless without the parent (a specific order's line items only make sense scoped to that order: /orders/{orderId}/items) and the child is always accessed IN that context. Do NOT nest when the resource has its own independent identity and is commonly accessed on its own (a specific order is meaningfully addressable as /orders/{orderId} directly, even though it also appears as one entry under /users/{userId}/orders); over-nesting (/users/{userId}/orders/{orderId}/items/{itemId}/reviews/{reviewId}) makes URLs unwieldy and couples every deep resource's addressability to knowing its entire ancestor chain, when a flatter /reviews/{reviewId} with the relationship expressed in the response body would serve most clients better.
Actions that are not plain CRUD. For a genuine action (cancel this order, publish this post), prefer a sub-resource action endpoint (POST /orders/{id}/cancel) over inventing a new HTTP verb or falling back to an RPC-style bare verb in the path; this keeps the resource-oriented structure while still allowing operations that do not map cleanly onto GET/PUT/PATCH/DELETE.
Avoiding endpoint explosion as relationships grow. As a domain accumulates more relationships (users have orders, orders have items, items have reviews, reviews have replies...), resist nesting every single one; instead, expose the DEEPLY nested resources at their own flat, top-level path (/reviews/{reviewId}, not /users/{userId}/orders/{orderId}/items/{itemId}/reviews/{reviewId}) and use query parameters or embedded links to express the parent relationship, rather than growing the URL structure to mirror the full entity-relationship diagram.
Trade-offs and pitfalls. The most common mistake is treating nesting depth as free; each additional level of nesting is a small but real tax on every client that has to construct or parse that URL, and a domain's relationships almost always grow faster than anyone expects when the API was first designed.
Unlock Full Question Bank
Get access to all 39 RESTful API Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.