API and Interface Design for Distributed Services Questions
Designing the contracts between services and clients: REST, gRPC, and GraphQL tradeoffs, versioning and backward compatibility, pagination, rate limiting, and idempotent endpoints. Covers request/response modeling, error contracts, and API gateway responsibilities. Focuses on the interface layer that ties distributed components together, not internal data schemas.
Describe how you would design API pagination and sync endpoints for a mobile app that must support partial offline sync. The endpoints should let the client reconcile local mutations with server state and fetch only deltas since the last sync point. Provide the high-level request/response contract and how you would surface conflicts to the caller.
Sample Answer
Direct answer
Split the contract into two endpoints threaded together by one opaque sync token: a push endpoint that accepts the client's local mutations tagged with the server version they were based on (so the server can detect conflicts), and a paginated pull endpoint that returns only the deltas since that token plus a new token to resume from. The client always pushes before it pulls, so its own pending changes are reflected in the server state it is about to reconcile against.
Contract shape
Push (client to server), sends local mutations:
POST /sync/push
{
"clientId": "device-abc",
"baseSyncToken": "tkn-123",
"mutations": [
{"localId": "c1", "type": "update", "resource": "note", "id": "srv-45", "baseVersion": 77, "payload": {"title": "New title"}}
]
}
Push response, tells the client what happened to each mutation:
{
"applied": [
{"localId": "c1", "serverId": "srv-45", "status": "applied", "serverVersion": 78}
],
"conflicts": [],
"newSyncToken": "tkn-124"
}
Pull (server to client), paginated deltas since the token:
GET /sync/pull?since=tkn-124&pageToken=null
{
"items": [
{"serverId": "srv-46", "resource": "note", "op": "upsert", "payload": {"title": "Meeting notes"}, "version": 12}
],
"nextPageToken": null,
"newSyncToken": "tkn-125"
}
Conflict signaling
Every mutation in the push request carries baseVersion, the server version the client last saw for that resource. The server compares it to the resource's current version before applying:
| Strategy | When to use | What the client sees |
|---|---|---|
| Reject and surface the conflict | Data loss risk is high (financial fields, anything a human should review) | {"localId": "c2", "reason": "version_mismatch", "serverState": {...}} in the conflicts array; client shows the current server value alongside the pending local change |
| Last-writer-wins by timestamp | Low-stakes fields where losing a rare concurrent edit is acceptable | The mutation applies silently; client sees status: "applied" even though its base version was stale |
| Field-level merge | Structured objects where two edits touch different fields | Server merges non-overlapping fields and returns the merged serverVersion; only overlapping-field conflicts surface |
Whichever strategy is chosen, baseVersion is what makes the push idempotent and safe to retry after a network drop: replaying the same push twice against an already-applied localId returns the same applied result rather than double-applying it.
Worked example
sequenceDiagram
participant C as Mobile client
participant S as Sync API
C->>S: POST /sync/push {baseSyncToken, mutations}
S-->>C: {applied, conflicts, newSyncToken}
C->>S: GET /sync/pull?since=newSyncToken&pageToken=null
S-->>C: {items, nextPageToken=abc, newSyncToken}
C->>S: GET /sync/pull?since=newSyncToken&pageToken=abc
S-->>C: {items, nextPageToken=null, newSyncToken}
Note over C: pull loop ends when nextPageToken is absent
Tracing the token through the exchange above: the push response's newSyncToken (tkn-124) becomes the since value on the first pull; each pull response echoes back the same newSyncToken (it only advances once a full pull pass completes) until nextPageToken comes back absent, which is the client's signal to stop paging and consider itself caught up as of that token.
Trade-offs and pitfalls
Deletes need a tombstone record (a marker saying "this resource was removed," not just its absence) so a client pulling deltas can distinguish "never existed" from "existed then got deleted." Tombstones cannot be kept forever: they need a retention window and a background cleanup pass that garbage-collects (permanently purges) tombstones older than the window, after which a client that reconnects past that window can no longer compute a delta and must fall back to a full resync instead of a partial pull.
A client that reconnects after a very long offline period is the sharpest edge case: if its baseSyncToken predates the server's retained history, the pull endpoint must detect that explicitly (rather than silently returning an incomplete delta) and tell the client to discard local state and re-fetch a full snapshot.
Monotonic version numbers per resource are enough for the common single-writer-per-record case. True multi-device, multi-writer scenarios (the same resource edited concurrently from two devices before either has synced) need richer causality tracking than a single version number can express; that machinery belongs to your consistency model, not this contract, so treat it as a known limitation to flag rather than something to solve inside the sync endpoints themselves.
Two clients read the same resource through your API, both submit an update, and the second write silently overwrites the first client's change without either client finding out. Design a request/response contract that prevents this kind of lost update, and walk through what the client sees when its update is rejected because the resource changed underneath it.
Sample Answer
Direct answer
Make every write conditional on the version of the resource the client actually read. The server attaches an ETag (entity tag: an opaque fingerprint that changes whenever the resource changes) to every representation it returns; the client must echo that value back in an If-Match request header on its update. If the resource changed since the client's read, the current ETag no longer matches, and the server rejects the write with 412 Precondition Failed instead of silently overwriting, so the client finds out immediately instead of losing the first writer's change.
Contract
GET /orders/42returns200 OKwith headerETag: "5"and the resource body.PUT /orders/42must include headerIf-Match: "5". If the resource's current version is still5, the server applies the write and returns200 OKwith the newETag: "6".- If someone else already wrote first (current version is now
6, not5), the server returns412 Precondition Failedand does not apply the write at all. - If the client omits
If-Matchentirely on an endpoint that requires it, return428 Precondition Requiredrather than allowing an unconditional overwrite by default. This closes the gap where a client that never learned about conditional writes would otherwise clobber silently, which is the exact bug this contract exists to prevent.
What the client sees on a rejected write
HTTP/1.1 412 Precondition Failed
Content-Type: application/json
{
"error": "precondition_failed",
"message": "Resource has changed since you last read it.",
"currentETag": "6",
"currentValue": {"orderId": "42", "status": "shipped", "quantity": 3}
}
The client now has everything it needs to recover on its own: reload the current value, decide whether to merge its pending change into it or discard it, and retry the write with the fresh If-Match: "6".
Worked example
sequenceDiagram
participant C as Client
participant A as API
participant D as Database
C->>A: GET /orders/42
A->>D: read order 42
D-->>A: value=v1, version=5
A-->>C: 200 OK, ETag: "5"
C->>A: PUT /orders/42 (If-Match: "5")
A->>D: compare-and-swap on version 5
D-->>A: mismatch (version now 6)
A-->>C: 412 Precondition Failed
Trace it through concretely: Client A reads at ETag: "5". Before Client A writes, Client B also read at ETag: "5" and successfully wrote, moving the server to ETag: "6". When Client A's PUT arrives with If-Match: "5", the server's compare-and-swap against the current version 6 fails, so Client A gets 412 Precondition Failed with the current state in the body, exactly the outcome that prevents its write from silently erasing Client B's change.
Trade-offs and pitfalls
ETag and If-Match are HTTP-native, so they interoperate for free with generic HTTP caches, CDNs, and client libraries that already understand conditional requests, unlike a bespoke version field buried only in the JSON body. A weak ETag, written as W/"5", signals "semantically equivalent," and is comparable with looser equality rules; a strong ETag (no W/ prefix) signals byte-for-byte identical representations. Pick weak ETags for most application-level version checks since exact byte identity is rarely what you actually need.
The sharpest pitfall is implementing the check and the write as two separate application-code steps (read current version, compare, then write) without an atomic compare-and-swap at the storage layer: that reintroduces a race window between the check and the write, which silently recreates the exact lost-update bug this pattern is meant to eliminate. The comparison and the write must be one atomic operation, typically a conditional update expression the database itself evaluates.
Optimistic concurrency control (detecting conflicts at write time rather than locking on read) works well for low-contention resources. Under heavy contention on the same resource, clients retrying after repeated 412 responses can thrash; that is a signal to consider a different concurrency strategy for that specific hot resource, not evidence that the ETag contract itself is wrong.
Design a REST API for listing and creating 'products' that supports pagination, filtering, sorting, and versioning. Specify the request/response shapes, your pagination strategy, your versioning approach, and how you would roll out a breaking change to both internal and external clients without a hard cutover.
Sample Answer
Direct answer
Model the collection with a resource-oriented URL (GET/POST /v1/products), cursor pagination for listing, structured filter and sort query parameters, and a URI-path major version (/v1/). To roll out a breaking change without a hard cutover, let the old and new response shapes run side by side under different version paths, keep the new field additive first, and retire the old version only once usage telemetry shows it is safe, not on a calendar date alone.
Framework
Endpoints and shapes
| Method & path | Purpose |
|---|---|
GET /v1/products?limit=20&after=<cursor>&sort=price:asc&status=active&category=tools | List products, paginated, filtered, sorted |
POST /v1/products | Create a product |
GET /v1/products/{id} | Fetch one product |
PATCH /v1/products/{id} | Partial update |
GET /v1/products?limit=2&sort=price:asc
200 OK
{
"data": [
{ "id": "p1", "name": "Widget", "price": 9.99, "status": "active" },
{ "id": "p2", "name": "Gadget", "price": 14.50, "status": "active" }
],
"page": { "limit": 2, "next_cursor": "<opaque-cursor-for-p2>", "has_more": true }
}
POST /v1/products
{ "name": "New Widget", "price": 19.99, "category": "tools" }
201 Created
Location: /v1/products/p3
{ "id": "p3", "name": "New Widget", "price": 19.99, "category": "tools", "status": "active" }
Pagination strategy: cursor-based, for the same reason as any list that can grow and be written to concurrently: stable under inserts, roughly constant query cost at depth, at the cost of not supporting a direct "jump to page 43." Offer limit and an opaque after cursor rather than page/offset.
Filtering and sorting: plain query parameters for common fields (status=active, category=tools), and a single sort=field:direction parameter, comma-separated for multiple fields, for example sort=price:asc,name:desc. The filterable field vocabulary is documented per field in the OpenAPI spec (a machine-readable API description) so clients know exactly what is filterable rather than guessing.
Versioning: the major version lives in the URL path (/v1/), because it is simple to route at a gateway or load balancer without inspecting headers, and it is trivially cacheable per version. Additive changes, new optional fields, new optional query parameters, ship inside v1 without a version bump; anything that removes a field, changes a field's type, or changes default sort or filter behavior goes into v2.
Rolling out a breaking change without a hard cutover: suppose v2 needs to rename price to unit_price and change it from a decimal amount to an integer number of cents.
- Ship
v2alongsidev1on the same deployment.v2's handler reads and writes the same underlying data asv1, so there is exactly one source of truth behind two response shapes. - Mark
v1'spricefield deprecated via a response header and a documented sunset date, while it keeps working exactly as before. Nothing breaks yet. - Instrument both versions: tag every request with which version served it, and track the fraction of traffic still hitting
v1. - Migrate internal clients first, since you can coordinate with them directly, then notify external and partner clients with a fixed migration window, for example 90 days, pointing at the
v2docs and a short code sample. - Remove
v1only once telemetry shows its remaining traffic has dropped to a level you have explicitly decided is safe to force-migrate, for example only a handful of clients you can contact individually, not on the calendar date alone. If a significant client is still onv1at the deadline, extend the window rather than break them, and treat that as a signal the migration tooling or communication needs work.
Worked example
Concretely, at the moment v2 ships, v1 carries 100% of traffic. After internal clients migrate in the first month, v1's share drops to 70%. Announcing the 90-day window to external partners brings it down further; by day 60, v1 is at 8%, all from three named partner integrations already contacted directly about their remaining migration steps. That 8%-and-named state, not the passage of 60 days on its own, is what tells you it is close to safe to set a hard removal date, once those three integrations confirm.
Trade-offs and pitfalls
- Renaming
pricetounit_priceand changing its unit, dollars to cents, in the same release conflates two changes into one migration. A client that only cared about the rename still has to handle the unit change, which raises the odds the migration is done wrong. Prefer landing one breaking change at a time when volume allows it. - A common pitfall is announcing deprecation only in documentation and not in the response itself. Clients that never read a changelog will not notice until the sunset date arrives, so a machine-readable response header, not prose, is what actually drives safe removal.
- Cursor pagination combined with
sortneeds care: the opaque cursor usually encodes the sort key's value, so changingsortmid-pagination, fetching page 1 by price then asking for page 2 by name, should be rejected or restarted from page 1, since a cursor from one sort order is meaningless under a different one.
Design the pagination contract for a feed API where new items are frequently inserted at the head while a client is mid-page or has filters and sorting applied. Explain how the cursor or page token stays valid across those inserts, how you avoid returning duplicate or skipped items, and what the response needs to signal so the client can handle the change smoothly.
Sample Answer
Direct answer
Anchor the pagination cursor to a specific item's identity in the ordering, never to a position count, so an insert at the head cannot shift what "the next page" means. Give the client explicit control over when new head items enter its view (a banner it opts into, not a silent reflow), and dedupe every incoming page against a client-side set of already-seen IDs so a race between a live push and a page fetch can never render the same item twice. This generalizes past feeds to any high-churn ordered list built on top of a mutable store.
Why offset-based paging breaks here
Offset pagination asks "give me the rows starting at position N," where position is counted fresh on every request. If a new row lands at the head between two calls, every existing row's position shifts by one: the client's next request for "position N" now returns a row it already saw (duplicate) while the row that used to sit at position N slides out of view entirely (skip). Cursor pagination instead asks "give me the rows after this specific row," so a head insert changes nothing about what "after this row" means.
| Aspect | Offset pagination (?page=3&size=20) | Cursor pagination (?cursor=<opaque>&limit=20) |
|---|---|---|
| What "position" means | The Nth row counted fresh on this call | Everything ordered after one specific, named row |
| Effect of a head insert | Shifts every row's position, causing duplicates or skips | No effect: the cursor still names the same row |
| Client must understand internals | Yes, page number is a count | No, the cursor is opaque |
| Typical failure mode under churn | Duplicate or missing items mid-scroll | Only fails if the anchor row itself is later deleted |
Cursor contract and handling head inserts
- The cursor encodes the last-seen row's ordering key (for example
created_atplusidas a tie-breaker) and is treated as opaque by the client. - Real-time inserts arrive over a separate channel (stream or websocket) carrying the full item payload. The client never merges them straight into the currently rendered list; it buffers them and shows a "N new" affordance the user can tap.
- When the client does fetch the next page, it filters incoming rows against a local set of IDs it has already rendered (from either the page fetch or the live stream) before appending, so a row delivered by both paths only renders once.
- If the anchor row referenced by a cursor is later deleted, the server cannot resolve "after this row" any more. Signal that explicitly (for example a
410 Gone-style error on that specific cursor) so the client knows to discard its cursor and reload the current view, rather than the server silently guessing a nearby row.
Worked example
Assume feed posts are labeled by creation order, oldest to newest: p1 ... p9 already exist when the client loads page one, and p10 arrives from the live stream while the client is browsing.
sequenceDiagram
participant C as Client
participant F as Feed API
participant R as Realtime stream
C->>F: GET /feed?cursor=null&limit=2
F-->>C: items=[p9,p8], next_cursor=after:p8
R-->>C: push new item p10 (inserted at head)
Note over C: buffer p10, show "1 new" banner, do not reorder current page
C->>F: GET /feed?cursor=after:p8&limit=2
F-->>C: items=[p7,p6], next_cursor=after:p6
Note over C: dedupe against seen-id set before appending
p10's arrival never touches the cursor after:p8: it still means exactly what it meant before the insert, so the second page correctly returns [p7, p6] with no repeat of p9 or p8 and no skip past p7.
Trade-offs and pitfalls
Cursor pagination trades away random access (you cannot jump straight to "page 7") for stability under churn, which is almost always the right trade for a live feed. A common mistake is auto-merging live-pushed items straight into the rendered list the instant they arrive: this yanks the scroll position under the user's finger and is a worse experience than a controllable banner, even though it feels "more real-time."
If filters or sort order change mid-session, the old cursor and seen-ID set are no longer meaningful and must be discarded, not reused; treat that as a fresh pagination session. For most feed use cases, accepting eventual consistency (the client may briefly be a few items behind, reconciled by the next page fetch or a manual refresh) is the right call over paying for a strict per-client consistent snapshot, which adds real operational cost for a benefit most users never notice.
Design an API and server-side protocol for resumable uploads supporting 10k concurrent clients, files up to 200MB, and intermittent mobile connectivity. Specify endpoints for initiating an upload, uploading chunks, resuming, validating integrity, and finalizing. Explain what server state you would persist, how idempotency applies at each step, and how you would handle abandoned uploads and abuse.
Sample Answer
Direct answer
Model the upload as a stateful resource, an upload session with its own id, and make idempotency work at two different levels: session creation is protected by a client-supplied Idempotency-Key so a retried "start upload" call cannot spin up two competing sessions for the same file, and each chunk is naturally idempotent because it is addressed by its byte range plus a checksum, so re-sending the same range is a safe no-op rather than a duplicate write. Everything the server needs to persist is small: the session's metadata and the set of byte ranges already received; the file bytes themselves live in object storage, not in the session record.
Structured elaboration
Endpoints:
POST /uploads(initiate). Body:filename,total_size,content_type. Headers:Idempotency-Key(recommended). Response201:{ upload_id, expires_at, chunk_size_hint }.PATCH /uploads/{upload_id}(upload a chunk). Headers:Content-Range: bytes {start}-{end}/{total}, a chunk checksum header. Response200:{ received_ranges: [...], next_offset }. Accepts chunks out of order.GET /uploads/{upload_id}(resume). Returnsreceived_rangesandnext_offsetso a reconnecting mobile client knows exactly what is missing.POST /uploads/{upload_id}/complete(finalize). Body: client-computed whole-file checksum. Server assembles/validates and returns the final file id.DELETE /uploads/{upload_id}(abandon explicitly).
Server state persisted:
- Session row:
upload_id,user_id,filename,total_size,content_type,state(initiated / in_progress / completed / aborted),created_at,last_activity_at,expires_at. - Received-ranges record: a compact list of
{start, end, checksum}entries per chunk actually stored, not the chunk bytes themselves (those go straight to object storage). - Nothing here requires holding an open connection between chunks; a mobile client can drop off the network for minutes and resume against the same
upload_id.
Idempotency at each step:
- Initiate:
Idempotency-Keyplus a fingerprint of(filename, total_size, content_type)maps to oneupload_idfor a bounded window (documented, e.g., 24 hours); replaying the same initiate call returns the same session instead of creating a second one. - Chunk upload: idempotent by construction through
Content-Range. If a range that was already received arrives again with a matching checksum, the server responds200and does nothing further (safe retry after a lost response). If the same range arrives with a different checksum, that is a real conflict (409), not a retry, because the client is now sending different bytes for a range it already committed. - Finalize: idempotent on
upload_id; callingcompletetwice after success returns the same final file id both times rather than re-assembling or erroring, so a client that couldn't tell whether its firstcompletecall landed can safely call it again.
Handling intermittent mobile connectivity:
- Small chunk sizes (server hints a size, e.g., a few hundred kilobytes to a few megabytes) so a dropped connection loses at most one chunk's worth of progress, not the whole upload.
GET /uploads/{upload_id}is the resume contract: the client asks what ranges are already received and only re-sends what's missing, rather than restarting from byte zero.
Abandoned uploads and abuse:
expires_aton the session (returned at initiate time so the client can see its own deadline); a background sweeper deletes sessions and their stored chunks only afterexpires_athas passed, never touching a session with recentlast_activity_at.- Per-user quotas (maximum concurrent sessions, maximum total bytes in flight) enforced at initiate time, returning
429when exceeded, the same documented-quota pattern used for any other rate-limited endpoint. total_sizeis validated against the declared 200MB ceiling at initiate time, and each chunk's checksum is verified on arrival so a corrupted or tampered chunk is rejected before it is ever assembled into a final file.
Worked example
Initiate:
// POST /uploads Idempotency-Key: up-init-9f21
{ "filename": "trip-video.mp4", "total_size": 52428800, "content_type": "video/mp4" }
// 201 Created
{ "upload_id": "up_7c14", "expires_at": "2026-07-20T09:00:00Z", "chunk_size_hint": 262144 }
Upload the first chunk (256KB, matching the hinted chunk size, so the byte range is 0 through 262143 of the 52428800-byte total):
PATCH /uploads/up_7c14
Content-Range: bytes 0-262143/52428800
// 200 OK
{ "received_ranges": [[0, 262143]], "next_offset": 262144 }
Client drops off the network, reconnects later, and asks what it still needs:
GET /uploads/up_7c14
{ "upload_id": "up_7c14", "received_ranges": [[0, 262143]], "next_offset": 262144, "total_size": 52428800 }
It resumes exactly at byte 262144, no re-upload of the first chunk. Finalize once all ranges are in:
// POST /uploads/up_7c14/complete
{ "checksum_algorithm": "sha256", "checksum": "<client-computed full-file digest>" }
// 200 OK
{ "file_id": "file_31a0", "state": "completed" }
Trade-offs and pitfalls
- Server-proxied chunk uploads (as shown) are simpler to validate per-chunk but cost server bandwidth; direct-to-object-store multipart uploads (client uploads parts straight to storage using short-lived credentials) save that bandwidth but push part-tracking and checksum bookkeeping onto a more complex client-storage handshake. For 10k concurrent clients, the bandwidth savings usually win, at the cost of that added complexity.
- Chunk size is a real trade-off: smaller chunks tolerate flaky mobile networks better (less to re-send after a drop) but add per-chunk request overhead at scale; larger chunks are more efficient per byte but riskier on an unreliable connection.
- A pitfall specific to idempotency here: treating "same range, different checksum" as a silent overwrite instead of a
409hides a real client bug (a chunk that got corrupted or a retry that picked up stale data) and can quietly assemble a corrupted final file. - Garbage-collecting by
expires_atalone, without checkinglast_activity_at, risks deleting a session that is still actively (if slowly) being uploaded to from a poor connection right at the boundary of its TTL; the sweeper should treat both signals together, the same discipline as garbage-collecting any other long-lived job resource.
Unlock Full Question Bank
Get access to all 34 API and Interface Design for Distributed Services interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.