API and Interface Design for Distributed Services Questions
Designing the contracts between services and clients: REST, gRPC, and GraphQL tradeoffs, versioning and backward compatibility, pagination, rate limiting, and idempotent endpoints. Covers request/response modeling, error contracts, and API gateway responsibilities. Focuses on the interface layer that ties distributed components together, not internal data schemas.
Compare cursor-based pagination with offset-based pagination for APIs. For a frequently-updated feed with a high insert rate at the head, explain which approach you would pick and why. Provide a sample response shape for a cursor-based page and mention how you would encode and expire cursors.
Sample Answer
Direct answer
For a feed with a high insert rate at the head, pick cursor-based (keyset) pagination. It anchors each page to the last item's sort key instead of a row count, so it does not skip or repeat rows when new items appear ahead of the client's read position. Offset-based pagination (OFFSET n LIMIT m) is simpler but silently shifts under concurrent inserts, which is exactly what a frequently-updated feed does continuously.
Framework
What each approach actually does
- Offset pagination: the client asks for "skip N, take M." The server re-runs the ordering on every request and counts N rows in before returning results. A page's identity is a position in a list, and that position can move.
- Cursor (keyset) pagination: the client sends the sort-key value(s) of the last item it saw. The server issues a seek query (
WHERE (created_at, id) < (:last_created_at, :last_id) ORDER BY created_at DESC, id DESC LIMIT M). A page's identity is anchored to a specific row's key, not a position.
| Dimension | Offset (OFFSET n LIMIT m) | Cursor / keyset |
|---|---|---|
| Correctness under head inserts | Breaks: new rows shift every row's position, causing duplicates or skipped rows on the next page | Stable: the seek predicate is relative to a fixed key and is unaffected by inserts elsewhere in the ordering |
| Query cost at depth | Grows with the offset (the database still has to skip N rows before returning M) | Roughly constant, an index seek to the key followed by M rows, no matter how deep the client has paged |
| Jump to an arbitrary page | Trivial (OFFSET 4000) | Not supported directly; you can only move forward or backward from a known cursor |
| Implementation complexity | Low | Higher: needs a stable, unique sort key (usually a tie-breaker column) and a cursor encoding scheme |
| Best fit | Small, mostly-static result sets, admin UIs that want page numbers | Feeds, timelines, anything with concurrent writes or deep paging |
Sorting by created_at alone is not enough if two rows can share a timestamp. The seek predicate needs a secondary, unique tie-breaker column (commonly id) so (created_at, id) is a total order with no ties.
Worked example
A concrete trace of why offset duplicates a row when an item is inserted at the head, and why cursor pagination does not.
Five existing rows, ordered newest-first by created_at:
| id | created_at |
|---|---|
| 105 | 2026-01-01T10:04:00Z |
| 104 | 2026-01-01T10:03:00Z |
| 103 | 2026-01-01T10:02:00Z |
| 102 | 2026-01-01T10:01:00Z |
| 101 | 2026-01-01T10:00:00Z |
Offset pagination, page size 2:
- Page 1:
OFFSET 0 LIMIT 2returns positions 1-2:[105, 104]. - A new row
106(created_at2026-01-01T10:05:00Z) is inserted at the head. Every existing row shifts down one position:105is now position 2,104is position 3, and so on. - Page 2:
OFFSET 2 LIMIT 2returns whatever now sits at positions 3-4, which is[104, 103]. The client already saw104on page 1, so it is duplicated, and the client has no way to notice.
Cursor pagination, same scenario:
- Page 1: no cursor supplied, returns
[105, 104];next_cursorencodes{"id": "104", "created_at": "2026-01-01T10:03:00Z"}. - Row
106is inserted at the head. - Page 2:
WHERE (created_at, id) < ('2026-01-01T10:03:00Z', 104) ORDER BY created_at DESC, id DESC LIMIT 2. Row106'screated_atis greater than104's, so it fails the<predicate and is correctly excluded. The query still returns[103, 102], exactly the rows that follow104, regardless of what was inserted ahead of it.
Sample cursor-based response shape (page 1 above):
{
"items": [
{ "id": "105", "created_at": "2026-01-01T10:04:00Z" },
{ "id": "104", "created_at": "2026-01-01T10:03:00Z" }
],
"next_cursor": "eyJpZCI6IjEwNCIsImNyZWF0ZWRfYXQiOiIyMDI2LTAxLTAxVDEwOjAzOjAwWiJ9.f0e3e3b16f56ab33",
"has_more": true
}
How that cursor was actually built (reproducible, step by step):
- Payload:
{"id": "104", "created_at": "2026-01-01T10:03:00Z"}. - Compact JSON serialization:
{"id":"104","created_at":"2026-01-01T10:03:00Z"}. - Base64url-encode the bytes, no padding:
eyJpZCI6IjEwNCIsImNyZWF0ZWRfYXQiOiIyMDI2LTAxLTAxVDEwOjAzOjAwWiJ9. - Sign the raw JSON bytes with a server-side hash-based message authentication code (HMAC-SHA256) using a secret only the server holds, and take the first 16 hex characters of the digest as a tamper check:
f0e3e3b16f56ab33. - Join encoded payload and signature with a period:
<base64url>.<signature>.
A production cursor payload would also carry an issued-at timestamp and a time-to-live (TTL): the server checks issued-at plus TTL on every request and, if the cursor has expired, rejects it (for example with a 400 response) and asks the client to restart from the first page, rather than trying to resume mid-sequence from stale coordinates.
Trade-offs and pitfalls
- Cursor pagination cannot jump to an arbitrary page. If the product genuinely needs "go to page 43," you need offset, or a hybrid where a search step lands the client near a position and cursor paging takes over from there.
- If the tie-breaker column is not unique, or is not indexed together with the primary sort key, the seek query degrades toward a scan instead of an index lookup.
- Signing the cursor matters for more than tamper-proofing: an unsigned or unencrypted cursor can leak internal ids and ordering, and a forged cursor could be used to skip access checks that were only applied on page 1.
- A common pitfall is anchoring the cursor to a mutable field, such as an "updated at" timestamp instead of creation time. An edit to any row changes its position under that ordering, so items can be revisited or skipped after an edit, not just after an insert. Prefer an immutable creation-time key plus a stable id tie-breaker.
A public API needs a new required field added to its request schema, and it's used by many client versions whose release cadence you don't control. Walk through how you would roll this out with zero downtime and minimal disruption to existing clients, including how you would eventually know it's actually safe to enforce the field as required.
Sample Answer
Direct answer
Ship the server change so it accepts requests both with and without the new field, defaulting it server-side when absent, so nothing breaks the moment the change deploys. Roll enforcement in behind a flag keyed by caller or client version, expanding gradually while watching a "requests missing this field" metric drop toward zero, and only flip the field to actually required once that metric shows it is safe for the population of clients you cannot update directly, not on a fixed calendar date.
Framework
Phase 0: ship tolerant, observe only
Deploy server code that accepts the field as optional, applies a sensible default when it is missing, and logs, tagged by caller and client version, every request that arrives without it. Nothing is rejected in this phase. This is what makes the eventual rollout zero-downtime: the deploy that adds validation never also has to add the field, because those are two separate deploys.
Phases 1-3: progressive enforcement
Behind a server-side feature flag keyed by caller id or client version, start rejecting requests missing the field for a small, known-safe slice of traffic first, for example your own internal callers, who you can fix immediately if they break. Widen the slice in stages, a percentage of external traffic, then all traffic from clients above a known-good version, watching the missing-field rate and the resulting error rate at each stage before widening further.
Phase 4: flip to required
Once the missing-field rate across all remaining traffic is at, or has been driven to, zero, or the only remaining non-compliant callers are ones identified by name who have either been fixed directly or have explicitly accepted the breakage, remove the tolerant path and the flag. The field is now genuinely required, not just documented as required while the server quietly still accepts its absence.
Worked example
Concretely: at Phase 0, telemetry shows 12% of requests lack the new field. Enabling strict validation for 1% of traffic, Phase 1, internal callers only, shows 0% errors, since internal callers were already updated first. Widening to 25% of external traffic, Phase 2, surfaces a 2% error rate concentrated in exactly 3 identifiable partner API keys still on an old integration version; those 3 keys are contacted directly and updated. Phase 3 widens to 100% traffic and the error rate returns to 0%, confirming it is now safe to make Phase 4's flip permanent and delete the tolerant path.
Rollback plan, defined before the rollout starts, not improvised during an incident: the fastest rollback is flipping the enforcement flag back to "accept missing field" for the affected slice, which takes effect without a new deployment. If the underlying server change itself is the problem, not just the validation strictness, redeploy the previous server version behind the same gradual-traffic mechanism used to roll it forward, rather than an all-at-once revert.
Trade-offs and pitfalls
- Splitting "accept the field" and "require the field" into two separate deploys costs an extra release cycle compared to shipping both at once, but it is what makes the rollout genuinely zero-downtime instead of a bet that every client updates before a single deploy lands.
- A common pitfall is widening the enforcement percentage on a fixed schedule, day 1 at 10%, day 7 at 50%, day 14 at 100%, instead of on the observed missing-field rate at each stage. A schedule that ignores what the data is saying can either drag on longer than necessary, or worse, widen past a real problem simply because a date arrived.
- Another pitfall is treating "0% errors so far" as proof of safety when the current traffic slice does not yet include the population most likely to be non-compliant, batch jobs that run monthly, or mobile clients on very old versions most users have already updated away from. Widen enforcement along dimensions that specifically include the slowest-moving clients before declaring the rollout complete.
Your microservices communicate using protobuf messages. Describe a safe process for evolving that schema over time: how you would add fields, remove fields, and handle a change that would otherwise break existing consumers, including the rollout sequence you would follow.
Sample Answer
Direct answer
Treat every protobuf message type as append-only: new fields get new, never-reused field numbers with sensible defaults, retiring fields get marked reserved (never deleted or renumbered) only after every consumer has stopped reading them, and any change that cannot be expressed additively goes through a new versioned message or endpoint instead of mutating the existing one in place.
Safe changes vs unsafe changes
| Change | Safe? | Why |
|---|---|---|
| Add a new field with a fresh field number | Yes | Old binaries silently ignore unknown field numbers; new binaries treat an unset field as its default until it is populated |
| Remove a field, then later reuse its number for something new | No | A newer message reusing field 3 for a different meaning looks like valid old data to a stale reader: silent data corruption, not a crash |
| Rename a field | Safe on the wire, risky elsewhere | The wire format only encodes field numbers, not names, but generated code and any JSON representation change, breaking anything that keyed off the old name |
Change a field's type (for example int32 to string) | No | The wire encoding differs by type; old readers misinterpret the bytes |
Change optional to repeated (or back) | No | The wire representation differs between the two |
Rollout sequence
- Add the new field to the
.protofile with a fresh field number; document it, but do not remove anything yet. - Ship the producer so it emits both the old and new fields simultaneously (a brief dual-write window).
- Migrate consumers to read the new field, verified through automated compatibility checks and a canary rollout before full deployment.
- Once telemetry shows zero consumers still reading the old field, stop populating it in the producer.
- After a sunset window with no incidents, mark the retired field
reservedin the schema so its number and name can never be reused by accident.
Worked example
// v1: original schema
message Order {
string order_id = 1;
string customer_id = 2;
int32 quantity = 3;
}
Step 2, additive change while quantity is being retired in favor of a field that supports multiple line items:
// v2: additive only, nothing removed yet
message Order {
string order_id = 1;
string customer_id = 2;
int32 quantity = 3 [deprecated = true]; // still populated during the migration window
string shipping_address = 4; // new field, new number
repeated int32 line_quantities = 5; // replaces "quantity" semantics going forward
}
Step 5, only after monitoring confirms field 3 has zero remaining readers:
// v3: retiring field 3, its number and name can never be reused
message Order {
reserved 3;
reserved "quantity";
string order_id = 1;
string customer_id = 2;
string shipping_address = 4;
repeated int32 line_quantities = 5;
}
reserved 3 and reserved "quantity" together stop the compiler from letting anyone accidentally reassign that number or that name to a future field, which is exactly the mistake row two of the table above warns against.
Tooling and CI safety net
Add a schema-compatibility check to continuous integration (CI, the automated pipeline that builds and tests every change before merge), such as buf breaking from the buf command-line tool, configured to fail the pull request if a diff removes a field number, changes a type, or otherwise breaks wire compatibility with the previous committed schema. This turns "did we just ship a breaking change" from a question answered in an incident review into one answered before merge.
Trade-offs and pitfalls
The dual-write, dual-read window costs real code complexity and, if the old field is expensive to keep populating, some resource cost too; it should have an explicit end date, not run indefinitely. Enum evolution deserves the same discipline: adding new enum values is safe only if every consumer treats an unrecognized value as its unspecified default rather than crashing, which is why proto3 conventions reserve zero as that default case.
Avro's evolution model differs in a way worth naming if the sub-area comes up: Avro compatibility is enforced by a schema registry comparing reader and writer schemas directly, and it requires an explicit default value for any field being added or removed to stay compatible, rather than protobuf's per-field-number independence. Confusing the two models (assuming Avro's "just add a default" rule applies to protobuf's numbered-field model, or vice versa) is a common mistake when a team works with both formats.
You are seeing intermittent duplicates and missing items in a paginated transaction list shown to clients, and the backend has concurrent writes. Describe a troubleshooting plan: which logs, traces, and client telemetry you would collect; synthetic tests to reproduce it; short-term mitigations you could ship quickly; and the long-term fix.
Sample Answer
Direct answer
Intermittent duplicates and missing rows in a paginated list under concurrent writes are almost always a symptom of a pagination contract that counts position (offset) rather than anchoring to a specific row's identity, reading a list that keeps moving underneath it. The fix path is: collect telemetry that lets you correlate a client's exact requests against what changed on the backend, reproduce it deterministically with a synthetic test rather than chasing it live, ship a contained mitigation, then replace the contract itself.
Telemetry to collect
- Per-request: the exact pagination parameters used (cursor or offset, limit), a request ID, and the item IDs actually returned, logged on both client and server so the two logs can be joined.
- Per-write: a monotonic write sequence number or commit timestamp for every insert, update, or delete against the transaction table, so you can reconstruct exactly which writes landed between two page fetches.
- Aggregate dashboards: rate of duplicate-ID and gap incidents (a page that skips an ID present in a request one page earlier) bucketed by client, endpoint, and page size, so you can tell whether the rate correlates with write volume or a specific client behavior (e.g. slow scrolling, background tab refetches).
Synthetic reproduction
The fastest way to confirm the hypothesis is a small, fully deterministic script rather than a live capture: insert a row between two page fetches and check whether the second fetch repeats or skips an item. This is directly runnable and needs no production access:
class Store:
def __init__(self, items):
self.items = items # newest first
def insert_at_head(self, item):
self.items.insert(0, item)
def page_by_offset(self, offset, limit):
return self.items[offset: offset + limit]
def page_by_cursor(self, after_id, limit):
start = 0 if after_id is None else self.items.index(after_id) + 1
return self.items[start: start + limit]
store = Store(["txn-5", "txn-4", "txn-3", "txn-2", "txn-1"])
page1 = store.page_by_offset(0, 2) # fetch page 1 before the insert
store.insert_at_head("txn-6") # concurrent write lands
page2 = store.page_by_offset(2, 2) # fetch page 2 after the insert
print("OFFSET page1:", page1)
print("OFFSET page2:", page2)
print("txn-4 in both pages:", "txn-4" in page1 and "txn-4" in page2)
print("txn-2 skipped entirely:", "txn-2" not in page1 and "txn-2" not in page2)
Running this prints OFFSET page1: ['txn-5', 'txn-4'] and OFFSET page2: ['txn-4', 'txn-3'], with both follow-up checks printing True: txn-4 appears in both pages (duplicate) and txn-2 never appears in either page (skipped entirely), reproduced with a handful of lines and no timing dependency.
Re-running the same scenario against a cursor-anchored implementation instead:
store2 = Store(["txn-5", "txn-4", "txn-3", "txn-2", "txn-1"])
page1c = store2.page_by_cursor(None, 2)
store2.insert_at_head("txn-6")
page2c = store2.page_by_cursor(page1c[-1], 2)
print("CURSOR page1:", page1c)
print("CURSOR page2:", page2c)
print("no overlap between pages:", set(page1c).isdisjoint(page2c))
This prints CURSOR page1: ['txn-5', 'txn-4'], CURSOR page2: ['txn-3', 'txn-2'], and no overlap between pages: True. This is the reproducible evidence that anchors the fix, not a guess.
Short-term mitigation vs long-term fix
| Timeframe | Action | What it buys | What it does not fix |
|---|---|---|---|
| Immediate (same day) | Shrink page size; log cursor/offset plus returned IDs on every request | Smaller inconsistency window, faster correlation when a user reports it | Root cause is untouched |
| Short-term (days) | Client-side de-dup by transaction ID; detect a gap and re-fetch the missing range | Masks the user-visible symptom quickly | Adds client complexity, still fragile under heavier write load |
| Long-term (weeks) | Replace offset with a cursor anchored on a deterministic key such as (commit_ts, id) | Removes the failure mode structurally, matches the pattern in the previous scenario's fix | Requires every consumer to migrate off page-number semantics |
Trade-offs and pitfalls
It is tempting to chase this as a backend consistency problem (replica lag, read quorum) rather than a pagination contract problem. If the storage layer is genuinely eventually consistent, that is worth knowing, but the API-level fix here does not require tuning consensus or replication internals: it requires the endpoint to commit to an explicit, anchored ordering contract and, if staleness is possible, to say so in the response (for example a snapshot or version marker the client can compare) so the client can detect it rather than silently rendering wrong data. Conflating "our replication is laggy" with "our pagination contract is unanchored" leads teams to spend weeks on the wrong layer.
A related pitfall: fixing this only on the read path while leaving deletes unhandled. A cursor anchored on a row that gets deleted between pages needs an explicit signal (do not silently skip past it or error opaquely) so the client knows to discard that cursor rather than retry it forever.
Design a REST API for listing and creating 'products' that supports pagination, filtering, sorting, and versioning. Specify the request/response shapes, your pagination strategy, your versioning approach, and how you would roll out a breaking change to both internal and external clients without a hard cutover.
Sample Answer
Direct answer
Model the collection with a resource-oriented URL (GET/POST /v1/products), cursor pagination for listing, structured filter and sort query parameters, and a URI-path major version (/v1/). To roll out a breaking change without a hard cutover, let the old and new response shapes run side by side under different version paths, keep the new field additive first, and retire the old version only once usage telemetry shows it is safe, not on a calendar date alone.
Framework
Endpoints and shapes
| Method & path | Purpose |
|---|---|
GET /v1/products?limit=20&after=<cursor>&sort=price:asc&status=active&category=tools | List products, paginated, filtered, sorted |
POST /v1/products | Create a product |
GET /v1/products/{id} | Fetch one product |
PATCH /v1/products/{id} | Partial update |
GET /v1/products?limit=2&sort=price:asc
200 OK
{
"data": [
{ "id": "p1", "name": "Widget", "price": 9.99, "status": "active" },
{ "id": "p2", "name": "Gadget", "price": 14.50, "status": "active" }
],
"page": { "limit": 2, "next_cursor": "<opaque-cursor-for-p2>", "has_more": true }
}
POST /v1/products
{ "name": "New Widget", "price": 19.99, "category": "tools" }
201 Created
Location: /v1/products/p3
{ "id": "p3", "name": "New Widget", "price": 19.99, "category": "tools", "status": "active" }
Pagination strategy: cursor-based, for the same reason as any list that can grow and be written to concurrently: stable under inserts, roughly constant query cost at depth, at the cost of not supporting a direct "jump to page 43." Offer limit and an opaque after cursor rather than page/offset.
Filtering and sorting: plain query parameters for common fields (status=active, category=tools), and a single sort=field:direction parameter, comma-separated for multiple fields, for example sort=price:asc,name:desc. The filterable field vocabulary is documented per field in the OpenAPI spec (a machine-readable API description) so clients know exactly what is filterable rather than guessing.
Versioning: the major version lives in the URL path (/v1/), because it is simple to route at a gateway or load balancer without inspecting headers, and it is trivially cacheable per version. Additive changes, new optional fields, new optional query parameters, ship inside v1 without a version bump; anything that removes a field, changes a field's type, or changes default sort or filter behavior goes into v2.
Rolling out a breaking change without a hard cutover: suppose v2 needs to rename price to unit_price and change it from a decimal amount to an integer number of cents.
- Ship
v2alongsidev1on the same deployment.v2's handler reads and writes the same underlying data asv1, so there is exactly one source of truth behind two response shapes. - Mark
v1'spricefield deprecated via a response header and a documented sunset date, while it keeps working exactly as before. Nothing breaks yet. - Instrument both versions: tag every request with which version served it, and track the fraction of traffic still hitting
v1. - Migrate internal clients first, since you can coordinate with them directly, then notify external and partner clients with a fixed migration window, for example 90 days, pointing at the
v2docs and a short code sample. - Remove
v1only once telemetry shows its remaining traffic has dropped to a level you have explicitly decided is safe to force-migrate, for example only a handful of clients you can contact individually, not on the calendar date alone. If a significant client is still onv1at the deadline, extend the window rather than break them, and treat that as a signal the migration tooling or communication needs work.
Worked example
Concretely, at the moment v2 ships, v1 carries 100% of traffic. After internal clients migrate in the first month, v1's share drops to 70%. Announcing the 90-day window to external partners brings it down further; by day 60, v1 is at 8%, all from three named partner integrations already contacted directly about their remaining migration steps. That 8%-and-named state, not the passage of 60 days on its own, is what tells you it is close to safe to set a hard removal date, once those three integrations confirm.
Trade-offs and pitfalls
- Renaming
pricetounit_priceand changing its unit, dollars to cents, in the same release conflates two changes into one migration. A client that only cared about the rename still has to handle the unit change, which raises the odds the migration is done wrong. Prefer landing one breaking change at a time when volume allows it. - A common pitfall is announcing deprecation only in documentation and not in the response itself. Clients that never read a changelog will not notice until the sunset date arrives, so a machine-readable response header, not prose, is what actually drives safe removal.
- Cursor pagination combined with
sortneeds care: the opaque cursor usually encodes the sort key's value, so changingsortmid-pagination, fetching page 1 by price then asking for page 2 by name, should be rejected or restarted from page 1, since a cursor from one sort order is meaningless under a different one.
Unlock Full Question Bank
Get access to all 35 API and Interface Design for Distributed Services interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.