API and Interface Design for Distributed Services Questions
Designing the contracts between services and clients: REST, gRPC, and GraphQL tradeoffs, versioning and backward compatibility, pagination, rate limiting, and idempotent endpoints. Covers request/response modeling, error contracts, and API gateway responsibilities. Focuses on the interface layer that ties distributed components together, not internal data schemas.
Design an API and backend workflow for long-running analytical queries that supports submit, query, cancel, and status endpoints, and can stream partial results. Explain how you would model job state, what you'd persist, how you'd bound concurrency per user, and how you would safely garbage-collect old jobs without affecting ones still running.
Sample Answer
Direct answer
Model each query as a durable resource with an explicit id and a small state machine, not as something the API tries to hold open. The client submits a query and gets back a job_id immediately; every other verb (status, partial results, cancel) operates on that id. Persist just enough to answer "what state is this job in and what has it produced so far" without persisting more than needed, bound concurrency at submission time as a documented quota rather than an invisible queue, and garbage-collect only jobs that are both in a terminal state and confirmed idle, never on a timer alone.
Structured elaboration
Endpoints:
POST /queriessubmits a query, returns202 Acceptedwith aquery_idand initial statequeued.GET /queries/{id}returns current state, progress, and (once available) a pointer to partial or final results.GET /queries/{id}/results?cursor=...streams results in pages; safe to call repeatedly, a pure read.POST /queries/{id}/cancelrequests cancellation; returns immediately, cancellation is best-effort, not a guarantee the job stops mid-instruction.
Job state model, as a state machine the API exposes to the caller:
stateDiagram-v2
[*] --> queued
queued --> running
running --> partial: first result page ready
partial --> partial: more pages
partial --> succeeded
running --> succeeded
running --> failed
running --> canceled: cancel accepted
partial --> canceled: cancel accepted
succeeded --> [*]
failed --> [*]
canceled --> [*]
What gets persisted (durable store, e.g. a relational table, not just in-memory):
query_id,user_id, submitted query text/params,state,progress(rows or bytes processed),created_at,last_heartbeat_at,result_manifest(list of completed result-page locations in object storage),expires_at.- Result data itself does not belong in the same store as job metadata: pages are written to object storage as the job produces them, and the job row only tracks pointers (a manifest), so the metadata store stays small and fast to query regardless of result size.
Bounding concurrency per user, as a contract, not just an internal limiter:
- At submission time, the server checks the caller's currently-
running/queuedjob count against a documented per-user limit (for example, a fixed number of concurrent queries per account tier). Over the limit, the API returns429with aRetry-Aftervalue and a body explaining which quota was hit, rather than silently queueing the request indefinitely. This gives the caller something to program against instead of guessing why the request never seems to make progress.
Streaming partial results:
- Once the query has produced at least one page, state moves to
partial, andGET /queries/{id}/resultsreturns whatever pages exist plus a cursor for the next page, the same pagination shape whether the job is still running or alreadysucceeded. A client can start rendering rows well before the whole query finishes.
Cancellation:
POST /queries/{id}/cancelis asynchronous and best-effort: the state moves to a transient "cancel-requested" condition, the worker checks for it between chunks of work and stops, and the final state the client observes could still legitimately besucceededif the job finished before it noticed the cancellation. The contract should say this plainly rather than pretend cancel is instantaneous.
Safe garbage collection without touching running jobs:
- Only terminal-state jobs (
succeeded,failed,canceled) are eligible for cleanup, and only onceexpires_athas passed ANDlast_heartbeat_atshows no recent activity, so a job that looks terminal in the metadata store but has a worker still flushing a last results page is not deleted out from under it. - Cleanup removes the object-storage pages referenced by the manifest and then the metadata row, in that order, so a crash mid-cleanup leaves an orphaned-but-harmless metadata row rather than a manifest pointing at deleted data.
expires_atis returned to the caller in the job status response, so a client polling late can see its results are about to disappear instead of getting a surprise 404.
Worked example
Submit and inspect a query:
// POST /queries
{ "sql": "SELECT region, SUM(revenue_cents) FROM orders WHERE order_date >= '2026-06-01' GROUP BY region", "priority": "normal" }
// 202 Accepted
{ "query_id": "q_5f31", "state": "queued", "submitted_at": "2026-07-18T10:00:00Z" }
// GET /queries/q_5f31 (mid-run)
{
"query_id": "q_5f31",
"state": "partial",
"progress_pct": 55,
"result_pages_ready": 2,
"expires_at": "2026-07-20T10:00:00Z"
}
// GET /queries/q_5f31/results?cursor=page_2
{
"rows": [
{ "region": "us-east", "revenue_cents": 48291000 },
{ "region": "us-west", "revenue_cents": 30112500 }
],
"next_cursor": "page_3",
"final": false
}
Once state becomes succeeded, the same results endpoint keeps working with final: true on the last page, so the client's polling and rendering code does not change between "still running" and "done."
Trade-offs and pitfalls
- A common mistake is coupling the results endpoint's shape to whether the job is done, forcing the client to special-case "partial vs final" results differently; keeping one paginated results contract regardless of job state avoids that branch entirely.
- Treating cancel as synchronous and guaranteed is a pitfall: a distributed worker may be mid-chunk when the cancel request arrives, so the honest contract is "cancellation requested, best-effort, check final state," not "cancelled means stopped immediately."
- Garbage collection racing a slow-to-report worker is the sharpest failure mode here: a heartbeat check in addition to a TTL is what prevents deleting a job's output while a worker is still writing to it, a pure clock-based TTL is not sufficient on its own.
- Per-user concurrency enforced only at submission time (not re-checked mid-run) is simpler but means a user who is throttled cannot "sneak in" more work by resubmitting after their first job finishes just below the limit; that is a deliberate simplicity trade-off worth naming rather than over-engineering a live-recheck.
Compare RESTful HTTP APIs and gRPC for both public-facing and internal service-to-service communication. For an internal microservices platform, explain when you would prefer one over the other, and what would change about your answer if the API needed to be called directly from a browser.
Sample Answer
Direct answer
For an internal microservices platform, prefer gRPC for service-to-service calls: its binary encoding and HTTP/2 transport give lower latency and smaller payloads at high call volumes, and its contract-first interface keeps many services in sync automatically. For anything a browser calls directly, prefer REST over plain HTTP and JSON, because a browser cannot open a native gRPC connection at all.
Framework
REST is an architectural style over HTTP that uses JSON bodies, standard HTTP methods, and standard status codes. gRPC is a contract-first remote procedure call framework that runs over HTTP/2 and encodes messages with Protocol Buffers (protobuf), a compact binary serialization format defined by a .proto schema file that generates client and server code in multiple languages.
| Dimension | REST (HTTP + JSON) | gRPC (HTTP/2 + protobuf) |
|---|---|---|
| Payload | Text (JSON), larger, human-readable | Binary (protobuf), smaller, not human-readable on the wire |
| Transport | HTTP/1.1 or HTTP/2, typically one request per response | HTTP/2 only, many requests multiplexed over one connection |
| Streaming | Bolted on via WebSockets or Server-Sent Events | Native client, server, and bidirectional streaming |
| Contract | Loose, documented separately (for example an OpenAPI spec) and easy to drift from the code | Strict: a .proto file generates both sides' code, harder to drift silently |
| Browser support | Native (fetch, XMLHttpRequest) | Not native; needs gRPC-Web plus a translating proxy |
| Caching and intermediaries | Works with standard HTTP caches, CDNs, and load balancers out of the box | Binary framing and multiplexing make generic HTTP caching not applicable |
| Best fit | Public APIs, browser clients, wide interoperability | Internal, high-volume, latency-sensitive service-to-service calls |
For the internal platform: gRPC, because service-to-service calls are usually high in volume, latency-sensitive, and made by services the team already controls, so the cost of generating and deploying stubs from a shared .proto file pays off, and the smaller binary payload plus multiplexed streams cut both bandwidth and per-connection overhead compared to many REST connections.
What changes for a browser: a browser's networking stack cannot speak native gRPC, because it cannot control HTTP/2 framing and trailers the way a gRPC client library does. Two ways to still reach the same backend from a browser: put a REST/JSON facade in front of the gRPC service (a translation layer sometimes called a backend for frontend), or use gRPC-Web, a variant browsers can speak that still needs a proxy in front of the gRPC service (commonly Envoy) to translate wire formats, and that does not support the full bidirectional streaming native gRPC does. In practice, most teams pick REST/JSON for anything a browser calls directly and keep gRPC strictly behind that boundary.
A third style, GraphQL, is not really the comparison this question is asking for. It trades either extreme for a single flexible query language over HTTP, and it would be the right answer if the real problem were a client needing to shape its own response across many resources at once, not the internal-platform protocol choice being asked about here.
Worked example
Suppose an internal order service calls an inventory service 50 times per incoming user request, once per line item, to check stock. With REST/JSON, each call opens or reuses an HTTP connection and pays JSON parsing cost per call. With gRPC, all 50 calls can multiplex over a single HTTP/2 connection to the inventory service, and each message is a compact protobuf encoding instead of a JSON object with repeated field-name strings, so the marginal cost per call is lower at that fan-out.
If that same platform later needs to expose an inventory check directly to a web storefront, the storefront calls a REST endpoint instead:
GET /v1/inventory/A1
200 OK
{ "sku": "A1", "in_stock": 42 }
The storefront is never handed a .proto file or a gRPC client; it gets ordinary JSON over HTTP.
Trade-offs and pitfalls
- gRPC's stricter schema is also a cost: every field change means regenerating and redeploying stubs across every consuming service, whereas REST/JSON tolerates an unexpected extra field without anyone regenerating anything.
- A common pitfall is exposing gRPC-Web directly to third-party public API consumers to save engineering time. It forces every external integrator to adopt protobuf tooling and a compatible proxy, which is a worse experience than JSON for most public consumers.
- Debugging gRPC's binary frames is not as simple as opening a browser's network tab or running curl the way it is with JSON; teams that skip investing in gRPC-aware tracing early often regret it once they have many services talking to each other.
Design API endpoints and backend read models for an analytics dashboard that needs aggregated counts and top-k lists, without pushing heavy aggregation onto the client. Discuss the trade-offs between freshness, storage cost, and query complexity in how you shape those endpoints.
Sample Answer
Expose two endpoint shapes, one for time-bucketed aggregated counts and one for top-k lists, and never let the client request an ungoverned ad-hoc aggregation. Every response carries an explicit freshness field (an asOf timestamp and a precision flag of exact or approximate) so the client-visible contract, not just the internal pipeline, states what freshness the caller is actually getting; the backend may route a request to a fast precomputed view or a slower exact path, but that routing choice is invisible to the caller except through this field.
Endpoints
GET /metrics/counts?metric=events&start=...&end=...&granularity=hour&dim=region: a time series of pre-aggregated buckets.GET /metrics/top?metric=clicks&k=10&start=...&end=...&dim=category: a ranked top-k list.- Both accept a
freshnessquery parameter as a request-side contract:freshness=fast(accept approximate or slightly stale data, get the lowest latency) versusfreshness=exact(route to an exact recompute, higher latency, and the contract documents a range cap or async fallback rather than letting the request hang indefinitely for a huge range).
The staleness contract, in the response itself
Every response includes asOf (the timestamp through which data is complete) and precision. This is the part of the design that belongs squarely to an API-contract discussion: it turns an internal trade-off (how fresh is the backing view) into an explicit, documented, client-visible field, so a caller building a dashboard can show an "as of 2 minutes ago" badge instead of silently trusting a number that might be stale.
| Client-visible option | What the caller gets | What it costs to offer |
|---|---|---|
freshness=fast (default) | Sub-second response, precision: "approximate" for top-k, asOf typically within the last minute | Requires precomputed rollups for common ranges; the contract must cap which dim combinations are supported, since only precomputed ones can be fast |
freshness=exact | Exact counts, precision: "exact", asOf reflects true request time | Higher, less predictable latency; the contract caps the allowed date range for this mode and documents a slower service-level agreement (SLA: a documented commitment about a specific response characteristic, here latency), so callers don't assume it is always fast |
| Arbitrary ad-hoc dimension or filter | Maximum flexibility | Not offered directly: the contract restricts dim to an enumerated, indexed set and returns a documented 400 Bad Request for an unsupported dimension, rather than silently running an expensive query |
Bounding and shaping the response
k is capped server-side (for example, a maximum of 100) regardless of what the client requests, and the response documents the cap via requestedK and returnedK, so a client cannot silently receive a truncated list without knowing it was truncated. The response is already the final shape the dashboard renders, bucketed counts or a ranked list; the client never receives raw event rows and aggregates them itself. That is the same overfetch-avoidance principle as any other read-heavy endpoint: the API promises a specific, small, pre-shaped payload, not a firehose the client has to post-process.
Worked example
GET /metrics/counts?metric=events&start=2026-07-18T00:00Z&end=2026-07-18T03:00Z&granularity=hour&freshness=fast
{
"metric": "events",
"granularity": "hour",
"asOf": "2026-07-18T03:00:42Z",
"precision": "approximate",
"buckets": [
{ "start": "2026-07-18T00:00Z", "count": 48210 },
{ "start": "2026-07-18T01:00Z", "count": 51330 },
{ "start": "2026-07-18T02:00Z", "count": 49980 }
]
}
GET /metrics/top?metric=clicks&k=5&start=2026-07-18T00:00Z&end=2026-07-18T03:00Z&dim=category&freshness=fast
{
"metric": "clicks",
"asOf": "2026-07-18T03:00:42Z",
"precision": "approximate",
"requestedK": 5,
"returnedK": 5,
"items": [
{ "key": "footwear", "count": 9120 },
{ "key": "outerwear", "count": 7040 },
{ "key": "accessories", "count": 5210 },
{ "key": "electronics", "count": 4880 },
{ "key": "home", "count": 3990 }
]
}
Trade-offs and pitfalls
The biggest interview-relevant trade-off is that offering freshness=exact at all is a contract commitment: once a caller can request it, some caller eventually will, on a huge date range, and the API needs a documented, bounded answer (a hard range cap, a queued and pollable async job, or an explicit rejection) rather than an unbounded query that degrades the whole service. A common wrong turn is exposing a single endpoint that "just returns the data" with no freshness or precision field, leaving the client to guess whether a number is authoritative; once a dashboard has shown a number without a staleness caveat, a later "actually that was approximate" correction reads as the API being wrong, when the real defect was an underspecified contract. Enumerating allowed dim values, rather than accepting any column name, trades flexibility for the ability to document, cap, and index every supported query shape; a genuinely open-ended analytics need should be pointed at a dedicated data-warehouse query interface, not bolted onto this API.
An API endpoint that aggregates data from several microservices for a single client request is making one call per related item, and latency balloons as the result set grows. Walk through how you would diagnose this and redesign the aggregation approach, and how you would weigh the trade-offs between the fixes available to you.
Sample Answer
Direct answer
This is the classic N+1 fan-out: fetching a list of related items triggers one call for the list plus one additional call per item, instead of a single call that fetches all the related data at once, so latency and downstream load scale with the size of the result set rather than staying flat. The fix is to collapse the per-item calls into a single batched call, or a request-scoped loader that performs that collapsing automatically, and only reach for a precomputed denormalized store when the read path is hot enough to justify the added consistency cost.
Diagnosing it
Confirm it is really N+1 and not just "a slow downstream service" by counting calls per request, not just measuring total time: instrument the aggregation layer to log how many downstream calls a single incoming request triggers, then check whether that count grows linearly with the size of the requested list. A request for 5 related items making 5 downstream calls and a request for 200 items making 200 calls is the signature; a flat call count regardless of list size means the problem is elsewhere.
Fix patterns
| Pattern | What it does | Best when | Trade-off |
|---|---|---|---|
| Single batched call | Replace N item calls with one call carrying all the IDs | The downstream service already supports (or can add) a bulk lookup endpoint | Requires that bulk endpoint to exist |
| Request-scoped loader (a "DataLoader"-style collector) | Collects every key requested during one incoming request's lifetime, then issues one batched call for all of them | Composition-heavy code (many independent resolvers each needing the same kind of lookup, common in GraphQL-style servers) | Only helps within a single request; needs its own coordination layer |
| Concurrency-capped fan-out | Keeps per-item calls but issues only a fixed number k at once | No bulk endpoint exists and cannot be added quickly | Still N total calls and downstream load, just paced instead of instantaneous |
| Denormalized read store | Precomputes the joined view ahead of time via events or change-data-capture, CDC (a stream of row-level changes from the source database) | Read-heavy, latency-critical path where some staleness is acceptable | Eventual consistency, plus a pipeline and storage to keep the copy in sync |
| Caching the aggregated response | Caches the whole assembled response or the individual entities | Same aggregation requested repeatedly with a low underlying write rate | Needs a real invalidation strategy; a wrong invalidation reintroduces staleness bugs |
A "DataLoader" here means a request-scoped utility, popularized by GraphQL server implementations, that defers execution until every resolver in the current request has registered the key it needs, then fires one batched downstream call instead of many individual ones.
Worked example
flowchart LR
subgraph Naive["Naive: one call per item"]
G1[API Gateway] --> I1[Item 1 call]
G1 --> I2[Item 2 call]
G1 --> I3[Item N call]
end
subgraph Batched["Fixed: single batched call"]
G2[API Gateway] --> B1[Batch call: all item IDs]
end
Reasoning about the scaling, not wall-clock timing (real timings depend on hardware and network and are not reproducible from an answer alone): if a request needs data for N related items and each downstream round trip costs a fixed unit c, a fully serial naive fan-out costs N×c total round-trip time. A single batched call collapses this to a constant cost of c regardless of N. A concurrency-capped fan-out with cap k costs ⌈N/k⌉×c: for example, with N=200 items and k=20 concurrent calls, that is ⌈200/20⌉=10 rounds, versus 1 round batched or 200 rounds fully serial. The batched call is the only option whose cost does not grow with the result set at all.
Trade-offs and pitfalls
Batching changes the failure model: a single bad ID in a batch call must not fail the entire batch. The batched response needs a shape that supports partial success (per-item status alongside a shared envelope), or one missing related record takes down an otherwise-successful response for every other item in the batch.
A request-scoped loader only helps inside one request's lifetime: if the same key is requested by two separate incoming requests seconds apart, the loader does not share that saved work across them, that is what a cache is for. Reaching for a denormalized store as the first fix, before trying batching or a loader, is a common overreach: it is the most expensive option operationally (a whole pipeline to build and keep in sync) and should be reserved for paths where the simpler fixes have already been tried and the traffic genuinely justifies it.
Design a contract-testing approach that catches a breaking API change before it reaches production, given that the consuming clients are owned by different teams than the API itself. Explain how you would automate this in CI and how a team would find out their contract test failed.
Sample Answer
Direct answer
Use consumer-driven contract testing: each consuming team publishes the exact requests and expected responses it actually depends on, and the provider team's own continuous integration (CI) pipeline verifies the real API against every published consumer contract before a change can merge or deploy. This inverts the usual failure mode, where a consumer discovers a break after the provider has already shipped; here the provider's own pipeline fails first, and the consumer team is notified automatically through the same system that stores the contracts, not by someone noticing something broke in production.
Structured elaboration
The contract and where it lives:
- Each consumer team writes a small set of interactions ("when I call this endpoint with this request, I expect a response shaped like this") and publishes them to a shared, versioned contract registry (a broker), tagged with which environment or release they apply to.
- This is deliberately narrower than a full schema: the consumer only asserts the exact fields and shapes it actually reads, so the contract reflects real usage rather than the provider's entire response shape.
Automating verification in the provider's CI:
- On every pull request to the provider's API code, a CI step fetches the latest contracts tagged for the relevant environment, spins up (or points at a running instance of) the provider, replays each consumer's recorded requests against it, and asserts the actual responses still satisfy what each consumer expects. Any mismatch fails the build and blocks the merge.
- Before an actual deploy (not just before merge), a second gate checks whether the version about to be deployed is compatible with whichever consumer versions are currently running in production, not just the newest consumer contract, since a provider change can be compatible with a consumer's latest contract while still breaking a consumer instance that hasn't picked up its own latest release yet.
How the affected team finds out:
- The contract broker's webhook posts directly to the consumer team's own channel (chat integration or issue tracker) whenever a provider verification against their specific contract fails, naming the exact interaction that broke and linking to a diff between what was expected and what the provider now returns. The team learns about the break from an automated, targeted message tied to their own contract, not from a shared build dashboard they'd have to be watching.
Worked example
A consumer's published contract (one interaction, in a contract-broker style format):
{
"consumer": { "name": "checkout-web" },
"provider": { "name": "orders-api" },
"interactions": [
{
"description": "fetching an order by id",
"request": { "method": "GET", "path": "/orders/500" },
"response": {
"status": 200,
"body": { "id": "500", "status": "PAID", "total_cents": 4200 }
}
}
]
}
Provider CI step (illustrative, tool-agnostic pseudocode) that fails the build on a mismatch:
verify-contracts:
script:
- fetch-contracts --provider orders-api --tag production
- contract-verify --provider-base-url http://localhost:8080 --contracts ./contracts/*.json
# non-zero exit code on any interaction mismatch blocks the merge
Suppose a developer on the provider team renames total_cents to totalCents in a refactor. The verification step replays the recorded request, gets back a body missing the field the consumer's contract asserted, and the build fails with a message like: "interaction 'fetching an order by id' failed: expected field total_cents, field not present in actual response." The broker then posts that same message, with a link to the failing interaction, straight to the checkout-web team's channel, so they know about the incompatibility before it ever reaches a shared environment.
Trade-offs and pitfalls
- Contract testing only catches what a consumer actually asserted. If a consumer's published contract never checked the
total_centsfield in the first place, removing it silently passes; the technique is only as good as how faithfully each contract reflects real usage, which puts real responsibility on consumer teams to keep their contracts current as their own code changes. - Verifying against the latest contract at merge time is necessary but not sufficient; gating the actual deploy against what is compatible with consumer versions currently running in production (not just their newest published contract) is what catches the case where two independently-safe changes combine badly because of deploy ordering.
- This approach is faster and cheaper than full end-to-end integration testing, but it is not a replacement for it: contract tests verify shape and basic behavior, not real-world concerns like load, timing, or business-logic correctness across a full user flow.
- A pitfall on the process side: if publishing and updating contracts is treated as optional or bolted on late, teams tend to let contracts drift stale, at which point the safety net silently degrades and nobody notices until a real break gets through anyway.
Unlock Full Question Bank
Get access to all 35 API and Interface Design for Distributed Services interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.