Backend-for-Frontend and Client-Specific API Design Questions
Shaping API responses to the needs of specific client types: the Backend-for-Frontend (BFF) pattern, one BFF per client versus a shared gateway, client-side composition or a generic API, what belongs in a BFF (including normalizing inconsistent downstream pagination and error formats for one client), and aggregating several backend calls into one client-shaped response. Covers payload shaping (sparse fieldsets and field selection against relational and document stores, partial responses, optional expansion of related collections, delta responses with deletions, compact serialization including binary wire formats), round-trip reduction and composite views for feeds and dashboards, adapting to mobile versus web bandwidth, battery and device constraints (device-appropriate image variants, prefetching and refresh frequency tuned to connection quality), and choosing GraphQL versus REST to fix over- and under-fetching, including incremental migration from REST to GraphQL. Authentication, general REST resource design, versioning, rate limiting, offline sync and real-time transport are covered elsewhere.
Three internal microservices each use a different auth scheme, pagination style and error format, and your web app has to render screens from all of them. Design the Backend-for-Frontend that sits in front. What belongs in it, what should stay out of it, and when would you rather normalize in the frontend instead?
Sample Answer
Direct answer
The BFF (Backend-for-Frontend, a server owned by the web team whose only customer is the web app) should absorb the differences between the services: authentication to each upstream, pagination, error format, timeouts and partial failure, plus the joining and trimming that one screen needs. It should not hold business rules, a database of record, or presentation concerns such as button labels. Normalize in the frontend only when the data comes from a single service, needs no secret, and is already shaped the way the UI wants it.
The three services in this scenario are a concrete example of the BFF's job: service A uses a bearer token (a credential sent in the Authorization: Bearer <token> header), offset pagination and {error: "..."}; service B uses an API key header, cursor pagination and {errors: [{code, detail}]}; service C uses a session cookie (a service session the BFF holds), page numbers and plain-text errors.
The three pagination styles ask for "the next page" in three different ways. Using the same lists as the runnable skeleton below:
| Style | Request for the next page | How the server says there is more |
|---|---|---|
| Offset (skip this many items) | GET /orders?offset=3&limit=3 | total: 7, so the client works out that 6 is still below 7 |
| Cursor (a bookmark the server hands back) | GET /tickets?cursor=2 | next_cursor: 4, or null on the last page |
| Page number | GET /invoices?page=2 | pages: 2, so page 2 is the last |
The UI wants one pager component, so the BFF turns all three into {items, nextCursor}. An opaque cursor is a string the client sends back unchanged without interpreting it; here "3" is an offset for orders while "2" is a page number for invoices, and the UI never needs to know.
What the BFF does for a single browser request to an account screen: (1) check the browser's session with the BFF; (2) start three upstream calls at once, each with its own credential and a deadline; (3) rewrite each reply into the shared item and error shapes; (4) return one JSON object with one panel per service, so a failed service marks only its own panel.
What belongs in the BFF
| Concern | What the BFF does | Why here |
|---|---|---|
| Upstream authentication | Holds the credential each service wants (token, key, session) and attaches it; the browser only ever holds its own session with the BFF. | B's API key and C's service session cookie are credentials issued to the BFF, not to any user. They cannot be shipped to a browser at all, because anything in page JavaScript can be read by every visitor, so this is not a style choice. |
| Pagination | Presents one shape to the UI, {items, nextCursor}, with an opaque cursor string; the adapter translates it to offset, cursor or page number. | The UI needs one pager component, not three. |
| Error format | Maps every upstream failure to {code, message, retryable}. The message is the BFF's own wording; the upstream text stays on the server error object for logs and never reaches the panel. | Upstream messages can leak internals; the UI needs one error component. |
| Fan-out and failure | Calls upstreams in parallel, each with a deadline, and returns per-panel status instead of failing the whole screen. | One slow or broken service should degrade one panel. |
| Shaping | Renames and drops fields so the payload matches the screen. | Saves bytes and removes mapping code from the browser. |
| Read caching | Caches upstream GETs briefly, and coalesces identical in-flight requests (when a second identical request arrives while the first is still waiting, it shares the first one's result instead of calling the service again). | Protects the services from a screen that fires the same call twice. |
What stays out
- Business rules and authorization decisions (is this refund allowed, what does this plan cost). The services own them; the BFF forwards the caller's identity and lets the service decide. If the BFF copies a rule, the next client re-implements it differently.
- Its own database of record. A cache is fine; the only copy of anything is not.
- Cross-service transactions. If two writes must succeed together, that is a service-level workflow, not something the web BFF orchestrates.
- Presentation text and layout, such as formatted strings for a button. Keep the payload semantic and let the UI format; otherwise every copy change is a backend deploy.
- Blind retries of writes. Retry only idempotent reads (an idempotent request has the same effect whether it is sent once or five times, as a read does); a repeated POST can double-charge.
A runnable BFF skeleton
Three fake upstreams with the three styles, adapters that normalize them, and a screen function that returns partial results. Run in a node:22 container (Node 22.23.3) with node bff-aggregator.mjs; there is nothing to install.
// Three upstream services with different auth, pagination and error formats,
// and a BFF that normalizes them into one screen payload.
import http from 'node:http';
// ---------- upstream services (stand-ins for internal microservices) ----------
const orders = Array.from({ length: 7 }, (_, i) => ({ oid: 100 + i, total: 10 * (i + 1) }));
const tickets = Array.from({ length: 5 }, (_, i) => ({ id: 't' + i, subject: 'Issue ' + i }));
const invoices = Array.from({ length: 4 }, (_, i) => ({ no: 9000 + i, amount_due: 5 * i }));
function listen(handler) {
return new Promise((res) => {
const s = http.createServer(handler).listen(0, '127.0.0.1', () => res(s));
});
}
const json = (r, code, body) => { r.writeHead(code, { 'content-type': 'application/json' }); r.end(JSON.stringify(body)); };
// Service A: bearer token, offset pagination, errors as {error:"..."}
const A = await listen((q, r) => {
const u = new URL(q.url, 'http://x');
if (q.headers.authorization !== 'Bearer svc-a-token') return json(r, 401, { error: 'bad token' });
const off = +(u.searchParams.get('offset') ?? 0), lim = +(u.searchParams.get('limit') ?? 3);
json(r, 200, { items: orders.slice(off, off + lim), total: orders.length });
});
// Service B: API key header, cursor pagination, errors as {errors:[{code,detail}]}
const B = await listen((q, r) => {
const u = new URL(q.url, 'http://x');
if (q.headers['x-api-key'] !== 'svc-b-key') return json(r, 403, { errors: [{ code: 'FORBIDDEN', detail: 'bad key' }] });
const c = +(u.searchParams.get('cursor') ?? 0);
const data = tickets.slice(c, c + 2);
json(r, 200, { data, next_cursor: c + 2 < tickets.length ? c + 2 : null });
});
// Service C: session cookie, page-number pagination, errors as plain text. Mode: 'ok', 'down' (503) or 'hang' (never answers).
let cMode = 'ok';
const C = await listen((q, r) => {
if (cMode === 'hang') return;
if (cMode === 'down') { r.writeHead(503, { 'content-type': 'text/plain' }); return r.end('pool exhausted on db-7 (10.0.3.9)'); }
const u = new URL(q.url, 'http://x');
if (q.headers.cookie !== 'sid=svc-c-session') { r.writeHead(401, { 'content-type': 'text/plain' }); return r.end('login required'); }
const p = +(u.searchParams.get('page') ?? 1), size = 2;
json(r, 200, { results: invoices.slice((p - 1) * size, p * size), page: p, pages: Math.ceil(invoices.length / size) });
});
const url = (s) => `http://127.0.0.1:${s.address().port}`;
// ---------- the BFF: adapters hide auth, pagination and error differences ----------
// One normalized error shape and one normalized page shape: {items, nextCursor}
class UpstreamError extends Error {
constructor(source, status, code, message) { super(message); Object.assign(this, { source, status, code }); }
get retryable() { return this.status >= 500 || this.status === 429; }
}
async function call(source, u, headers) {
const ctl = AbortSignal.timeout(500); // every upstream call has a deadline
const res = await fetch(u, { headers, signal: ctl });
const text = await res.text();
if (!res.ok) {
let code = 'UPSTREAM_' + res.status, message = text;
try { const j = JSON.parse(text); message = j.error ?? j.errors?.[0]?.detail ?? text; code = j.errors?.[0]?.code ?? code; } catch {}
throw new UpstreamError(source, res.status, code, message);
}
return JSON.parse(text);
}
const adapters = {
orders: async (cursor) => { // cursor is the offset, encoded as a string
const off = cursor ? +cursor : 0;
const j = await call('orders', `${url(A)}/orders?offset=${off}&limit=3`, { authorization: 'Bearer svc-a-token' });
const next = off + j.items.length;
return { items: j.items.map((o) => ({ id: String(o.oid), amount: o.total })), nextCursor: next < j.total ? String(next) : null };
},
tickets: async (cursor) => {
const j = await call('tickets', `${url(B)}/tickets?cursor=${cursor ?? 0}`, { 'x-api-key': 'svc-b-key' });
return { items: j.data.map((t) => ({ id: t.id, title: t.subject })), nextCursor: j.next_cursor === null ? null : String(j.next_cursor) };
},
invoices: async (cursor) => {
const p = cursor ? +cursor : 1;
const j = await call('invoices', `${url(C)}/invoices?page=${p}`, { cookie: 'sid=svc-c-session' });
return { items: j.results.map((i) => ({ id: String(i.no), due: i.amount_due })), nextCursor: p < j.pages ? String(p + 1) : null };
},
};
// The screen endpoint: one browser request, three parallel upstream calls, partial results allowed.
async function accountScreen() {
const names = Object.keys(adapters);
const settled = await Promise.allSettled(names.map((n) => adapters[n]()));
const panels = {};
settled.forEach((s, i) => {
// Only UpstreamError is a normalized failure. A timeout (DOMException) or a refused connection (TypeError) is not,
// and a DOMException has its own numeric `code` (23 for a timeout), so check the type instead of reading `.code`.
const e = s.reason;
panels[names[i]] = s.status === 'fulfilled'
? { status: 'ok', ...s.value }
: { status: 'error', error: e instanceof UpstreamError
? { code: e.code, message: `${names[i]} service failed (HTTP ${e.status})`, retryable: e.retryable } // e.message (the upstream text) stays server-side, for logs
: { code: 'UNAVAILABLE', message: 'upstream did not answer', retryable: true } };
});
return panels;
}
console.log('--- all upstreams healthy ---');
console.log(JSON.stringify(await accountScreen()));
console.log('--- page 2 of orders through the same normalized cursor ---');
console.log(JSON.stringify(await adapters.orders('3')));
console.log('--- invoices service down ---');
cMode = 'down';
const degraded = await accountScreen();
console.log(JSON.stringify({ orders: degraded.orders.status, tickets: degraded.tickets.status, invoices: degraded.invoices }));
console.log('--- invoices service hangs: the 500 ms deadline fires ---');
cMode = 'hang';
const t0 = Date.now();
const hung = await accountScreen();
console.log(JSON.stringify({ orders: hung.orders.status, invoices: hung.invoices }), 'waited under 1 s:', Date.now() - t0 < 1000);
console.log('--- wrong credential on service A: the thrown error keeps the upstream detail for server logs ---');
try { await call('orders', `${url(A)}/orders`, { authorization: 'Bearer nope' }); } catch (e) { console.log(e.source, e.status, e.code, e.message, 'retryable=' + e.retryable); }
[A, B, C].forEach((s) => { s.closeAllConnections(); s.close(); });
It prints:
--- all upstreams healthy ---
{"orders":{"status":"ok","items":[{"id":"100","amount":10},{"id":"101","amount":20},{"id":"102","amount":30}],"nextCursor":"3"},"tickets":{"status":"ok","items":[{"id":"t0","title":"Issue 0"},{"id":"t1","title":"Issue 1"}],"nextCursor":"2"},"invoices":{"status":"ok","items":[{"id":"9000","due":0},{"id":"9001","due":5}],"nextCursor":"2"}}
--- page 2 of orders through the same normalized cursor ---
{"items":[{"id":"103","amount":40},{"id":"104","amount":50},{"id":"105","amount":60}],"nextCursor":"6"}
--- invoices service down ---
{"orders":"ok","tickets":"ok","invoices":{"status":"error","error":{"code":"UPSTREAM_503","message":"invoices service failed (HTTP 503)","retryable":true}}}
--- invoices service hangs: the 500 ms deadline fires ---
{"orders":"ok","invoices":{"status":"error","error":{"code":"UNAVAILABLE","message":"upstream did not answer","retryable":true}}} waited under 1 s: true
--- wrong credential on service A: the thrown error keeps the upstream detail for server logs ---
orders 401 UPSTREAM_401 bad token retryable=false
Reading the code: AbortSignal.timeout(500) gives each upstream call a signal that aborts the fetch after 500 ms (MDN documents the abort reason as a TimeoutError), so a hung service cannot hold the screen. Promise.allSettled waits for all three adapter promises and returns one {status: 'fulfilled' | 'rejected'} entry per input, in input order, without rejecting itself when one input fails; that is why one dead service becomes one error panel instead of a failed screen. Only an UpstreamError is a normalized failure. A timeout rejects with a DOMException and a refused connection with a TypeError; neither is an UpstreamError, so the screen function checks the type with instanceof and labels them UNAVAILABLE, retryable: true, with a fixed message so no raw internals reach the UI. Reading reason.code instead would be a bug: a timeout's DOMException carries the legacy numeric code 23, which would leak into the panel as "code":23.
Reading the output: the screen came back from one function call with three panels, each carrying its own status. With service C returning 503 (its body, pool exhausted on db-7 (10.0.3.9), is internal detail), orders and tickets are still ok and invoices carries a normalized, retryable error whose message is the BFF's own invoices service failed (HTTP 503); the upstream text does not reach the panel. With service C hanging instead, the 500 ms deadline fires and the panel comes back as UNAVAILABLE in well under a second while orders stays ok. A wrong credential on service A (401) is an UpstreamError with retryable=false, because a client retry cannot fix a bad token; the last line prints bad token, the upstream detail that stays on the server-side error for logging, while the screen function would give the panel the BFF's own message. The normalized cursor for orders ("3") is the BFF's own opaque token; the UI never learns it is an offset.
When pagination crosses services
Each panel paginates itself, so no cross-service ordering is needed. If a screen needs one merged, sorted list across services, the BFF fetches a page from each, merges, and returns a cursor that encodes every source's position. Do that only when the product needs it: it multiplies upstream calls per page and makes the cursor fragile if any source reorders.
When to normalize in the frontend instead
Choose frontend normalization when all of these hold:
- The screen reads one service, so there is nothing to aggregate.
- The service needs no secret the browser cannot hold (public data, or the user's own login cookie where the browser and the service share a site, so the browser attaches the cookie itself and no secret sits in page JavaScript; this is different from a service credential such as C's session above).
- The difference is presentational: sorting, grouping or formatting data the UI already has, or interaction state such as an optimistic list update (showing a newly added row immediately, before the server has confirmed it).
- The BFF would add a hop and an owner for no gain, for example a prototype or a team without capacity to run one more service.
Anything that needs a secret, a join, or a unified error contract across services goes in the BFF.
Trade-offs and pitfalls
- The BFF is a new single point of failure for the web app. Mitigate with deadlines per upstream, partial responses (as shown), and health checks per adapter.
- Adapters are the part that rots: when service B changes its error format the BFF breaks. Cover each adapter with contract tests (tests that feed the adapter a recorded or fake copy of the service's reply format and fail when the format changes), as the fake upstreams here do in miniature.
- Keep the BFF thin. When a feature request starts with "the BFF should decide", check whether the rule belongs in a service.
- Do not leak adapter differences through the contract: a panel field called
nextOffsetin one panel andnextin another defeats the point.
Your analytics dashboard renders a dozen widgets, and each one currently calls its own backend endpoint, so users stare at spinners for seconds even though each call is fast. What would you change on the API side to cut perceived latency, and how would those changes affect what the frontend can render first?
Sample Answer
Direct answer
Replace the dozen independent calls with one aggregation endpoint (a server-side endpoint that fans out to the backends and returns a screen-shaped result) that streams its answer as newline-delimited JSON: first a tiny layout line describing every widget and its expected size, then one line per widget the instant that widget's data is ready, each with its own status. Fast calls feeding a slow screen usually point at the request pattern, not the services: the calls queue behind each other in the browser, start late because the page has to boot and discover the widgets first, and each pays connection, authentication and cross-origin overhead (a request to a different origin, meaning a different scheme, host or port, triggers the browser's cross-origin permission checks, covered in item 3 below). One request that fans out inside the data centre removes those costs, and streaming means the first widgets paint while the slowest is still loading.
Why a dozen fast calls still feel slow
Check in this order, using the browser DevTools Network panel's waterfall (the timeline of every request):
- Start times: do the widget requests all begin at once, or only after a config or layout call returns? A late start is a dependency chain, and no amount of per-call speed fixes it.
- Queueing: bars with a long grey "Queueing" or "Stalled" segment mean the browser is holding requests back. Over HTTP/1.1 browsers open only a handful of parallel connections per domain (MDN describes 6 as the common figure), so twelve requests run in two waves. HTTP/2 multiplexes (carries many requests and responses at the same time over one connection), so this particular cause disappears when the server speaks HTTP/2.
- Per-request overhead: repeated connection setup, token validation, and
OPTIONSpreflights (the browser's permission check before a cross-origin request) on every widget call. - Shared bottleneck behind the calls: twelve "fast" calls hitting the same database can serialise there (end up running one after another instead of side by side, for example waiting for the same lock or a small connection pool).
The measurement that matters is time to first meaningful widget and time to the last one, not average call latency. Time to first meaningful widget is the time from navigation until the first widget shows real data instead of a skeleton. Record both with the PerformanceObserver API (a browser API that lets page JavaScript subscribe to performance entries, such as paint times and its own named marks, as the browser records them) or your real-user monitoring (collecting these timings from the browsers of real visitors instead of from a lab test), per release.
The endpoint design
Request: GET /dashboard?widgets=kpi_revenue,kpi_orders,trend_chart,... (or a saved dashboard id the server expands).
Response, application/x-ndjson, one JSON object per line:
{"type":"layout","widgets":[{"id":"kpi_revenue","rows":1},{"id":"trend_chart","rows":30},...]}
{"type":"widget","id":"kpi_revenue","status":"ok","data":{...}}
{"type":"widget","id":"cohort_table","status":"error","error":"upstream 503"}
Design decisions:
- Fan-out happens on the server in parallel, with a per-widget deadline (1 s in the demo; choose it from the slowest acceptable widget, not the average). A slow widget becomes
status: "error"or"timeout"for that widget instead of stalling the page. - Partial failure is part of the contract. The HTTP status is 200 once the stream starts, because the headers are already sent; failure is reported per widget. The frontend renders an inline error with a retry button on that widget only, and retries it through a single-widget call (
/dashboard?widgets=cohort_table). - The layout line enables skeleton-first rendering. With row counts known up front the client draws correctly sized placeholders, so content arriving later does not shift the page (unexpected layout shift hurts the Cumulative Layout Shift score, CLS, a Web Vitals metric that adds up how far visible content jumps around while the page loads; lower is better).
- Widget payloads are shaped for the widget: pre-aggregated series (30 points for a trend chart) rather than raw rows the browser has to reduce.
What this changes for the frontend
- The page can render the shell and skeletons after the first line, and each widget flips from skeleton to content as its line arrives, so "first widget" depends on the fastest backend rather than the slowest.
- State becomes per-widget (
loading | ok | error) instead of one page-level flag. The component for a widget receives its own slice and knows nothing about the others. - On the layout line the client draws one skeleton box per widget, with height taken from
rows; on each widget line it finds that widget's box and swaps in the chart or number, or an error with a retry button whenstatusiserror. The client parses the stream line by line withfetchand the response body reader. Anything that buffers the whole response (some proxies and compression layers) turns streaming back into one big wait, so verify it end to end, not only on localhost. If a proxy or CDN in front does buffer, the options are to turn response buffering off for this route in its configuration, to exclude the route from layers that wait for the whole body, or to fall back to the composite JSON response.
Worked example
This program starts a server with seven simulated widgets (one fails, one exceeds the deadline) and a client that reads the stream and prints each line as it arrives. The simulated delays are inputs, spaced at least 75 ms apart so the arrival order is stable.
import http from 'node:http';
// Simulated downstream services: delay in ms (spaced at least 75 ms apart so the arrival order is stable)
const widgets = {
kpi_revenue: { delay: 50, rows: 1 },
kpi_orders: { delay: 200, rows: 1 },
trend_chart: { delay: 350, rows: 30 },
top_products: { delay: 500, rows: 10 },
funnel: { delay: 650, rows: 5 },
geo_map: { delay: 4000, rows: 50 }, // slower than the per-widget budget
cohort_table: { delay: 125, fail: true }, // downstream error
};
const BUDGET_MS = 1000; // per-widget deadline
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function load(id) {
const w = widgets[id];
const work = (async () => { await sleep(w.delay); if (w.fail) throw new Error('upstream 503'); return { rows: w.rows }; })();
const deadline = sleep(BUDGET_MS).then(() => { throw new Error('timeout'); });
return Promise.race([work, deadline]);
}
const server = http.createServer(async (req, res) => {
const ids = new URL(req.url, 'http://x').searchParams.get('widgets').split(',');
res.writeHead(200, { 'content-type': 'application/x-ndjson' });
// Line 1 returns at once: the layout the client can draw skeletons from
res.write(JSON.stringify({ type: 'layout', widgets: ids.map((id) => ({ id, rows: widgets[id]?.rows ?? 0 })) }) + '\n');
await Promise.all(ids.map(async (id) => { // all upstream calls start together
let line;
try { line = { type: 'widget', id, status: 'ok', data: await load(id) }; }
catch (e) { line = { type: 'widget', id, status: 'error', error: e.message }; }
res.write(JSON.stringify(line) + '\n'); // flushed the moment this widget settles
}));
res.end();
}).listen(0);
const url = `http://127.0.0.1:${server.address().port}/dashboard?widgets=${Object.keys(widgets).join(',')}`;
const t0 = Date.now();
const res = await fetch(url);
const dec = new TextDecoder(); let buf = '', tFirst = null, tLast = 0;
for await (const chunk of res.body) {
buf += dec.decode(chunk, { stream: true });
let i;
while ((i = buf.indexOf('\n')) >= 0) {
const m = JSON.parse(buf.slice(0, i)); buf = buf.slice(i + 1);
if (m.type === 'layout') console.log('layout -> skeletons for', m.widgets.map((w) => `${w.id}(${w.rows})`).join(' '));
else {
tFirst ??= Date.now() - t0; tLast = Date.now() - t0;
console.log(`widget -> ${m.id.padEnd(13)} ${m.status}${m.error ? ' (' + m.error + ')' : ''}`);
}
}
}
// Thresholds, not raw milliseconds, so the output is stable from run to run
console.log(`first widget painted before 200 ms: ${tFirst < 200}`);
console.log(`last line waited for the ${BUDGET_MS} ms deadline: ${tLast >= BUDGET_MS - 10}`);
server.close();
It prints (the arrival order is the same on every run):
layout -> skeletons for kpi_revenue(1) kpi_orders(1) trend_chart(30) top_products(10) funnel(5) geo_map(50) cohort_table(0)
widget -> kpi_revenue ok
widget -> cohort_table error (upstream 503)
widget -> kpi_orders ok
widget -> trend_chart ok
widget -> top_products ok
widget -> funnel ok
widget -> geo_map error (timeout)
first widget painted before 200 ms: true
last line waited for the 1000 ms deadline: true
Reading the code: load starts the simulated work and a deadline timer and uses Promise.race, so whichever settles first wins; geo_map needs 4,000 ms against the 1,000 ms budget, so the deadline wins and it becomes timeout. On the server, res.write sends the layout line immediately and then one line per widget as soon as that widget settles; Promise.all only decides when res.end() closes the stream. In the client loop, for await (const chunk of res.body) receives the body as byte chunks as they arrive. TextDecoder with { stream: true } turns bytes into text while keeping a character that was split across two chunks intact. A chunk can hold half a line or several lines, so buf accumulates text, indexOf('\n') finds each complete line, JSON.parse reads it, and buf.slice keeps the remainder for the next chunk.
The two final lines are what show streaming: the first widget line reaches the client well before 200 ms even though the last one only arrives at the 1,000 ms deadline. If the server held every line until Promise.all finished (a composite response), the order of the lines would not change but the first line would arrive at about 1,000 ms and the first check would print false. Widgets arrive in order of their backend latency, the failed widget reports immediately without blocking the rest, and the stuck map is cut off at the deadline instead of holding the page.
Alternatives and trade-offs
| Option | Strength | Weakness |
|---|---|---|
| One composite JSON response | Simplest client, easy to cache | Every widget waits for the slowest; one failure needs a partial-failure convention anyway |
| Streaming aggregation (recommended) | First widget paints early, per-widget errors | Needs streaming-capable infrastructure and a small stream parser |
| Keep twelve calls, move to HTTP/2 or edge (edge caching stores individual widget responses on CDN servers close to the users) | Zero API change, fixes connection queueing | Does not fix the late start or per-call auth overhead; each widget still pays its own round trip |
| Batch endpoint that takes ids and returns one JSON map | Middle ground | Same slowest-wins behaviour as the composite |
Recommendation: streaming aggregation for this dashboard. Fall back to the composite JSON if your gateway cannot stream, and take HTTP/2 regardless since it is nearly free. What would flip the choice: if all twelve widgets are cheap and come from one database, a composite is enough; if widgets are independently cacheable and personalised data is small, edge caching of individual widget calls can beat aggregation.
Pitfalls
- Aggregating in the browser's critical path (a client-side loop that waits for all promises) recreates the composite's wait with extra requests.
- Failing the whole call on one error: the dashboard goes blank because of one broken widget.
- No deadline: a hung dependency holds the connection and the layout open indefinitely.
- Verifying only on a fast laptop: test with DevTools network throttling and a single widget artificially slowed so the progressive rendering is actually exercised.
Running the code
docker run --rm --ulimit core=0 -v "$PWD":/w -w /w node:22 sh -c 'timeout 60 node dash.mjs'
Save the listing as dash.mjs (Node 22 ES module, no install step, uses the built-in fetch).
What are over-fetching and under-fetching in client-server APIs, and how do REST and GraphQL each deal with them? Describe what a frontend or mobile client feels from each approach.
Sample Answer
Direct answer
Over-fetching means the response carries more data than the screen uses. Under-fetching means one response is not enough, so the client must make follow-up requests to assemble a screen. Both waste the client's time, data allowance and battery. REST fixes them by shaping resources and endpoints on the server: sparse fieldsets (a fields= parameter that lists the only fields the client wants, shown below), embedded related data, or screen-specific endpoints. GraphQL, a query language where the server publishes a schema (a typed list of every object and field a client may ask for), fixes them by letting the client name exactly the fields it wants in a single request, at the price of a heavier server and a harder caching story.
What each problem looks like
Take a feed screen. Each row shows a post title, the author's name and a thumbnail.
- Over-fetching.
GET /postsreturns the whole post: body text, tags, view counts, edit history. The row uses three fields. If each post is 2,000 bytes and a row needs about 100, a page of 20 rows downloads 20 x 2,000 = 40,000 bytes to render 20 x 100 = 2,000 bytes of content, a factor of 20. The phone also parses all of it as JSON (JavaScript Object Notation) and holds it in memory. - Under-fetching.
GET /postsreturnsauthorIdbut not the name. If the 20 posts have 3 distinct authors, the client makes 1 + 3 = 4 requests (the pattern is called N+1: one request for the list, then one more per distinct related item). The author lookups cannot start until the list has arrived, so the screen waits for at least two sequential network round trips (list, then authors), and each extra hop on a weak cellular link is a chance to stall.
How REST deals with it
REST (resources addressed by URL, standard HTTP verbs) has no built-in query language, so you shape responses with conventions:
- Sparse fieldsets:
GET /posts?fields=id,title,author.name,thumbnailUrl. - Embedding or inclusion:
GET /posts?embed=authororinclude=author, so related data rides along (JSON:API, a public REST convention, definesfields[TYPE]andincludeand requires400 Bad Requestwhen a server does not supportinclude). - Screen-shaped endpoints:
GET /screens/feed, which belongs to a Backend-for-Frontend (BFF, a server layer owned by the client team that returns exactly what one client needs).
How GraphQL deals with it
GraphQL exposes one typed schema. The client sends a query that names fields, and the response mirrors its shape:
{ posts(first: 20) { id title author { name } thumbnailUrl(width: 96) } }
Over-fetching disappears because only named fields come back. Under-fetching disappears because the nested author { name } is resolved on the server in the same request (a resolver is the server function that produces the value of one field, here the author's name). Per the GraphQL over HTTP documentation, servers must accept POST for queries and mutations (mutations are the GraphQL operations that write data) and may accept GET for queries only. A response can hold both data and errors, and a response with non-null data should still use a 2xx status, because HTTP has no status for partial success.
What the client feels
| Concern | REST with fieldsets / BFF | GraphQL |
|---|---|---|
| Bytes per screen | Small if the server implements shaping; fixed endpoints over-fetch for other screens | Small by construction, per query |
| Round trips | 1 with embeds or a screen endpoint, otherwise 1 + N (one list request plus one per related item) | 1 per screen query |
| Client CPU | Parses a small JSON; no query building | Parses a small JSON; a client library may normalize results into a cache (split each response into one record per object, keyed by id, so an author shared by many posts is stored once), and that splitting and merging costs CPU on low-end devices |
| CDN (content delivery network) caching | Easy: URL is the cache key, GET is cacheable | Hard with POST; possible with GET plus persisted query ids (short identifiers that replace long query text, which also avoids URL length limits) |
| Developer velocity | Frontend waits for a new field or endpoint if the shape is missing | Frontend adds a field to its own query, no backend ticket, if the schema already exposes it |
| Versioning | New fields are additive; breaking shapes need /v2 or negotiation | Add fields freely and mark old ones with the built-in @deprecated directive (an annotation on a schema field that tells tools and developers the field is going away); usage is visible per field |
| Error handling | HTTP status per request tells the story | Often HTTP 200 with an errors array and partial data, so clients must inspect the body |
Which screens feel it most
- Personalized feed: many small rows, several related entities per row (author, counts, reaction). Under-fetching dominates, so one composed request (GraphQL query or BFF endpoint) matters most.
- Settings screen: one small object, rarely changing. Plain REST is fine and caches well; GraphQL adds little.
- Analytics sync: the client uploads or downloads bulk data in the background. The cost is bytes and radio wake-ups (every time the phone's cellular radio leaves idle to send a request it draws extra battery, so many small requests cost more energy than one batched request), not flexible field choice. Batch it into one request and compress it; a query language is not the lever.
Pitfalls
- GraphQL does not remove over-fetching on the server: a resolver (the function that produces one field) may still load the whole database row. Field selection saves wire bytes, not necessarily database work.
- A permissive
fields=parameter can leak internal columns. Validate it against an allowlist. - Choosing GraphQL to fix under-fetching, then discovering that each nested field issues its own database query (the N+1 problem), swaps a client problem for a server one.
That is every published Backend-for-Frontend and Client-Specific API Design question for Frontend Developer so far. Browse the other topics in this category, or practice this one interactively.