Backend-for-Frontend and Client-Specific API Design Questions
Shaping API responses to the needs of specific client types: the Backend-for-Frontend (BFF) pattern, one BFF per client versus a shared gateway, client-side composition or a generic API, what belongs in a BFF (including normalizing inconsistent downstream pagination and error formats for one client), and aggregating several backend calls into one client-shaped response. Covers payload shaping (sparse fieldsets and field selection against relational and document stores, partial responses, optional expansion of related collections, delta responses with deletions, compact serialization including binary wire formats), round-trip reduction and composite views for feeds and dashboards, adapting to mobile versus web bandwidth, battery and device constraints (device-appropriate image variants, prefetching and refresh frequency tuned to connection quality), and choosing GraphQL versus REST to fix over- and under-fetching, including incremental migration from REST to GraphQL. Authentication, general REST resource design, versioning, rate limiting, offline sync and real-time transport are covered elsewhere.
Your analytics dashboard renders a dozen widgets, and each one currently calls its own backend endpoint, so users stare at spinners for seconds even though each call is fast. What would you change on the API side to cut perceived latency, and how would those changes affect what the frontend can render first?
Sample Answer
Direct answer
Replace the dozen independent calls with one aggregation endpoint (a server-side endpoint that fans out to the backends and returns a screen-shaped result) that streams its answer as newline-delimited JSON: first a tiny layout line describing every widget and its expected size, then one line per widget the instant that widget's data is ready, each with its own status. Fast calls feeding a slow screen usually point at the request pattern, not the services: the calls queue behind each other in the browser, start late because the page has to boot and discover the widgets first, and each pays connection, authentication and cross-origin overhead (a request to a different origin, meaning a different scheme, host or port, triggers the browser's cross-origin permission checks, covered in item 3 below). One request that fans out inside the data centre removes those costs, and streaming means the first widgets paint while the slowest is still loading.
Why a dozen fast calls still feel slow
Check in this order, using the browser DevTools Network panel's waterfall (the timeline of every request):
- Start times: do the widget requests all begin at once, or only after a config or layout call returns? A late start is a dependency chain, and no amount of per-call speed fixes it.
- Queueing: bars with a long grey "Queueing" or "Stalled" segment mean the browser is holding requests back. Over HTTP/1.1 browsers open only a handful of parallel connections per domain (MDN describes 6 as the common figure), so twelve requests run in two waves. HTTP/2 multiplexes (carries many requests and responses at the same time over one connection), so this particular cause disappears when the server speaks HTTP/2.
- Per-request overhead: repeated connection setup, token validation, and
OPTIONSpreflights (the browser's permission check before a cross-origin request) on every widget call. - Shared bottleneck behind the calls: twelve "fast" calls hitting the same database can serialise there (end up running one after another instead of side by side, for example waiting for the same lock or a small connection pool).
The measurement that matters is time to first meaningful widget and time to the last one, not average call latency. Time to first meaningful widget is the time from navigation until the first widget shows real data instead of a skeleton. Record both with the PerformanceObserver API (a browser API that lets page JavaScript subscribe to performance entries, such as paint times and its own named marks, as the browser records them) or your real-user monitoring (collecting these timings from the browsers of real visitors instead of from a lab test), per release.
The endpoint design
Request: GET /dashboard?widgets=kpi_revenue,kpi_orders,trend_chart,... (or a saved dashboard id the server expands).
Response, application/x-ndjson, one JSON object per line:
{"type":"layout","widgets":[{"id":"kpi_revenue","rows":1},{"id":"trend_chart","rows":30},...]}
{"type":"widget","id":"kpi_revenue","status":"ok","data":{...}}
{"type":"widget","id":"cohort_table","status":"error","error":"upstream 503"}
Design decisions:
- Fan-out happens on the server in parallel, with a per-widget deadline (1 s in the demo; choose it from the slowest acceptable widget, not the average). A slow widget becomes
status: "error"or"timeout"for that widget instead of stalling the page. - Partial failure is part of the contract. The HTTP status is 200 once the stream starts, because the headers are already sent; failure is reported per widget. The frontend renders an inline error with a retry button on that widget only, and retries it through a single-widget call (
/dashboard?widgets=cohort_table). - The layout line enables skeleton-first rendering. With row counts known up front the client draws correctly sized placeholders, so content arriving later does not shift the page (unexpected layout shift hurts the Cumulative Layout Shift score, CLS, a Web Vitals metric that adds up how far visible content jumps around while the page loads; lower is better).
- Widget payloads are shaped for the widget: pre-aggregated series (30 points for a trend chart) rather than raw rows the browser has to reduce.
What this changes for the frontend
- The page can render the shell and skeletons after the first line, and each widget flips from skeleton to content as its line arrives, so "first widget" depends on the fastest backend rather than the slowest.
- State becomes per-widget (
loading | ok | error) instead of one page-level flag. The component for a widget receives its own slice and knows nothing about the others. - On the layout line the client draws one skeleton box per widget, with height taken from
rows; on each widget line it finds that widget's box and swaps in the chart or number, or an error with a retry button whenstatusiserror. The client parses the stream line by line withfetchand the response body reader. Anything that buffers the whole response (some proxies and compression layers) turns streaming back into one big wait, so verify it end to end, not only on localhost. If a proxy or CDN in front does buffer, the options are to turn response buffering off for this route in its configuration, to exclude the route from layers that wait for the whole body, or to fall back to the composite JSON response.
Worked example
This program starts a server with seven simulated widgets (one fails, one exceeds the deadline) and a client that reads the stream and prints each line as it arrives. The simulated delays are inputs, spaced at least 75 ms apart so the arrival order is stable.
import http from 'node:http';
// Simulated downstream services: delay in ms (spaced at least 75 ms apart so the arrival order is stable)
const widgets = {
kpi_revenue: { delay: 50, rows: 1 },
kpi_orders: { delay: 200, rows: 1 },
trend_chart: { delay: 350, rows: 30 },
top_products: { delay: 500, rows: 10 },
funnel: { delay: 650, rows: 5 },
geo_map: { delay: 4000, rows: 50 }, // slower than the per-widget budget
cohort_table: { delay: 125, fail: true }, // downstream error
};
const BUDGET_MS = 1000; // per-widget deadline
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function load(id) {
const w = widgets[id];
const work = (async () => { await sleep(w.delay); if (w.fail) throw new Error('upstream 503'); return { rows: w.rows }; })();
const deadline = sleep(BUDGET_MS).then(() => { throw new Error('timeout'); });
return Promise.race([work, deadline]);
}
const server = http.createServer(async (req, res) => {
const ids = new URL(req.url, 'http://x').searchParams.get('widgets').split(',');
res.writeHead(200, { 'content-type': 'application/x-ndjson' });
// Line 1 returns at once: the layout the client can draw skeletons from
res.write(JSON.stringify({ type: 'layout', widgets: ids.map((id) => ({ id, rows: widgets[id]?.rows ?? 0 })) }) + '\n');
await Promise.all(ids.map(async (id) => { // all upstream calls start together
let line;
try { line = { type: 'widget', id, status: 'ok', data: await load(id) }; }
catch (e) { line = { type: 'widget', id, status: 'error', error: e.message }; }
res.write(JSON.stringify(line) + '\n'); // flushed the moment this widget settles
}));
res.end();
}).listen(0);
const url = `http://127.0.0.1:${server.address().port}/dashboard?widgets=${Object.keys(widgets).join(',')}`;
const t0 = Date.now();
const res = await fetch(url);
const dec = new TextDecoder(); let buf = '', tFirst = null, tLast = 0;
for await (const chunk of res.body) {
buf += dec.decode(chunk, { stream: true });
let i;
while ((i = buf.indexOf('\n')) >= 0) {
const m = JSON.parse(buf.slice(0, i)); buf = buf.slice(i + 1);
if (m.type === 'layout') console.log('layout -> skeletons for', m.widgets.map((w) => `${w.id}(${w.rows})`).join(' '));
else {
tFirst ??= Date.now() - t0; tLast = Date.now() - t0;
console.log(`widget -> ${m.id.padEnd(13)} ${m.status}${m.error ? ' (' + m.error + ')' : ''}`);
}
}
}
// Thresholds, not raw milliseconds, so the output is stable from run to run
console.log(`first widget painted before 200 ms: ${tFirst < 200}`);
console.log(`last line waited for the ${BUDGET_MS} ms deadline: ${tLast >= BUDGET_MS - 10}`);
server.close();
It prints (the arrival order is the same on every run):
layout -> skeletons for kpi_revenue(1) kpi_orders(1) trend_chart(30) top_products(10) funnel(5) geo_map(50) cohort_table(0)
widget -> kpi_revenue ok
widget -> cohort_table error (upstream 503)
widget -> kpi_orders ok
widget -> trend_chart ok
widget -> top_products ok
widget -> funnel ok
widget -> geo_map error (timeout)
first widget painted before 200 ms: true
last line waited for the 1000 ms deadline: true
Reading the code: load starts the simulated work and a deadline timer and uses Promise.race, so whichever settles first wins; geo_map needs 4,000 ms against the 1,000 ms budget, so the deadline wins and it becomes timeout. On the server, res.write sends the layout line immediately and then one line per widget as soon as that widget settles; Promise.all only decides when res.end() closes the stream. In the client loop, for await (const chunk of res.body) receives the body as byte chunks as they arrive. TextDecoder with { stream: true } turns bytes into text while keeping a character that was split across two chunks intact. A chunk can hold half a line or several lines, so buf accumulates text, indexOf('\n') finds each complete line, JSON.parse reads it, and buf.slice keeps the remainder for the next chunk.
The two final lines are what show streaming: the first widget line reaches the client well before 200 ms even though the last one only arrives at the 1,000 ms deadline. If the server held every line until Promise.all finished (a composite response), the order of the lines would not change but the first line would arrive at about 1,000 ms and the first check would print false. Widgets arrive in order of their backend latency, the failed widget reports immediately without blocking the rest, and the stuck map is cut off at the deadline instead of holding the page.
Alternatives and trade-offs
| Option | Strength | Weakness |
|---|---|---|
| One composite JSON response | Simplest client, easy to cache | Every widget waits for the slowest; one failure needs a partial-failure convention anyway |
| Streaming aggregation (recommended) | First widget paints early, per-widget errors | Needs streaming-capable infrastructure and a small stream parser |
| Keep twelve calls, move to HTTP/2 or edge (edge caching stores individual widget responses on CDN servers close to the users) | Zero API change, fixes connection queueing | Does not fix the late start or per-call auth overhead; each widget still pays its own round trip |
| Batch endpoint that takes ids and returns one JSON map | Middle ground | Same slowest-wins behaviour as the composite |
Recommendation: streaming aggregation for this dashboard. Fall back to the composite JSON if your gateway cannot stream, and take HTTP/2 regardless since it is nearly free. What would flip the choice: if all twelve widgets are cheap and come from one database, a composite is enough; if widgets are independently cacheable and personalised data is small, edge caching of individual widget calls can beat aggregation.
Pitfalls
- Aggregating in the browser's critical path (a client-side loop that waits for all promises) recreates the composite's wait with extra requests.
- Failing the whole call on one error: the dashboard goes blank because of one broken widget.
- No deadline: a hung dependency holds the connection and the layout open indefinitely.
- Verifying only on a fast laptop: test with DevTools network throttling and a single widget artificially slowed so the progressive rendering is actually exercised.
Running the code
docker run --rm --ulimit core=0 -v "$PWD":/w -w /w node:22 sh -c 'timeout 60 node dash.mjs'
Save the listing as dash.mjs (Node 22 ES module, no install step, uses the built-in fetch).
Your large REST-based product wants to adopt GraphQL without a rewrite. Propose a migration path that keeps existing clients working while screens move over one at a time, and say how you would decide when the old endpoints can be retired.
Sample Answer
Direct answer
Put a GraphQL layer in front of the existing REST services as a facade (a thin new front door that translates the new query language into calls to the old system) and migrate the product screen by screen. Existing clients keep calling REST untouched, the facade reuses the same services and business rules, and each screen switches behind a feature flag (a configuration switch, read at run time, that turns the new code path on or off for chosen users without a deploy) once its GraphQL version is proven equivalent. Retire a REST endpoint only when measured traffic to it has been zero for longer than your slowest client can take to upgrade.
Phase 1: facade over the existing services
- Stand up one GraphQL server (the facade) whose resolvers (the functions that fetch each field) call the existing REST endpoints or the service layer behind them. No data model rewrite and no REST change, so there is nothing for old clients to notice.
- Model the schema around what screens need, not a 1:1 mirror of the REST resources. This is where the over-fetching (receiving fields the screen ignores) and under-fetching (needing extra calls for missing fields) problems are actually fixed.
- Pass the caller's credentials through and keep authorisation in the services, so the facade cannot become a bypass.
- Guard against the classic resolver cost: a list of 50 orders whose resolver calls the customer service once per order issues 50 calls (the "N+1" problem). Batch lookups per request with a batching helper such as DataLoader, and set query depth and complexity limits, or allow only pre-registered ("persisted") queries from your own apps.
Phase 2: move screens one at a time
- Pick a low-risk, high-pain first screen (one with many REST calls, so the win is visible).
- Build it against GraphQL behind a feature flag, and run a parity check in staging: load the same data through the REST path and the GraphQL path and diff what the screen renders. For example, for order 1042 the REST path renders the total
$84.50and the GraphQL path renders$84.5; the diff flags that one field, and the fix is a formatting rule or a missing field in the schema, not a change to the check. - Roll out by canary (named after the canary in the coal mine: a small slice of users gets the change first, so a problem shows up on few people): for example 1 percent, 10 percent, 50 percent, 100 percent, holding at each step long enough to see error and latency data. Rollback is the flag, not a deploy.
- Compare the same signals for both paths: error rate, 95th percentile latency (the time that 95 percent of requests beat; only the slowest 1 in 20 take longer), payload bytes, and backend calls per page. Stop the rollout if any regresses.
- Repeat for the next screen. New screens are GraphQL from the start, which stops the legacy surface from growing.
Mark REST-shaped parts of the schema that you want people to leave with GraphQL's built-in @deprecated(reason: "...") directive, which the GraphQL documentation describes as the way to annotate deprecated fields; tooling then warns developers.
Gateway, stitching, or one facade
- Single facade owned by one team is the recommendation to start: one schema, one deploy, simple debugging.
- Schema stitching or federation (two ways of combining several teams' separate GraphQL schemas into one graph that clients query as a single API; a gateway routes each field to the team that owns it) pays off when multiple teams own different parts and cannot share one repository. It adds a gateway layer to operate, so adopt it when team boundaries demand it, not on day one.
- A BFF per client type (a backend-for-frontend, a server tailored to one client such as web or mobile) is an alternative when clients differ sharply; GraphQL lets one server serve both because each client chooses its fields.
Keeping existing clients working
REST contracts are frozen: no breaking changes while GraphQL adoption proceeds. Old mobile app versions stay in the wild for months, so the legacy API must outlive the web migration. Add a client-version tag to requests now (a header such as an app version) if one does not exist, since the retirement decision depends on it.
Deciding when to retire an endpoint
Retire per endpoint, never as one big switch. Evidence to require:
- Traffic: request count per endpoint, broken down by client and app version, shows zero from supported versions for a full window. Choose the window to cover your slowest release cadence plus your forced-upgrade policy (for example the longest period a mobile version stays supported); that number is a business decision, not a technical constant. Illustrative policy: each app version stays supported for 12 months after the release that replaced it, so an endpoint last used by version 7 can be retired only once version 7's support has ended (12 months after version 8 shipped), plus a 30-day margin, with zero supported-version traffic in the final 30 days.
- Remaining callers identified: any non-zero traffic maps to a named owner (internal job, partner, old app version), each with a migration path or an agreed end date.
- Replacement proven: the GraphQL path has run at full traffic through at least one complete business cycle (month end, sale peaks) with no regression.
- Staged shutdown: announce, then log warnings and email owners of the remaining callers, then run short scheduled outages of the endpoint to flush out hidden dependants, then remove. The scheduled outages convert silent dependencies into visible, reversible failures.
Communicate throughout: a published timeline, a migration guide per endpoint, a single dashboard showing traffic left, and a named contact for partners. Incentives matter for internal teams: new features ship only on GraphQL, and their performance wins are reported back to them.
Worked example
A mobile "order details" screen currently makes 4 REST calls: /orders/{id}, /customers/{id}, /orders/{id}/items, /shipments?order={id}. The facade exposes one order(id) query asking for 9 of the roughly 40 fields across those responses. In staging, a parity test fetches the screen's data both ways for 200 sampled order ids and fails if any rendered field differs. The screen then ships at 1 percent of users behind a flag; the team watches error rate and latency against the REST cohort for a week before raising it. Illustrative figures, assuming a 150 ms round trip from the phone, 30 ms of server time per call, a 5 ms internal link from the facade to the services, and about 300 bytes per field: the REST screen needs two sequential waves (the order first, then customer, items and shipments in parallel), about 2 x (150 + 30) = 360 ms and 40 x 300 = 12,000 bytes; the GraphQL screen needs one round trip plus the facade's own two waves, about 150 + 2 x (5 + 30) = 220 ms and 9 x 300 = 2,700 bytes. The four REST routes stay alive for the previous app versions; their per-version traffic chart is what later justifies removal.
Trade-offs and pitfalls
- Facade becomes the bottleneck: every GraphQL request now depends on it, so give it its own scaling, timeouts and per-resolver circuit breaking (when one downstream service keeps failing, the resolver stops calling it for a short time and returns an error or a fallback at once, instead of letting every request wait on a service that is down).
- HTTP caching is harder: GraphQL queries are typically
POST, so URL-based caches and CDNs (content delivery networks) do less; plan persisted queries or response-level caching. - Two systems to run during migration: budget for the overlap period and set the end date up front, or the legacy layer lives forever.
- Mirroring REST 1:1 in the schema keeps the old problems with a new syntax.
- What flips the recommendation: if the problem is only one or two chatty screens, a purpose-built BFF endpoint is cheaper than adopting GraphQL across the product.
Clients can ask your resource endpoints for only some of the fields. How would you implement that field selection on the server, against a relational database and against a document store, so it actually reduces backend work and not only response size?
Sample Answer
Direct answer
Parsing ?fields=title,status and deleting keys from the JSON before sending only shrinks the response. To reduce backend work, the field list must change what the data layer is asked to do: build the SQL SELECT list (or the document-store projection, which is the part of a query that names the columns or fields to return) from a whitelist (an explicit list of allowed field names; anything not on it is refused), skip the joins that unrequested fields would need, and design indexes so common field sets can be answered from the index alone. An index is a separate sorted structure, like the index of a book: the database finds an entry by value without scanning every row. A "covered" query is one answered entirely from an index without reading the table rows or documents, because the index already contains every column the query needs. Picture 1,000 articles with a 5,000-character body each. An index on (status, title) holds only those two columns plus the row id. A query for title and status is answered from that small structure; a query that also wants body must follow each index entry back to the full row and read its 5,000 characters, 20 rows x 5,000 = 100,000 characters for one page.
Design
- Whitelist, not pass-through. Map each public field name to how it is fetched: a column, a joined column, or a computed value with its dependencies. Reject unknown names with
400, testing membership on the whitelist's own keys only (in JavaScriptObject.hasOwn, or aMap, sinceinalso matches inherited names likeconstructor). Never interpolate the client string into SQL. - Always include the key (
id) so the client can identify rows, and apply a default field set whenfieldsis absent so old clients keep working. - Relational database: generate the
SELECTlist and add aJOINonly if a requested field needs it. Wide columns (large text, JSON blobs) are the most valuable to leave out. In PostgreSQL, when a row is wider than about 2 kB its large values are compressed and moved out of line (TOAST, The Oversized-Attribute Storage Technique), and the documentation states that "the big values of TOASTed attributes will only be pulled out (if selected at all) at the time the result set is sent to the client", so not selectingbodyavoids reading it. - Document store (MongoDB): pass the same whitelist as a projection document such as
{ _id: 1, title: 1, status: 1 }. Projection limits what is returned and cuts network and serialization; storage reads drop only for a covered query, which MongoDB defines as one satisfied "entirely using an index" without examining any document. That requires every queried and returned field to be in one index, and_id: 0in the projection unless_idis in the index. Traced example: with an index on{ status: 1, title: 1, _id: 1 }, the query{ status: "published" }with projection{ title: 1, status: 1 }returns only fields the index holds, so no document is read. The same query withbody: 1added must open every matching document. With an index on{ status: 1, title: 1 }alone, the projection must say_id: 0, because_idis returned by default and is not in that index. A query that is not covered has to examine documents, so for it the saving is mainly network transfer and serialization. - Computed fields declare their inputs (
fullNameneedsfirstandlast), so requesting one selects those columns. - Nested relations (
author.name): resolve in one batched query or join, never one query per row. - Guard the combinations. Each distinct field set is a different query plan. Cap the number of fields and nesting depth, and list the few high-traffic field sets in the index design rather than trying to cover all combinations.
Runnable demonstration (relational side)
An in-memory SQLite database (the node:sqlite module that ships with Node 22) with 1,000 articles, each with a 5,000-character body, and a users table. The whitelist builds the SQL; EXPLAIN QUERY PLAN (a SQLite command that prints how the database would run a query, without running it) shows what the database will do. In the plan text, SEARCH means the engine jumps straight to matching entries using an index, as opposed to SCAN, which reads every row; COVERING INDEX means the index alone supplies every column the query reads. Run in a node:22 container (Node 22.23.3) with node fields.mjs; there is nothing to install. SQLite prints an experimental-feature warning on stderr that is not part of the output below.
import { DatabaseSync } from 'node:sqlite';
const db = new DatabaseSync(':memory:');
db.exec(`
CREATE TABLE users(id INTEGER PRIMARY KEY, name TEXT);
CREATE TABLE articles(id INTEGER PRIMARY KEY, title TEXT, status TEXT, author_id INTEGER, body TEXT, updated_at INTEGER);
CREATE INDEX idx_articles_status_title ON articles(status, title);
`);
const ins = db.prepare('INSERT INTO articles VALUES (?,?,?,?,?,?)');
db.exec('BEGIN');
for (let i = 1; i <= 1000; i++) ins.run(i, 'Title ' + i, i % 2 ? 'published' : 'draft', (i % 10) + 1, 'x'.repeat(5000), 1000 + i);
for (let u = 1; u <= 10; u++) db.prepare('INSERT INTO users VALUES (?,?)').run(u, 'User ' + u);
db.exec('COMMIT');
// Whitelist: API field name -> how to fetch it. Anything not listed is rejected, never interpolated.
// Membership is checked with Object.hasOwn: `in` would also accept inherited names such as constructor or __proto__.
const FIELDS = {
id: { select: 'a.id' },
title: { select: 'a.title' },
status: { select: 'a.status' },
body: { select: 'a.body' },
updatedAt: { select: 'a.updated_at AS updatedAt' },
'author.name': { select: 'u.name AS authorName', join: 'JOIN users u ON u.id = a.author_id' },
};
const DEFAULT = ['id', 'title', 'status'];
function buildQuery(fieldsParam) {
const wanted = fieldsParam ? fieldsParam.split(',') : DEFAULT;
const bad = wanted.filter((f) => !Object.hasOwn(FIELDS, f));
if (bad.length) throw Object.assign(new Error('unknown fields: ' + bad.join(',')), { status: 400 });
const cols = [...new Set(['id', ...wanted])]; // id is always kept so the client can key rows
const joins = [...new Set(cols.map((c) => FIELDS[c].join).filter(Boolean))];
const sql = `SELECT ${cols.map((c) => FIELDS[c].select).join(', ')} FROM articles a ${joins.join(' ')} WHERE a.status = ? ORDER BY a.title LIMIT 20`;
return sql;
}
const plan = (sql) => db.prepare('EXPLAIN QUERY PLAN ' + sql).all('published').map((r) => r.detail).join(' | ');
const size = (sql) => Buffer.byteLength(JSON.stringify(db.prepare(sql).all('published')));
const cases = {
'no fields param (default)': undefined,
'fields=title,status': 'title,status',
'fields=title,body': 'title,body',
'fields=title,author.name': 'title,author.name',
};
for (const [label, f] of Object.entries(cases)) {
const sql = buildQuery(f);
console.log(`${label}\n SQL : ${sql}\n plan: ${plan(sql)}\n response bytes: ${size(sql)}`);
}
console.log('SELECT * for comparison, response bytes:', size('SELECT * FROM articles WHERE status = ? LIMIT 20'));
for (const attempt of ['title;DROP TABLE users', 'title,constructor', '__proto__']) { // inherited names such as constructor must not count as whitelisted
try { buildQuery(attempt); } catch (e) { console.log('rejected', JSON.stringify(attempt), '->', e.status, e.message); }
}
// Document store: the same whitelist drives a projection object.
const toMongoProjection = (fieldsParam) => {
const wanted = (fieldsParam ? fieldsParam.split(',') : DEFAULT);
if (wanted.some((f) => !Object.hasOwn(FIELDS, f))) throw new Error('unknown field');
return Object.fromEntries([...new Set(['id', ...wanted])].map((f) => [f === 'id' ? '_id' : f, 1]));
};
console.log('Mongo projection for fields=title,status:', JSON.stringify(toMongoProjection('title,status')));
It prints:
no fields param (default)
SQL : SELECT a.id, a.title, a.status FROM articles a WHERE a.status = ? ORDER BY a.title LIMIT 20
plan: SEARCH a USING COVERING INDEX idx_articles_status_title (status=?)
response bytes: 1033
fields=title,status
SQL : SELECT a.id, a.title, a.status FROM articles a WHERE a.status = ? ORDER BY a.title LIMIT 20
plan: SEARCH a USING COVERING INDEX idx_articles_status_title (status=?)
response bytes: 1033
fields=title,body
SQL : SELECT a.id, a.title, a.body FROM articles a WHERE a.status = ? ORDER BY a.title LIMIT 20
plan: SEARCH a USING INDEX idx_articles_status_title (status=?)
response bytes: 100813
fields=title,author.name
SQL : SELECT a.id, a.title, u.name AS authorName FROM articles a JOIN users u ON u.id = a.author_id WHERE a.status = ? ORDER BY a.title LIMIT 20
plan: SEARCH a USING INDEX idx_articles_status_title (status=?) | SEARCH u USING INTEGER PRIMARY KEY (rowid=?)
response bytes: 1056
SELECT * for comparison, response bytes: 101876
rejected "title;DROP TABLE users" -> 400 unknown fields: title;DROP TABLE users
rejected "title,constructor" -> 400 unknown fields: constructor
rejected "__proto__" -> 400 unknown fields: __proto__
Mongo projection for fields=title,status: {"_id":1,"title":1,"status":1}
What each case shows:
- Default and
fields=title,statusproduce the same query and a plan ofCOVERING INDEX idx_articles_status_title. The index was created on(status, title)and a SQLite table keyed byINTEGER PRIMARY KEYalso stores thatidin every index entry, soid,titleandstatusare all in the index: the database answers from the index and never touches the table rows. That is backend work saved, not only bytes. fields=title,bodystill uses the index for the filter but must read the table row forbody(the index holds nobody, which is why the plan loses the wordCOVERING), and the response grows from 1,033 to 100,813 bytes.fields=title,author.nameadds theusersjoin only for this request; the other plans have no join. The joined value comes back asauthorNameand the API layer nests it asauthor.name.SELECT *returns 101,876 bytes for the same 20 rows, about 99 times the default field set.- The injection attempt is rejected with
400before any SQL is built. So are inherited object names such asconstructorand__proto__: the membership test isObject.hasOwn(FIELDS, f), becausef in FIELDSwould accept them and produce a malformedSELECT a.id, FROM ...that fails as a server error instead of a400.
SQLite's plan wording is its own, but the mechanism carries over to other relational engines: a column list that the index contains lets the engine skip the row lookup.
The last line of the output shows the document-store side: the same whitelist yields the projection {"_id":1,"title":1,"status":1}. The projection includes _id, so under the MongoDB documentation's covered-query rules it is covered only when the index contains _id (otherwise the projection needs _id: 0, which the API then has to re-add if clients need ids). The covered-query rules are those in the MongoDB documentation's query-optimization page: every queried field and every returned field is in one index, and no queried field equals null.
How to verify it reduced backend work
- Compare
EXPLAIN(orEXPLAIN ANALYZEin PostgreSQL) for the default field set and a heavy one, as in the output above; in MongoDB, runexplain()on the projected query to confirm it is covered, meaning no documents are examined. - Load-test the default field set before and after: the metrics that should move are rows or buffers read (buffers are the fixed-size pages of table or index data the database reads from memory or disk; PostgreSQL's
EXPLAIN (ANALYZE, BUFFERS)reports them) and database CPU, not only response bytes. - Add a test that every public field name resolves, so a new field cannot be added without a data-layer mapping.
Trade-offs and pitfalls
- Index explosion. Covering every field combination is impossible; cover the few common sets (list views) and let rare combinations hit the table.
- Cache fragmentation. Every distinct
fieldsvalue is a separate cache key. Normalize the parameter (sort names, remove duplicates) before it becomes a key, and consider named field sets (view=summary) for the hot paths. - Computed fields that fetch more than they return. A field that quietly needs a join per row cancels the saving; declare dependencies and check the plan.
- GraphQL does not do this for you. A resolver that loads the full record and returns the requested fields trims only the response. The resolver (or a look-ahead, which means the resolver inspects which sub-fields the query selected before it fetches anything) has to pass the selection to the data layer.
- Authorization. Field selection must never expose a field the caller may not read; apply field-level permission checks after parsing and before building the query.
What are over-fetching and under-fetching in client-server APIs, and how do REST and GraphQL each deal with them? Describe what a frontend or mobile client feels from each approach.
Sample Answer
Direct answer
Over-fetching means the response carries more data than the screen uses. Under-fetching means one response is not enough, so the client must make follow-up requests to assemble a screen. Both waste the client's time, data allowance and battery. REST fixes them by shaping resources and endpoints on the server: sparse fieldsets (a fields= parameter that lists the only fields the client wants, shown below), embedded related data, or screen-specific endpoints. GraphQL, a query language where the server publishes a schema (a typed list of every object and field a client may ask for), fixes them by letting the client name exactly the fields it wants in a single request, at the price of a heavier server and a harder caching story.
What each problem looks like
Take a feed screen. Each row shows a post title, the author's name and a thumbnail.
- Over-fetching.
GET /postsreturns the whole post: body text, tags, view counts, edit history. The row uses three fields. If each post is 2,000 bytes and a row needs about 100, a page of 20 rows downloads 20 x 2,000 = 40,000 bytes to render 20 x 100 = 2,000 bytes of content, a factor of 20. The phone also parses all of it as JSON (JavaScript Object Notation) and holds it in memory. - Under-fetching.
GET /postsreturnsauthorIdbut not the name. If the 20 posts have 3 distinct authors, the client makes 1 + 3 = 4 requests (the pattern is called N+1: one request for the list, then one more per distinct related item). The author lookups cannot start until the list has arrived, so the screen waits for at least two sequential network round trips (list, then authors), and each extra hop on a weak cellular link is a chance to stall.
How REST deals with it
REST (resources addressed by URL, standard HTTP verbs) has no built-in query language, so you shape responses with conventions:
- Sparse fieldsets:
GET /posts?fields=id,title,author.name,thumbnailUrl. - Embedding or inclusion:
GET /posts?embed=authororinclude=author, so related data rides along (JSON:API, a public REST convention, definesfields[TYPE]andincludeand requires400 Bad Requestwhen a server does not supportinclude). - Screen-shaped endpoints:
GET /screens/feed, which belongs to a Backend-for-Frontend (BFF, a server layer owned by the client team that returns exactly what one client needs).
How GraphQL deals with it
GraphQL exposes one typed schema. The client sends a query that names fields, and the response mirrors its shape:
{ posts(first: 20) { id title author { name } thumbnailUrl(width: 96) } }
Over-fetching disappears because only named fields come back. Under-fetching disappears because the nested author { name } is resolved on the server in the same request (a resolver is the server function that produces the value of one field, here the author's name). Per the GraphQL over HTTP documentation, servers must accept POST for queries and mutations (mutations are the GraphQL operations that write data) and may accept GET for queries only. A response can hold both data and errors, and a response with non-null data should still use a 2xx status, because HTTP has no status for partial success.
What the client feels
| Concern | REST with fieldsets / BFF | GraphQL |
|---|---|---|
| Bytes per screen | Small if the server implements shaping; fixed endpoints over-fetch for other screens | Small by construction, per query |
| Round trips | 1 with embeds or a screen endpoint, otherwise 1 + N (one list request plus one per related item) | 1 per screen query |
| Client CPU | Parses a small JSON; no query building | Parses a small JSON; a client library may normalize results into a cache (split each response into one record per object, keyed by id, so an author shared by many posts is stored once), and that splitting and merging costs CPU on low-end devices |
| CDN (content delivery network) caching | Easy: URL is the cache key, GET is cacheable | Hard with POST; possible with GET plus persisted query ids (short identifiers that replace long query text, which also avoids URL length limits) |
| Developer velocity | Frontend waits for a new field or endpoint if the shape is missing | Frontend adds a field to its own query, no backend ticket, if the schema already exposes it |
| Versioning | New fields are additive; breaking shapes need /v2 or negotiation | Add fields freely and mark old ones with the built-in @deprecated directive (an annotation on a schema field that tells tools and developers the field is going away); usage is visible per field |
| Error handling | HTTP status per request tells the story | Often HTTP 200 with an errors array and partial data, so clients must inspect the body |
Which screens feel it most
- Personalized feed: many small rows, several related entities per row (author, counts, reaction). Under-fetching dominates, so one composed request (GraphQL query or BFF endpoint) matters most.
- Settings screen: one small object, rarely changing. Plain REST is fine and caches well; GraphQL adds little.
- Analytics sync: the client uploads or downloads bulk data in the background. The cost is bytes and radio wake-ups (every time the phone's cellular radio leaves idle to send a request it draws extra battery, so many small requests cost more energy than one batched request), not flexible field choice. Batch it into one request and compress it; a query language is not the lever.
Pitfalls
- GraphQL does not remove over-fetching on the server: a resolver (the function that produces one field) may still load the whole database row. Field selection saves wire bytes, not necessarily database work.
- A permissive
fields=parameter can leak internal columns. Validate it against an allowlist. - Choosing GraphQL to fix under-fetching, then discovering that each nested field issues its own database query (the N+1 problem), swaps a client problem for a server one.
You need one API strategy that serves mobile clients on poor connectivity and web clients that need rich interactivity. Would you serve both from the same API, a shared gateway, or separate client-specific backends? How would the protocol, payload shape and caching differ per client class?
Sample Answer
Direct answer
Run one set of shared domain services behind a thin gateway, with a separate BFF for each client class (a mobile BFF and a web BFF). Serving both from one generic API forces a compromise: either mobile pays for the web's richness in bytes and round trips, or web pays for mobile's frugality in chattiness. A shared gateway alone does not fix this, because it has no knowledge of any one client's screens. The BFFs differ in protocol tuning, payload shape and caching, while the business rules stay in the shared services.
Why the link, not the device, drives the split
A first-order model of one request on an open connection: time = RTT + payload bits / bandwidth, where RTT (round trip time) is how long a request and the first byte of its answer take. The link figures are assumptions (weak cellular 400 ms and 1 Mbit/s; office Wi-Fi 30 ms and 50 Mbit/s), and the model ignores slow start (a new connection begins by sending only a little data per round trip and speeds up as the network proves it can cope), packet loss (data lost in transit that must be sent again) and server time. Run in a node:22 container (Node 22.23.3) with node transfer.mjs:
// Time for one request/response on an already-open connection:
// rtt (request out, first byte back) + payload bits / bandwidth.
// Ignores TCP/QUIC slow start, packet loss and server time: a first-order model with ASSUMED link figures.
const links = {
'weak cellular (RTT 400 ms, 1 Mbit/s)': { rttMs: 400, mbps: 1 },
'office Wi-Fi (RTT 30 ms, 50 Mbit/s)': { rttMs: 30, mbps: 50 },
};
const ms = (link, bytes) => link.rttMs + (bytes * 8) / (link.mbps * 1e6) * 1000;
const plans = [
{ name: 'one composite response, 120 kB', calls: [120_000] },
{ name: 'one lean response, 30 kB', calls: [30_000] },
{ name: 'six sequential calls, 20 kB each (120 kB)', calls: Array(6).fill(20_000) },
{ name: 'lean 30 kB, then lazy 90 kB for below the fold', calls: [30_000, 90_000], firstOnly: true },
];
for (const [ln, link] of Object.entries(links)) {
console.log(ln);
for (const p of plans) {
const total = p.calls.reduce((t, b) => t + ms(link, b), 0);
const first = ms(link, p.calls[0]);
console.log(' ', p.name.padEnd(50), p.firstOnly ? `first content ${Math.round(first)} ms, all ${Math.round(total)} ms` : `all ${Math.round(total)} ms`);
}
}
It prints:
weak cellular (RTT 400 ms, 1 Mbit/s)
one composite response, 120 kB all 1360 ms
one lean response, 30 kB all 640 ms
six sequential calls, 20 kB each (120 kB) all 3360 ms
lean 30 kB, then lazy 90 kB for below the fold first content 640 ms, all 1760 ms
office Wi-Fi (RTT 30 ms, 50 Mbit/s)
one composite response, 120 kB all 49 ms
one lean response, 30 kB all 35 ms
six sequential calls, 20 kB each (120 kB) all 199 ms
lean 30 kB, then lazy 90 kB for below the fold first content 35 ms, all 79 ms
What this shows:
- On weak cellular, both terms hurt: six small sequential calls cost 3,360 ms because each pays 400 ms of RTT, and one 120 kB response costs 1,360 ms because 120 kB at 1 Mbit/s is 960 ms of transfer. The lean 30 kB response needs 640 ms.
- On Wi-Fi, bytes are nearly free (120 kB is 19 ms of transfer) and round trips dominate (six calls take 199 ms against 49 ms for one).
- Splitting a screen into a lean first response plus a lazy second one costs 1,760 ms in total on cellular, 400 ms more than a single response, but the user sees first content at 640 ms instead of 1,360 ms. That is the right trade for perceived speed on the mobile client.
So the mobile BFF optimizes for fewer bytes and fewer, resumable requests, and the web BFF optimizes for fewer round trips and interactivity; neither is a copy of the other.
What differs per client class
| Dimension | Mobile BFF (poor connectivity) | Web BFF (rich interactivity) |
|---|---|---|
| Protocol | HTTP/3 where available (the newest HTTP version, which runs over QUIC, a transport protocol built on UDP instead of TCP). In plain terms: stream multiplexing means many requests share one connection at the same time, and the point is that on QUIC one lost packet delays only the stream it belongs to. RFC 9114 maps HTTP onto QUIC, whose features include stream multiplexing and low-latency connection establishment, while the same RFC notes that over HTTP/2 on TCP a lost or reordered packet stalls all active transactions, which is the failure mode that hurts most on lossy cellular links (lossy meaning packets are often dropped). The practical rule is: send few requests, keep connections open, and let the platform negotiate the best protocol. | HTTP/2 or HTTP/3 through the CDN, many parallel requests are cheap (one connection carries them all); add streaming or push channels only where the product needs them (choosing that transport is its own design decision). |
| Payload shape | One composite response per screen, sparse fields, short lists with small first pages, image URLs chosen for the device, compact enums instead of repeated labels. | Resource-oriented or query-style endpoints that let the page fetch exactly what an interaction needs, richer objects (previews, permissions, related items) to avoid follow-up calls. |
| Request pattern | Idempotent GETs (reads, which are safe to repeat) with retry and backoff (wait longer after each failed attempt, for example 1 s, then 2 s, then 4 s, so a struggling server is not hammered); a cursor that survives app restart; first content separated from below-the-fold content (the part of the screen the user has to scroll to reach). | Prefetch on hover or route intent (start loading the next page's data when the pointer hovers a link or the router predicts the visit); optimistic updates (show the result of an action immediately and undo it if the server refuses); fine-grained revalidation (re-check only the one resource that may have changed). |
| Caching | ETag (a version tag for a response) and If-None-Match (the client sends the tag back) so a refresh can return 304 with no body when nothing changed; stale content shown immediately (stale-while-revalidate, which lets a cache reuse a stale response while it revalidates); private for per-user screens. | CDN for shared resources with s-maxage (a Cache-Control value that sets how long shared caches such as a CDN may reuse a response), browser HTTP cache for assets, and a client-side server-state cache (a library layer in the page that keeps fetched API data by key, serves it instantly and refetches it in the background) for API reads. |
| Compression | Compress text bodies; avoid sending fields at all rather than relying on compression, which cannot save the round trips. | Compress text bodies; field selection matters less than avoiding extra calls. |
| Versioning of the contract | Must support old app versions for months, so keep the contract additive. | Deployed with the page; the contract can change in step with the frontend release. |
Example of the same entity, a product card, as each BFF returns it:
{ "mobile": { "id": "pr_1", "t": "Trail shoe", "pc": 8900, "img": "https://img.example/pr_1/w640.jpg", "rt": 4.6 } }
{ "web": { "id": "pr_1", "title": "Trail shoe", "priceCents": 8900, "currency": "USD",
"images": [{ "url": "https://img.example/pr_1/w1290.jpg", "w": 1290, "h": 968 }],
"rating": { "average": 4.6, "count": 812 },
"variants": [{ "sku": "pr_1_42", "size": 42, "inStock": true }],
"permissions": { "canReview": true } } }
The mobile card uses short keys for a small, stable object. Short keys save bytes only when the response is large and compression is off; with compression the saving shrinks because repeated key names compress well, so use them where a profile of the real payload shows it matters, and never at the cost of a contract nobody can read.
Same API, shared gateway or separate backends
| Option | Fits when | Breaks when |
|---|---|---|
| Same API for both | One small team, similar screens, early product | Mobile starts needing a different shape and every change is negotiated with web |
| Shared gateway only | Edge concerns (auth, limits) are the main problem | Screen-specific shaping lands in a team-shared codebase |
| Gateway plus a BFF per client class | Different link quality, release cadence and screen needs, as here | Only one client exists, or no team can own a second backend |
The recommendation is the third row. What would flip it: if usage analytics show that mobile and web request nearly identical shapes, collapse to a single BFF.
Pitfalls
- Treating "mobile" as one thing: a phone on office Wi-Fi does not need the cellular profile. Let the client send a coarse profile name (such as lean or balanced) and keep the server choice overridable.
- Letting the two BFFs fork business logic. The shared services decide; the BFFs only shape.
- Optimizing the byte count and ignoring depth: the cellular numbers above show round trips costing as much as the payload.
- Caching a per-user screen publicly. Mark it
privateand cache upstream fragments instead.
Unlock Full Question Bank
Get access to all 18 Backend-for-Frontend and Client-Specific API Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.