API Contracts and Schema Design Questions
Defining the interface contract between producers and consumers: request/response payload shapes, data models, field-level validation, nullability, and enum/typing decisions (including safely evolving an enum's allowed values without breaking clients). Covers contract-first design with OpenAPI and JSON Schema (authoring specs, generating SDKs and mock servers, catching breaking changes in CI), mapping internal domain models to external DTOs as both evolve independently, and designing a stable error-response contract (structured error codes, correlation IDs, retryable classification). Also covers the leadership and behavioral practice of establishing and governing contract standards across teams. The contract is the durable artifact clients depend on.
Explain normalization versus denormalization in your API's response shape, from the consumer's perspective. When would you prefer a normalized response, and when would you intentionally duplicate data into a read model or a single response payload instead?
Sample Answer
From the consumer's side, normalization means the response gives you references (an ID) that you must resolve with a further call, while denormalization means the response embeds the related data directly so no follow-up call is needed. The right choice depends on how the client is actually going to use the data, not on how the backend's database happens to be structured.
When a normalized response is the better choice
Prefer references when the related data is large, changes independently of the primary resource, or is rarely needed by the majority of callers. Returning an order with just a customer_id (instead of the customer's full profile embedded) keeps the response small for the common case where a caller only needs the order itself, and avoids serving stale customer data that has since changed, since the client fetches the customer separately only when it actually needs it.
When denormalizing into the response is the right call
Prefer embedding data directly, effectively building a small denormalized read model into the response, when the consuming client needs it on nearly every call and a follow-up round trip would be a real cost: a mobile app rendering an order list screen that always shows the customer's name benefits from the API embedding customer_name directly, trading a slightly larger and slightly less "pure" response for avoiding an extra network round trip on every single request.
Worked example
An orders API serving two very different consumers:
- An internal reconciliation job that only needs order totals and IDs: a normalized response (just
customer_id, not the customer object) keeps the payload small across millions of rows and avoids duplicating customer data that changes independently. - A mobile order-history screen that always displays the customer's name and the first product's thumbnail: a denormalized response embedding
customer_nameandfirst_item_thumbnail_urldirectly saves that client two round trips it would otherwise make on every single screen load.
The same underlying order table can back both: the choice is made at the API-response layer, not at the storage layer, precisely because the two consumers have different access patterns.
Trade-offs and pitfalls
Denormalizing into a response means that data can go stale between the time it was embedded and the time the client reads it, so anything embedded needs a story for how staleness is acceptable (or how the client gets told an embedded value might be out of date). Normalizing everything, on the other hand, pushes N+1-style round trips onto every consumer regardless of their actual access pattern, penalizing the common case for the sake of theoretical purity. A frequent middle-ground pattern is offering both: a lean, normalized response by default, with an explicit expansion parameter (?expand=customer) that lets a caller opt into the denormalized fields only when it actually needs them.
Your team wants to add a computed field to an existing API response, but the value is expensive to calculate and could materially change the response time. How would you decide whether to compute it synchronously, cache it, materialize it in storage ahead of time, or expose it through a separate endpoint instead?
Sample Answer
The decision hinges on two independent questions: how often is this field actually read relative to how often the underlying data changes, and how expensive is the computation relative to the latency budget of the endpoint it would live on. Getting the pairing right (cheap-and-frequent versus expensive-and-rare) is what separates a good answer from a checklist of four generic options.
The four options, and when each wins
Compute synchronously, inline in the response. Only viable if the computation is genuinely cheap relative to the endpoint's existing latency budget; adding a moderately expensive computation directly into a hot-path response risks making every caller pay a cost that only some of them actually need.
Cache the computed value. A strong fit when the underlying data changes far less often than the field is read (a product's aggregate rating computed from reviews that arrive far less frequently than the product page is viewed): compute once, serve many times, with a clear invalidation trigger (recompute when a new review lands, or on a time-based TTL (time-to-live) if slight staleness is acceptable).
Materialize it in storage ahead of time. The right choice when the computation is too expensive to redo per-cache-miss and the read pattern is frequent and predictable enough to justify pre-computing and storing the result as data, updated by a background job or an event trigger rather than computed on any request path at all.
Expose it through a separate endpoint. The right choice when only a MINORITY of callers actually need the field, since folding an expensive computation into the primary response penalizes every caller (including the majority who never wanted it) to serve the few who do; a separate, explicitly-named endpoint lets the cost be paid only by the consumers who ask for it.
Worked example
A product page's response currently includes basic fields (name, price, description) served in under 50ms. The team wants to add a computed similar_products field requiring a moderately expensive similarity computation across the catalog, roughly 200ms.
- If most callers of this endpoint (analytics jobs, internal tooling) never use
similar_products, but the customer-facing web page always renders it: expose it through a separate endpoint (GET /products/{id}/similar) so the majority of callers keep their fast response, and the web page issues a second, parallel request specifically for the expensive field. - If the underlying catalog data changes only a few times a day but the product page is viewed millions of times: cache the computed value, recomputing on a catalog-change event rather than per-request, so the 200ms cost is paid rarely instead of on every page view.
Trade-offs and pitfalls
Folding an expensive field directly into the primary response "for convenience" is the most common mistake: it looks simpler in the short term (one endpoint, one call) but silently taxes every caller of that endpoint with the new field's cost, including callers who will never read it. The opposite mistake, splitting out a separate endpoint for a field nearly every caller actually needs, just adds an extra round trip for the common case; the decision genuinely depends on measuring who calls this endpoint and how they use the response, not on a general preference for either simplicity or slimness.
Describe a time you mentored an engineer who struggled with API design or data modeling. What was the gap, how did you coach them, and what told you they had internalized the lesson rather than just fixing one review comment?
Sample Answer
The strongest mentoring stories in this space are specific about the GAP (not "they needed to learn API design" but the exact recurring mistake, like designing responses around what the database table looked like instead of what the consumer actually needed) and specific about the evidence that the lesson actually transferred to a NEW situation, not just the one review comment that prompted the conversation.
What a strong answer covers
Naming the gap precisely. Vague framing ("they weren't very senior yet") is weak; a strong answer names the actual recurring pattern: consistently exposing internal field names verbatim in the API response, or treating every field as required because that was simpler to implement, or never considering what happens when a field needs to be removed later.
The coaching approach, not just the correction. Fixing the one pull request yourself, or leaving a detailed review comment on that one instance, teaches the engineer that this specific PR had a problem; walking through the underlying PRINCIPLE (why does field-shape decoupling matter, what breaks later if you skip it) with a concrete example from their own code is what actually builds the judgment to catch it themselves next time.
The evidence of internalization. This is the part that separates a real mentoring story from a generic one: what did you actually observe, later, that told you the lesson had stuck? Seeing them catch the same category of issue in someone ELSE's review, unprompted, or seeing them independently flag a design risk in their own next feature before you had to say anything, is much stronger evidence than "they said they understood" or "the next PR was fine" (which could just mean the next task happened not to touch that particular failure mode).
Worked example shape
"An engineer kept designing API responses that mirrored our internal database schema almost one-to-one, including internal-only status codes that made sense to us but meant nothing to an external consumer. Rather than rewriting their schema in review, I walked through one specific consequence with them: our internal status codes had already changed twice for operational reasons unrelated to the API's meaning, and each time, external clients who happened to be reading that raw code had broken. We renamed the field to a small, stable, contract-owned enum decoupled from the internal representation. Two features later, on an unrelated endpoint, they came to me proactively and said 'I want to make sure this new field isn't just mirroring our internal representation, can you sanity check the contract I've drafted before I start implementing?' That was the signal the lesson had generalized past the original example, not just fixed the one instance."
Trade-offs and pitfalls
A common weak spot in this kind of story is skipping straight to "and now they're great at this," without the intermediate evidence; interviewers are listening for the specific moment or signal that told you the coaching had worked, not just an assertion that it did. Equally common: over-indexing on a single correction rather than the underlying principle, which produces an engineer who avoids that ONE mistake in the future but hasn't developed the more general judgment to catch a different symptom of the same root problem.
Walk me through a time you helped your team improve API or data-model consistency by introducing a review practice, template, or standard. How did you get buy-in, and what measurable outcome showed the change was worth it?
Sample Answer
The strongest version of this story treats "getting buy-in" as its own real problem to solve, not an afterthought after the standard was written: a review checklist or template that nobody asked for and that adds friction to every pull request will be quietly ignored, no matter how sound its content is, unless the people adopting it were part of shaping it or can see a concrete problem it actually prevented.
What a strong answer walks through
The recurring problem that motivated the standard. Name a specific, repeated pain point (three separate incidents where a field was renamed without warning, or a recurring pattern of inconsistent error shapes across services making client-side error handling brittle), not an abstract desire for "more consistency."
How you got buy-in, specifically. Did you pilot the template on your own team first and bring results, not just a proposal, to a wider forum? Did you involve the engineers who would actually use it in drafting it, so it reflected real workflow rather than an ideal imposed from outside? Buy-in earned through a concrete, already-demonstrated win is far more durable than buy-in secured by a mandate from above.
The measurable outcome. A credible answer names something you could actually point to: a drop in a specific class of production incident, a measured reduction in review-comment cycles for API design discussions, or an increase in some adoption metric (percentage of new endpoints using the reviewed template) tracked over a real time window, not a vague "things got better."
Worked example shape
"We had three separate incidents in one quarter where a field was silently renamed or retyped in an API response, each caught only after a client broke in production. I proposed a lightweight schema-change checklist as part of the pull request template, but rather than mandating it top-down, I piloted it on my own team's PRs for a month first, tracking how many schema-affecting changes it actually caught before merge. It caught two would-be breaking changes in that pilot month alone. I brought that concrete result, not just the proposal, to the broader engineering review, and adoption followed because people could see it had already prevented two real incidents rather than being a hypothetical improvement. Six months after wider rollout, breaking-change incidents traced to unreviewed schema changes dropped from three in the prior quarter to zero."
Trade-offs and pitfalls
A common weak point is describing the STANDARD in detail while glossing over how buy-in was actually secured, which is usually the harder and more interesting half of this story; a checklist's content is rarely the reason adoption succeeds or fails. Another common weak spot: an outcome metric that is really just "adoption happened" rather than a downstream result (fewer incidents, faster reviews) the standard was actually meant to produce; adoption is a leading indicator, not the outcome itself.
Tell me about a time you had to push back on an API or data-model design that seemed simpler for short-term delivery but would have created long-term pain for downstream consumers. How did you make the case, and what was the final decision?
Sample Answer
The strongest version of this story is not "I was right and they were wrong," it's showing the concrete, specific downstream cost the simpler design would have created, translated into terms the other side of the table actually cared about (a delivery deadline, a support burden, a migration cost six months out), and then describing how the actual decision got made once that cost was visible to everyone.
What a strong answer walks through
The situation: name the specific design choice under pressure (a schema shortcut, a field reused for two purposes, a response shape copy-pasted from an unrelated endpoint) and the delivery pressure driving it, honestly, without caricaturing the other side's position as simply wrong.
The case you made: the strongest version of this case is concrete, not principled in the abstract. "This will be hard to maintain" rarely moves a deadline; "this specific field reuse means every future consumer has to special-case whether this response came from path A or path B, and we already have three planned features that will need to tell them apart" is something a stakeholder can actually weigh against the deadline.
How the decision actually got made: did you get the design changed outright, negotiate a smaller fix with a follow-up ticket to do it properly, or lose the argument and later have to deal with the consequence you predicted? All three are legitimate answers; the weakest version of this story claims a clean win with no friction, which reads as either an easy problem or a polished retelling.
The outcome: what happened afterward, concretely, that validates (or complicates) the case you made at the time.
Worked example shape
"We were under a two-week deadline to ship a partner integration, and the plan was to reuse an existing internal user_id field on the response as the partner-facing identifier, saving us from adding a new field and updating a few internal services. I pushed back because I could see two features already on the roadmap that would need a distinct partner-facing identifier separate from the internal one (partner-scoped API keys, and a planned data-residency requirement that needed to know which records were partner-visible). I proposed adding a new partner_ref field instead, which took two extra days but meant we did not have to do a breaking migration eight weeks later when the partner-scoped-API-keys feature landed and genuinely needed that separation. The team agreed to the two extra days once the specific upcoming conflict was visible, not because of a general 'good practice' argument."
Trade-offs and pitfalls
A common weak point in this story is framing it as pure technical purism ("clean code" or "best practices") rather than a concrete, forecastable cost; interviewers are listening for whether you can translate a technical concern into terms a non-technical stakeholder, or a deadline-driven engineering lead, would actually act on. Equally common: telling this story as an unqualified win when the honest answer is more nuanced (you got a smaller compromise, or the team shipped the shortcut anyway and you were right six months later); either honest version is stronger than a suspiciously frictionless win.
Unlock Full Question Bank
Get access to all 8 API Contracts and Schema Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.