API Contracts and Schema Design Questions
Defining the interface contract between producers and consumers: request/response payload shapes, data models, field-level validation, nullability, and enum/typing decisions (including safely evolving an enum's allowed values without breaking clients). Covers contract-first design with OpenAPI and JSON Schema (authoring specs, generating SDKs and mock servers, catching breaking changes in CI), mapping internal domain models to external DTOs as both evolve independently, and designing a stable error-response contract (structured error codes, correlation IDs, retryable classification). Also covers the leadership and behavioral practice of establishing and governing contract standards across teams. The contract is the durable artifact clients depend on.
Compare JSON and Protocol Buffers (protobuf) as serialization formats for APIs. Discuss trade-offs across latency, payload size, schema evolution, developer ergonomics, client diversity, human-readability, and tooling. Provide scenarios where you would choose JSON for a public REST API and where you would choose protobuf for internal, high-throughput communication, and why.
Sample Answer
A strong first answer: pick JSON when you need human-readable, browser-friendly, loosely-coupled contracts (a public REST API, a webhook payload, anything a developer might read in a curl response), and pick Protobuf when you control both ends of the wire (the wire = the literal bytes sent over the network, not the source code or an in-memory value) and need small, fast, strongly-typed messages (internal gRPC calls [gRPC: a common RPC framework built on Protobuf], high-throughput or high-volume traffic between your own services).
The axes that actually decide it
Payload size and latency. Protobuf is a binary, tag-length-value format: field names never travel on the wire, only small integer field numbers do, and values are packed (varints -- a compact encoding that uses fewer bytes for smaller integers -- for integers, no quoting for strings). JSON repeats every field name as a string on every message. For small, frequent messages the difference in serialization/deserialization cost and wire size adds up.
Schema evolution. Schema evolution is where the format choice matters most, because API contracts keep changing long after they first ship, and the two formats handle that change very differently. Protobuf ties every field to a stable field NUMBER, not to its position or name: a reader that does not recognize a field number simply skips it, so adding a new field is safe by construction, and renaming a field is free (the name is compile-time only, never on the wire). The one hard rule is that a field number must never be reused for a different meaning once it has shipped. JSON has no built-in schema at all; "schema evolution" for JSON is really "JSON Schema evolution," and safety depends entirely on the discipline the team layers on top (never treating a field as positionally required, always tolerating unknown fields).
Developer ergonomics and human-readability. JSON wins here outright: you can read it in a browser network tab, log it, curl it, and paste it into a bug report with zero tooling. Protobuf messages are opaque bytes without the compiled descriptor; debugging typically means adding a debug JSON-encoding path or using a tool that understands the .proto file.
Client diversity and tooling. JSON needs nothing beyond a standard library in essentially every language ever shipped, which matters when the API is public and you do not control the client. Protobuf needs the compiler and generated stubs for each client language, which is a real integration cost for external, unknown consumers but a small one-time cost inside a service mesh you already control.
Worked example: what the format choice actually costs on the wire
Take a tiny, realistic message: {"id": 42, "name": "widget", "in_stock": true}.
As compact JSON (no extra whitespace), this is:
{"id":42,"name":"widget","in_stock":true}
That is 41 bytes.
Encoding the same three fields as Protobuf (field 1 = id, varint; field 2 = name, length-delimited; field 3 = in_stock, varint), by hand-rolling the tag-length-value bytes. Every field starts with a tag byte, and that byte is not arbitrary: it packs the field number together with a wire type (a small numeric code -- 0 for varint, 2 for length-delimited -- that tells the decoder how many bytes to read and how to interpret them) using the formula tag = (field_number << 3) | wire_type. You can verify each one by hand below:
- Field 1 (varint, wire type 0): tag = (1 << 3) | 0 = 8 =
0x08; value 42 fits in one varint byte0x2A-> 2 bytes. - Field 2 (length-delimited, wire type 2): tag = (2 << 3) | 2 = 18 =
0x12; length byte0x06, then the 6 raw ASCII bytes of "widget" -> 8 bytes. - Field 3 (varint, wire type 0): tag = (3 << 3) | 0 = 24 =
0x18; valuetrueencodes as0x01-> 2 bytes.
Total: 12 bytes (08 2a 12 06 77 69 64 67 65 74 18 01).
reduction=(1−4112)×100%≈70.7%
That 70.7% is specific to this exact 3-field, mostly-numeric message; it is not a universal constant. Industry write-ups on protobuf vs. JSON commonly report a broader range (roughly 50 to 85% smaller, 3 to 10 times faster to parse) across many message shapes, which is consistent in direction with this worked example but is a separately-reported range, not something this specific calculation proves on its own.
When the axes actually point in different directions
For an internal, high-throughput RPC path exchanging large model metadata or batched inputs between services you own end to end, Protobuf's size and schema-evolution guarantees dominate and the tooling cost is a one-time investment. For a public REST API serving a heterogeneous client base (mobile, web, and IoT devices you do not control), JSON's zero-tooling accessibility usually outweighs the wire-format savings, unless payloads are large enough (media, bulk exports) that the size difference becomes the bottleneck.
You need to transform internal domain models into external API response objects (DTOs) and back. Describe patterns for this mapping (hand-written mappers, code generation, annotation-based), how you validate the DTO independently of the domain model, and how you handle optional or deprecated fields in the DTO as the underlying domain model evolves.
Sample Answer
The core idea: never let the external API contract be a direct, unmediated reflection of the internal domain model. A dedicated Data Transfer Object (DTO) layer sits between them, and every mapping pattern below exists to answer one question: when the domain model changes, does the contract have to change too, or can the mapping layer absorb it?
Mapping patterns, from least to most automated
Hand-written mappers. A function (or small class) that explicitly reads fields off the domain object and constructs the DTO field by field. Slower to write, but every mapping decision is visible in code review, which matters most when the domain model and the DTO genuinely need to diverge (a domain field that should never be exposed, or a computed field the DTO needs that the domain object doesn't have).
Code generation. A build step generates the mapper from a declarative description (often the same schema used for the OpenAPI spec), trading hand-written flexibility for consistency and less boilerplate across dozens of similar DTOs. Works best when most mappings really are one-to-one field copies and the exceptions are rare enough to special-case.
Annotation-based mapping. Framework-level annotations on the domain class or DTO class (common in typed languages) tell a mapping library how to convert between them at runtime or compile time, which is fast to set up but couples the domain model's annotations to the API's concerns, a coupling the whole DTO pattern usually exists to avoid.
Validating the DTO independently of the domain model
Validation belongs on the DTO, not (only) on the domain model, because the DTO is what an untrusted client actually sent: a required-field check, a string-length limit, or an enum constraint enforced on the DTO catches a malformed request before it ever reaches domain logic. If validation only lives on the domain model, a malformed DTO can be silently coerced into a domain object that never should have existed.
Handling optional or deprecated fields as the domain model evolves
This is where the mapping layer earns its cost. When the domain model adds a new internal field, the DTO does not need to change at all unless that field is meant to become part of the public contract, which keeps internal refactoring from becoming an API-breaking event. When a domain field the DTO exposes gets deprecated internally, the mapper can keep populating the DTO field from a fallback or a computed value during a transition window, so the external contract stays stable even while the internal implementation moves out from under it.
Worked example
A domain Order object internally tracks internalRiskScore (never exposed) and legacyStatusCode (being replaced by a new status enum). The mapper: copies orderId, items, and totalAmount straight across; omits internalRiskScore entirely; and populates the DTO's public status field by translating legacyStatusCode through a small lookup table, so a client sees the contract's stable enum values (pending, shipped, delivered) even while the backend migrates its own internal status representation over several releases without a single external-facing change.
Trade-offs and pitfalls
The most common failure is skipping the mapping layer entirely for speed ("just serialize the domain object directly") and discovering months later that every internal refactor is now a potential breaking API change, because clients had been silently depending on whatever internal fields happened to serialize. The dedicated DTO layer costs real development time up front; it pays that back the first time the domain model needs to change in a way the API contract must not reflect.
Design the API contract between frontend dashboards and a metrics microservice: the request/response shape for a paginated tile of metrics, how you would evolve that contract over time without breaking existing dashboard clients, and what you would classify as an additive change versus a breaking one.
Sample Answer
The contract between a dashboard frontend and a metrics microservice needs two things nailed down explicitly: a stable shape for a single paginated "tile" of metrics data, and a documented rule for what counts as additive versus breaking, since dashboards tend to be built by a different team than the metrics service and neither wants a silent shape change to break the other's release schedule.
The tile response shape
A single tile response should carry: the metric's identifier and display label, the paginated data points themselves (each with a timestamp and value), and pagination metadata (a cursor or page token -- an opaque value representing "where you left off" in the result set, unlike a page number the client cannot compute or guess on its own -- and a flag for whether more data exists). Keeping the SHAPE of one tile's response consistent across every metric type (rather than a different shape per metric) means the dashboard's rendering code does not need a special case per metric.
Evolving the contract without breaking dashboard clients
Additive, and therefore safe without coordination: adding a new optional field to a data point (a confidence interval, a data-quality flag), adding a new tile type that reuses the existing envelope shape, adding new optional query parameters for filtering.
Breaking, and therefore requiring coordination and a version bump: renaming or removing an existing field the dashboard already renders, changing a field's type (a value going from a number to a string), changing what the pagination cursor means in a way that invalidates cursors a client might have cached.
Worked example
A metrics service wants to add a rolling 7-day trend indicator alongside each metric's current value. Because this is a NEW optional field (trend_direction: "up" | "down" | "flat", absent unless computed) added to the existing tile shape, dashboard clients that do not yet render it simply ignore the extra field and keep working unmodified; dashboard clients that DO want to render it can start doing so on their own release schedule, entirely decoupled from the metrics service's deploy. Contrast that with changing the existing value field from a bare number to an object with value and unit nested inside: that is a breaking type change to a field every existing dashboard already parses, and would need a new tile-schema version with a migration window, not a same-day deploy.
Trade-offs and pitfalls
The most common failure mode on a contract like this is a metrics team treating "the dashboard team hasn't complained yet" as evidence a change was safe, when the real signal (a stated additive-versus-breaking policy, checked in CI (continuous integration)) was never actually enforced. Without an explicit, written rule for what counts as additive on THIS specific tile contract, both teams end up relying on tribal knowledge about what "should" be safe, which breaks down the moment either team has new engineers who were not there when the informal rule was established.
Your team wants to add a computed field to an existing API response, but the value is expensive to calculate and could materially change the response time. How would you decide whether to compute it synchronously, cache it, materialize it in storage ahead of time, or expose it through a separate endpoint instead?
Sample Answer
The decision hinges on two independent questions: how often is this field actually read relative to how often the underlying data changes, and how expensive is the computation relative to the latency budget of the endpoint it would live on. Getting the pairing right (cheap-and-frequent versus expensive-and-rare) is what separates a good answer from a checklist of four generic options.
The four options, and when each wins
Compute synchronously, inline in the response. Only viable if the computation is genuinely cheap relative to the endpoint's existing latency budget; adding a moderately expensive computation directly into a hot-path response risks making every caller pay a cost that only some of them actually need.
Cache the computed value. A strong fit when the underlying data changes far less often than the field is read (a product's aggregate rating computed from reviews that arrive far less frequently than the product page is viewed): compute once, serve many times, with a clear invalidation trigger (recompute when a new review lands, or on a time-based TTL (time-to-live) if slight staleness is acceptable).
Materialize it in storage ahead of time. The right choice when the computation is too expensive to redo per-cache-miss and the read pattern is frequent and predictable enough to justify pre-computing and storing the result as data, updated by a background job or an event trigger rather than computed on any request path at all.
Expose it through a separate endpoint. The right choice when only a MINORITY of callers actually need the field, since folding an expensive computation into the primary response penalizes every caller (including the majority who never wanted it) to serve the few who do; a separate, explicitly-named endpoint lets the cost be paid only by the consumers who ask for it.
Worked example
A product page's response currently includes basic fields (name, price, description) served in under 50ms. The team wants to add a computed similar_products field requiring a moderately expensive similarity computation across the catalog, roughly 200ms.
- If most callers of this endpoint (analytics jobs, internal tooling) never use
similar_products, but the customer-facing web page always renders it: expose it through a separate endpoint (GET /products/{id}/similar) so the majority of callers keep their fast response, and the web page issues a second, parallel request specifically for the expensive field. - If the underlying catalog data changes only a few times a day but the product page is viewed millions of times: cache the computed value, recomputing on a catalog-change event rather than per-request, so the 200ms cost is paid rarely instead of on every page view.
Trade-offs and pitfalls
Folding an expensive field directly into the primary response "for convenience" is the most common mistake: it looks simpler in the short term (one endpoint, one call) but silently taxes every caller of that endpoint with the new field's cost, including callers who will never read it. The opposite mistake, splitting out a separate endpoint for a field nearly every caller actually needs, just adds an extra round trip for the common case; the decision genuinely depends on measuring who calls this endpoint and how they use the response, not on a general preference for either simplicity or slimness.
You need to add new enum values to a public API response without breaking older clients that may error or crash on an unrecognized value. Propose both server-side and client-side strategies to make this safe, and describe how you would migrate the schema and communicate the change to minimize client impact.
Sample Answer
The safest approach combines both sides: design the client to treat unrecognized enum values as a known, handled case (never a crash), and design the server rollout so the new value only starts appearing after clients have had a chance to deploy that handling.
Server-side strategies
Prefer strings over closed enums where the value set is expected to grow. If the API's own schema does not hard-enforce a closed enum (or uses an "open" pattern that documents known values without technically restricting the wire format to them), adding a new value is a purely additive change with nothing to coordinate on the server side.
Stage the rollout. Add the new value to the documentation and to a small percentage of test traffic before it appears broadly, giving client teams a window to deploy defensive handling against a real, already-defined value rather than a hypothetical future one.
Version the schema explicitly if the enum is truly closed (a formal JSON Schema enum constraint with no forward-compatibility clause): adding a value is then treated as a breaking change requiring a new API version, which is the conservative, slower, but bulletproof path.
Client-side strategies
Always include a default/unknown branch. Any code that switches on the enum's value needs an explicit "unrecognized value" branch, treated as a real code path to test, not an afterthought: log it, degrade gracefully (show a generic label instead of crashing), and move on.
Treat the enum as open by default, even if it currently only has three values, unless the contract explicitly documents it as closed. Writing an exhaustive switch statement with no default case is the single most common way this breaks in practice.
Worked example: migrating an order-status enum
An order-status field currently has values pending, shipped, delivered. The team wants to add returned.
- Client audit first: confirm (via API usage logs, or a client-side telemetry event on the "unknown status" branch) that consuming clients actually have a default/unknown handler; if telemetry shows zero hits on an unknown-value path today, that is itself informative (either nobody has an unknown handler, worth flagging before rollout, or the handler exists and has simply never fired yet).
- Document and communicate
returnedas a new value, giving client teams a fixed window (for example, 30 days) to confirm their unknown-value handling covers it, even though nothing about the wire contract has changed yet. - Start emitting
returnedin production, initially on a small fraction of traffic (a single test account, then a small tenant cohort) to catch a client with no real default handling before it hits every consumer at once. - Roll out to all traffic, monitoring the same "unknown status" telemetry sinks to see the metric spike briefly then fall to zero as clients pick up the new value in their own next deploys.
Trade-offs and pitfalls
The most common mistake is treating "unknown enum value" handling as a theoretical concern until the first time it actually happens in production, at which point some client somewhere is crashing on the very first real-world new value, with no staged rollout to catch it earlier. The conservative alternative, always requiring a new major version for any new enum value, is bulletproof but slow, and if applied to every genuinely additive value, teams tend to route around it by not modeling values as enums at all and losing the schema's documentation value entirely.
Unlock Full Question Bank
Get access to all 15 API Contracts and Schema Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.