API Contracts and Schema Design Questions
Defining the interface contract between producers and consumers: request/response payload shapes, data models, field-level validation, nullability, and enum/typing decisions (including safely evolving an enum's allowed values without breaking clients). Covers contract-first design with OpenAPI and JSON Schema (authoring specs, generating SDKs and mock servers, catching breaking changes in CI), mapping internal domain models to external DTOs as both evolve independently, and designing a stable error-response contract (structured error codes, correlation IDs, retryable classification). Also covers the leadership and behavioral practice of establishing and governing contract standards across teams. The contract is the durable artifact clients depend on.
You need to add new enum values to a public API response without breaking older clients that may error or crash on an unrecognized value. Propose both server-side and client-side strategies to make this safe, and describe how you would migrate the schema and communicate the change to minimize client impact.
Sample Answer
The safest approach combines both sides: design the client to treat unrecognized enum values as a known, handled case (never a crash), and design the server rollout so the new value only starts appearing after clients have had a chance to deploy that handling.
Server-side strategies
Prefer strings over closed enums where the value set is expected to grow. If the API's own schema does not hard-enforce a closed enum (or uses an "open" pattern that documents known values without technically restricting the wire format to them), adding a new value is a purely additive change with nothing to coordinate on the server side.
Stage the rollout. Add the new value to the documentation and to a small percentage of test traffic before it appears broadly, giving client teams a window to deploy defensive handling against a real, already-defined value rather than a hypothetical future one.
Version the schema explicitly if the enum is truly closed (a formal JSON Schema enum constraint with no forward-compatibility clause): adding a value is then treated as a breaking change requiring a new API version, which is the conservative, slower, but bulletproof path.
Client-side strategies
Always include a default/unknown branch. Any code that switches on the enum's value needs an explicit "unrecognized value" branch, treated as a real code path to test, not an afterthought: log it, degrade gracefully (show a generic label instead of crashing), and move on.
Treat the enum as open by default, even if it currently only has three values, unless the contract explicitly documents it as closed. Writing an exhaustive switch statement with no default case is the single most common way this breaks in practice.
Worked example: migrating an order-status enum
An order-status field currently has values pending, shipped, delivered. The team wants to add returned.
- Client audit first: confirm (via API usage logs, or a client-side telemetry event on the "unknown status" branch) that consuming clients actually have a default/unknown handler; if telemetry shows zero hits on an unknown-value path today, that is itself informative (either nobody has an unknown handler, worth flagging before rollout, or the handler exists and has simply never fired yet).
- Document and communicate
returnedas a new value, giving client teams a fixed window (for example, 30 days) to confirm their unknown-value handling covers it, even though nothing about the wire contract has changed yet. - Start emitting
returnedin production, initially on a small fraction of traffic (a single test account, then a small tenant cohort) to catch a client with no real default handling before it hits every consumer at once. - Roll out to all traffic, monitoring the same "unknown status" telemetry sinks to see the metric spike briefly then fall to zero as clients pick up the new value in their own next deploys.
Trade-offs and pitfalls
The most common mistake is treating "unknown enum value" handling as a theoretical concern until the first time it actually happens in production, at which point some client somewhere is crashing on the very first real-world new value, with no staged rollout to catch it earlier. The conservative alternative, always requiring a new major version for any new enum value, is bulletproof but slow, and if applied to every genuinely additive value, teams tend to route around it by not modeling values as enums at all and losing the schema's documentation value entirely.
Design an error contract for an API that aggregates calls to multiple third-party services. The contract should expose meaningful high-level errors to consumers while masking internal or third-party-sensitive details. Include how you would categorize transient versus permanent errors and propagate a correlation ID for debugging.
Sample Answer
An aggregator sitting in front of several third-party services needs its own STABLE error vocabulary, translated from whatever each third party actually returns, so that a change to a partner's internal error format never leaks through as a breaking change to the aggregator's own consumers.
Categorizing transient versus permanent errors
Transient (the caller should retry, possibly after a delay): a third-party timeout, a rate limit from the partner, a temporary partner outage. Permanent (retrying will never help without a code change or user action): the partner rejected the request as fundamentally invalid, an authentication failure with the partner's credentials, a resource that genuinely does not exist. This classification, not the raw partner error, is what the aggregator's OWN error contract should expose, because a consumer of the aggregator should not need to know which specific third party was involved to decide whether retrying makes sense.
Masking internal and third-party-sensitive detail
The aggregator's public error response should never leak a partner's internal error codes, stack traces, or account-specific detail verbatim; those get logged internally (tied to the correlation ID) for debugging, while the public-facing error exposes only the aggregator's own stable vocabulary (upstream_timeout, upstream_rejected, and so on) plus a correlation ID a consumer can hand back for support escalation.
Propagating a correlation ID for debugging
A single correlation ID, generated when the aggregator receives the original request, should be threaded through every downstream call to every third party and included in every log line on both sides of the boundary. When something goes wrong three services deep, that one ID is what lets an engineer reconstruct the whole call chain instead of correlating timestamps across three different systems' logs by hand.
Worked example
The aggregator calls a shipping-rate partner whose gateway times out during a transient outage, returning a partner-specific error like ERR_GATEWAY_TIMEOUT with an internal partner request ID in the body. The aggregator's response to ITS OWN consumer never repeats that partner error verbatim; instead it returns:
{
"error": {
"code": "upstream_unavailable",
"category": "transient",
"message": "A shipping provider is temporarily unavailable. Please retry.",
"correlation_id": "req_a91f2b3c"
}
}
Internally, the aggregator's logs (searchable by req_a91f2b3c) retain the full partner error detail (including the raw ERR_GATEWAY_TIMEOUT code and the partner's own internal request ID) for an engineer investigating the incident, while the consumer only ever sees the stable, categorized, non-sensitive shape. Contrast this with a DIFFERENT partner error, ERR_CARRIER_ACCT_SUSPENDED (the aggregator's own account with that carrier has been suspended over a billing dispute): even though it also arrives from a third party, it belongs in the PERMANENT bucket, not transient, because no amount of client-side retrying resolves a suspended account. That failure should map to a distinct code (upstream_rejected) with "category": "permanent" and a message that does not invite a retry, since telling a client to retry a failure that only a human resolving a billing dispute can fix wastes capacity on both sides and delays anyone noticing the real, unretryable problem.
Trade-offs and pitfalls
Masking too aggressively can leave consumers unable to distinguish genuinely different failure modes that they need to handle differently (treating every upstream failure as one generic "something went wrong" code removes the transient-versus-permanent signal that makes the categorization useful in the first place). The opposite failure, passing partner error detail through unmodified "to be helpful," ties the aggregator's own contract to every partner's internal error format, so a partner changing their error codes becomes a breaking change for the aggregator's consumers even though nothing about the aggregator's own contract changed on purpose.
Write a minimal OpenAPI 3.1 specification (YAML) for a POST /orders endpoint on an e-commerce service. Include the request body schema (items: an array of objects with product_id and quantity), a 201 success response schema (order_id, total_amount), 400 and 401 error responses, and an example request and response embedded in the spec.
Sample Answer
The spec needs a request body schema for the array of order items, a 201 response schema for the created order, and explicit 400 and 401 error responses, each with an embedded example so a reader can see the exact shape without inferring it from the schema alone.
The specification
openapi: 3.1.0
info:
title: Orders API
version: "1.0.0"
paths:
/orders:
post:
operationId: createOrder
requestBody:
required: true
content:
application/json:
schema:
type: object
required: [items]
properties:
items:
type: array
minItems: 1
items:
type: object
required: [product_id, quantity]
properties:
product_id:
type: string
quantity:
type: integer
minimum: 1
example:
items:
- product_id: "sku_1001"
quantity: 2
- product_id: "sku_2044"
quantity: 1
responses:
"201":
description: Order created
content:
application/json:
schema:
type: object
required: [order_id, total_amount]
properties:
order_id:
type: string
total_amount:
type: number
example:
order_id: "ord_7f3a"
total_amount: 64.50
"400":
description: Validation error
content:
application/json:
schema:
type: object
required: [error_code, message]
properties:
error_code:
type: string
message:
type: string
example:
error_code: "invalid_quantity"
message: "quantity must be at least 1"
"401":
description: Missing or invalid credentials
content:
application/json:
schema:
type: object
properties:
error_code:
type: string
example:
error_code: "unauthorized"
This document was parsed and structurally checked with a YAML parser: it parses cleanly as valid YAML, declares openapi: 3.1.0, defines the /orders path with the post operation described above, and every example object matches the shape its sibling schema declares (the 201 example has both order_id and total_amount, matching the required list; the 400 example has both error_code and message).
Design decisions worth calling out
itemsrequiresminItems: 1, so an order with zero line items is rejected at the schema level rather than reaching business logic.quantityhasminimum: 1, catching a zero-or-negative quantity the same way, at the contract layer.- The 400 response has a structured shape (
error_codeplusmessage), not a bare string, so a client can branch onerror_codeprogrammatically instead of pattern-matching on human-readable text.
Trade-offs and pitfalls
Keeping this spec minimal (one endpoint, three responses) makes it easy to read in an interview or a code review, but a production version would need to decide, and document, what OTHER error shapes exist (a 404 if referencing a nonexistent product_id, a 409 for a stock conflict) so the response schema doesn't quietly grow undocumented shapes over time as edge cases get patched in. A second real trade-off: embedding literal examples directly in the spec (as done here) is excellent for readability and for generating realistic mocks, but every example needs to be kept in sync with the schema by hand or by a linter; an example that drifts from its own schema is worse than no example, because it actively misleads a reader.
Draft a concise API contract (endpoints, request/response examples, and audit behavior) for an internal feature-flagging service used by multiple teams. Include endpoints for creating a flag, evaluating a flag on a low-latency path, toggling a flag, and retrieving audit logs. Note the performance and safety considerations your schema choices need to account for.
Sample Answer
A feature-flag service's contract centers on one asymmetry: the evaluate path is read constantly, on a hot path, by every service that checks a flag, while create/toggle/audit are low-volume management operations. The contract should reflect that asymmetry directly in its shape, not just in the implementation behind it.
The four endpoints
Create a flag (POST /flags): accepts a flag key, a human-readable description, and a default value, and returns the created flag's identifier and current state. This is a low-frequency, administrative operation, so its response can afford to be verbose (full flag metadata) without any latency concern.
Evaluate a flag (GET /flags/{key}/evaluate?context=...): the hot path. The request carries an evaluation context (user ID, environment, or other targeting attributes); the response should be as small as possible, essentially just the resolved boolean or variant value, because this call happens on every request path that checks the flag and any extra payload weight is paid millions of times over.
Toggle a flag (PATCH /flags/{key}): flips a flag on or off (or changes its targeting rules), and should return the new state plus a version or timestamp so a caller can confirm the toggle actually took effect and see what it changed from.
Retrieve audit logs (GET /flags/{key}/audit): returns a paginated history of who changed the flag, when, and what the before/after state was, which is the accountability trail for a service that can change application behavior without a code deploy.
Worked example: the evaluate response shape
{
"flag_key": "new_checkout_flow",
"value": true,
"reason": "targeting_rule_match"
}
Compare that to a create response, which can be much richer:
{
"flag_key": "new_checkout_flow",
"description": "Enables the redesigned checkout flow for eligible users",
"default_value": false,
"created_at": "2026-07-28T12:00:00Z",
"version": 1
}
The evaluate response is deliberately minimal (three small fields); the create response can afford ten times the payload because it runs orders of magnitude less often.
Performance and safety considerations the schema has to account for
- Evaluate must be cacheable and fast to parse, so its schema stays flat and small; nesting a full targeting-rule explanation into every evaluate call would add latency to the busiest path in the system for the sake of a debugging convenience almost nobody needs on every call.
- Toggle must be auditable by construction, not as an afterthought: every toggle response including a version number means a caller (and the audit log) can always answer "what did this flag look like immediately before this specific change."
- The evaluation context in the evaluate request needs a stable, documented shape (which attributes are supported for targeting) so that adding a new targeting dimension later is an additive, backward-compatible schema change, not a breaking one for every existing caller.
Trade-offs and pitfalls
The most common mistake is putting all four operations behind one uniform, verbose response schema for consistency's sake, which quietly taxes the highest-traffic endpoint (evaluate) to make the lowest-traffic ones (create, audit) marginally easier to read. A feature-flag contract earns its keep specifically by treating "how often is this called" as a first-class input into how big its response is allowed to be.
Write a JSON Schema (draft 2020-12 or later) for a POST endpoint /v1/apps/{app_id}/events that accepts an event payload with properties: eventType (enum), timestamp (ISO 8601 string), actor (object with id and type), metadata (optional object with string values), and items (array of objects with id and quantity). Include validation rules, required fields, and example success and validation-error responses.
Sample Answer
The schema needs five things: a type constraint on every field, an enum for eventType so only known event kinds validate, a format constraint on timestamp, an object shape for actor with its own required sub-fields, and an array shape for items with per-item structure. Below is a complete, valid JSON Schema (draft 2020-12) for the event.
The schema
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://api.example.com/schemas/app-event.json",
"title": "AppEvent",
"type": "object",
"properties": {
"eventType": {
"type": "string",
"enum": ["item_created", "item_updated", "item_deleted"]
},
"timestamp": {
"type": "string",
"format": "date-time"
},
"actor": {
"type": "object",
"properties": {
"id": { "type": "string" },
"type": { "type": "string", "enum": ["user", "service"] }
},
"required": ["id", "type"],
"additionalProperties": false
},
"metadata": {
"type": "object",
"additionalProperties": { "type": "string" }
},
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": { "type": "string" },
"quantity": { "type": "integer", "minimum": 1 }
},
"required": ["id", "quantity"],
"additionalProperties": false
},
"minItems": 1
}
},
"required": ["eventType", "timestamp", "actor", "items"],
"additionalProperties": false
}
Design decisions worth calling out
metadatais genuinely optional (absent fromrequired), and itsadditionalPropertiesconstraint restricts VALUES to strings without naming the specific keys, which keeps the schema honest about what "optional object with string values" means without over-constraining future keys.itemsrequires at least one entry (minItems: 1): an event with an empty items array is treated as invalid here, a deliberate validation rule rather than an oversight.additionalProperties: falseon the top level and on each nested object rejects unrecognized fields, which catches client typos early but is itself a compatibility decision: it means adding a genuinely new field later requires updating this schema in lockstep with any producer that starts sending it, rather than silently tolerating it.
Worked example: validating example payloads against this schema
A valid event:
{
"eventType": "item_created",
"timestamp": "2026-07-28T12:00:00Z",
"actor": { "id": "usr_123", "type": "user" },
"metadata": { "source": "mobile_app" },
"items": [{ "id": "sku_1", "quantity": 2 }]
}
This validates cleanly: every required field is present, eventType is one of the three enum values, and items has one well-formed entry.
An invalid event (empty items):
{
"eventType": "item_created",
"timestamp": "2026-07-28T12:00:00Z",
"actor": { "id": "usr_123", "type": "user" },
"items": []
}
This is rejected, specifically because items has zero entries and the schema's minItems: 1 constraint requires at least one. Both examples above were run against this exact schema with a JSON Schema validator; the valid example passed and the empty-items example failed with the minItems violation, confirming the schema behaves as written.
Example success response: {"event_id": "evt_9a1b", "accepted": true} (HTTP 201). Example validation-error response: {"error_code": "validation_error", "message": "items must contain at least 1 item"} (HTTP 400).
Trade-offs and pitfalls
additionalProperties: false is a real trade-off, not a free safety net: it makes typos in client requests fail loudly (good), but it also means the schema and every producer must be updated together the moment a genuinely new, backward-compatible field needs to be added, which is slower than a looser schema that simply ignores unknown fields. Choosing strict validation here is defensible for an internal event contract with few, coordinated producers; a public-facing schema with many independent producers might reasonably relax this to allow additive fields to pass through unvalidated.
Unlock Full Question Bank
Get access to all 17 API Contracts and Schema Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.