API and Contract Testing Questions
Testing services and their interfaces directly. Covers REST and other API testing, request/response and schema validation, status and error handling, and contract testing between producers and consumers. Includes service-level and integration testing without a UI.
Your team has an SLA that a given API endpoint should respond within a set time for the large majority of requests. How would you write an automated test that checks this without it becoming flaky in CI? Think about what tolerance you'd allow and whether a violation should fail the build outright or just raise a warning.
Sample Answer
Direct answer
Asserting a response-time SLA without making the test flaky means testing against a statistical tolerance, not a single hard cutoff on a single request, and treating the result as a warning until you have enough samples to trust it as a real signal.
Structured elaboration
Why a single-request threshold is flaky. A single request's response time is noisy: network jitter, a GC pause, a neighboring process briefly using CPU on the CI runner, any of these can push one request over a hard threshold without the service actually being slow in any way that matters. Asserting response_time < 500ms on one request will occasionally fail for reasons that have nothing to do with the SLA actually being violated.
Percentile-based assertions instead. The SLA itself is usually stated as a percentile ("95% of requests under 500ms"), so the test should match that shape: send a batch of requests (enough to make a percentile calculation meaningful, tens rather than one or two), compute the actual p95 from that batch, and assert the p95 is under the threshold, tolerating individual outliers the way the SLA itself does.
Tolerance windows. Even a percentile calculated from a modest batch size has some natural variance run to run. Building in an explicit tolerance (asserting p95 under 550ms rather than exactly 500ms, for a documented 500ms SLA) absorbs that natural variance without either being so loose it never catches a real regression or so tight it fails on noise.
Fail vs. warn. A reasonable policy: a violation of the tolerance-adjusted threshold on a single CI run raises a warning, visible but non-blocking, while a violation that's CONSISTENT across several runs (say, 3 consecutive CI runs, or a sustained trend visible over a day) escalates to actually failing the build. This distinguishes a one-off noisy run from a genuine regression without requiring a human to manually triage every borderline case.
Minimizing noise in how the test runs. Running the timing batch against a dedicated, unshared instance of the service (rather than a busy shared staging environment where other tests' load bleeds into your timing measurements) removes one major source of noise, and running the batch as a tight, back-to-back burst rather than spread across a longer window reduces exposure to any transient environmental blip.
Worked example
import statistics
def assert_p95_under_sla(response_times_ms: list[float], sla_ms: float, tolerance_pct: float = 0.10):
# n=100 splits the data into 99 cut points, one for each percentile from 1 to 99.
# Index 94 (0-based) is the 95th of those 99 cut points, i.e. the 95th percentile.
p95 = statistics.quantiles(response_times_ms, n=100)[94]
threshold = sla_ms * (1 + tolerance_pct)
assert p95 <= threshold, f"p95={p95:.1f}ms exceeds tolerance-adjusted threshold={threshold:.1f}ms (SLA={sla_ms}ms)"
return p95
samples_ok = [420, 430, 410, 405, 440, 415, 425, 460, 480, 400] * 10 # p95 well under 550ms
p95 = assert_p95_under_sla(samples_ok, sla_ms=500)
print(f"passed, p95={p95:.1f}ms")
Output:
passed, p95=480.0ms
Trade-offs and pitfalls
A percentile-based assertion needs enough samples to be statistically meaningful; asserting a p95 from a batch of 5 requests isn't really measuring the 95th percentile at all (there's no clean way to compute a meaningful p95 from 5 points), it's an unreliable dressed-up version of the same single-outlier flakiness this approach is meant to fix. Enforcing a minimum sample size (tens of requests at least) is part of the design, not an optional nicety.
Explain consumer-driven contract testing: what problem it solves, how it prevents integration regressions between a service and its consumers, and what the consumer and provider sides are each responsible for in CI. What tooling typically supports it, and what pitfalls do teams commonly run into when adopting it?
Sample Answer
Direct answer
Consumer-driven contract testing (CDC) verifies that a service (the provider) still honors the expectations of the services that call it (the consumers) without spinning up the whole system. Each consumer records what it expects from the provider as a machine-readable "contract." The provider then replays that contract against its own code, in its own pipeline, to confirm it still satisfies it. Nobody has to run both services at once.
Structured elaboration
Why this exists. A full integration test where every dependent service is actually running catches real bugs, but it's slow, flaky, and gets worse as the number of services grows. Contract testing answers a narrower, cheaper question: does the provider still return what this specific consumer needs? That narrower question is fast to check and can run on every commit.
Division of responsibility.
- The consumer writes a test against a mock of the provider that asserts what it needs (a field, a status code, a shape). Running that test generates the contract as a byproduct, so the contract is always in sync with what the consumer's code actually depends on, not with what someone remembers to document.
- The provider takes that contract and replays it against its real implementation: it sets up whatever preconditions the contract's scenario needs (a "provider state," like "a user with ID 42 exists"), serves the request the contract describes, and checks the response matches.
- CI wiring on both sides is what makes this useful: the consumer's pipeline publishes new contracts when they change, and the provider's pipeline pulls and verifies against the latest contracts before it's allowed to deploy. A provider that would break a consumer fails its own build, not the consumer's.
Typical infrastructure. A contract-testing tool (Pact is the best known) plus a broker: a small service that stores contracts, tracks which provider version has verified which consumer version, and can answer "is it safe for me to deploy this provider version given what's currently in production?"
Benefits. Fast (no real network calls, no shared environment), and it fails on the side that actually caused the problem, since the provider's own CI is where a broken contract shows up.
Trade-offs and pitfalls
Contract tests are not a substitute for integration or end-to-end tests. They confirm the pieces individually honor their agreements; they don't prove the whole system behaves correctly when wired together, and they say nothing about things a contract can't easily express, like timing, ordering across multiple calls, or business-logic correctness. Common early-adoption pitfalls: contracts that are too strict (asserting on fields the consumer doesn't actually use, so unrelated provider changes fail verification for no reason) and contracts that are too loose (missing something the consumer genuinely depends on, so a real break slips through). A useful discipline is to only assert on what the consumer's code actually reads, nothing more.
Consumer-driven vs. provider-driven. In the consumer-driven model described above, the consumer's own tests generate the contract, so it's automatically grounded in real usage. A provider-driven model runs the other way: the provider maintains its own suite of contract tests representing what it believes its consumers need, without a consumer necessarily being involved. That's easier to set up when consumers can't or won't participate, but the contract can silently drift out of sync with what a consumer actually needs, since nothing forces it to be re-derived from real consumer code. Consumer-driven is preferable whenever you can get consumer teams to participate; provider-driven is a fallback when a consumer is external, unresponsive, or a system you don't control.
You're serving a machine learning model behind an HTTP API that accepts a JSON feature payload and returns a probability distribution with some metadata. Design contract and schema tests for it: what would the schema need to specify, and how would you test that the API stays backward compatible and fails gracefully when an optional field is missing?
Sample Answer
Direct answer
Contract and schema tests for a model-serving API need to cover the same ground as any REST API's contract tests, request/response shape, status codes, but with two additions specific to ML serving: the response is a probability distribution (which has its own structural constraints beyond "valid JSON") and the model itself can legitimately evolve in ways an ordinary CRUD API's schema doesn't.
Structured elaboration
Schema for the request. The feature payload's schema needs to specify required versus optional features, their types, and, where meaningful, valid ranges (a feature representing an age shouldn't validate as -50, even though "any number" would pass a plain type check).
Schema for the response. Beyond "the response is valid JSON with the right field names and types," a probability distribution has structural properties worth asserting explicitly: the probabilities should sum to (approximately) 1, each individual probability should be in [0, 1], and the number of classes returned should match what the model is documented to predict. A response that's syntactically valid JSON but has probabilities summing to 1.3 is a real bug a plain JSON Schema check (which only validates types and structure, not this kind of numeric invariant) would miss entirely, worth a dedicated assertion beyond the schema tool.
Metadata fields. If the response includes metadata (a model version identifier, a confidence score, timing information), the schema should specify what's guaranteed to be present versus optional, the same discipline as any other API's optional-field handling.
Backward compatibility. A model-serving API evolving, a retrained model, a new feature added to the input schema, shouldn't break existing callers making the same request shape they always have. Testing this means confirming an OLDER, still-valid request shape continues to be served correctly (with reasonable output) after a model update, not just that the NEW shape works.
Graceful handling of missing optional fields. If a feature is genuinely optional (the model has a documented default or imputation strategy for it (imputation means the model or pipeline fills in a reasonable stand-in value for the missing feature, for example the average value seen during training, rather than failing outright)), a request omitting that feature should still succeed with a sensible response, not silently degrade to garbage output or fail with an unclear error. This needs testing explicitly, since "the model still runs" and "the model produces a reasonable response with a missing optional feature" are different claims.
Worked example (schema and invariant check outline)
def validate_prediction_response(response: dict, expected_num_classes: int):
probs = response.get("probabilities")
assert probs is not None, "missing probabilities field"
assert len(probs) == expected_num_classes, (
f"expected {expected_num_classes} classes, got {len(probs)}"
)
assert all(0.0 <= p <= 1.0 for p in probs), f"probability out of [0,1] range: {probs}"
total = sum(probs)
assert abs(total - 1.0) < 1e-3, f"probabilities sum to {total}, expected ~1.0"
Trade-offs and pitfalls
The distribution-sums-to-one check is easy to skip if you're only reaching for a generic JSON Schema validator, since it's a numeric invariant across multiple fields, not a per-field type or range constraint, and most schema tools don't express that natively. It needs its own explicit assertion outside the schema check, and it's exactly the kind of bug (a normalization step dropped somewhere in a model update) that would otherwise ship silently, since the response would still look completely valid to a schema validator that only checks structure and types.
Write Postman test scripts (using pm.test and pm.expect) that check three things: a login request returns a token with HTTP 200, creating a resource returns HTTP 201 with the expected fields, and an unauthorized request returns HTTP 401. Also explain how you'd handle the token as an environment variable between requests.
Sample Answer
Direct answer
Below are Postman test scripts (added to the Tests tab of each request) covering a login returning a token with 200, a resource-creation call returning 201 with the expected fields, and an unauthorized request returning 401.
Structured elaboration
Postman test scripts run in a sandboxed JavaScript environment after the request completes, using the pm object to access the response and to persist values (like a token) into the environment for later requests to use. pm.test wraps each individual assertion so failures are reported per-check rather than as one opaque script failure.
Worked example
Request 1: POST /login
pm.test("login returns 200 with a token", function () {
pm.response.to.have.status(200);
const body = pm.response.json();
pm.expect(body).to.have.property("token");
pm.expect(body.token).to.be.a("string").and.not.empty;
// Store the token for subsequent requests in this collection run.
pm.environment.set("auth_token", body.token);
});
Request 2: POST /resources (using {{auth_token}} as a Bearer token in this request's Authorization header)
pm.test("creating a resource returns 201 with expected fields", function () {
pm.response.to.have.status(201);
const body = pm.response.json();
pm.expect(body).to.have.property("id");
pm.expect(body.id).to.be.a("number");
pm.expect(body).to.have.property("created_at");
});
Request 3: POST /resources (deliberately sent with NO Authorization header, to exercise the unauthorized path)
pm.test("unauthorized request is rejected with 401", function () {
pm.response.to.have.status(401);
});
Environment-variable handling. pm.environment.set("auth_token", body.token) after the login request is what makes the token available to later requests without hardcoding it: the resource-creation request's Authorization header is configured as Bearer {{auth_token}}, a template reference Postman resolves from the environment at request time. This keeps the token entirely out of the collection's static configuration, since it only exists at runtime, generated fresh by request 1 of the actual run.
Executed via Newman against a local fixture server implementing exactly this login/create/unauthorized behavior:
POST /login
login returns 200 with a token ✓ (2 assertions)
POST /resources
creating a resource returns 201 with expected fields ✓ (3 assertions)
POST /resources
unauthorized request is rejected with 401 ✓ (1 assertion)
3 requests, 0 failures
Trade-offs and pitfalls
The unauthorized-request test needs to be its own separate request in the collection, deliberately configured without the Authorization header, rather than a variant assertion on the authorized request, Postman doesn't have a clean way to "temporarily remove" a header for one assertion within an otherwise-authorized request, so this is naturally two distinct requests representing two distinct scenarios, not one request with conditional logic.
You have a Postman collection that represents your team's API test suite. Describe, step by step, how you would run that collection automatically in CI: handling environment variables and secrets, deciding what should fail the build, and producing a report a dashboard could consume.
Sample Answer
Direct answer
Newman is Postman's own official command-line runner: it executes a saved Postman collection outside the desktop app, which is what makes a Postman-based test suite runnable on a CI machine that has no GUI at all. Running a Postman collection in CI with Newman means three things beyond just installing it: getting secrets and environment values into the run without hardcoding them into the collection file, deciding what failure behavior should mean for the build, and producing a report format your CI and dashboards can actually consume.
Structured elaboration
Environment variables and secrets. A Postman collection references variables (a base URL, an API key) rather than hardcoding them, and Newman accepts an environment file or individual --env-var flags to supply those values at run time. Secrets specifically should never be committed as part of the environment file itself, they're injected from the CI system's own secret store as environment variables that get passed through to Newman's invocation, keeping the actual secret values out of version control entirely.
Pass/fail behavior. Newman exits non-zero if any test assertion in the collection fails, which is what a CI pipeline checks to decide whether the build passes. It's worth being deliberate about which failures should actually block the build versus which should be visible but non-blocking, a flaky or known-broken low-priority check might be tagged and excluded from the pass/fail gate while still running and reporting, rather than either silently skipping it or letting it block every deploy.
Machine-readable reporting. Newman supports several reporters; for CI dashboards, the JUnit XML reporter is usually the most broadly compatible choice, since most CI systems already know how to parse and display JUnit XML natively (per-test pass/fail, timing, failure messages) without custom tooling. A JSON reporter is the better choice when you're feeding a custom internal dashboard that wants the full structured detail rather than the JUnit-shaped summary.
Worked example (CI step, illustrative)
- name: Run Postman API tests
run: |
npx newman run collection.json \
-e ci-environment.json \
--env-var "api_key=$API_KEY" \
--reporters cli,junit \
--reporter-junit-export results/newman-results.xml
with $API_KEY supplied by the CI system's own secret store, never committed to ci-environment.json. The cli reporter gives readable console output for a human scanning the build log; the junit reporter's XML output is what the CI dashboard actually parses to render pass/fail status and historical trends.
Trade-offs and pitfalls
A collection that references a variable Newman was never given a value for fails at request-build time with an unresolved-variable error, worth explicitly testing that every variable the collection needs is actually supplied by the CI environment file (or the secret injection), rather than discovering a missing one only when the pipeline runs for the first time in a new environment. It's also worth keeping the environment file itself checked into version control (with real secrets excluded, only placeholder or non-sensitive values), so the collection's expected variable shape stays visible and reviewable even though the actual secret values live elsewhere.
Unlock Full Question Bank
Get access to all 47 API and Contract Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.