API and Contract Testing Questions
Testing services and their interfaces directly. Covers REST and other API testing, request/response and schema validation, status and error handling, and contract testing between producers and consumers. Includes service-level and integration testing without a UI.
Explain consumer-driven contract testing: what problem it solves, how it prevents integration regressions between a service and its consumers, and what the consumer and provider sides are each responsible for in CI. What tooling typically supports it, and what pitfalls do teams commonly run into when adopting it?
Sample Answer
Direct answer
Consumer-driven contract testing (CDC) verifies that a service (the provider) still honors the expectations of the services that call it (the consumers) without spinning up the whole system. Each consumer records what it expects from the provider as a machine-readable "contract." The provider then replays that contract against its own code, in its own pipeline, to confirm it still satisfies it. Nobody has to run both services at once.
Structured elaboration
Why this exists. A full integration test where every dependent service is actually running catches real bugs, but it's slow, flaky, and gets worse as the number of services grows. Contract testing answers a narrower, cheaper question: does the provider still return what this specific consumer needs? That narrower question is fast to check and can run on every commit.
Division of responsibility.
- The consumer writes a test against a mock of the provider that asserts what it needs (a field, a status code, a shape). Running that test generates the contract as a byproduct, so the contract is always in sync with what the consumer's code actually depends on, not with what someone remembers to document.
- The provider takes that contract and replays it against its real implementation: it sets up whatever preconditions the contract's scenario needs (a "provider state," like "a user with ID 42 exists"), serves the request the contract describes, and checks the response matches.
- CI wiring on both sides is what makes this useful: the consumer's pipeline publishes new contracts when they change, and the provider's pipeline pulls and verifies against the latest contracts before it's allowed to deploy. A provider that would break a consumer fails its own build, not the consumer's.
Typical infrastructure. A contract-testing tool (Pact is the best known) plus a broker: a small service that stores contracts, tracks which provider version has verified which consumer version, and can answer "is it safe for me to deploy this provider version given what's currently in production?"
Benefits. Fast (no real network calls, no shared environment), and it fails on the side that actually caused the problem, since the provider's own CI is where a broken contract shows up.
Trade-offs and pitfalls
Contract tests are not a substitute for integration or end-to-end tests. They confirm the pieces individually honor their agreements; they don't prove the whole system behaves correctly when wired together, and they say nothing about things a contract can't easily express, like timing, ordering across multiple calls, or business-logic correctness. Common early-adoption pitfalls: contracts that are too strict (asserting on fields the consumer doesn't actually use, so unrelated provider changes fail verification for no reason) and contracts that are too loose (missing something the consumer genuinely depends on, so a real break slips through). A useful discipline is to only assert on what the consumer's code actually reads, nothing more.
Consumer-driven vs. provider-driven. In the consumer-driven model described above, the consumer's own tests generate the contract, so it's automatically grounded in real usage. A provider-driven model runs the other way: the provider maintains its own suite of contract tests representing what it believes its consumers need, without a consumer necessarily being involved. That's easier to set up when consumers can't or won't participate, but the contract can silently drift out of sync with what a consumer actually needs, since nothing forces it to be re-derived from real consumer code. Consumer-driven is preferable whenever you can get consumer teams to participate; provider-driven is a fallback when a consumer is external, unresponsive, or a system you don't control.
How would you design tests for a real-time API built on WebSockets or gRPC streaming? Think about what a test harness for this looks like, how you'd get deterministic message ordering, how you'd verify reconnect and resume behavior, and how you'd test many concurrent streams at once.
Sample Answer
Direct answer
A real-time API test harness needs to think in terms of a persistent connection and a stream of messages over time, rather than the single request/response pair a normal API test assumes, which changes what "deterministic" and "correct" even mean for the test.
Structured elaboration
Harness architecture. Instead of a single request-then-assert, the harness opens a connection (a WebSocket, or a gRPC streaming call), sends and/or receives a SEQUENCE of messages over that connection, and needs to buffer and inspect that sequence rather than a single response. A typical shape: connect, send a subscribe or handshake message, collect incoming messages for a bounded window or until a specific terminal message arrives, then assert on the collected sequence.
Deterministic message sequences and IDs. Real-time systems often don't guarantee message arrival timing, so a test needs its assertions to be based on message CONTENT and IDs, not wall-clock timing. Giving each message a client-assigned correlation ID (or asserting on a server-assigned one) lets the test match "the response to message 3" reliably even if messages arrive close together or slightly out of the order they were logically generated.
Ordering guarantees. If the protocol promises ordering (messages for a given channel arrive in the order they were sent), the test should explicitly verify that promise rather than assume it, sending a sequence of numbered messages and asserting they're received in the same order is a direct way to test this.
Reconnect and resume. A realistic test deliberately drops the connection mid-stream and reconnects, then asserts on what happens next: does the client receive messages it missed while disconnected (if the protocol supports resume-from-a-point), or does it need to re-subscribe from scratch? This needs to be tested explicitly since it's exactly the scenario a naive implementation is most likely to get wrong, and it's the scenario most likely to occur in production intermittently, exactly when it's hardest to debug after the fact.
Backpressure. For a stream where the client can't keep up with the volume of incoming messages, the test verifies the system's actual backpressure behavior, whether it buffers, drops, or slows the sender, matches what's documented, rather than silently overwhelming the client and hiding the failure mode until it happens in production under real load.
Testing N concurrent streams. Beyond a single connection's correctness, testing many simultaneous connections verifies the system's behavior under realistic concurrency: does message delivery to one connection ever leak into another, does the system maintain per-connection state correctly, and does overall latency and throughput hold up as connection count grows.
Worked example
A minimal correlation-ID trace makes the "assert on content and IDs, not timing" guidance concrete. A client subscribes, then the server streams three numbered updates, and the test asserts on the sequence it actually received:
client sends: {"id": 1, "type": "subscribe", "channel": "prices"}
server sends: {"id": 1, "type": "ack"}
server sends: {"channel": "prices", "seq": 1, "value": 101.2}
server sends: {"channel": "prices", "seq": 2, "value": 101.5}
server sends: {"channel": "prices", "seq": 3, "value": 101.4}
The test does not assert anything about WHEN these arrive, only that message id=1 got an ack (confirming the subscribe was accepted before anything else happened), and that the three seq values arrived as [1, 2, 3] in that order with no gap and no repeat. A reconnect-mid-stream variant of the same trace would drop the connection after seq=2, reconnect, and then assert on whatever the protocol promises happens next: either seq=3 still arrives (a resume-from-a-point protocol) or the client re-subscribes and a fresh sequence starts from whatever the server considers current (a re-subscribe-from-scratch protocol). The test's job is to confirm which one the system actually does, not to assume either.
Trade-offs and pitfalls
The single biggest trap in testing real-time systems is asserting on WALL-CLOCK timing ("the response should arrive within 100ms") as if it were a hard correctness property. That kind of assertion is inherently flaky in CI, where resource contention can introduce timing variance that has nothing to do with whether the system is actually correct. Assertions should target message CONTENT, ORDER, and DELIVERY GUARANTEES, treating raw latency as a separate, explicitly-labeled performance concern with its own tolerance bands, not folded into the same test as correctness.
Write a script that compares two API contract manifests (each describing endpoints, HTTP methods, required parameters, and response fields) and reports the differences between them. What normalization would you need to do first, and how does the approach scale as the manifests grow large?
Sample Answer
Direct answer
Below is a script that loads two API contract manifests and reports the differences between them: endpoints only in one, methods that changed, required parameters added or removed, and response fields added or removed.
Structured elaboration
The core idea is normalizing both manifests into the same comparable structure (a dictionary keyed by (path, method)) before diffing, so the comparison logic doesn't care what order the manifest's entries happened to be listed in.
Worked example
def normalize_manifest(manifest: list[dict]) -> dict:
"""Key each endpoint entry by (path, method) for order-independent comparison."""
return {(entry["path"], entry["method"]): entry for entry in manifest}
def diff_manifests(old_manifest: list[dict], new_manifest: list[dict]) -> dict:
old = normalize_manifest(old_manifest)
new = normalize_manifest(new_manifest)
old_keys = set(old.keys())
new_keys = set(new.keys())
report = {
"added_endpoints": sorted(new_keys - old_keys),
"removed_endpoints": sorted(old_keys - new_keys),
"changed_endpoints": {},
}
for key in sorted(old_keys & new_keys):
old_entry, new_entry = old[key], new[key]
old_params, new_params = set(old_entry.get("params", [])), set(new_entry.get("params", []))
old_resp, new_resp = set(old_entry.get("response", [])), set(new_entry.get("response", []))
changes = {}
if old_params != new_params:
changes["params_added"] = sorted(new_params - old_params)
changes["params_removed"] = sorted(old_params - new_params)
if old_resp != new_resp:
changes["response_fields_added"] = sorted(new_resp - old_resp)
changes["response_fields_removed"] = sorted(old_resp - new_resp)
if changes:
report["changed_endpoints"][f"{key[1]} {key[0]}"] = changes
return report
if __name__ == "__main__":
old_manifest = [
{"path": "/users", "method": "GET", "params": ["q"], "response": ["id", "name"]},
{"path": "/orders", "method": "POST", "params": ["userId"], "response": ["orderId"]},
]
new_manifest = [
{"path": "/users", "method": "GET", "params": ["q", "limit"], "response": ["id", "name", "email"]},
{"path": "/products", "method": "GET", "params": [], "response": ["id", "price"]},
]
import json
print(json.dumps(diff_manifests(old_manifest, new_manifest), indent=2))
Executed output:
{
"added_endpoints": [["/products", "GET"]],
"removed_endpoints": [["/orders", "POST"]],
"changed_endpoints": {
"GET /users": {
"params_added": ["limit"],
"params_removed": [],
"response_fields_added": ["email"],
"response_fields_removed": []
}
}
}
Normalization. Keying by (path, method) rather than comparing the manifests as raw lists is what makes the comparison correct regardless of ordering, without it, a manifest with the same endpoints listed in a different order would show up as entirely different.
Trade-offs and pitfalls
Complexity at scale. As written, this is O(n) in the number of endpoints, each manifest is normalized into a dict once (O(n)), and the diff itself is set operations over the keys, so it scales linearly and stays fast even for a manifest with thousands of endpoints. The part that DOESN'T scale as cleanly is a manifest where params or response fields are deeply nested objects rather than flat lists of names, this script's set-based comparison only detects a field being added or removed at the top level; a genuinely nested schema diff (a field's own sub-structure changing) needs a recursive comparison, which is real added complexity worth scoping in deliberately rather than assuming this flat version already handles it.
You consume webhooks from an external vendor that signs each payload with HMAC SHA-256 and includes a timestamp to guard against replay. Write a Postman pre-request script (or describe the equivalent code) that generates the correct signature header for a test webhook request, and describe the automated tests you'd write on the receiver side to verify signature validation, timestamp freshness, and replay protection.
Sample Answer
Direct answer
Generating the signature is a few lines: HMAC (hash-based message authentication code) SHA-256 over the timestamp and raw body, using the shared secret. Verifying it correctly on the receiver side is the part that actually matters, and it has three genuinely separate checks: the signature is valid, the timestamp is fresh, and this exact request hasn't been processed before.
Structured elaboration
Signing (sender side, what the Postman pre-request script does). The vendor's convention here is standard: concatenate the timestamp and the raw request body with a separator, then HMAC-SHA256 that combined string with the shared secret, and send the result as a header alongside the timestamp itself. The receiver has to sign the exact same bytes the same way to check it, so the pre-request script and the receiver's verification logic must agree on the signed-content format down to the separator character.
Verifying (receiver side), three checks, each catching a different failure:
- Timestamp freshness. Reject anything outside a tolerance window (a few minutes is typical) before doing anything else. This is what limits how long a captured request stays replayable even before the replay check runs, and it's cheap to check first so an obviously stale request doesn't cost a cryptographic comparison.
- Signature validation. Recompute the expected signature from the secret, the timestamp, and the body, and compare it to the header using a constant-time comparison, never a plain string equality, which can leak timing information about how many leading bytes matched and make the secret guessable byte by byte over many attempts.
- Replay protection. Even a validly-signed, fresh request should only be accepted once. Track signatures (or a vendor-supplied event ID, if one exists) already processed, and reject a repeat.
Worked example
The pre-request script that generates the signature, using CryptoJS, the crypto library Postman's sandbox exposes as a global:
const secret = 'test-webhook-secret-shared-with-vendor';
const timestamp = Math.floor(Date.now() / 1000).toString();
const body = pm.request.body.raw;
const signedContent = timestamp + '.' + body;
const signature = CryptoJS.HmacSHA256(signedContent, secret).toString(CryptoJS.enc.Hex);
pm.request.headers.upsert({ key: 'X-Webhook-Timestamp', value: timestamp });
pm.request.headers.upsert({ key: 'X-Webhook-Signature', value: signature });
And the receiver-side verification, in Python:
import hmac, hashlib, time
def verify_webhook(secret, timestamp, body, signature_header, seen_signatures, max_age_seconds=300):
ts = int(timestamp)
if abs(time.time() - ts) > max_age_seconds:
return False, "stale timestamp"
signed_content = timestamp.encode() + b"." + body
expected = hmac.new(secret.encode(), signed_content, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, signature_header):
return False, "signature mismatch"
if signature_header in seen_signatures:
return False, "replay detected"
seen_signatures.add(signature_header)
return True, "accepted"
Verified end to end with a minimal receiver, not just the isolated function above. Wiring verify_webhook into a small Flask endpoint and driving it with Flask's test client (real HTTP request/response objects, no mocking) confirms the three cases the tests need to cover:
from flask import Flask, request, jsonify
SECRET = "test-webhook-secret-shared-with-vendor"
seen_signatures = set()
app = Flask(__name__)
@app.route("/webhook", methods=["POST"])
def webhook():
timestamp = request.headers.get("X-Webhook-Timestamp", "")
signature = request.headers.get("X-Webhook-Signature", "")
body = request.get_data()
ok, reason = verify_webhook(SECRET, timestamp, body, signature, seen_signatures)
if ok:
return jsonify({"status": reason}), 200
if reason == "replay detected":
return jsonify({"error": reason}), 409
return jsonify({"error": reason}), 400
Driving it with three real round trips through app.test_client():
fresh first-time request -> 200 {'status': 'accepted'}
replayed request -> 409 {'error': 'replay detected'}
stale (1hr old) request -> 400 {'error': 'stale timestamp'}
Exactly the three outcomes the acceptance criteria need: a fresh, correctly-signed request accepted, the identical request replayed and rejected as a duplicate, and a correctly-signed but hour-old request rejected as stale.
Trade-offs and pitfalls
A real gotcha found while building this, worth knowing before you hit it live. The obvious "modernization" of the pre-request script, replacing the bare CryptoJS global with const CryptoJS = require('crypto-js'), actually breaks in the current Postman sandbox: CryptoJS is already bound as a global, and redeclaring it throws SyntaxError: Identifier 'CryptoJS' has already been declared. The bare global form is deprecated (Postman's own console warns about it) but is still the one that actually works today; don't "fix" a deprecation warning by introducing a naming collision that breaks the script outright, verify a replacement actually runs before trusting a deprecation notice's suggested fix.
The signed-content format (timestamp, a separator, then the raw body, in that exact order and byte form) has to match the vendor's real convention exactly; a mismatch anywhere (a different separator, a parsed-and-re-serialized body instead of the raw bytes, a different byte encoding) makes every signature fail to verify even though the logic is otherwise correct, so the first thing to check against a real vendor's docs, not assume, is the exact signed-content construction.
What does a solid test-data strategy for API and service tests actually look like in practice? Cover how you'd decide between generating data on the fly versus using fixed fixtures, when synthetic data is good enough versus when you need something closer to real production data, and how you keep tests from interfering with each other's data.
Sample Answer
Direct answer
A solid test-data strategy for API and service tests rests on four decisions made deliberately rather than by default: generate-on-the-fly versus fixed fixtures, synthetic versus production-derived data, how tests stay isolated from each other's data, and what actually needs cleanup versus what can be safely shared.
Structured elaboration
Generated on the fly vs. fixed fixtures. A fixture generated fresh for each test (a new user created at the start of the test, torn down at the end) gives strong isolation, no test can be affected by another test's leftover state, at the cost of some setup time per test. A fixed, shared fixture (a small set of known reference records seeded once) is faster to use repeatedly but risks one test's assumptions about that shared data becoming invalid if another test modifies it. The practical split: generate fresh data for anything a test needs to MUTATE, and use shared fixtures only for read-only reference data nothing is expected to change.
Synthetic vs. closer-to-production data. Purely synthetic data (templated names, generated emails) is simplest and carries no privacy risk, but can miss bugs that only show up with realistic value distributions or realistic correlations between fields. Production-derived data (masked or anonymized) is more realistic but requires real discipline around what's actually safe to use in a test environment, and is usually reserved for a smaller number of higher-fidelity tests (load testing, a final pre-release check) rather than the bulk of the day-to-day suite, where synthetic data's speed and simplicity matter more than its lower realism.
Isolation between tests. However the data is generated, tests running concurrently need to not collide: namespacing (a unique prefix or identifier per test run), using ephemeral per-test-run resources (a fresh database schema, a dedicated test namespace), or simply ensuring every generated record's identifiers are unique enough that two tests' data can coexist without either noticing the other.
Cleanup. Data a test creates should be removed when the test finishes, via the test's own teardown when things go well, and via a backstop (a scheduled sweep for anything matching a test-data naming pattern and older than some threshold) for the cases where a test crashes before its own teardown runs. Shared, read-only reference data doesn't need this kind of cleanup at all, that's part of why the mutate-vs-read-only split above matters: it determines which data even needs a cleanup story.
Trade-offs and pitfalls
The mistake that causes the most downstream pain is treating "test data strategy" as one uniform decision applied to everything, generate everything fresh, or share everything, rather than recognizing that different data has different needs. A shared reference fixture that's occasionally, accidentally mutated by a test that was supposed to only read it is one of the more confusing categories of flaky test to debug, since the failure shows up in a DIFFERENT, unrelated test that happened to run afterward and inherited the corrupted shared state, not in the test that actually caused the problem.
Unlock Full Question Bank
Get access to all 29 API and Contract Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.