API and Contract Testing Questions
Testing services and their interfaces directly. Covers REST and other API testing, request/response and schema validation, status and error handling, and contract testing between producers and consumers. Includes service-level and integration testing without a UI.
A service you depend on keeps shipping changes that break your integration with it, and the breakage is only caught after it reaches production. Propose both a technical and a process fix: what would you put in place so a breaking change is caught before it ships, who should own the tests that catch it, and how would you pilot the change and show it actually reduced regressions?
Sample Answer
Direct answer
The fix has to work on two tracks at once: a technical track that catches a breaking change before it ships, and a process track that makes catching it everyone's default rather than something only found in a postmortem. Neither alone is durable, technical checks without ownership get bypassed under deadline pressure, and process without automation depends on someone remembering to look.
Structured elaboration
Technical track. Add consumer-driven contract tests to the dependency relationship: the depending service's real expectations get encoded as a contract, and the upstream service's CI verifies against it before every deploy. This is what actually blocks a breaking change from shipping, rather than catching it after the fact in a bug report. Where a contract test isn't practical to set up quickly, an automated compatibility check that diffs the API's schema or response shape between versions is a cheaper first step that still catches the most common failure mode: a field silently removed, renamed, or retyped.
CI policy. The upstream service's pipeline should fail the build, not just warn, when provider verification against a downstream consumer's contract fails. Placement matters: this check belongs in the upstream service's own CI, where the change originates, not bolted onto the downstream service's pipeline as an afterthought.
Process track. Even a working technical check needs an owner. A PR template or review checklist item that asks "does this change anything a documented consumer depends on" catches the cases contract tests haven't been written for yet. Clear ownership, who's accountable when a contract test fails, and who's accountable for keeping the downstream service's contract accurate as its own needs evolve, is what prevents the check from silently rotting once nobody's paying attention to it.
Retrofitting onto an org with no prior governance. When there's no existing discipline at all, start with a migration plan rather than a mandate: pick the highest-incident-risk dependency pairs first (the ones that have actually caused regressions), prove contract testing catches something real there, and use that as the case for expanding rather than rolling it out everywhere simultaneously with no working example.
Piloting and measuring. Run the new discipline on one or two service pairs first. Track something concrete before and after: the number of contract-related production incidents on those pairs, or the number of contract-test failures caught in CI versus caught in production. That comparison is what actually demonstrates the pilot worked, rather than assuming it did because the process feels more rigorous.
Trade-offs and pitfalls
The most common way this fails is treating it as a one-time cleanup instead of an ongoing discipline: teams write contract tests once, the regressions stop for a while, and the checklist item quietly gets skipped under the next deadline because nothing enforces it anymore. Ownership needs to be durable, not a one-time assignment, and the CI gate needs to actually block a deploy, not just post a warning that's easy to ignore.
Implement a reusable function that performs an HTTP GET with retry and exponential backoff for transient failures (server errors and network errors), with configurable attempt count and base delay. What do you need to be careful about if this function is used concurrently by many tests at once?
Sample Answer
Direct answer
Below is a reusable HTTP GET function with retry and exponential backoff for transient server errors and network errors, with configurable attempt count and base delay.
Structured elaboration
The function distinguishes what's worth retrying (a timeout, a connection error, a 5xx that's likely transient) from what isn't (a 4xx, which means the request itself is wrong and retrying an unmodified request will just fail the same way again), and backs off exponentially between attempts so a struggling server isn't hit with an immediate retry storm.
Worked example
import time
import requests
def resilient_get(url, max_attempts=5, initial_backoff=0.5, backoff_factor=2,
retry_statuses=(500, 502, 503, 504)):
last_exception = None
delay = initial_backoff
for attempt in range(1, max_attempts + 1):
try:
resp = requests.get(url, timeout=5)
if resp.status_code not in retry_statuses:
return resp # success, or a non-retryable error (e.g. 4xx): return as-is
last_exception = None
except (requests.ConnectionError, requests.Timeout) as e:
last_exception = e
resp = None
if attempt == max_attempts:
if last_exception:
raise last_exception
return resp # exhausted retries on a retryable status; return the last response
time.sleep(delay)
delay *= backoff_factor
raise RuntimeError("unreachable") # defensive; loop always returns or raises above
Thread-safety. As written, resilient_get has no shared mutable state at all: delay, attempt, and last_exception are all local to each call, so many test threads calling it concurrently don't interact with each other in any way. The one thing worth being deliberate about in concurrent use is the underlying requests session: this version uses the module-level requests.get, which creates a new connection per call and is safe under concurrency but doesn't reuse connections. If you switch to a shared requests.Session() for connection pooling (a reasonable optimization under high concurrency), the Session object itself needs to be either one per thread or explicitly documented as thread-safe for your use case, since requests.Session is not guaranteed thread-safe for concurrent use by requests' own documentation.
Verified with a fixture that fails twice with a 503 and then succeeds:
call_count = [0]
def flaky_get(url, timeout):
call_count[0] += 1
class FakeResp:
status_code = 503 if call_count[0] <= 2 else 200
return FakeResp()
requests.get = flaky_get # monkeypatched into requests.get for this test
resp = resilient_get("http://fake/x", initial_backoff=0.01)
print(f"attempts made: {call_count[0]}")
print(f"final status_code: {resp.status_code}")
Running resilient_get against this fixture (with initial_backoff=0.01 to keep the test fast) returns a 200 after exactly 3 attempts, confirming the retry loop and the eventual-success path both work:
attempts made: 3
final status_code: 200
Trade-offs and pitfalls
A 4xx status code deliberately does NOT trigger a retry in this implementation: retrying an unmodified request that the server has already rejected as invalid wastes time and, in the worst case, can look like an attempted abuse pattern to the server (repeated requests to an endpoint that keeps rejecting them). If a caller genuinely wants to retry a 429 (rate-limited) specifically, that status needs to be added to retry_statuses deliberately, and ideally the delay should respect a Retry-After header if the server provides one, rather than blindly following the function's own generic backoff schedule.
Design automated tests that validate JWT-based authentication for a set of realistic failure cases: a missing token, a malformed token, an expired token, a token with the wrong issuer or audience, and a token that was signed with a key that has since been revoked. For each case, what status code do you expect, and how would you generate the test tokens deterministically?
Sample Answer
Direct answer
Each JWT failure case needs its own deterministic token and its own expected status code: a missing token and a malformed token both typically return 401, but they represent different bugs if the wrong one shows up, so testing them separately, rather than treating "any 401" as sufficient, actually matters.
Structured elaboration
Missing token. No Authorization header at all. Expect 401.
Malformed token. A string that isn't even a structurally valid JWT (wrong number of segments, invalid base64). Expect 401, and specifically confirm the server doesn't throw an unhandled parsing exception that would otherwise leak as a 500, a malformed token is exactly the kind of input a naive parser might not defensively handle.
Expired token. A structurally valid, correctly signed token whose exp claim is in the past. Expect 401 with an error that's ideally distinguishable from "invalid signature" (a client's retry logic often differs: an expired token means "refresh and retry," an invalid signature means "something is fundamentally wrong, don't retry with the same flow").
Wrong issuer or audience. A token that's validly signed and unexpired, but whose iss or aud claim doesn't match what this API expects, this is the case that catches a token meant for a DIFFERENT service being replayed against this one. Expect 401.
Revoked signing key. A token signed with a key that was valid when issued but has since been revoked (rotated out, compromised). Expect 401. This is the case most likely to be missed in testing, because it requires the test infrastructure to actually track key rotation, not just accept anything with a valid-looking signature.
Worked example
Generating deterministic test tokens (using a library like PyJWT) rather than depending on a real token-issuing flow:
import jwt
import time
SIGNING_KEY = "test-secret-for-fixtures-only"
REVOKED_KEY = "an-old-key-that-has-since-been-rotated-out"
def make_token(claims_override=None, key=SIGNING_KEY, algorithm="HS256"):
now = int(time.time())
claims = {
"sub": "user-123",
"iss": "https://auth.example.com",
"aud": "https://api.example.com",
"iat": now,
"exp": now + 3600,
}
if claims_override:
claims.update(claims_override)
return jwt.encode(claims, key, algorithm=algorithm)
expired_token = make_token({"exp": int(time.time()) - 60})
wrong_issuer_token = make_token({"iss": "https://not-your-auth-server.com"})
revoked_key_token = make_token(key=REVOKED_KEY)
malformed_token = "not.a.valid.jwt.at.all"
print("expired:", expired_token[:20], "...")
print("wrong issuer:", wrong_issuer_token[:20], "...")
print("revoked-key-signed:", revoked_key_token[:20], "...")
Executed output:
expired: eyJhbGciOiJIUzI1NiIs ...
wrong issuer: eyJhbGciOiJIUzI1NiIs ...
revoked-key-signed: eyJhbGciOiJIUzI1NiIs ...
Each variant is generated deterministically from a known signing key and a known set of claim overrides, which is what makes this reproducible: no dependency on a real, time-sensitive login flow, and no waiting around for a token to actually expire in real time.
Trade-offs and pitfalls
The revoked-key case is easy to test wrong: if your test's "revoked key" is simply a key your test suite never registered as valid in the first place, you're really testing "unknown signing key," not "a key that WAS valid and got revoked." The distinction matters because revocation is usually handled by a different code path (checking a key against a revocation list or rotation timestamp) than simple signature validation, and a test that conflates the two can pass while the real revocation-checking logic is broken.
Write a script that compares two API contract manifests (each describing endpoints, HTTP methods, required parameters, and response fields) and reports the differences between them. What normalization would you need to do first, and how does the approach scale as the manifests grow large?
Sample Answer
Direct answer
Below is a script that loads two API contract manifests and reports the differences between them: endpoints only in one, methods that changed, required parameters added or removed, and response fields added or removed.
Structured elaboration
The core idea is normalizing both manifests into the same comparable structure (a dictionary keyed by (path, method)) before diffing, so the comparison logic doesn't care what order the manifest's entries happened to be listed in.
Worked example
def normalize_manifest(manifest: list[dict]) -> dict:
"""Key each endpoint entry by (path, method) for order-independent comparison."""
return {(entry["path"], entry["method"]): entry for entry in manifest}
def diff_manifests(old_manifest: list[dict], new_manifest: list[dict]) -> dict:
old = normalize_manifest(old_manifest)
new = normalize_manifest(new_manifest)
old_keys = set(old.keys())
new_keys = set(new.keys())
report = {
"added_endpoints": sorted(new_keys - old_keys),
"removed_endpoints": sorted(old_keys - new_keys),
"changed_endpoints": {},
}
for key in sorted(old_keys & new_keys):
old_entry, new_entry = old[key], new[key]
old_params, new_params = set(old_entry.get("params", [])), set(new_entry.get("params", []))
old_resp, new_resp = set(old_entry.get("response", [])), set(new_entry.get("response", []))
changes = {}
if old_params != new_params:
changes["params_added"] = sorted(new_params - old_params)
changes["params_removed"] = sorted(old_params - new_params)
if old_resp != new_resp:
changes["response_fields_added"] = sorted(new_resp - old_resp)
changes["response_fields_removed"] = sorted(old_resp - new_resp)
if changes:
report["changed_endpoints"][f"{key[1]} {key[0]}"] = changes
return report
if __name__ == "__main__":
old_manifest = [
{"path": "/users", "method": "GET", "params": ["q"], "response": ["id", "name"]},
{"path": "/orders", "method": "POST", "params": ["userId"], "response": ["orderId"]},
]
new_manifest = [
{"path": "/users", "method": "GET", "params": ["q", "limit"], "response": ["id", "name", "email"]},
{"path": "/products", "method": "GET", "params": [], "response": ["id", "price"]},
]
import json
print(json.dumps(diff_manifests(old_manifest, new_manifest), indent=2))
Executed output:
{
"added_endpoints": [["/products", "GET"]],
"removed_endpoints": [["/orders", "POST"]],
"changed_endpoints": {
"GET /users": {
"params_added": ["limit"],
"params_removed": [],
"response_fields_added": ["email"],
"response_fields_removed": []
}
}
}
Normalization. Keying by (path, method) rather than comparing the manifests as raw lists is what makes the comparison correct regardless of ordering, without it, a manifest with the same endpoints listed in a different order would show up as entirely different.
Trade-offs and pitfalls
Complexity at scale. As written, this is O(n) in the number of endpoints, each manifest is normalized into a dict once (O(n)), and the diff itself is set operations over the keys, so it scales linearly and stays fast even for a manifest with thousands of endpoints. The part that DOESN'T scale as cleanly is a manifest where params or response fields are deeply nested objects rather than flat lists of names, this script's set-based comparison only detects a field being added or removed at the top level; a genuinely nested schema diff (a field's own sub-structure changing) needs a recursive comparison, which is real added complexity worth scoping in deliberately rather than assuming this flat version already handles it.
You consume webhooks from an external vendor that signs each payload with HMAC SHA-256 and includes a timestamp to guard against replay. Write a Postman pre-request script (or describe the equivalent code) that generates the correct signature header for a test webhook request, and describe the automated tests you'd write on the receiver side to verify signature validation, timestamp freshness, and replay protection.
Sample Answer
Direct answer
Generating the signature is a few lines: HMAC (hash-based message authentication code) SHA-256 over the timestamp and raw body, using the shared secret. Verifying it correctly on the receiver side is the part that actually matters, and it has three genuinely separate checks: the signature is valid, the timestamp is fresh, and this exact request hasn't been processed before.
Structured elaboration
Signing (sender side, what the Postman pre-request script does). The vendor's convention here is standard: concatenate the timestamp and the raw request body with a separator, then HMAC-SHA256 that combined string with the shared secret, and send the result as a header alongside the timestamp itself. The receiver has to sign the exact same bytes the same way to check it, so the pre-request script and the receiver's verification logic must agree on the signed-content format down to the separator character.
Verifying (receiver side), three checks, each catching a different failure:
- Timestamp freshness. Reject anything outside a tolerance window (a few minutes is typical) before doing anything else. This is what limits how long a captured request stays replayable even before the replay check runs, and it's cheap to check first so an obviously stale request doesn't cost a cryptographic comparison.
- Signature validation. Recompute the expected signature from the secret, the timestamp, and the body, and compare it to the header using a constant-time comparison, never a plain string equality, which can leak timing information about how many leading bytes matched and make the secret guessable byte by byte over many attempts.
- Replay protection. Even a validly-signed, fresh request should only be accepted once. Track signatures (or a vendor-supplied event ID, if one exists) already processed, and reject a repeat.
Worked example
The pre-request script that generates the signature, using CryptoJS, the crypto library Postman's sandbox exposes as a global:
const secret = 'test-webhook-secret-shared-with-vendor';
const timestamp = Math.floor(Date.now() / 1000).toString();
const body = pm.request.body.raw;
const signedContent = timestamp + '.' + body;
const signature = CryptoJS.HmacSHA256(signedContent, secret).toString(CryptoJS.enc.Hex);
pm.request.headers.upsert({ key: 'X-Webhook-Timestamp', value: timestamp });
pm.request.headers.upsert({ key: 'X-Webhook-Signature', value: signature });
And the receiver-side verification, in Python:
import hmac, hashlib, time
def verify_webhook(secret, timestamp, body, signature_header, seen_signatures, max_age_seconds=300):
ts = int(timestamp)
if abs(time.time() - ts) > max_age_seconds:
return False, "stale timestamp"
signed_content = timestamp.encode() + b"." + body
expected = hmac.new(secret.encode(), signed_content, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, signature_header):
return False, "signature mismatch"
if signature_header in seen_signatures:
return False, "replay detected"
seen_signatures.add(signature_header)
return True, "accepted"
Verified end to end with a minimal receiver, not just the isolated function above. Wiring verify_webhook into a small Flask endpoint and driving it with Flask's test client (real HTTP request/response objects, no mocking) confirms the three cases the tests need to cover:
from flask import Flask, request, jsonify
SECRET = "test-webhook-secret-shared-with-vendor"
seen_signatures = set()
app = Flask(__name__)
@app.route("/webhook", methods=["POST"])
def webhook():
timestamp = request.headers.get("X-Webhook-Timestamp", "")
signature = request.headers.get("X-Webhook-Signature", "")
body = request.get_data()
ok, reason = verify_webhook(SECRET, timestamp, body, signature, seen_signatures)
if ok:
return jsonify({"status": reason}), 200
if reason == "replay detected":
return jsonify({"error": reason}), 409
return jsonify({"error": reason}), 400
Driving it with three real round trips through app.test_client():
fresh first-time request -> 200 {'status': 'accepted'}
replayed request -> 409 {'error': 'replay detected'}
stale (1hr old) request -> 400 {'error': 'stale timestamp'}
Exactly the three outcomes the acceptance criteria need: a fresh, correctly-signed request accepted, the identical request replayed and rejected as a duplicate, and a correctly-signed but hour-old request rejected as stale.
Trade-offs and pitfalls
A real gotcha found while building this, worth knowing before you hit it live. The obvious "modernization" of the pre-request script, replacing the bare CryptoJS global with const CryptoJS = require('crypto-js'), actually breaks in the current Postman sandbox: CryptoJS is already bound as a global, and redeclaring it throws SyntaxError: Identifier 'CryptoJS' has already been declared. The bare global form is deprecated (Postman's own console warns about it) but is still the one that actually works today; don't "fix" a deprecation warning by introducing a naming collision that breaks the script outright, verify a replacement actually runs before trusting a deprecation notice's suggested fix.
The signed-content format (timestamp, a separator, then the raw body, in that exact order and byte form) has to match the vendor's real convention exactly; a mismatch anywhere (a different separator, a parsed-and-re-serialized body instead of the raw bytes, a different byte encoding) makes every signature fail to verify even though the logic is otherwise correct, so the first thing to check against a real vendor's docs, not assume, is the exact signed-content construction.
Unlock Full Question Bank
Get access to all 18 API and Contract Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.