Code Quality, Error Handling, and Defensive Programming Questions
Writing robust, high-quality code that fails safely. Covers defensive programming, input validation, error handling and fault tolerance, logging for diagnosability, and general engineering-quality standards. Includes anticipating failure modes and making code resilient to bad inputs and unexpected states.
You need to verify that a payment endpoint is safe to retry, meaning a client that times out and retries doesn't create a duplicate charge. How would you test this, and what would the endpoint need to implement to make it testable?
Sample Answer
Direct answer
The endpoint needs an idempotency key, a client-generated unique id sent with the request, so the server recognizes a retry as "the same request" and returns the original result instead of processing it twice. To test it, simulate the exact failure, the server processes the request but the response is lost, and the client retries with the same key, then assert only one charge exists.
Structured elaboration
- Implementation requirement: the server stores (idempotency key, request hash, result) and on a repeat key returns the stored result instead of re-executing side effects, typically with a TTL and a check that the retried body matches the original.
- Test 1, happy-path retry: send request A with key K, capture the result; resend identical request A with key K; assert the second call returns the same result and no second charge is created.
- Test 2, concurrent retry: fire two requests with the same key K at nearly the same time, simulating a client retrying before the first response returns; assert exactly one charge is created. This tests that the check is atomic, not "check then insert" with a race window.
- Test 3, key reuse with a different payload: send key K with amount 10, then key K with amount 20; assert the defined behavior (reject as a conflict) rather than silently using either amount.
- Test 4, TTL expiry: confirm the deliberately chosen, tested behavior when the same key is reused after the idempotency record has expired.
Worked example
A checkout service processes a 50-unit charge under key "order-4471-attempt-1". The client's connection drops after the charge succeeds but before the response arrives, so it retries with the same key. The test asserts the payment gateway shows exactly one 50-unit charge and the API returns the same transaction id both times, not a new one.
Trade-offs and pitfalls
The most common bug is a "check if key exists, then insert" pattern that isn't atomic; it passes sequential tests but fails under real concurrent retries, which is why Test 2 is the one that actually matters and the one teams skip. Storing idempotency records forever is a data-growth problem, but too short a TTL reopens the double-charge window during exactly the retry storms it exists to prevent.
What the interviewer probes next
Whether the candidate reaches for the concurrency test unprompted, since a sequential-only idempotency suite gives false confidence.
How would you verify, before an incident happens, that your circuit breakers and timeouts actually work as intended, rather than trusting they do because the code was reviewed?
Sample Answer
Direct answer
Deliberately inject the failure the mechanism is meant to handle, kill a dependency, add latency, drop connections, in a controlled environment and observe whether the circuit breaker actually trips, the timeout actually fires, and the system actually recovers when the dependency comes back. Code review confirms the logic looks right; only exercising the real failure path confirms it behaves right under conditions like connection-pool exhaustion that are hard to reason about statically.
Structured elaboration
- Fault injection tools: introduce latency, errors, or drops at the network layer, using a service mesh's fault injection or a tool like Toxiproxy, between the service and its dependency, rather than mocking the dependency, so the test exercises the real timeout and retry code paths.
- What to assert: the breaker opens within the expected failure count or window, requests fail fast (not hang) once it's open, and it attempts to close again (half-open) once the dependency recovers, without immediately re-overwhelming a barely-recovered dependency.
- Start in staging, graduate to production: run the same fault injection in staging first, then a controlled game-day exercise in production during low-traffic hours with the team watching dashboards, before trusting it to hold up during a real incident.
- Regression protection: once verified, add it to a scheduled chaos test so a future refactor that accidentally breaks the timeout configuration gets caught automatically instead of at the next real incident.
Worked example
A service has a 2-second timeout and a breaker configured to open after 5 consecutive failures. Injecting 10 seconds of latency on the downstream dependency, the test asserts requests time out around 2 seconds rather than hanging, the breaker opens after the 5th failure so the 6th request fails immediately, and after removing the injected latency and waiting the reset window, the next request succeeds and the breaker closes.
Trade-offs and pitfalls
Running fault injection in production carries real risk, so it needs a blast-radius limit, a single instance or a low-traffic canary window, and a fast abort mechanism. A common pitfall is testing that the breaker opens but never testing the half-open recovery behavior, which is where thundering-herd bugs, every waiting client retrying the instant the breaker closes, tend to hide.
What the interviewer probes next
Whether the candidate would actually test the recovery path, not just the failure-triggering path, since that's the part teams usually skip.
How would you build automated tests that catch when an upstream API or data source silently changes its response schema, before that bad data reaches production consumers?
Sample Answer
Direct answer
Add a schema-validation step at the ingestion boundary, using something like JSON Schema or Pydantic, that runs on every real response as part of the pipeline, not just at test time, so drift fails fast with a clear error instead of silently flowing downstream. Pair that with a scheduled contract test against the live upstream, not just mocked fixtures, so the team learns about a breaking change before a batch job does.
Structured elaboration
- Define the contract explicitly: required fields, types, and value constraints, not just "the response was 200 OK."
- Validate at the boundary: check every real response against the schema as it enters your system, and quarantine records that fail rather than letting a type mismatch propagate through several transformations before it's hard to trace back.
- Contract-test against the real upstream, scheduled independently of your own deploys: mocked fixtures will happily keep passing forever after the real API changes, because the mock never changes.
- Version and alert on schema changes, even for fields you don't currently use, since today's ignored field can become tomorrow's dependency.
Worked example
A pipeline ingests a partner's product feed where "price" was always a decimal like 19.99. The partner changes their API to a nested object with an amount in cents and a currency code. A schema check asserting price is a number fails immediately with "price: expected number, got object" and quarantines the batch, instead of the pipeline coercing the object to a missing value and silently loading zero-priced products.
Trade-offs and pitfalls
An overly strict schema that rejects any unknown extra field creates noisy failures every time upstream adds something unrelated; a good contract distinguishes fields you depend on from fields you ignore. Mock-only contract tests give false confidence because the mock and the real API can drift apart silently for months.
What the interviewer probes next
Whether the candidate distinguishes a test that checks your own parsing code from a test that checks the upstream contract itself.
A boundary check validates that a value (an index, an offset, a size) falls within the range the code actually handles correctly, and it routinely catches real production bugs before they cause damage. Pick three DIFFERENT kinds of boundary bugs you've seen or can construct realistically, and for each: describe the bug it would cause if unchecked, the specific defensive check you'd add, and a unit test that would catch a regression if the check were later removed.
Sample Answer
Direct answer
A boundary check catches a specific class of bug (accessing an index, offset, or value outside the range the code actually handles correctly) at the moment it happens, instead of letting it silently produce wrong output or crash somewhere unrelated later; three concrete examples: array/list indexing, pagination offsets, and numeric limits.
Structured elaboration and worked examples
- Array indexing: the bug is an off-by-one or attacker-controlled index reading past the end of a buffer or list. The defensive check: validate
0 <= index < len(array)before accessing, raising a clearIndexError/custom exception instead of either crashing with a cryptic native error or, in an unsafe language, reading adjacent memory. A unit test:assert_raises(IndexError, get_item, [1,2,3], 5). - Pagination offsets: the bug is a negative or absurdly large
offset/limitfrom a client, which can either error confusingly deep in a SQL driver or, worse, silently return zero rows and look like 'no data' rather than 'bad request'. The defensive check: clamp or rejectoffset < 0and caplimitto a sane maximum (say 1000) before it reaches the query layer. A unit test:assert paginate(items, offset=-5, limit=10) raises ValueError. - Numeric limits: the bug is an integer overflow or an out-of-domain value (a negative quantity in an order, a percentage over 100) silently producing a nonsensical result instead of an error. The defensive check: validate the value's range explicitly before using it in a calculation. A unit test:
assert_raises(ValueError, apply_discount, price=100, percent=150).
Trade-offs and pitfalls
Each of these checks is cheap individually, but the value comes from applying them CONSISTENTLY at every place the boundary is actually crossed (every array access from external input, not just the ones you happen to remember); a single unguarded pagination endpoint added six months later by someone who didn't see this pattern reintroduces the exact bug class. Treat these as patterns to lint for or wrap in a shared utility function, not as one-off checks to remember individually.
Write a small JavaScript function safeParseJSON(jsonString) that returns an object {success: boolean, value: any, error: string|null}. It must not throw, must handle invalid JSON gracefully, and must not allow prototype pollution (it must avoid assigning to proto or constructor). Also write one example unit test in Jest for an invalid input.
Sample Answer
Direct answer
A safe JSON parser must never throw on bad input, must report success or failure through its return value instead, and must specifically guard against prototype pollution, which is JSON.parse's own quiet trap: a payload like {"__proto__": {"polluted": true}} can corrupt the global Object prototype for the entire process if you are not careful about how you consume the parsed result.
Structured elaboration
Never throw. JSON.parse throws a SyntaxError on malformed input; a "safe" wrapper catches that and returns a structured result instead, so callers do not need a try/catch at every call site.
Prototype pollution. JSON.parse itself does not mutate Object.prototype (unless you additionally use a reviver or later merge the result unsafely into another object with a naive deep-merge). The real risk in a "safe parse" utility is what happens AFTER parsing: if calling code later does something like Object.assign(defaults, parsed) or a recursive merge without checking keys, an attacker-controlled __proto__ or constructor.prototype key can walk up to the shared prototype and add or override properties on every object in the process. The defensive fix at the parse boundary is twofold: use a JSON.parse reviver function that strips dangerous keys (__proto__, constructor, prototype) as soon as they are encountered, and additionally verify with Object.prototype.hasOwnProperty.call that the top-level parsed object does not carry one of those keys before returning it, since a reviver alone can miss certain nested shapes depending on how the object is later traversed.
Structured return, not an exception. Return {success, value, error} so a caller writes if (!result.success) { ...handle... } instead of a try/catch, which keeps the calling code linear and makes "parsing failed" an ordinary value instead of a special control-flow path.
Worked example
function safeParseJSON(jsonString) {
if (typeof jsonString !== 'string') {
return { success: false, value: null, error: 'input is not a string' };
}
let parsed;
try {
parsed = JSON.parse(jsonString, (key, value) => {
if (key === '__proto__' || key === 'constructor' || key === 'prototype') return undefined;
return value;
});
} catch (err) {
return { success: false, value: null, error: err.message };
}
if (parsed && typeof parsed === 'object') {
if (Object.prototype.hasOwnProperty.call(parsed, '__proto__') ||
Object.prototype.hasOwnProperty.call(parsed, 'constructor')) {
return { success: false, value: null, error: 'disallowed key detected' };
}
}
return { success: true, value: parsed, error: null };
}
Executed (Node.js, verified): safeParseJSON('{"a": 1, "b": [1,2,3]}') returns {success: true, value: {a: 1, b: [1,2,3]}, error: null}. safeParseJSON('{not valid json') returns success: false with the underlying SyntaxError message, and does not throw. safeParseJSON(42) returns success: false without ever calling JSON.parse on a non-string. Critically, safeParseJSON('{"__proto__": {"polluted": true}}') was run and confirmed that ({}).polluted is undefined afterward, meaning the global Object prototype was NOT polluted, and the parsed value has no __proto__ own-key surviving in it; the same holds for a constructor.prototype variant of the attack.
One Jest unit test for the invalid-input case:
test('returns a failure result instead of throwing on malformed JSON', () => {
const result = safeParseJSON('{not valid json');
expect(result.success).toBe(false);
expect(result.value).toBeNull();
expect(typeof result.error).toBe('string');
});
Trade-offs and pitfalls
The reviver-based key strip handles the common shape of the attack, but a defense-in-depth mindset says: never rely on parse-time stripping alone if you also deep-merge untrusted objects elsewhere in the codebase, because a different merge utility downstream might not go through this parser at all. Prefer Object.create(null) or a Map for any object you build from untrusted keys if you can avoid prototype-based objects entirely. A common mistake is checking for __proto__ as an enumerable key with a plain for...in loop, which will not find it, since __proto__ set via the object literal syntax is an accessor, not an own enumerable property in the usual sense; using hasOwnProperty.call directly, as above, avoids that gap.
Unlock Full Question Bank
Get access to all 11 Code Quality, Error Handling, and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.