Code Quality, Error Handling, and Defensive Programming Questions
Writing robust, high-quality code that fails safely. Covers defensive programming, input validation, error handling and fault tolerance, logging for diagnosability, and general engineering-quality standards. Includes anticipating failure modes and making code resilient to bad inputs and unexpected states.
A backend service builds a SQL query by concatenating a user-supplied search string directly into the query text. What's wrong with this from a defensive-programming standpoint, and what would you check for in code review to catch this class of bug at scale?
Sample Answer
Direct answer
Concatenating untrusted input directly into a SQL string lets an attacker change the query's structure, not just its data, which is SQL injection and can expose or modify anything the database connection can reach. The fix is parameterized queries, where input is always bound as data and can never be interpreted as SQL syntax; at scale, catching this means a static-analysis rule that flags any string concatenation feeding a query-execution call, not relying on a human to spot every instance.
Structured elaboration
- Why it's dangerous: if the search string is a SQL fragment like a statement terminator followed by a destructive command, a concatenated query executes it as SQL, because the database can't distinguish "data the user typed" from "query the developer wrote" once they're merged into one string.
- The fix: parameterized queries or an ORM's query builder, where the driver sends the query text and parameters separately, so the database always treats parameters as literal values.
- Scaling the catch beyond code review: a static-analysis rule, via a linter or a tool like Semgrep, that flags any database call built from string concatenation or interpolation, run in CI so it blocks the pull request automatically.
- Defense in depth: least-privilege database credentials, so the app's database user doesn't have permission to drop tables at all, shrinking the blast radius if a check is ever bypassed.
Worked example
Vulnerable: building the query text as SELECT * FROM users WHERE name = ' plus the raw user input plus '. Safe: SELECT * FROM users WHERE name = ? with the input passed as a bound parameter. With the vulnerable version, an input like x' OR '1'='1 turns the WHERE clause into always-true, returning every row instead of matching one name.
Trade-offs and pitfalls
Parameterization doesn't cover every injection surface; dynamically chosen table or column names can't be parameterized the same way and still need explicit allow-list validation, not string interpolation. A common pitfall is fixing the obvious query but missing a second one built the same way in a reporting script or admin tool that doesn't go through the same review path.
What the interviewer probes next
Whether the candidate reaches for "parameterize it" immediately, then goes further into how they'd find every other instance of the pattern across a large codebase, not just this one.
Compare and contrast graceful degradation and fail-fast design approaches for production systems. For each approach, explain a typical use case (for example a customer-facing API versus an internal pipeline), the operational trade-offs, how you would instrument each approach with metrics, logs, and traces, and how you would communicate degraded functionality to clients or downstream systems.
Sample Answer
Direct answer
Fail-fast stops the operation immediately and surfaces the error the moment something is wrong, trading availability for correctness and a clear signal; graceful degradation keeps the system partially functional by falling back to reduced capability, trading some correctness or completeness for continued availability. The right choice depends on whether a wrong or incomplete answer is worse than no answer at all for this specific system.
Structured elaboration
Fail-fast: when correctness matters more than availability. An internal data pipeline computing financial reconciliation numbers should fail fast and loudly the moment its inputs look wrong, because a wrong number that looks plausible and gets used in a report is a much worse outcome than the pipeline simply not running today. Fail-fast systems are also easier to operate: a hard failure with a clear error is diagnosable immediately, whereas a system that silently degrades can mask a real problem for a long time before anyone notices the quality of its output has quietly dropped.
Graceful degradation: when partial availability beats a hard stop. A customer-facing product page that depends on a recommendation service should degrade to a generic, non-personalized set of recommendations if that service is slow or down, rather than showing the customer an error page, because a slightly-worse-but-functional page is a much better outcome for both the customer and the business than a hard failure on a page that otherwise works fine.
Instrumentation differs by approach. A fail-fast system needs strong alerting on the failure itself, since the failure IS the signal: an error rate spike, a specific exception type, or a circuit breaker opening. A gracefully-degrading system needs the opposite kind of visibility: a metric or log line specifically for "we are currently in degraded mode", because the degraded path, by design, does not look like a failure to a simple error-rate dashboard, and a team that isn't specifically tracking degraded-mode usage can be running in a permanently degraded state for months without noticing.
Communicating degraded functionality. For a user-facing system, this usually means a visible but non-alarming UI signal ("Showing popular items while personalized recommendations are unavailable") rather than silence, since silent degradation erodes trust once a user notices the quality difference without being told why. For a downstream service-to-service dependency, this means an explicit field or header in the response indicating degraded mode, so the calling service can make its own informed choice about whether to also degrade or to fail.
Worked example
An internal pipeline vs. a downstream API dependency, side by side: the internal pipeline computing quarterly revenue numbers for a financial report should fail fast and halt if a required upstream table is empty or a row count sanity check fails, alerting the on-call data engineer immediately, because publishing a subtly wrong number in a financial report is far worse than the report being late. The customer-facing recommendation widget on the same company's storefront, dependent on a separate ML service, should instead catch a timeout from that service and immediately serve a cached "most popular this week" list, log a degraded_mode=true metric tagged with the reason, and continue serving the page, because an incomplete page for one widget is a minor UX cost, not a correctness failure the business needs to halt over.
Trade-offs and pitfalls
The most common mistake is applying the wrong default to a whole system uniformly: treating every dependency as fail-fast produces a fragile product where one non-critical service outage takes down an entire page, while treating every dependency as gracefully-degradable risks quietly serving wrong financial or safety-relevant data with no alert ever firing. The decision should be made dependency by dependency, based on whether being wrong is worse than being unavailable for that specific piece of functionality, not applied as a single system-wide policy.
What items should a code-review checklist contain to enforce production-quality, defensive-programming standards across distributed teams? Draft a prioritized checklist of at least eight review items, and for each one explain why it directly impacts production reliability or operability.
Sample Answer
Direct answer
A code-review checklist for defensive, production-quality standards should be short enough that reviewers actually use it every time, and should focus on the handful of items that correlate most directly with real production incidents: error handling, observability, resource leaks, secrets management, idempotency, and input validation, each with a concrete one-line test a reviewer can actually apply while reading a diff.
Structured elaboration
1. Error handling. Does every external call (network, database, file system) have explicit failure handling, and is there no bare, silent catch-and-ignore? This matters because a swallowed exception is one of the most common root causes of "the system silently stopped working and nobody noticed for days".
2. Observability. Does this change add or preserve logging/metrics for its new failure paths, not just its happy path? A new code path with no visibility into whether it's failing in production is effectively unmonitored the moment it ships.
3. Resource leaks. Are file handles, database connections, and locks acquired in this diff guaranteed to be released even when an exception occurs (via try/finally, a context manager, or the language's equivalent)? This matters because a resource leak in an error path specifically (the path least likely to be exercised in normal testing) is a classic source of a slow production degradation that only appears under sustained load or over a long uptime.
4. Secrets management. Does this diff introduce any hardcoded credential, API key, or token, or log anything that could contain one? This is a fast, mechanical check (often automatable via a pre-commit secret scanner) but still worth a human's attention, since scanners miss secrets embedded in less obvious places like a debug log statement.
5. Idempotency. If this diff adds or touches an operation that could be retried (by a client, a queue redelivery, or an internal retry mechanism), is that operation actually safe to run more than once? This matters because a non-idempotent operation that silently becomes retriable somewhere in the call stack is a duplicate-side-effect bug waiting to happen, often not caught until production traffic patterns exercise the retry path that testing never did.
6+. Input validation and the remaining prioritized items. Does every externally-supplied input reaching this code get validated at a clear boundary, rather than trusted implicitly? Beyond these top items, a fuller checklist includes: test coverage for the new failure paths specifically (not just the happy path), whether any deprecated or discouraged pattern was introduced, and whether the change includes a rollback plan for anything touching a schema or a stateful migration.
Why these six, and why prioritized. Each is chosen because it maps directly to a common, real production-incident root cause, and they are ordered so a reviewer under time pressure who only gets through the first three still caught the highest-impact categories.
Worked example
A pull request adds a new endpoint that calls an internal payments service and writes a record to a local database. Applying the checklist: (1) error handling: the diff has a bare except: pass around the payments call, flagged; (2) observability: no log line exists for the payments-call failure path, flagged; (3) resource leaks: the database connection is correctly used inside a context manager, passes; (4) secrets: no hardcoded credentials found, passes; (5) idempotency: the endpoint is a POST that creates a payment record with no idempotency key, and the client-facing API documentation doesn't mention retry safety, flagged as a real production risk given payments-adjacent code specifically; (6) input validation: the request body is validated via a shared schema, passes. Three of six items are flagged, and the two most severe (the swallowed exception and the missing idempotency key on a payments-adjacent endpoint) block the merge, while the missing log line is a required fix but not necessarily a hard blocker if paired with a fast-follow commitment.
Trade-offs and pitfalls
A checklist with thirty items reliably gets skimmed rather than actually applied under normal review-time pressure; keeping it to the highest-impact handful, with a concrete one-line test for each, is what makes it something a reviewer genuinely runs through on every diff rather than something referenced once and then forgotten. The most common failure mode for a checklist like this is treating it as a one-time training exercise rather than something enforced consistently: without periodic reinforcement (referencing it explicitly in review comments, tracking how often flagged items actually get raised) it tends to fade from active use within a few months of being introduced.
A boundary check validates that a value (an index, an offset, a size) falls within the range the code actually handles correctly, and it routinely catches real production bugs before they cause damage. Pick three DIFFERENT kinds of boundary bugs you've seen or can construct realistically, and for each: describe the bug it would cause if unchecked, the specific defensive check you'd add, and a unit test that would catch a regression if the check were later removed.
Sample Answer
Direct answer
A boundary check catches a specific class of bug (accessing an index, offset, or value outside the range the code actually handles correctly) at the moment it happens, instead of letting it silently produce wrong output or crash somewhere unrelated later; three concrete examples: array/list indexing, pagination offsets, and numeric limits.
Structured elaboration and worked examples
- Array indexing: the bug is an off-by-one or attacker-controlled index reading past the end of a buffer or list. The defensive check: validate
0 <= index < len(array)before accessing, raising a clearIndexError/custom exception instead of either crashing with a cryptic native error or, in an unsafe language, reading adjacent memory. A unit test:assert_raises(IndexError, get_item, [1,2,3], 5). - Pagination offsets: the bug is a negative or absurdly large
offset/limitfrom a client, which can either error confusingly deep in a SQL driver or, worse, silently return zero rows and look like 'no data' rather than 'bad request'. The defensive check: clamp or rejectoffset < 0and caplimitto a sane maximum (say 1000) before it reaches the query layer. A unit test:assert paginate(items, offset=-5, limit=10) raises ValueError. - Numeric limits: the bug is an integer overflow or an out-of-domain value (a negative quantity in an order, a percentage over 100) silently producing a nonsensical result instead of an error. The defensive check: validate the value's range explicitly before using it in a calculation. A unit test:
assert_raises(ValueError, apply_discount, price=100, percent=150).
Trade-offs and pitfalls
Each of these checks is cheap individually, but the value comes from applying them CONSISTENTLY at every place the boundary is actually crossed (every array access from external input, not just the ones you happen to remember); a single unguarded pagination endpoint added six months later by someone who didn't see this pattern reintroduces the exact bug class. Treat these as patterns to lint for or wrap in a shared utility function, not as one-off checks to remember individually.
When an authorization check throws an unexpected error, should the system fail open and allow the action, or fail closed and deny it? Walk me through how you'd decide, using a concrete example.
Sample Answer
Direct answer
For security-sensitive decisions like authorization, default to fail closed, deny the action when the check can't complete, because the cost of a false negative (briefly blocking a legitimate user) is almost always smaller than the cost of a false positive (an unauthorized action succeeding). The exception is when failing closed itself creates a worse safety or availability problem, which is a deliberate, domain-specific judgment call, not a blanket rule.
Structured elaboration
- Default posture: authorization and permission checks fail closed. If the permission service times out, deny the request with a clear error; don't silently treat "unknown" as "allowed."
- Where fail-open is sometimes deliberate: low-stakes, availability-critical paths where an outage of the check itself would be worse than the risk it guards against, and the exposure window is bounded and monitored.
- Make the decision explicit and documented per check, not an accident of how the code happens to be written; a try/catch that swallows the exception and falls through is an accidental fail-open.
- Instrument it: a fail-closed denial caused by an internal error should be logged and alerted distinctly from a legitimate permission denial, so an outage in the auth path shows up as an incident, not just a spike in 403s.
Worked example
An admin panel calls a permissions service to check if a user can delete a record. The service times out. If the code reads let allowed = true; try { allowed = permissionCheck(); } catch (e) { logger.warn(e); } followed by if (allowed) { allowDelete(); }, the caught exception is logged but never rethrown, so allowed is left at its optimistic default of true and the timeout results in the delete going through anyway, an accidental fail-open hiding inside code that looks defensive because it has a try/catch. The fix flips the default and the exception path: let allowed = false; try { allowed = permissionCheck(); } catch (e) { return deny(); }, so any failure to get a definitive answer denies the action.
Trade-offs and pitfalls
Fail-closed can turn a partial outage of a dependency into a full outage of everything gated behind it, so it needs to be paired with making that dependency itself highly available (caching the last known-good permission, short timeouts, circuit breakers) rather than accepting cascading denial as the cost of security. Fail-open, chosen without discussion, is the classic way an authorization bug ships silently, because everything still "works" in testing.
What the interviewer probes next
Whether the candidate can articulate the asymmetry between a false allow and a false deny, rather than reciting "fail closed is always right," and whether they'd catch an accidental fail-open hiding inside ordinary-looking exception handling.
Unlock Full Question Bank
Get access to all 15 Code Quality, Error Handling, and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.