Code Quality, Error Handling, and Defensive Programming Questions
Writing robust, high-quality code that fails safely. Covers defensive programming, input validation, error handling and fault tolerance, logging for diagnosability, and general engineering-quality standards. Includes anticipating failure modes and making code resilient to bad inputs and unexpected states.
Design a small set of custom static-analysis checks to detect common defensive-programming anti-patterns that make debugging harder, for example swallowing exceptions, broad try/catch blocks, and empty catch blocks. For each check, explain how you would implement it and give a sample warning message.
Sample Answer
Direct answer
A custom static-analysis check for defensive-programming anti-patterns works by pattern-matching the SHAPE of a code construct in an abstract syntax tree, not by textual pattern matching, and each check should target one specific, well-defined anti-pattern (an empty catch block, a catch that only logs and swallows, a catch of the broadest possible exception type) with a clear, actionable warning message rather than trying to build one general-purpose "bad error handling" detector.
Structured elaboration
Empty catch blocks. The simplest and highest-confidence check: find any catch/except block whose body is empty, or contains only a comment, or contains only a pass/no-op statement. This is almost never intentional and almost always indicates an error that was silently discarded during debugging and never properly handled afterward. Implementation: walk the AST for exception-handler nodes and check whether the handler's body block contains zero meaningful statements.
Broad exception catching. Flag a catch clause that catches the broadest possible exception type (except Exception, catch (Exception e), a bare except: with no type at all) UNLESS it is the outermost handler in an entry point specifically responsible for preventing a whole process crash (a top-level request handler, a main loop), which is a legitimate and common exception to the rule; implementation-wise, this means the check needs a configurable allowlist of file paths or function names where broad catching is intentional, rather than flagging every instance uniformly.
Catch-log-swallow. A subtler variant of the empty-catch problem: a catch block that logs the exception (so it isn't literally silent) but then continues execution as if nothing happened, without either re-raising, returning an error indicator to the caller, or taking an explicit recovery action. This requires slightly more sophisticated analysis than the empty-catch case: the check looks for a catch block whose only statements are a logging call, with no re-raise, no return of an error value, and no explicit recovery logic.
Sample warning messages. For an empty catch block: "Empty catch block silently discards <ExceptionType>. Either handle it explicitly or, if this really is intentional, add a comment explaining why swallowing this specific exception is safe here." For overly-broad catching: "Catching the base Exception type here will also catch bugs unrelated to what this block is meant to handle (a NullPointerException from an unrelated coding mistake, for instance). Catch the specific exception type(s) you actually expect from this call."
Integrating into developer workflow without excessive noise. Run as a fast, incremental check on changed files in CI (not a full-codebase scan on every commit, which would be slow and would re-flag pre-existing violations that aren't part of this change), and provide a narrow, explicit suppression mechanism (an inline comment like # noqa: empty-except with a required justification) for the genuine, rare cases where the pattern is intentional, rather than either blocking every legitimate exception or having no escape hatch at all.
Worked example
A simplified AST-based check for empty catch blocks, sketched in Python-like pseudocode operating over a parsed syntax tree:
def check_empty_except(tree):
violations = []
for node in tree.walk():
if node.type == "except_handler":
body_statements = [s for s in node.body if s.type not in ("comment", "pass")]
if len(body_statements) == 0:
violations.append({
"line": node.line,
"message": f"Empty except block for {node.exception_type or 'bare except'}; "
f"handle it explicitly or document why swallowing it is safe."
})
return violations
Applied to try: risky_call()\nexcept Exception:\n pass, this flags line 2 with a message naming the caught type and asking for either real handling or an explicit, documented justification, rather than a generic "bad code" warning that gives the developer no specific next step.
Trade-offs and pitfalls
A check that flags EVERY broad exception catch with no allowlist mechanism will immediately generate noisy, unwanted warnings on legitimate top-level handlers (a web framework's outermost request handler, which SHOULD catch broadly to prevent one request's unexpected error from crashing the whole server process), and a team that gets flooded with warnings on code that's actually correct will quickly learn to ignore the tool entirely, defeating its purpose; a configurable allowlist for known-legitimate broad-catch locations is what keeps the signal-to-noise ratio high enough that the warnings are still trusted. The most common implementation mistake is building this as a simple TEXT-based regex search over source files rather than an AST-based check: a regex looking for except Exception will miss a semantically-identical broad catch written slightly differently (a caught type stored in a variable, or a multi-line catch clause), and will also produce false positives on the string appearing inside a comment or a string literal that has nothing to do with actual exception-handling code.
You are developing firmware for an IoT device that sometimes loses power mid-write and ends up with corrupted state. Describe defensive strategies at the software and firmware level to ensure state consistency and recovery: journaling, atomic writes, checksums, transactional updates, wear-leveling, and graceful degradation for the case where the state cannot be repaired.
Sample Answer
Direct answer
An IoT device that can lose power mid-write needs its persistent state design to assume that ANY write can be interrupted at an arbitrary point, and to guarantee that after a power loss the device recovers to either the state before the write or the state after it, never a corrupted state in between; this is achieved through atomic writes, journaling, checksums, and a defined recovery procedure that runs on every boot.
Structured elaboration
Atomic writes. Never overwrite a live state file in place. Write the new state to a temporary location, then use a single atomic operation (a rename, on most embedded filesystems) to make it the active file; a power loss during the write leaves either the old file intact (rename never happened) or the new file fully intact (rename completed), never a half-written file masquerading as valid.
Journaling. Before making a change, write a small journal entry describing the intended change (a write-ahead log), then apply the change, then mark the journal entry complete. On boot, if an incomplete journal entry is found, the device can either replay it to completion or roll it back, rather than trusting whatever partial state exists on disk.
Checksums. Every persisted state block includes a checksum computed over its own contents. On boot, before trusting any stored state, recompute and compare the checksum; a mismatch means this block was only partially written before power was lost, and the device should fall back to its last known-good state (kept in a separate, previously-checksummed location) rather than trusting a block that failed its own integrity check.
Transactional updates. Where multiple related values must change together (updating both a counter and a corresponding flag), group them into a single atomic unit (write them as one block with one checksum, or use the journaling pattern above) rather than writing them as two separate operations, since a power loss between the two writes would otherwise leave them inconsistent with each other.
Wear-leveling. Flash storage has a limited number of write cycles per cell; spreading writes across the storage medium rather than always rewriting the same physical location extends the device's operational lifetime, and is usually handled by the flash translation layer or filesystem rather than application code, but application code should avoid patterns that defeat it, like writing extremely frequently to a fixed-size, unchanging file.
Graceful degradation when state cannot be repaired. If, after a checksum failure and a fallback to last-known-good state also fails (both copies corrupted, an unlikely but real scenario for a device that lost power during the write of BOTH), the device should fall back to a safe, known factory-default state rather than attempting to guess at a repair, and should surface this event clearly (a status flag, a log entry sent on next successful connectivity) so it is not silently operating on defaults indefinitely without anyone knowing.
Worked example
A smart thermostat storing its current schedule: instead of write(schedule_file, new_schedule) directly, the device writes new_schedule to schedule_file.tmp, computes and appends a checksum, then calls rename(schedule_file.tmp, schedule_file). If power is lost mid-write to .tmp, on reboot the device finds schedule_file unchanged (the rename never happened) and schedule_file.tmp incomplete or checksum-invalid, discards the .tmp file, and boots with the previous, still-valid schedule; the user loses only the update that was in progress, not the whole schedule. If power is lost DURING the rename itself (a much narrower window on most filesystems, since rename is designed to be atomic), the filesystem guarantees the rename either did or did not happen, never a partially-renamed file.
Trade-offs and pitfalls
Atomic-write-via-temp-file-and-rename assumes the underlying filesystem actually implements rename atomically, which is true for most embedded Linux filesystems but is a real assumption to verify for the specific hardware and filesystem in use, not something to take for granted. Checksums and journaling add both storage overhead and write amplification (more total bytes written per logical update), which directly interacts with the wear-leveling concern above; on a severely storage- or write-cycle-constrained device, this is a genuine engineering trade-off between corruption-safety and device lifetime that needs to be made deliberately, not defaulted into.
Explain patterns for handling missing or null values in strongly-typed languages like Java and dynamically-typed languages like Python or JavaScript. Include examples of Option/Maybe-style types, exceptions, and sentinel values, and explain when you would use an assertion compared to throwing a recoverable error.
Sample Answer
Direct answer
Strongly-typed languages let you make "this can be absent" part of the type itself (an Optional, Maybe, or nullable type), which forces every caller to handle the absent case at compile time; dynamically-typed languages have no such enforcement, so the same discipline has to be applied by convention, through explicit checks, sentinel values, or exceptions, and it is far easier to forget.
Structured elaboration
Option/Maybe types (Java's Optional<T>, Kotlin's T?, Rust's Option<T>). These make "might not have a value" visible in the type signature itself. A function returning Optional<User> cannot be called and have its result used as a User without the caller explicitly unwrapping it (via .get(), .orElse(default), or a null check), so the compiler catches the case where a developer forgot that the value might be absent.
Sentinel values. A special value from within the same type used to mean "nothing" (returning -1 for an index not found, or an empty string). These predate Optional types and are still common, particularly in older or lower-level codebases, but they are a real hazard: a sentinel is indistinguishable from a legitimate value of the same type unless every caller remembers to check for it, and nothing enforces that they do. indexOf returning -1 is the classic example: if a caller forgets to check and uses the result directly as an array index, it silently wraps or throws far from the actual bug.
Exceptions. Appropriate when absence represents an actual error condition the caller must react to (a required config value is missing), not merely a normal possible outcome (a user has no middle name). Throwing for something that is a completely normal case forces every caller into try/catch for ordinary control flow, which is a sign the wrong tool was chosen.
Dynamically-typed languages (Python, JavaScript). There is no compiler to force a null check, so None/null/undefined handling depends entirely on discipline: explicit is not None checks at the boundary where a value enters the system, defensive defaults (value = data.get("key", default)), and, where the codebase uses type hints, tools like mypy can catch some cases statically even though the language itself does not enforce them at runtime.
Assertions versus recoverable errors. An assertion says "this should be logically impossible given my own code's invariants; if it happens, my code has a bug", and is appropriate for catching a developer error early (an internal invariant that should never be violated if the code upstream is correct). A recoverable error (an exception or an Optional/error-result) is for a condition that is possible in the outside world regardless of whether the code is correct (a user did not provide their middle name; a file does not exist). Do not use an assertion for something a real caller can legitimately trigger, since assertions can be stripped in optimized production builds in several languages and are not guaranteed to run.
Worked example
A Java method Optional<User> findById(String id) forces every caller to write findById(id).map(User::getName).orElse("unknown") or similar, and the compiler will not let a caller treat the return value as a bare User. The equivalent Python function find_by_id(user_id) might return None on a miss, and nothing stops a caller from writing find_by_id(user_id).name and getting an AttributeError: 'NoneType' object has no attribute 'name' at runtime, potentially in a code path that only executes rarely, long after the function was written and long after the original author has moved to another project.
Trade-offs and pitfalls
Overuse of Optional/Maybe wrapping for values that are realistically always present adds ceremony without benefit; reserve it for genuinely-optional data. The most damaging mistake in dynamically-typed languages specifically is treating None/null handling as optional discipline rather than a hard rule at every boundary where external data enters the system (an API response, a database read, a config file): that is precisely where a missing null-check turns into a production incident, because it is exactly the boundary where the type system (if any) has the least information about what's actually there.
A boundary check validates that a value (an index, an offset, a size) falls within the range the code actually handles correctly, and it routinely catches real production bugs before they cause damage. Pick three DIFFERENT kinds of boundary bugs you've seen or can construct realistically, and for each: describe the bug it would cause if unchecked, the specific defensive check you'd add, and a unit test that would catch a regression if the check were later removed.
Sample Answer
Direct answer
A boundary check catches a specific class of bug (accessing an index, offset, or value outside the range the code actually handles correctly) at the moment it happens, instead of letting it silently produce wrong output or crash somewhere unrelated later; three concrete examples: array/list indexing, pagination offsets, and numeric limits.
Structured elaboration and worked examples
- Array indexing: the bug is an off-by-one or attacker-controlled index reading past the end of a buffer or list. The defensive check: validate
0 <= index < len(array)before accessing, raising a clearIndexError/custom exception instead of either crashing with a cryptic native error or, in an unsafe language, reading adjacent memory. A unit test:assert_raises(IndexError, get_item, [1,2,3], 5). - Pagination offsets: the bug is a negative or absurdly large
offset/limitfrom a client, which can either error confusingly deep in a SQL driver or, worse, silently return zero rows and look like 'no data' rather than 'bad request'. The defensive check: clamp or rejectoffset < 0and caplimitto a sane maximum (say 1000) before it reaches the query layer. A unit test:assert paginate(items, offset=-5, limit=10) raises ValueError. - Numeric limits: the bug is an integer overflow or an out-of-domain value (a negative quantity in an order, a percentage over 100) silently producing a nonsensical result instead of an error. The defensive check: validate the value's range explicitly before using it in a calculation. A unit test:
assert_raises(ValueError, apply_discount, price=100, percent=150).
Trade-offs and pitfalls
Each of these checks is cheap individually, but the value comes from applying them CONSISTENTLY at every place the boundary is actually crossed (every array access from external input, not just the ones you happen to remember); a single unguarded pagination endpoint added six months later by someone who didn't see this pattern reintroduces the exact bug class. Treat these as patterns to lint for or wrap in a shared utility function, not as one-off checks to remember individually.
Implement bool add_will_overflow(int32_t a, int32_t b) in C++ that returns true if a + b would overflow a 32-bit signed integer. Do not use a 64-bit type. Include unit tests for edge cases such as INT_MAX + 0, INT_MAX + 1, and negative overflows, and explain your approach.
Sample Answer
Direct answer
Detecting whether a + b would overflow a 32-bit signed integer, without widening to 64 bits, means reasoning about the operation BEFORE it happens using only the bounds of int32_t itself: check whether b is positive and a is already close enough to the maximum that adding b would exceed it, and symmetrically for a negative b against the minimum.
Structured elaboration
Why you can't just compute a + b and check the result. Computing the sum first and then checking whether it looks wrong is undefined behavior for signed integer overflow in C++, meaning the compiler is permitted to assume overflow never happens and can optimize the check away entirely, silently producing incorrect results specifically in the case you were trying to detect. The check has to be done using only values that are guaranteed to be representable, before the actual addition occurs.
The two symmetric cases. If b is positive, overflow happens when a is already greater than INT32_MAX - b (equivalently, adding b would push past the maximum); this comparison, a > INT32_MAX - b, is always computable without overflow since INT32_MAX - b cannot itself overflow when b is positive. If b is negative, overflow (underflow past the minimum) happens when a is less than INT32_MIN - b; note INT32_MIN - b is safe to compute here specifically because b is negative, making this subtraction move away from, not toward, the boundary.
The zero and boundary cases. b == 0 never overflows regardless of a, and the two comparisons above naturally handle this correctly without a special case, since a > INT32_MAX - 0 is simply a > INT32_MAX, which is never true for a valid int32_t value of a.
Worked example
bool add_will_overflow(int32_t a, int32_t b) {
if (b > 0 && a > std::numeric_limits<int32_t>::max() - b) return true;
if (b < 0 && a < std::numeric_limits<int32_t>::min() - b) return true;
return false;
}
Executed and verified (g++, -Wall -Wextra): add_will_overflow(INT32_MAX, 0) is false (no overflow); add_will_overflow(INT32_MAX, 1) is true (the classic overflow case); add_will_overflow(INT32_MAX - 1, 1) is false (exactly at the boundary, still valid); add_will_overflow(INT32_MIN, -1) is true (the symmetric underflow case); add_will_overflow(INT32_MIN, 0) is false; add_will_overflow(INT32_MIN + 1, -1) is false (exactly at the boundary on the negative side); ordinary values like add_will_overflow(100, 200) and add_will_overflow(-100, -200) are both false; and add_will_overflow(INT32_MAX/2 + 1, INT32_MAX/2 + 1) is true, confirming the check also catches an overflow that occurs from two moderately-large positive values rather than only from a value already at the exact boundary.
Trade-offs and pitfalls
The single most common mistake is writing the intuitive-looking but broken version, int32_t sum = a + b; if (sum < a) return true; (checking whether the result "wrapped around" to something smaller than one of the inputs): this relies on signed overflow actually wrapping, which is undefined behavior in C++ and not guaranteed to behave that way at all, especially under compiler optimizations that are explicitly permitted to assume signed overflow never occurs and can eliminate the check entirely. A second, more subtle mistake is getting the comparison direction backwards for the negative-b case (checking a < INT32_MIN + b instead of a < INT32_MIN - b), which happens to work correctly by luck for some inputs and silently fails for others; testing both boundary directions explicitly, as in the worked example, is what catches this class of subtle sign error.
Unlock Full Question Bank
Get access to all 10 Code Quality, Error Handling, and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.