Code Quality, Error Handling, and Defensive Programming Questions
Writing robust, high-quality code that fails safely. Covers defensive programming, input validation, error handling and fault tolerance, logging for diagnosability, and general engineering-quality standards. Includes anticipating failure modes and making code resilient to bad inputs and unexpected states.
Explain patterns for handling missing or null values in strongly-typed languages like Java and dynamically-typed languages like Python or JavaScript. Include examples of Option/Maybe-style types, exceptions, and sentinel values, and explain when you would use an assertion compared to throwing a recoverable error.
Sample Answer
Direct answer
Strongly-typed languages let you make "this can be absent" part of the type itself (an Optional, Maybe, or nullable type), which forces every caller to handle the absent case at compile time; dynamically-typed languages have no such enforcement, so the same discipline has to be applied by convention, through explicit checks, sentinel values, or exceptions, and it is far easier to forget.
Structured elaboration
Option/Maybe types (Java's Optional<T>, Kotlin's T?, Rust's Option<T>). These make "might not have a value" visible in the type signature itself. A function returning Optional<User> cannot be called and have its result used as a User without the caller explicitly unwrapping it (via .get(), .orElse(default), or a null check), so the compiler catches the case where a developer forgot that the value might be absent.
Sentinel values. A special value from within the same type used to mean "nothing" (returning -1 for an index not found, or an empty string). These predate Optional types and are still common, particularly in older or lower-level codebases, but they are a real hazard: a sentinel is indistinguishable from a legitimate value of the same type unless every caller remembers to check for it, and nothing enforces that they do. indexOf returning -1 is the classic example: if a caller forgets to check and uses the result directly as an array index, it silently wraps or throws far from the actual bug.
Exceptions. Appropriate when absence represents an actual error condition the caller must react to (a required config value is missing), not merely a normal possible outcome (a user has no middle name). Throwing for something that is a completely normal case forces every caller into try/catch for ordinary control flow, which is a sign the wrong tool was chosen.
Dynamically-typed languages (Python, JavaScript). There is no compiler to force a null check, so None/null/undefined handling depends entirely on discipline: explicit is not None checks at the boundary where a value enters the system, defensive defaults (value = data.get("key", default)), and, where the codebase uses type hints, tools like mypy can catch some cases statically even though the language itself does not enforce them at runtime.
Assertions versus recoverable errors. An assertion says "this should be logically impossible given my own code's invariants; if it happens, my code has a bug", and is appropriate for catching a developer error early (an internal invariant that should never be violated if the code upstream is correct). A recoverable error (an exception or an Optional/error-result) is for a condition that is possible in the outside world regardless of whether the code is correct (a user did not provide their middle name; a file does not exist). Do not use an assertion for something a real caller can legitimately trigger, since assertions can be stripped in optimized production builds in several languages and are not guaranteed to run.
Worked example
A Java method Optional<User> findById(String id) forces every caller to write findById(id).map(User::getName).orElse("unknown") or similar, and the compiler will not let a caller treat the return value as a bare User. The equivalent Python function find_by_id(user_id) might return None on a miss, and nothing stops a caller from writing find_by_id(user_id).name and getting an AttributeError: 'NoneType' object has no attribute 'name' at runtime, potentially in a code path that only executes rarely, long after the function was written and long after the original author has moved to another project.
Trade-offs and pitfalls
Overuse of Optional/Maybe wrapping for values that are realistically always present adds ceremony without benefit; reserve it for genuinely-optional data. The most damaging mistake in dynamically-typed languages specifically is treating None/null handling as optional discipline rather than a hard rule at every boundary where external data enters the system (an API response, a database read, a config file): that is precisely where a missing null-check turns into a production incident, because it is exactly the boundary where the type system (if any) has the least information about what's actually there.
For a public API, design a policy that decides what error detail is safe to return to CLIENTS versus what stays only in internal logs. Include examples of safe client-facing error formats, how to include a correlation id without leaking internals, and whether/when to include a stack trace in a log versus an API response. Propose an automated test that ensures no sensitive field ever leaks into a client-facing response.
Sample Answer
Direct answer
Decide what error detail reaches a client by defaulting to the minimum that's actually actionable for THAT client (a correlation id and a stable error code, always; a human-readable message only if it's genuinely safe and useful; never a stack trace or internal identifiers), keeping the full detail in internal logs correlated by the same id.
Structured elaboration
- Safe client-facing format:
{error_code, message, correlation_id}at minimum; themessageshould describe what went wrong from the CLIENT's perspective ("the email field is required") never from the server's internal perspective ("NullPointerException in UserValidator.java line 42"). - Correlation ids without leaking internals: a correlation id is safe to expose (it's an opaque token, not information about your system) and is exactly what lets support/engineering find the FULL internal detail later, without the client ever seeing that detail directly.
- Localized user messages: keep the machine-readable
error_codestable and English-invariant; localize the human-readablemessageseparately based on the client's locale, so client code branching onerror_codenever breaks when message wording/translation changes. - Automated tests for no leakage: a test suite that deliberately triggers every known internal exception type and asserts the CLIENT-FACING response contains none of a blocklist of sensitive patterns (stack trace markers, internal hostnames, SQL fragments, raw exception class names for internal errors) catches this class of leak before it ships, since manual review alone reliably misses it under time pressure.
Worked example
An internal psycopg2.OperationalError: could not connect to server: Connection refused... host "10.2.4.19" must never reach a client; the sanitized response is {"error_code": "internal_error", "message": "Something went wrong on our end. Please try again.", "correlation_id": "7f3e-9c"}, while the full raw exception (including the internal hostname) is logged server-side, findable by an engineer searching for correlation_id: 7f3e-9c.
Trade-offs and pitfalls
The hardest cases are 5xx errors that ARE genuinely useful for the client to know more about (a specific downstream service being down, which the client's own retry logic might want to know about specifically); resist the urge to pass through the raw exception message even here, and instead define a small, deliberate set of STRUCTURED, safe detail fields ({"error_code": "dependency_unavailable", "dependency": "payment_gateway"}) rather than either a blanket generic message or a raw leak.
Production just exhausted its error budget due to cascading 5xx errors triggered by a downstream change, and you must ship defensive changes quickly to prevent a repeat. Which mitigations do you prioritize first and why: request timeouts, retries with backoff and jitter, circuit breakers, bulkheads/isolated thread pools, backpressure, or graceful degradation? Explain how you would measure whether each change is actually working.
Sample Answer
Direct answer
Under active error-budget exhaustion from cascading 5xx errors, prioritize the mitigation that stops the cascade fastest with the least new risk: circuit breakers and request timeouts first (they cut the feedback loop immediately and are usually already-tested code paths), then backpressure/bulkheads to protect what's left, with retries-with-backoff and graceful degradation as the follow-up once the bleeding has stopped, not the first move.
Structured elaboration
- Circuit breakers, first: if a downstream change is causing cascading failures, the fastest way to stop the cascade is to stop CALLING the failing dependency; a circuit breaker (or a manual, config-driven kill switch if none exists yet for this path) halts the cascade immediately, faster than any code change can ship.
- Timeouts, immediately after: if calls to the failing dependency are hanging rather than failing fast, tightening the timeout (even a temporary, aggressive config change) frees up resources (threads, connections) being held hostage, which is often what's actually driving the cascade beyond the original failing dependency.
- Bulkheads/isolated thread pools: if the resource exhaustion has already spread to starve unrelated requests (see the bulkhead survivor), isolating pools limits further blast radius, though retrofitting a bulkhead mid-incident is a bigger, riskier change than flipping an existing breaker or timeout config.
- Retries with backoff and jitter: valuable for RECOVERY once the dependency is coming back, but retries added or left ENABLED during the active cascade make it worse, not better, by adding more load onto an already-struggling dependency; this is often the first thing to actively DISABLE, not add, during the incident.
- Backpressure: shedding load at the edge (rejecting a percentage of incoming requests outright, with a clear 503) protects the system's remaining capacity for the traffic it CAN serve, at the direct, visible cost of intentionally failing some requests.
- Graceful degradation: the longer-term fix (serve a fallback/cached response instead of failing) usually requires a code change that can't ship instantly during an active incident, so it's the follow-up hardening work, not the immediate mitigation.
- Measuring effectiveness: track the downstream dependency's own error rate and the error BUDGET burn rate in real time as each mitigation is applied; a mitigation is working if the burn rate visibly slows within minutes, not hours.
Worked example
During the incident: (1) immediately flip the circuit breaker for the failing downstream to force-open (or disable retries against it if no breaker exists) to stop the cascade; (2) tighten the client timeout for that dependency from 30s to 2s to stop threads from being held hostage; (3) if capacity is still degraded, enable load shedding (reject 20% of lowest-priority traffic) to protect the rest; (4) once the dependency confirms recovery, re-enable the breaker and retries gradually, watching the error rate as you do, rather than flipping everything back on at once.
Trade-offs and pitfalls
The instinctive first move during many outages is to add MORE retries ('the requests are failing, let's retry them harder'), which is exactly backwards during a cascading-failure incident: more retries onto an already-overloaded dependency deepens the cascade. The discipline that prevents this: stop calling the failing thing first, THEN worry about graceful recovery.
Compare and contrast graceful degradation and fail-fast design approaches for production systems. For each approach, explain a typical use case (for example a customer-facing API versus an internal pipeline), the operational trade-offs, how you would instrument each approach with metrics, logs, and traces, and how you would communicate degraded functionality to clients or downstream systems.
Sample Answer
Direct answer
Fail-fast stops the operation immediately and surfaces the error the moment something is wrong, trading availability for correctness and a clear signal; graceful degradation keeps the system partially functional by falling back to reduced capability, trading some correctness or completeness for continued availability. The right choice depends on whether a wrong or incomplete answer is worse than no answer at all for this specific system.
Structured elaboration
Fail-fast: when correctness matters more than availability. An internal data pipeline computing financial reconciliation numbers should fail fast and loudly the moment its inputs look wrong, because a wrong number that looks plausible and gets used in a report is a much worse outcome than the pipeline simply not running today. Fail-fast systems are also easier to operate: a hard failure with a clear error is diagnosable immediately, whereas a system that silently degrades can mask a real problem for a long time before anyone notices the quality of its output has quietly dropped.
Graceful degradation: when partial availability beats a hard stop. A customer-facing product page that depends on a recommendation service should degrade to a generic, non-personalized set of recommendations if that service is slow or down, rather than showing the customer an error page, because a slightly-worse-but-functional page is a much better outcome for both the customer and the business than a hard failure on a page that otherwise works fine.
Instrumentation differs by approach. A fail-fast system needs strong alerting on the failure itself, since the failure IS the signal: an error rate spike, a specific exception type, or a circuit breaker opening. A gracefully-degrading system needs the opposite kind of visibility: a metric or log line specifically for "we are currently in degraded mode", because the degraded path, by design, does not look like a failure to a simple error-rate dashboard, and a team that isn't specifically tracking degraded-mode usage can be running in a permanently degraded state for months without noticing.
Communicating degraded functionality. For a user-facing system, this usually means a visible but non-alarming UI signal ("Showing popular items while personalized recommendations are unavailable") rather than silence, since silent degradation erodes trust once a user notices the quality difference without being told why. For a downstream service-to-service dependency, this means an explicit field or header in the response indicating degraded mode, so the calling service can make its own informed choice about whether to also degrade or to fail.
Worked example
An internal pipeline vs. a downstream API dependency, side by side: the internal pipeline computing quarterly revenue numbers for a financial report should fail fast and halt if a required upstream table is empty or a row count sanity check fails, alerting the on-call data engineer immediately, because publishing a subtly wrong number in a financial report is far worse than the report being late. The customer-facing recommendation widget on the same company's storefront, dependent on a separate ML service, should instead catch a timeout from that service and immediately serve a cached "most popular this week" list, log a degraded_mode=true metric tagged with the reason, and continue serving the page, because an incomplete page for one widget is a minor UX cost, not a correctness failure the business needs to halt over.
Trade-offs and pitfalls
The most common mistake is applying the wrong default to a whole system uniformly: treating every dependency as fail-fast produces a fragile product where one non-critical service outage takes down an entire page, while treating every dependency as gracefully-degradable risks quietly serving wrong financial or safety-relevant data with no alert ever firing. The decision should be made dependency by dependency, based on whether being wrong is worse than being unavailable for that specific piece of functionality, not applied as a single system-wide policy.
Implement a React ErrorBoundary component that logs errors to a provided logger (for example, an error-tracking service like Sentry) and displays a localized fallback UI when a child component throws during render. Then write a React Testing Library test that asserts the logger was called and that the fallback text is rendered. Explain what an ErrorBoundary will and will not catch, and discuss the trade-off between showing a retry UI and surfacing the raw error to the user.
Sample Answer
Direct answer
A React error boundary catches rendering errors thrown by any component in its subtree during render, in lifecycle methods, and in constructors, logs the error with enough context to diagnose it, and shows a fallback UI instead of leaving the user with a blank screen or React's own default error overlay; it does not catch errors in event handlers, asynchronous code, or errors thrown in the boundary component itself.
Structured elaboration
What it catches. Errors thrown during the render phase of any component below the boundary in the tree, including errors in lifecycle methods (componentDidMount, etc.) and in constructors. This is React's mechanism for preventing one broken component from crashing the entire application.
What it does NOT catch, and why that matters. Event handlers (a click handler that throws is a normal JavaScript exception, not a React rendering error, and needs its own try/catch); asynchronous code (a .then() callback or an async function's rejection happens outside React's render cycle entirely); server-side rendering errors; and errors thrown by the error boundary component itself (a boundary cannot catch its own failures, which is why the boundary component should be kept as simple as possible, with minimal logic that could itself throw).
Logging to an error-tracking service. componentDidCatch(error, info) receives both the error object and a componentStack describing which component tree led to the failure; sending both to a service like Sentry, tagged with any available user or session context, turns "a customer reported a blank page" into "we can see exactly which component threw, with what stack, for which user" without waiting for the customer to describe what they were doing.
Retry UI versus surfacing the raw error. A "try again" button that resets the boundary's state and re-attempts rendering the subtree is appropriate when the failure might be transient (a component that failed because of a momentary bad prop from a slow API response); it is misleading for a deterministic bug that will fail identically on every retry, where a generic "something went wrong, we've been notified" message (with no false promise that retrying will help) is more honest to the user, even though it is less satisfying than a button that appears to offer control.
Worked example
class ErrorBoundary extends React.Component {
constructor(props) { super(props); this.state = { hasError: false }; }
static getDerivedStateFromError(error) { return { hasError: true }; }
componentDidCatch(error, info) {
if (this.props.logger) this.props.logger.logError(error, info.componentStack);
}
render() {
if (this.state.hasError) {
return <div role="alert">{this.props.fallbackText || 'Something went wrong.'}</div>;
}
return this.props.children;
}
}
Executed (React Testing Library, verified): rendering <ErrorBoundary logger={logger} fallbackText="We hit a snag. Please retry."><Boom /></ErrorBoundary>, where Boom throws during render, confirms screen.getByRole('alert') shows the fallback text and logger.logError was called exactly once with the thrown Error object. A second test confirms that when no child throws, the boundary renders its children normally and logger.logError is never called, so the boundary is confirmed to be transparent in the non-error case, not just functional in the error case.
Trade-offs and pitfalls
A single application-wide error boundary at the root catches everything but takes down the ENTIRE page for a failure in one small, non-critical widget; placing boundaries around individual independent sections (a sidebar widget, a comments section) means one broken component degrades gracefully to just that section showing a fallback, while the rest of the page keeps working, which is almost always the better default for anything with multiple independent sections. The most common mistake is assuming an error boundary catches an async data-fetching failure inside a useEffect: it does not, since that error occurs outside the render phase entirely, and needs its own explicit error state managed by the component, separate from the boundary mechanism.
Unlock Full Question Bank
Get access to all 10 Code Quality, Error Handling, and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.