Error Handling and Defensive Programming Questions
Making code robust against bad input and failure: exceptions versus error returns, input validation, guard clauses, graceful degradation, and designing for the unhappy path. Covers where to handle versus propagate errors and how to fail safely without hiding bugs. A recurring probe of production maturity.
Design a general (non-ML) CI/CD strategy to catch regressions in error handling and observability before they reach production: unit, integration, and contract tests, failure-injection (chaos) tests, gating deploys on test results, and post-deploy observability checks that can trigger an automatic rollback.
Sample Answer
Direct answer
A general (non-ML) CI/CD strategy for catching error-handling regressions layers unit tests for individual exception paths, integration tests with mocked failure modes, and failure-injection/chaos tests against real dependencies in a staging environment, gating deploys on all three passing plus post-deploy observability checks that can trigger an automatic rollback if the new version's error rate regresses in production.
Structured elaboration
- Unit tests for exception paths: test that a function raises the RIGHT exception type for a given bad input, not just that it doesn't crash on good input; a codebase's exception-path test coverage is often much lower than its happy-path coverage, and this gap is exactly where regressions like an accidentally-removed validation check hide.
- Integration tests with mocked failure modes: mock a downstream dependency to fail (timeout, 500, connection reset) and assert the calling service handles it as designed (retries appropriately, falls back, or fails cleanly with the right error), catching a regression in the FAILURE-HANDLING logic specifically, which a happy-path-only integration test never exercises.
- Contract tests: verify that your integration tests' assumptions about a downstream dependency's actual failure-mode behavior (which status code it returns for which failure, what shape an error response takes) still match the dependency's REAL current behavior, run periodically against the dependency's staging environment or a shared, versioned contract; without this tier, an integration test's mock can silently drift out of sync with what the real dependency actually does, and every test keeps passing the whole time even though the assumption it exercises is now wrong.
- Failure-injection/chaos tests: in staging, actually inject a real failure (kill a dependency's pod, add network latency) rather than mocking it, catching integration-level bugs a mock can't (a timeout configuration that's technically set but doesn't actually apply where you think it does, say).
- Gating deploys: require all three test tiers to pass before a deploy proceeds; a common structure is unit and integration tests gate the BUILD, while a scheduled/periodic (not necessarily per-deploy) chaos test suite gates confidence in staging more broadly, given chaos tests are typically slower and more expensive to run on every single commit.
- Post-deploy observability checks with auto-rollback: after a deploy reaches production, compare the NEW version's error rate/latency against the previous version's baseline over a short bake window; a significant regression triggers an automatic rollback, closing the loop for regressions that somehow passed every pre-deploy test tier but still manifest under real production traffic patterns.
Worked example
A PR accidentally removes a null check that used to convert a specific malformed input into a clean 400 response; a unit test asserting validate_payload(malformed_input) raises ValidationError would have caught this directly; if that specific test didn't exist, an integration test that sends a malformed request through the whole stack and asserts a 400 (not a 500 or a crash) provides a second layer of defense; if BOTH somehow miss it, the post-deploy observability check notices the new version's 500-error rate climbing relative to baseline within minutes of rollout and triggers an automatic rollback before most users are affected.
Trade-offs and pitfalls
The post-deploy auto-rollback safety net is valuable precisely BECAUSE pre-deploy tests can never achieve 100% coverage of every real-world failure mode, but it's not a substitute for investing in the pre-deploy layers; a team that leans entirely on post-deploy rollback as its only defense will ship (and briefly expose users to) many more regressions than a team that catches most of them earlier, cheaper, and before any real user is affected.
Given code that catches a broad exception and silently swallows it (try: ... except Exception: pass, or except Exception: print('error'), or a health-check that catches everything and always returns 200), explain why this is dangerous in production, and refactor it: catch specific exception types, log with full context and preserve the stack trace, decide when to re-raise versus recover, instrument a metric, and ensure resource cleanup still happens. Describe the unit tests you would add to prove failures are now visible.
Sample Answer
Direct answer
Refactor a broad, silent exception-swallowing pattern by catching only the specific exception types you actually anticipate, logging with full context (never a bare pass or an unlabeled print), deciding explicitly whether to re-raise (if the caller needs to know) or recover (if this layer has a genuine fallback), and instrumenting a metric so the failure is visible in aggregate even to someone not reading logs line-by-line.
Structured elaboration
- Why swallowing is dangerous:
except Exception: pass(orexcept: return Nonewith no logging) converts a loud, debuggable failure into complete silence; the bug the exception represented is still there, but now nobody knows it's firing until its downstream SYMPTOM (wrong data, a confused user, a slow accumulation of corruption) surfaces somewhere else entirely, disconnected from its actual cause. - Catch specific types: replace a bare
except Exceptionwith the SPECIFIC exception types this code path genuinely expects and can meaningfully handle (except (ValueError, KeyError):), letting anything else (a genuine bug, an unanticipated failure) propagate rather than being silently absorbed alongside the anticipated cases. - Log with context, always: even when the decision is to recover (not re-raise), log the full exception (type, message, and enough surrounding context to investigate) at an appropriate level (WARNING if recovered gracefully, ERROR if it represents a real problem worth attention).
- Instrument a metric: a counter incremented on every caught-and-recovered exception makes the failure rate VISIBLE on a dashboard, catching a sudden increase (a new bug, a data-quality regression) even for someone who never reads the raw logs.
Worked example (executed; refactor demonstrated to make failures visible where the original silently hid them)
# BEFORE (the anti-pattern):
try:
data = load_data()
process(data)
save_results(data)
except Exception:
pass
# AFTER:
import logging
logger = logging.getLogger(__name__)
processing_errors = Counter("processing_errors_total")
try:
data = load_data()
except (FileNotFoundError, json.JSONDecodeError) as e:
logger.error("failed to load data: %s", e, exc_info=True)
processing_errors.inc()
raise # a missing/corrupt data file is not something THIS layer can recover from
try:
process(data)
except ValidationError as e:
logger.warning("skipping invalid record: %s", e)
processing_errors.inc()
return # a genuinely recoverable, anticipated case: skip this record, don't crash the whole run
save_results(data) # left un-wrapped deliberately: an unexpected save failure should propagate loudly, not be caught by a broad handler meant for the two specific cases above
The refactor separates THREE distinct failure points that the original single broad try block conflated into one indistinguishable 'something failed' outcome, giving each its own specific exception type, its own recovery decision (re-raise vs skip-and-continue), and its own logged, metric-counted visibility.
Trade-offs and pitfalls
Splitting one broad try/except into several narrower ones (as shown) is more code, and it's tempting to 'simplify' back toward a single broad catch under time pressure; resist that, since the narrower version is precisely what lets a reviewer (and a future on-call engineer) understand exactly which failures this code anticipated and how each is meant to be handled, rather than a black box that swallows everything uniformly.
Implement a configuration loader (JSON or YAML, in Python or Go) that reads a config file, decodes it into a typed structure, validates required fields with clear, specific error messages (listing every missing or malformed field, not just the first), supports environment-variable overrides and sane defaults, and distinguishes failure modes explicitly (file not found returns defaults, a permission error raises a dedicated exception, a parse error is logged and does not silently succeed). Do not catch a broad Exception. Explain how you'd keep the loader maintainable as the schema evolves.
Sample Answer
Direct answer
A robust config loader distinguishes its failure modes explicitly (file missing returns defaults, a permission error raises a specific typed exception, a malformed file logs clearly and returns defaults rather than crashing outright), supports environment-variable overrides layered on top of the file, and never catches a broad Exception that would mask a genuinely unexpected failure mode alongside the anticipated ones.
Structured elaboration
- File not found: often a legitimate, expected case (no config file means 'use defaults'), so return defaults rather than raising, UNLESS the caller has explicitly indicated the file is required.
- Permission error: distinct from 'not found' and usually indicates a genuine environment misconfiguration worth surfacing loudly (a specific
ConfigAccessError), since silently falling back to defaults here could mask a real deployment problem (the service account lacks the permissions it's supposed to have). - Malformed JSON/YAML: log a clear, specific error (which file, what parse error) and fall back to defaults rather than crash the whole service on startup over a config typo, UNLESS the config is safety-critical enough that starting with defaults would itself be dangerous, in which case failing to start (loudly) is the safer choice; this is a genuine judgment call that depends on the config's actual criticality.
- Environment-variable overrides: apply AFTER loading the file, so overrides always win over file values, supporting the common deployment pattern of a baseline config file plus per-environment overrides via env vars, without needing separate config files per environment.
- Never catch broad Exception: catching only the SPECIFIC anticipated exception types (per-format parse errors,
PermissionError,FileNotFoundError) means a genuinely unexpected failure (a bug in the loader itself) still surfaces loudly rather than being silently absorbed alongside the anticipated cases.
Worked example (executed; all assertions passed)
class ConfigAccessError(Exception): pass
def load_json_config(path, required_fields, env_overrides=None):
if not os.path.exists(path):
raise FileNotFoundError(f"config not found: {path}")
try:
with open(path) as f:
cfg = json.load(f)
except json.JSONDecodeError as e:
raise ConfigAccessError(f"malformed config at {path}: {e}") from e
missing = [f for f in required_fields if f not in cfg]
if missing:
raise ConfigAccessError(f"missing required fields: {missing}")
if env_overrides:
for k, v in env_overrides.items():
if k in cfg:
cfg[k] = v
return cfg
Verified: a valid config with an env override correctly applies the override (port: 9090 overriding the file's port: 8080); a config missing a required field raises ConfigAccessError naming the specific missing field; a genuinely absent file raises FileNotFoundError distinctly; a malformed (unparseable) JSON file raises ConfigAccessError rather than an unhandled json.JSONDecodeError leaking a less-actionable, format-library-specific error type to the caller.
Trade-offs and pitfalls
This implementation deliberately RAISES FileNotFoundError rather than silently defaulting, which is a design choice worth being explicit about (some callers genuinely want 'file optional, use defaults', others want 'file required, fail loudly if absent'); a production version would likely take an explicit required: bool parameter rather than baking in one behavior, since conflating the two use cases into a single hardcoded choice is a common source of surprise for a caller with the opposite expectation.
You're leading a team with recurring bugs caused by poor error handling and sparse tests (or balancing shipping new features against investing time in defensive engineering). How would you introduce team-level practices to improve this over a quarter: code-review rules, linters, templates, testing quotas, and a phased rollout that gets buy-in from product? Describe your prioritization framework and how you'd measure success.
Sample Answer
Direct answer
Introduce team-level defensive-coding practices over a quarter through a phased rollout (define the practices, socialize and get buy-in, enforce via review/tooling, measure results) rather than a single mandate, since a practice imposed without buy-in or automated enforcement reliably decays back to old habits within weeks.
Structured elaboration
- Phase 1 (weeks 1-2): define and socialize: write down the SPECIFIC practices (not 'write better error handling' but concrete rules: no bare except, every public function validates its inputs, every resource-acquiring block uses a context manager) and discuss them WITH the team, incorporating their pushback, rather than presenting a finished mandate top-down.
- Testing quotas: pair the code-review rules with a lightweight, enforceable minimum (no PR touching error-handling code merges without at least one new test covering the failure path), tracked via a coverage-delta check in CI rather than left to reviewer memory, since 'sparse tests' was named alongside poor error handling as one of the two root problems and needs its own concrete lever, not just an assumption that better error-handling rules will incidentally produce more tests.
- Phase 2 (weeks 3-6): tooling and templates: back the practices with automated enforcement where possible (a lint rule catching the worst offenders) and a PR template/checklist reminding reviewers to check for the rest, so the practice doesn't rely purely on every individual remembering it every time.
- Phase 3 (weeks 7-10): review-driven enforcement: make the practices an explicit part of code review, with the manager (you) modeling the review comments initially so the team sees the calibration (how strict is too strict) rather than each reviewer independently guessing.
- Phase 4 (weeks 11-13): measure and adjust: track a concrete metric (incidents traced to the target failure classes, or a code-quality proxy like lint-rule violation trend) and share the result with the team AND with product, closing the loop on whether the investment paid off.
- Getting buy-in from product: frame the ask in terms product cares about (fewer firefighting-driven schedule disruptions, more predictable delivery) rather than purely as an engineering-quality initiative, and be explicit about the SHORT-TERM velocity cost (code review will be slightly slower initially) versus the medium-term payoff (fewer incidents pulling engineers off roadmap work).
Worked example
Week 1: propose 4 specific rules to the team in a design discussion, incorporating feedback that 2 of the originally-proposed rules were too strict for legacy code paths and should apply to new code only, and agree on the testing quota (one new test per error-handling fix); week 3: ship a lint rule catching bare-except patterns and a CI coverage-delta check enforcing the testing quota, both added as warnings (not yet blocking); week 7: flip the lint rule and the coverage-delta check to blocking for new code, with review-checklist backing for the harder-to-automate rules; week 13: present to the team and to product that incidents in the target category dropped from 3/month to 0/month over the quarter, with review turnaround time increasing by a measured (and acceptable) 10%, framing this as a concrete trade the data supports continuing.
Trade-offs and pitfalls
A rollout that skips the socialization/buy-in phase and goes straight to enforcement generates resentment and passive resistance (reflexive suppression comments, grudging compliance without real behavior change); a rollout that socializes endlessly without ever reaching enforcement never actually changes anything, since good intentions alone don't survive contact with a looming deadline. The phased structure exists specifically to avoid both failure modes.
Design user-facing error messaging for a web or ML-serving application when a request fails (temporary server overload, a validation failure, a corrupted model file, a GPU OOM). The message must be helpful (what happened, what the user can do) without exposing internal details; provide platform variants (mobile vs desktop, screen-reader accessibility) and describe retry UX (automatic retry, a 'try again' button with backoff). For each failure case, give a one-line example of the safe user-facing message and note what extra diagnostic metadata belongs only in the logs.
Sample Answer
Direct answer
Design user-facing error messaging that tells the user what happened and what they can do, without exposing internal details, adapted per platform (mobile vs desktop) and accessibility need (screen readers), while capturing full diagnostic detail separately in logs/telemetry so engineering retains everything needed to investigate even though the user only sees the safe, helpful version.
Structured elaboration and worked examples per failure case
- Temporary server overload: user-facing copy: "We're experiencing high demand right now. Please try again in a moment." with an automatic retry (with backoff) attempted silently before showing this message at all, and a visible 'Try again' button as the manual fallback; screen-reader accessible via an ARIA live region so the message is announced, not just visually shown.
- Validation failure: user-facing copy: field-specific, actionable ("Email address must include an @ symbol") rather than generic ("Invalid input"); shown inline near the specific field, not just as a top-of-page banner a screen-reader user might miss the association for.
- Corrupted model file / internal error: user-facing copy: a generic, honest "Something went wrong on our end. Our team has been notified." (never surfacing the internal cause), paired with a correlation id shown in small text for support purposes if the user needs to report it.
- GPU out-of-memory / capacity issue: user-facing copy, if user-visible at all (often this should be handled entirely server-side via graceful degradation before the user ever sees an error): "This request is taking longer than usual. We're working on it." with a longer, honest wait-time expectation rather than a generic failure.
- Platform variants (mobile vs desktop): on mobile, prefer a compact inline banner or toast that doesn't force a modal dismiss (screen space is scarce and a blocking modal is more disruptive on a small screen), with the 'Try again' action reachable by a single thumb-friendly tap; on desktop, a slightly more detailed inline message near the affected control (with room for a visible correlation id and a 'report this' link) is acceptable given the extra screen real estate and typically lower time-pressure; both platforms route the SAME underlying message text and correlation id through their respective native error/toast components, and both honor the screen-reader accessibility requirement (an ARIA live region on web, the equivalent native accessibility announcement API on mobile) rather than treating accessibility as a desktop-only concern.
- Retry UX: automatic retry for genuinely transient conditions (attempted silently, user only sees a brief loading state, not an error, if the retry succeeds quickly); a manual 'Try again' button with its own backoff (disabled briefly after a click to prevent a frustrated user from hammering it) for cases where automatic retry has been exhausted.
- Diagnostic capture: every shown error, regardless of how generic the user-facing copy is, is paired with a FULL internal log entry (exact exception, stack trace, request context) tagged with the SAME correlation id shown to the user, so a support ticket referencing that id gives engineering the complete picture even though the user-facing message deliberately said very little.
Trade-offs and pitfalls
A message that's TOO generic across every failure type ('something went wrong' for everything, including a simple validation error) frustrates users who could have easily self-corrected a validation mistake if told specifically what was wrong; calibrate specificity to the failure type: maximally specific and actionable for CLIENT-caused failures (validation), deliberately generic (but honest and reassuring) for SERVER-caused failures where detail would either leak internals or simply not help the user do anything differently.
Unlock Full Question Bank
Get access to all Error Handling and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.