Code Quality, Error Handling, and Defensive Programming Questions
Writing robust, high-quality code that fails safely. Covers defensive programming, input validation, error handling and fault tolerance, logging for diagnosability, and general engineering-quality standards. Includes anticipating failure modes and making code resilient to bad inputs and unexpected states.
A backend service builds a SQL query by concatenating a user-supplied search string directly into the query text. What's wrong with this from a defensive-programming standpoint, and what would you check for in code review to catch this class of bug at scale?
Sample Answer
Direct answer
Concatenating untrusted input directly into a SQL string lets an attacker change the query's structure, not just its data, which is SQL injection and can expose or modify anything the database connection can reach. The fix is parameterized queries, where input is always bound as data and can never be interpreted as SQL syntax; at scale, catching this means a static-analysis rule that flags any string concatenation feeding a query-execution call, not relying on a human to spot every instance.
Structured elaboration
- Why it's dangerous: if the search string is a SQL fragment like a statement terminator followed by a destructive command, a concatenated query executes it as SQL, because the database can't distinguish "data the user typed" from "query the developer wrote" once they're merged into one string.
- The fix: parameterized queries or an ORM's query builder, where the driver sends the query text and parameters separately, so the database always treats parameters as literal values.
- Scaling the catch beyond code review: a static-analysis rule, via a linter or a tool like Semgrep, that flags any database call built from string concatenation or interpolation, run in CI so it blocks the pull request automatically.
- Defense in depth: least-privilege database credentials, so the app's database user doesn't have permission to drop tables at all, shrinking the blast radius if a check is ever bypassed.
Worked example
Vulnerable: building the query text as SELECT * FROM users WHERE name = ' plus the raw user input plus '. Safe: SELECT * FROM users WHERE name = ? with the input passed as a bound parameter. With the vulnerable version, an input like x' OR '1'='1 turns the WHERE clause into always-true, returning every row instead of matching one name.
Trade-offs and pitfalls
Parameterization doesn't cover every injection surface; dynamically chosen table or column names can't be parameterized the same way and still need explicit allow-list validation, not string interpolation. A common pitfall is fixing the obvious query but missing a second one built the same way in a reporting script or admin tool that doesn't go through the same review path.
What the interviewer probes next
Whether the candidate reaches for "parameterize it" immediately, then goes further into how they'd find every other instance of the pattern across a large codebase, not just this one.
Write a small JavaScript function safeParseJSON(jsonString) that returns an object {success: boolean, value: any, error: string|null}. It must not throw, must handle invalid JSON gracefully, and must not allow prototype pollution (it must avoid assigning to proto or constructor). Also write one example unit test in Jest for an invalid input.
Sample Answer
Direct answer
A safe JSON parser must never throw on bad input, must report success or failure through its return value instead, and must specifically guard against prototype pollution, which is JSON.parse's own quiet trap: a payload like {"__proto__": {"polluted": true}} can corrupt the global Object prototype for the entire process if you are not careful about how you consume the parsed result.
Structured elaboration
Never throw. JSON.parse throws a SyntaxError on malformed input; a "safe" wrapper catches that and returns a structured result instead, so callers do not need a try/catch at every call site.
Prototype pollution. JSON.parse itself does not mutate Object.prototype (unless you additionally use a reviver or later merge the result unsafely into another object with a naive deep-merge). The real risk in a "safe parse" utility is what happens AFTER parsing: if calling code later does something like Object.assign(defaults, parsed) or a recursive merge without checking keys, an attacker-controlled __proto__ or constructor.prototype key can walk up to the shared prototype and add or override properties on every object in the process. The defensive fix at the parse boundary is twofold: use a JSON.parse reviver function that strips dangerous keys (__proto__, constructor, prototype) as soon as they are encountered, and additionally verify with Object.prototype.hasOwnProperty.call that the top-level parsed object does not carry one of those keys before returning it, since a reviver alone can miss certain nested shapes depending on how the object is later traversed.
Structured return, not an exception. Return {success, value, error} so a caller writes if (!result.success) { ...handle... } instead of a try/catch, which keeps the calling code linear and makes "parsing failed" an ordinary value instead of a special control-flow path.
Worked example
function safeParseJSON(jsonString) {
if (typeof jsonString !== 'string') {
return { success: false, value: null, error: 'input is not a string' };
}
let parsed;
try {
parsed = JSON.parse(jsonString, (key, value) => {
if (key === '__proto__' || key === 'constructor' || key === 'prototype') return undefined;
return value;
});
} catch (err) {
return { success: false, value: null, error: err.message };
}
if (parsed && typeof parsed === 'object') {
if (Object.prototype.hasOwnProperty.call(parsed, '__proto__') ||
Object.prototype.hasOwnProperty.call(parsed, 'constructor')) {
return { success: false, value: null, error: 'disallowed key detected' };
}
}
return { success: true, value: parsed, error: null };
}
Executed (Node.js, verified): safeParseJSON('{"a": 1, "b": [1,2,3]}') returns {success: true, value: {a: 1, b: [1,2,3]}, error: null}. safeParseJSON('{not valid json') returns success: false with the underlying SyntaxError message, and does not throw. safeParseJSON(42) returns success: false without ever calling JSON.parse on a non-string. Critically, safeParseJSON('{"__proto__": {"polluted": true}}') was run and confirmed that ({}).polluted is undefined afterward, meaning the global Object prototype was NOT polluted, and the parsed value has no __proto__ own-key surviving in it; the same holds for a constructor.prototype variant of the attack.
One Jest unit test for the invalid-input case:
test('returns a failure result instead of throwing on malformed JSON', () => {
const result = safeParseJSON('{not valid json');
expect(result.success).toBe(false);
expect(result.value).toBeNull();
expect(typeof result.error).toBe('string');
});
Trade-offs and pitfalls
The reviver-based key strip handles the common shape of the attack, but a defense-in-depth mindset says: never rely on parse-time stripping alone if you also deep-merge untrusted objects elsewhere in the codebase, because a different merge utility downstream might not go through this parser at all. Prefer Object.create(null) or a Map for any object you build from untrusted keys if you can avoid prototype-based objects entirely. A common mistake is checking for __proto__ as an enumerable key with a plain for...in loop, which will not find it, since __proto__ set via the object literal syntax is an accessor, not an own enumerable property in the usual sense; using hasOwnProperty.call directly, as above, avoids that gap.
When an authorization check throws an unexpected error, should the system fail open and allow the action, or fail closed and deny it? Walk me through how you'd decide, using a concrete example.
Sample Answer
Direct answer
For security-sensitive decisions like authorization, default to fail closed, deny the action when the check can't complete, because the cost of a false negative (briefly blocking a legitimate user) is almost always smaller than the cost of a false positive (an unauthorized action succeeding). The exception is when failing closed itself creates a worse safety or availability problem, which is a deliberate, domain-specific judgment call, not a blanket rule.
Structured elaboration
- Default posture: authorization and permission checks fail closed. If the permission service times out, deny the request with a clear error; don't silently treat "unknown" as "allowed."
- Where fail-open is sometimes deliberate: low-stakes, availability-critical paths where an outage of the check itself would be worse than the risk it guards against, and the exposure window is bounded and monitored.
- Make the decision explicit and documented per check, not an accident of how the code happens to be written; a try/catch that swallows the exception and falls through is an accidental fail-open.
- Instrument it: a fail-closed denial caused by an internal error should be logged and alerted distinctly from a legitimate permission denial, so an outage in the auth path shows up as an incident, not just a spike in 403s.
Worked example
An admin panel calls a permissions service to check if a user can delete a record. The service times out. If the code reads let allowed = true; try { allowed = permissionCheck(); } catch (e) { logger.warn(e); } followed by if (allowed) { allowDelete(); }, the caught exception is logged but never rethrown, so allowed is left at its optimistic default of true and the timeout results in the delete going through anyway, an accidental fail-open hiding inside code that looks defensive because it has a try/catch. The fix flips the default and the exception path: let allowed = false; try { allowed = permissionCheck(); } catch (e) { return deny(); }, so any failure to get a definitive answer denies the action.
Trade-offs and pitfalls
Fail-closed can turn a partial outage of a dependency into a full outage of everything gated behind it, so it needs to be paired with making that dependency itself highly available (caching the last known-good permission, short timeouts, circuit breakers) rather than accepting cascading denial as the cost of security. Fail-open, chosen without discussion, is the classic way an authorization bug ships silently, because everything still "works" in testing.
What the interviewer probes next
Whether the candidate can articulate the asymmetry between a false allow and a false deny, rather than reciting "fail closed is always right," and whether they'd catch an accidental fail-open hiding inside ordinary-looking exception handling.
For a public API, design a policy that decides what error detail is safe to return to CLIENTS versus what stays only in internal logs. Include examples of safe client-facing error formats, how to include a correlation id without leaking internals, and whether/when to include a stack trace in a log versus an API response. Propose an automated test that ensures no sensitive field ever leaks into a client-facing response.
Sample Answer
Direct answer
Decide what error detail reaches a client by defaulting to the minimum that's actually actionable for THAT client (a correlation id and a stable error code, always; a human-readable message only if it's genuinely safe and useful; never a stack trace or internal identifiers), keeping the full detail in internal logs correlated by the same id.
Structured elaboration
- Safe client-facing format:
{error_code, message, correlation_id}at minimum; themessageshould describe what went wrong from the CLIENT's perspective ("the email field is required") never from the server's internal perspective ("NullPointerException in UserValidator.java line 42"). - Correlation ids without leaking internals: a correlation id is safe to expose (it's an opaque token, not information about your system) and is exactly what lets support/engineering find the FULL internal detail later, without the client ever seeing that detail directly.
- Localized user messages: keep the machine-readable
error_codestable and English-invariant; localize the human-readablemessageseparately based on the client's locale, so client code branching onerror_codenever breaks when message wording/translation changes. - Automated tests for no leakage: a test suite that deliberately triggers every known internal exception type and asserts the CLIENT-FACING response contains none of a blocklist of sensitive patterns (stack trace markers, internal hostnames, SQL fragments, raw exception class names for internal errors) catches this class of leak before it ships, since manual review alone reliably misses it under time pressure.
Worked example
An internal psycopg2.OperationalError: could not connect to server: Connection refused... host "10.2.4.19" must never reach a client; the sanitized response is {"error_code": "internal_error", "message": "Something went wrong on our end. Please try again.", "correlation_id": "7f3e-9c"}, while the full raw exception (including the internal hostname) is logged server-side, findable by an engineer searching for correlation_id: 7f3e-9c.
Trade-offs and pitfalls
The hardest cases are 5xx errors that ARE genuinely useful for the client to know more about (a specific downstream service being down, which the client's own retry logic might want to know about specifically); resist the urge to pass through the raw exception message even here, and instead define a small, deliberate set of STRUCTURED, safe detail fields ({"error_code": "dependency_unavailable", "dependency": "payment_gateway"}) rather than either a blanket generic message or a raw leak.
Design a Java logging helper that redacts common PII, such as email addresses, Social Security numbers, and credit-card numbers, from log messages before they are written. State your assumptions, show the use of compiled regular-expression patterns, discuss the performance considerations, explain how you would configure the helper to extend the redaction patterns, and describe how you would test and validate it at scale. Also cover what structured fields you would include (for example a correlation ID and a job or request identifier), what you would log at INFO versus DEBUG level, and how the approach differs for a nightly batch scoring job versus a real-time service.
Sample Answer
Direct answer
A PII-safe logging helper redacts common sensitive patterns (emails, Social Security numbers, credit card numbers) from a message before it is ever written to a log, using compiled regular expressions applied consistently across every log call, while also enforcing structured fields (a correlation ID, a job or request identifier), log-level discipline, and an extensibility point for adding new redaction patterns as new sensitive-data types are identified.
Structured elaboration
Redaction via compiled regex patterns. A small set of well-tested regular expressions, compiled once and reused (not recompiled on every log call, which would be wasteful), match common PII shapes: an email address pattern, a Social Security number pattern (\d{3}-\d{2}-\d{4}), and a credit-card-like sequence of 13 to 16 digits (allowing spaces or dashes as separators, since real card numbers are often written with them). Each match is replaced with a fixed [REDACTED] marker before the message is passed to the underlying logging framework.
Performance considerations. Compiling patterns once at startup (not per-call) and running them against every log message adds a small, roughly-constant regex-matching cost per log call; for a very high-throughput logging path this is measurable but usually acceptable, since the alternative (an actual PII leak into a log aggregation system with far broader read access than the original data source) is a materially worse outcome, and the specific cost can be validated by simply measuring log throughput with and without the redaction step in a realistic load test.
Configuration to extend redaction patterns. New PII patterns identified over time (an internal account-number format specific to the business, for instance) should be addable without modifying the core logging class, via a constructor parameter or configuration file accepting additional patterns, so the redaction logic can grow as new sensitive-data types are identified in practice, not require a code change and redeploy of the core logging utility itself for every new pattern.
Testing and validating at scale. Beyond unit tests confirming each pattern redacts correctly and that non-sensitive messages pass through unchanged, validating "at scale" means running the redaction logic against a genuinely large, realistic sample of actual (or realistically-synthetic) log messages and manually or statistically auditing a sample of the OUTPUT for anything that looks like it should have been redacted but wasn't, since regex patterns can have false negatives on real-world data that a small, hand-written unit test suite won't surface (an international phone number format, a differently-formatted SSN with no dashes).
Structured fields, log levels, and INFO versus DEBUG. Beyond redaction, the logging helper should ensure every log entry carries a correlation ID (tying related log lines from the same request or job together) and a job/run identifier where relevant (for a batch context specifically). INFO-level logging should capture what a normal operator needs to see to understand system behavior (a job started, a job completed, a summary count); DEBUG-level logging can include more granular detail useful only during active troubleshooting, but should still be redacted with exactly the same discipline as INFO, since a DEBUG log accidentally left enabled in production is a very common real-world path by which sensitive data actually ends up in logs.
Differences for a batch job versus a real-time service. A nightly batch scoring job's logging strategy centers on a per-run summary (start time, record count processed, error count, completion status) tagged with a job_id and run_id, since a human reviewing a batch job's logs the next morning wants an overview, not necessarily a line per record. A real-time service's logging centers on a per-REQUEST correlation ID and typically much higher log volume, and needs sampling strategies for very high-traffic endpoints (logging a representative fraction of successful requests at INFO, while still logging every failure) to keep log volume and cost manageable without losing visibility into failures specifically.
Worked example
public class PiiSafeLogger {
private static final Pattern EMAIL = Pattern.compile("[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}");
private static final Pattern SSN = Pattern.compile("\\b\\d{3}-\\d{2}-\\d{4}\\b");
private static final Pattern CREDIT_CARD = Pattern.compile("\\b(?:\\d[ -]*?){13,16}\\b");
private final List<Pattern> patterns = new ArrayList<>(List.of(EMAIL, SSN, CREDIT_CARD));
public PiiSafeLogger(List<Pattern> extraPatterns) { patterns.addAll(extraPatterns); }
public String redact(String message) {
if (message == null) return null;
String result = message;
for (Pattern p : patterns) result = p.matcher(result).replaceAll("[REDACTED]");
return result;
}
}
Verification note: a Java Development Kit was not available in this execution sandbox, so the exact class above was not compiled directly; the identical regular-expression patterns and replacement logic were instead executed against an equivalent Python translation (Python's re module uses the same pattern syntax for these specific expressions) and confirmed: an email address embedded in a sentence is fully redacted; a dashed Social Security number is fully redacted; a spaced 16-digit credit-card-like number is fully redacted; a message containing no PII passes through completely unchanged; and a null input returns null rather than throwing. This confirms the REGEX LOGIC is correct; it does not confirm Java-specific behavior (for example, java.util.regex.Pattern's exact semantics versus Python's re), which is a disclosed limitation of this verification, not a claim of full Java execution.
Trade-offs and pitfalls
Regex-based redaction is a strong first line of defense but is not exhaustive: a credit card number with unusual formatting, a non-US identification number format, or a PII field that doesn't match any recognizable pattern at all (a person's name in a free-text field, which has no distinguishing shape a regex can reliably catch) can pass through unredacted; this is exactly why validating "at scale" against a realistic log sample, not just a handful of unit tests, matters, and why some organizations pair regex-based redaction with an explicit ALLOWLIST discipline (only log fields you've deliberately reviewed) for the highest-sensitivity data flows, rather than relying on a denylist-style redaction pattern to catch everything. The most common real-world failure mode is a DEBUG-level log statement, written during development and forgotten, that logs an entire request or record object directly (bypassing the redaction helper entirely, since it wasn't routed through it) rather than going through this centralized logging path, which is why enforcing that ALL logging goes through a single, redacting logger (not calling a raw print or an unwrapped logging framework call directly) is as important as the redaction logic itself.
Unlock Full Question Bank
Get access to all 10 Code Quality, Error Handling, and Defensive Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.