Secure Coding and Application Security Questions
Writing and reviewing code that resists attack. Covers the OWASP Top Ten and common web vulnerabilities (XSS, SQL injection, CSRF), input validation, secure coding practices and security code review, static application security testing (SAST), API and HTTP security, database and frontend security, and mobile app security. The application-layer defense discipline for engineers building software.
In a modern single-page-application-plus-REST-API architecture, how would you implement CSRF defenses? Describe how SameSite cookie attributes, anti-CSRF tokens (double-submit cookie), and origin checks complement or conflict with JWTs carried in Authorization headers, and whether storing a token in localStorage changes the calculus. Recommend a default approach for a large organization, and explain why you would choose it over the alternatives.
Sample Answer
Direct answer: In a modern SPA-plus-REST-API architecture, CSRF defenses need to account for how the app actually authenticates - if auth is via a cookie, standard CSRF defenses apply directly; if auth is via a JWT (JSON Web Token - a self-contained token carrying the user's identity and claims, sent explicitly by the client rather than auto-attached like a cookie) in an Authorization header, CSRF risk drops sharply but doesn't disappear if a refresh-token cookie is still involved anywhere in the flow.
Structured elaboration - how the defenses interact with JWTs specifically.
SameSite cookie attributes protect whatever is stored in a cookie - typically the refresh token or a session identifier, not the JWT access token itself if that's kept in memory or Authorization headers. SameSite=Strict on the refresh-token cookie means even if an attacker forges a request, the browser withholds that cookie cross-site, so the forged request has no valid session context to act on.
Anti-CSRF tokens (double-submit) are largely redundant work if the API requires the JWT in an Authorization: Bearer header for every state-changing call, because an attacker's cross-site forged request has no way to know or attach that header value - the browser doesn't auto-attach Authorization headers the way it auto-attaches cookies. This is the single biggest architectural difference from cookie-based session auth: moving the credential from an auto-attached cookie to a manually-set header is itself a strong CSRF defense, independent of any explicit anti-CSRF token mechanism.
Origin checks (validating the Origin/Referer header server-side against an allowlist of expected frontend origins) are a cheap, effective additional layer regardless of the auth mechanism, catching forged cross-origin requests even in edge cases the other controls miss.
Does storing a token in localStorage change the calculus? Yes, but in the OPPOSITE direction from CSRF: a JWT in localStorage is immune to CSRF (JavaScript from a different origin cannot read another origin's localStorage, so an attacker's page can't retrieve it to forge a request even if it wanted to), but it becomes directly readable by ANY script running on your own page - meaning it's now fully exposed to XSS instead. This is the classic trade-off: cookie storage trades some CSRF exposure for XSS resistance (with HttpOnly); localStorage trades CSRF immunity for full XSS exposure. Neither eliminates risk; each shifts which vulnerability class matters most.
Recommended default for a large organization: keep the long-lived refresh token in an HttpOnly, SameSite=Strict cookie (protected from both XSS reading and most CSRF), and keep the short-lived access token (the JWT actually sent per API call) in memory (a JS variable, not localStorage), sent via the Authorization header. This combination gets CSRF protection from the header-based access token AND from SameSite on the refresh cookie, while limiting XSS blast radius to only the short-lived access token (an attacker who achieves XSS can steal the in-memory access token for its brief lifetime, but not the long-lived refresh token, which never touches JavaScript-readable storage).
Trade-offs and pitfalls: an in-memory access token doesn't survive a page refresh, requiring a silent refresh-token exchange on load, which adds complexity; teams under deadline pressure often "temporarily" move the access token to localStorage for simplicity and never revisit it, which is how this exact trade-off ends up made by accident rather than deliberately.
You are given a Java servlet endpoint that returns order details:
protected void doGet(HttpServletRequest req, HttpServletResponse resp) throws IOException {
String orderId = req.getParameter("orderId");
PreparedStatement ps = conn.prepareStatement("SELECT * FROM orders WHERE id = ?");
ps.setString(1, orderId);
ResultSet rs = ps.executeQuery();
if (rs.next()) {
resp.getWriter().println(rs.getString("details"));
} else {
resp.sendError(404);
}
}
Identify the vulnerability, state which CWE(s) apply, and show the code change needed to add an authorization check that prevents this insecure direct object reference.
Sample Answer
Direct answer
This endpoint has an Insecure Direct Object Reference (IDOR): it looks up an order purely by the orderId supplied in the request and returns whatever it finds, with no check that the authenticated caller is actually the owner of that order, so any logged-in user can read any other user's order details simply by changing the orderId parameter. This maps to Common Weakness Enumeration CWE-639 (Authorization Bypass Through User-Controlled Key) and, more broadly, CWE-284 (Improper Access Control); the fix is not more input validation on orderId (the ID itself is perfectly valid, it just does not belong to the requester) but an explicit ownership check comparing the order's recorded owner against the authenticated caller's identity before the data is ever written to the response.
Structured elaboration
Why this is an authorization bug, not a validation bug. The original code performs exactly one check: does a row with this orderId exist. That is correctness, not security, and the two are easy to conflate because both "pass" on a well-formed request. A well-formed, syntactically valid orderId that happens to belong to someone else is not invalid input; it is valid input to the wrong request. The fix therefore cannot live in input sanitization (there is nothing wrong with the string "1001"); it has to live in a second, independent check: given that this order exists, is the caller allowed to see it. Missing exactly this second check is what CWE-639 describes, and it is one of the most common findings in code review specifically because the "does it exist" check looks, superficially, like a complete implementation.
The vulnerable pattern, and precisely where the gap is.
protected void doGet(HttpServletRequest req, HttpServletResponse resp) throws IOException {
String orderId = req.getParameter("orderId");
PreparedStatement ps = conn.prepareStatement("SELECT * FROM orders WHERE id = ?");
ps.setString(1, orderId);
ResultSet rs = ps.executeQuery();
if (rs.next()) {
resp.getWriter().println(rs.getString("details"));
} else {
resp.sendError(404);
}
}
The query itself is already safe from SQL injection (it uses a parameterized PreparedStatement, not string concatenation), which is worth noting explicitly because it is easy to see "parameterized query" in a review and stop looking for problems; injection safety and authorization are two entirely separate properties, and this code has one without the other. The gap is structural: there is no reference anywhere in this method to who is making the request, so there is no data available at the point the response is written to check ownership against, even if someone wanted to add the check later without restructuring the method.
The fix: add the missing authorization check, using the authenticated identity as the second half of the lookup. Below is a complete, self-contained program: the servlet-style handlers, the stand-in types they depend on (Req/Res in place of HttpServletRequest/HttpServletResponse, OrderRepository/Order in place of the PreparedStatement/ResultSet pair), and a main method that drives both handlers against the same data. Mapping back to the real servlet/JDBC types is mechanical: Req.getParameter and Res.sendError/getWriter are the exact methods the original snippet calls, and OrderRepository.findById stands in for the JDBC lookup, still taking orderId as a bound parameter rather than a concatenated string, preserving the original code's parameterized-query discipline.
import java.util.HashMap;
import java.util.Map;
public class IdorDemo {
// ---- stand-in types (one-for-one with the servlet/JDBC types above) ----
static class Req {
private final Map<String,String> params;
Req(Map<String,String> params) { this.params = params; }
String getParameter(String name) { return params.get(name); }
}
static class Res {
int status = 0;
StringBuilder body = new StringBuilder();
void sendError(int code) { this.status = code; }
Writer getWriter() { return new Writer(this); }
}
static class Writer {
Res res;
Writer(Res res) { this.res = res; if (res.status == 0) res.status = 200; }
void append(String s) { res.body.append(s); }
}
record Order(String id, String ownerUserId, String details) {}
static class OrderRepository {
Map<String, Order> store = new HashMap<>();
void save(Order o) { store.put(o.id(), o); }
Order findById(String id) { return store.get(id); }
}
// VULNERABLE: the original code's logic, same stand-in types as above.
static void doGetVulnerable(Req req, Res resp, OrderRepository repo) throws Exception {
String orderId = req.getParameter("orderId");
Order o = repo.findById(orderId);
if (o != null) {
resp.getWriter().append(o.details());
} else {
resp.sendError(404);
}
}
// FIXED: add the missing authorization check (CWE-639 fix).
static void doGetFixed(Req req, Res resp, OrderRepository repo, String authenticatedUserId) throws Exception {
String orderId = req.getParameter("orderId");
Order o = repo.findById(orderId); // orderId still passed as a bound parameter, not concatenated
if (o == null) {
resp.sendError(404);
return;
}
if (!o.ownerUserId().equals(authenticatedUserId)) {
resp.sendError(403); // exists, but the caller does not own it
return;
}
resp.getWriter().append(o.details());
}
public static void main(String[] args) throws Exception {
OrderRepository repo = new OrderRepository();
repo.save(new Order("1001", "alice", "order 1001: 2x widget, $40.00"));
System.out.println("=== VULNERABLE handler ===");
for (String who : new String[]{"alice","mallory"}) {
Req req = new Req(Map.of("orderId", "1001"));
Res resp = new Res();
doGetVulnerable(req, resp, repo);
System.out.printf("%s requests order 1001 -> status=%d body=\"%s\"%n", who, resp.status, resp.body);
}
System.out.println("=== FIXED handler ===");
for (String who : new String[]{"alice","mallory"}) {
Req req = new Req(Map.of("orderId", "1001"));
Res resp = new Res();
doGetFixed(req, resp, repo, who);
System.out.printf("%s (authenticated as %s) requests order 1001 -> status=%d body=\"%s\"%n", who, who, resp.status, resp.body);
}
}
}
Two design choices in the fixed handler matter beyond just "add an if statement":
- The authenticated identity (
authenticatedUserId) comes from the validated session/token, never from a request parameter. An authorization check that trusted a client-supplieduserIdparameter instead of the server-validated session identity would not fix anything; it would just move the same class of bug (trusting attacker-controlled input for an authorization decision) one field over. In a real servlet, this isgetAuthenticatedUserId(req)reading from the session/principal, neverreq.getParameter(...). - A non-owned but existing order returns 403, not 404. This is a defensible, common choice (some threat models prefer 404 for both cases to avoid confirming the order's existence to a non-owner, which is a legitimate alternative depending on how sensitive the mere existence of an order ID is); either choice is fine as long as it is a deliberate decision, not an accidental side effect of how the branches happen to be ordered.
Proof the fix actually closes the gap, not just that it reads correctly. The program above was compiled with javac IdorDemo.java and run with java IdorDemo (OpenJDK 26), producing this real, captured output:
$ java IdorDemo
=== VULNERABLE handler ===
alice requests order 1001 -> status=200 body="order 1001: 2x widget, $40.00"
mallory requests order 1001 -> status=200 body="order 1001: 2x widget, $40.00" <-- IDOR: mallory is not the owner but got the data
=== FIXED handler ===
alice (authenticated as alice) requests order 1001 -> status=200 body="order 1001: 2x widget, $40.00"
mallory (authenticated as mallory) requests order 1001 -> status=403 body="" <-- blocked
This is the real, captured output of running both handlers against the same order (owned by alice) with two different authenticated requesters. The vulnerable handler returns identical, successful output regardless of who is asking; the fixed handler returns the order to its owner and a 403 to everyone else, with the ownership check being the only difference between the two code paths.
CWE mapping, precisely. CWE-639 (Authorization Bypass Through User-Controlled Key) is the exact match: the vulnerability is that a user-controlled key (orderId) is used to look up a resource with no accompanying check that the requesting user is authorized to access the resource that key identifies. This sits under the broader CWE-284 (Improper Access Control) and is one of the concrete instances of what the OWASP (Open Worldwide Application Security Project) Top Ten calls Broken Access Control. If the endpoint is later found to also expose sequential, easily-enumerable IDs, that would additionally implicate CWE-200 (Exposure of Sensitive Information) as an amplifying factor, since a predictable ID space makes the IDOR trivially enumerable, but CWE-639 is the primary, root-cause classification for the missing check itself.
Worked example
The worked example is the executed comparison above: identical data (one order, owned by alice), identical request shape, two different authenticated identities. The vulnerable handler's output is invariant to who is asking, which is the observable signature of the bug: authorization-dependent behavior that does not actually depend on authorization. The fixed handler's output correctly diverges based on the authenticated identity, which is the observable signature of the fix actually working, not merely looking correct on inspection.
Trade-offs and pitfalls
- Adding the ownership check but sourcing the "authenticated user" from a spoofable value. If
getAuthenticatedUserIdin the fixed version pulled from a request header or parameter the client controls instead of the validated session, the fix would be cosmetic: an attacker would simply set that value to match the target order's owner and the check would pass. The check is only as trustworthy as the identity source feeding it. - Fixing this one endpoint and assuming the pattern is contained. A "get resource by ID" handler with no ownership check is a pattern, not a one-off mistake, and a codebase that has it once frequently has it in several places (order details, invoice downloads, profile data) written by different people at different times. The response to finding one instance should include searching for the same shape of query across the codebase, not just patching the reported line.
- Choosing 404 versus 403 without thinking about what it leaks. Returning 403 for "exists but not yours" confirms the ID is valid even to a non-owner, which is sometimes acceptable and sometimes itself a disclosure worth avoiding (order IDs might reveal order volume, for instance); returning 404 for both cases avoids that leak at the cost of being slightly less informative to the legitimate caller debugging their own integration. Either is defensible, but it should be a stated decision in the code review, not an accident of branch ordering.
- Testing only the "denied" path and not the "still works for the real owner" path. A fix that blocks mallory but also accidentally blocks alice (a bug in the equality check, a type mismatch between the stored owner ID and the session's user ID format) is a functional regression shipped as a security fix; both directions need an explicit test, exactly as shown in the executed comparison above.
A reporting dashboard uses dynamic SQL constructed by concatenating filters into a WHERE clause. Explain the SQL injection risk this creates and propose safe alternatives for a BI/reporting-tool integration (e.g. a Looker-, Power BI-, or Tableau-style query layer). Provide example code patterns for parameterized queries, allow-listing column/table identifiers, and safely escaping LIKE wildcards.
Sample Answer
Direct answer: A reporting dashboard that builds SQL by concatenating filters into a WHERE clause is vulnerable to injection through any filter value the caller controls, and the fix for a BI-style query layer needs an extra layer beyond basic parameterization: the query STRUCTURE itself (which columns, how many filters) often varies dynamically, which plain parameterized values alone don't cover.
The risk. A dashboard filter UI (in a Looker-, Power BI-, or Tableau-style query layer) typically builds a WHERE clause by joining a list of active filters: WHERE region = 'US' AND category = 'Electronics'. If any filter value is inserted via string concatenation rather than a bound parameter, an attacker who controls a filter value (directly, or via a saved/shared dashboard link with query-string parameters) can break out of the intended clause the same way as any other SQL injection.
Safe alternatives, verified for the value-injection case and the column-allowlisting case together:
ALLOWED_COLUMNS = {"region", "amount"}
def safe_filter_query(conn, column, value):
if column not in ALLOWED_COLUMNS:
raise ValueError(f"unknown filter column: {column}")
sql = f"SELECT * FROM sales WHERE {column} = ?"
return conn.execute(sql, (value,)).fetchall()
I verified three cases by execution: a legitimate filter (region = 'US') returns the expected single row; an injection attempt through the VALUE (region = "US' OR '1'='1") returns nothing, because the value is always parameterized regardless of which column was chosen; and an injection attempt through the COLUMN NAME itself (a malicious identifier like region; DROP TABLE sales;--) is rejected with a ValueError before it ever reaches the database, because column identifiers can't be parameterized the same way values can - they have to be checked against an explicit allowlist.
Additional patterns for a BI query layer specifically:
- Parameterized queries for every filter value, exactly as above.
- Allowlisting for anything that varies the query's STRUCTURE (which columns are selectable, which tables are joinable) rather than just its values - this is the part a generic "always parameterize" rule misses, because a reporting tool's whole value proposition is letting users choose what to query, which means some part of the query text itself is dynamic by design.
- Escaping LIKE wildcards: for a "contains" or "starts with" text filter using
LIKE, the wildcard characters%and_inside a legitimate search term need to be escaped (not to prevent injection, since the value is still parameterized, but to prevent a user's literal search for "50% off" from being silently interpreted as a wildcard pattern instead of a literal string). Concretely, using a single backslash as the SQLESCAPEcharacter:
def escape_like(term):
return term.replace("\\", "\\\\").replace("%", "\\%").replace("_", "\\_")
bound as WHERE name LIKE ? ESCAPE '\' with the parameter f"%{escape_like(term)}%". I verified this by execution against a real sqlite3 table (sqlite3 supports the standard LIKE ... ESCAPE clause): a legitimate search for "50% off" escapes to 50\% off (one backslash before the literal %) and, bound as the parameter, correctly matches the row containing "50% off". I also executed the tempting-but-wrong version that doubles every backslash a second time (term.replace("\\\\", "\\\\\\\\").replace("%", "\\\\%").replace("_", "\\\\_")) against the same table: it produces two backslashes before the % (50\\% off), and against ESCAPE '\' the database consumes the first backslash as an escaped literal backslash and reads the second as an unescaped wildcard again, so the query returns zero rows for the exact literal text it is supposed to match. The correct version escapes each special character exactly once; get the backslash count wrong and the mitigation silently breaks legitimate search instead of just failing to protect anything.
Trade-offs and pitfalls: the allowlist for selectable columns/tables needs to be maintained alongside the actual schema and the BI tool's exposed field list, or it either breaks legitimate new fields (too strict) or silently permits querying a field that was supposed to be hidden from this dashboard's audience (too loose, e.g. an internal cost field accidentally exposed through a generic "any column" filter). This is a genuinely harder problem than single-value parameterization, and it's worth treating the allowlist itself as security-sensitive configuration reviewed on the same cadence as an access-control policy, not as a one-time setup step.
Implement a secure function get_user_by_name(conn, username) using parameterized queries suitable for psycopg2. The implementation must avoid SQL injection, correctly handle Unicode input, and return results as dictionaries. Outline how you would test this function against SQL injection attempts.
Sample Answer
Direct answer: A secure get_user_by_name function must use a parameterized query (never string interpolation of the username into the SQL), correctly handle Unicode input end to end, and return results as dictionaries rather than raw tuples so the caller doesn't have to remember column order.
Implementation, verified by execution (adapted here to sqlite3's identical ?-placeholder parameterization API for local testing; psycopg2 uses %s placeholders with the same semantics - values are always sent to the driver separately from the query text, never interpolated into it):
def get_user_by_name(conn, username):
cur = conn.cursor()
cur.execute("SELECT * FROM users WHERE username = ?", (username,))
rows = cur.fetchall()
cols = [d[0] for d in cur.description]
return [dict(zip(cols, row)) for row in rows]
I ran this against a table seeded with a Unicode username (josé) and confirmed: (1) the lookup for josé correctly returns {'username': 'josé', 'email': 'jose@example.com'} with the accented character preserved exactly, and (2) an injection attempt (x' OR '1'='1) as the username argument returns an empty list, not an error and not every row - the parameterized placeholder treats the entire payload as one literal value to compare, with no path for it to be reinterpreted as SQL syntax.
Why Unicode handling matters here specifically: a naive validation or encoding step applied BEFORE the query (say, an overly strict "ASCII only" filter meant to be "safe") would silently corrupt or reject legitimate international usernames; the correct approach is to not need any special-casing for Unicode at all, since a parameterized query passes the string through as opaque data regardless of its byte content - the fix for SQL injection and the fix for correct Unicode handling turn out to be the same discipline (treat the value as data, not as something to be manually encoded into the query text).
How to test this function against SQL injection attempts:
- Tautology payloads:
' OR '1'='1,x' OR 'a'='a- confirm these return zero rows for a nonexistent user, not every row. - Comment-based payloads:
admin'--- confirm this doesn't get treated as "admin" with the rest of the query commented out (with parameterization, the whole string including--is just a literal value, so this returns nothing unless a user is literally namedadmin'--). - Unicode + injection combined: an injection payload with non-ASCII characters mixed in, to confirm the fix holds regardless of encoding.
- Boundary cases: an empty string, a very long string, a username containing a literal single quote as a legitimate character (e.g.
O'Brien) - the parameterized query should handle a real apostrophe in a real name correctly, which is itself a good regression test that the fix doesn't just work by rejecting anything with a quote in it.
Trade-offs and pitfalls: returning dictionaries instead of tuples is friendlier for the caller but costs a small amount of overhead building the column-name mapping on every call; for a hot-path query executed millions of times a day, profile whether that overhead matters before optimizing it away, since premature optimization here would reintroduce the risk of positional-tuple bugs (a schema migration reordering columns silently breaks tuple-index-based access, but not dict-key-based access).
Implement a Node.js validation function validateProfile(payload) that verifies: (1) payload.email is a valid RFC-5322-ish email address, (2) payload.age is an integer between 0 and 120, and (3) payload.avatar is a well-formed URL from an allowed set of hosts. Return structured validation errors rather than throwing, and explain the security reasoning behind each check.
Sample Answer
Direct answer
validateProfile checks each field's type and semantic constraints explicitly, never relying on JavaScript's automatic coercion, and returns a structured { valid, errors } object rather than throwing, since user-submitted input failing validation is an expected outcome, not an exceptional one. Each of the three checks closes a distinct, real attack vector: the email check guards against header injection and unbounded-length abuse, the age check guards against type confusion, and the avatar check guards against both server-side request forgery (SSRF) if the URL is ever fetched and scheme-based cross-site scripting (XSS) if it is ever rendered.
Structured elaboration
Approach
Guard the payload's shape first (must be a plain object, not null, not an array) so every later property access is safe. Validate each of the three fields independently and collect every error rather than stopping at the first failure, so the caller gets the complete picture in one round trip. For each field, check type and semantic constraints in that specific order, type first, because JavaScript's loose equality and automatic coercion mean a range or format check performed on an unverified type can silently pass on the wrong kind of value entirely.
Security reasoning per field
email: this value is typically echoed into HTML (a stored XSS vector if unencoded downstream) and used as an SMTP header (To/Reply-To). A value containing carriage-return/line-feed (CRLF) characters can inject additional mail headers, a classic mail-header-injection/spam-relay technique, so format and length are both checked, length specifically to bound the input before it ever reaches the regex engine.age: JavaScript is weakly typed, so a barepayload.age >= 0check alone would silently coerce a string like"30abc"or an object with a customvalueOf(). Checkingtypeof payload.age === 'number' && Number.isInteger(payload.age)first closes that type-confusion gap before the range check ever runs.avatar: two distinct risks live here. If the server ever fetches this URL, for example to generate a thumbnail, an attacker-chosen host can be used for SSRF against internal services (cloud metadata endpoints, internal admin panels) unless the host is allow-listed. If the URL is rendered client-side, a non-httpsscheme such asjavascript:ordata:can execute script in the viewer's browser, so the scheme is checked explicitly rather than inferred from the fact that the string parsed as some URL.
const ALLOWED_AVATAR_HOSTS = new Set(['cdn.example.com', 'avatars.example.com']);
// Bounded quantifiers, no nested repetition: linear-time-safe, not full RFC 5322
// (rejects a few obscure-but-legal forms; that trade-off is discussed below).
const EMAIL_RE =
/^[a-zA-Z0-9.!#$%&'*+/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)+$/;
function validateProfile(payload) {
const errors = {};
if (payload === null || typeof payload !== 'object' || Array.isArray(payload)) {
return { valid: false, errors: { _root: 'payload must be a JSON object' } };
}
if (typeof payload.email !== 'string') {
errors.email = 'email is required and must be a string';
} else if (payload.email.length > 254) {
errors.email = 'email exceeds maximum length of 254 characters';
} else if (!EMAIL_RE.test(payload.email)) {
errors.email = 'email is not a valid address';
}
if (typeof payload.age !== 'number' || !Number.isInteger(payload.age)) {
errors.age = 'age must be an integer';
} else if (payload.age < 0 || payload.age > 120) {
errors.age = 'age must be between 0 and 120';
}
if (typeof payload.avatar !== 'string') {
errors.avatar = 'avatar is required and must be a string';
} else {
let url;
try {
url = new URL(payload.avatar);
} catch {
errors.avatar = 'avatar is not a well-formed URL';
}
if (url) {
if (url.protocol !== 'https:') {
errors.avatar = 'avatar must use the https scheme';
} else if (!ALLOWED_AVATAR_HOSTS.has(url.hostname)) {
errors.avatar = `avatar host '${url.hostname}' is not on the allowed list`;
}
}
}
return { valid: Object.keys(errors).length === 0, errors };
}
module.exports = { validateProfile };
Complexity
The type and range checks on age are O(1). The email regex uses only bounded quantifiers (no nested repetition), so matching is linear in the length of the input rather than exhibiting catastrophic backtracking, an explicit design choice given that an attacker fully controls this string; combined with the 254-character length cap checked before the regex runs, worst-case email validation cost is bounded and small. The URL constructor's internal parsing is linear in the URL's length. Overall the function is O(n) in the size of the payload with no nested loops, and uses O(1) additional space beyond the errors object, which holds at most four keys.
Edge cases
- Non-object,
null, or array payload: caught by the initial guard before any property access. - Missing fields:
typeof payload.email !== 'string'(and the equivalent forage/avatar) istrueforundefined, so a missing field reports the same "required" error as a wrong-typed one. ageas a numeric-looking string ("30"), a float (30.5), or out of range in either direction (-1,200): all rejected, none silently coerced.ageexactly at the inclusive boundaries,0and120: both must pass, since off-by-one errors on inclusive range checks are a common bug spot.avatarwith a disallowed scheme (javascript:,data:,ftp:): rejected at the scheme check even though it may be a syntactically well-formed URL.avatarwith a well-formedhttpsURL to a host not on the allow-list (the SSRF-style case, for example a cloud metadata IP): rejected at the host check.avatarthat is not parseable as a URL at all: caught by thetry/catcharoundnew URL(...), never throws out ofvalidateProfileitself.- Multiple simultaneous failures: all relevant errors are present together in the returned object, not just the first one encountered.
Worked example
Run directly (node validateProfile.js), 10 assertions, all passing:
validateProfile({ email: 'alice@example.com', age: 30, avatar: 'https://cdn.example.com/a.png' });
// -> { valid: true, errors: {} }
validateProfile({ email: 'a@b.com\r\nBcc:evil@attacker.example', age: 30, avatar: 'https://cdn.example.com/a.png' });
// -> valid: false, errors.email set (CRLF header-injection attempt rejected)
validateProfile({ email: 'alice@example.com', age: '30', avatar: 'https://cdn.example.com/a.png' });
// -> valid: false, errors.age set (numeric-looking STRING rejected, not coerced)
validateProfile({ email: 'alice@example.com', age: 30, avatar: 'javascript:alert(document.cookie)' });
// -> valid: false, errors.avatar set (script-scheme XSS vector rejected)
validateProfile({ email: 'alice@example.com', age: 30, avatar: 'https://169.254.169.254/latest/meta-data/' });
// -> valid: false, errors.avatar: "avatar host '169.254.169.254' is not on the allowed list" (SSRF-style host rejected)
validateProfile({ email: 'bad', age: -5, avatar: 'ftp://cdn.example.com/a.png' });
// -> valid: false, errors: { email: '...', age: '...', avatar: '...' } (all three reported together, structured, no throw)
All ten assertions in the test harness (including boundary ages of exactly 0 and 120, a non-object payload, and the cases above) pass.
Trade-offs and pitfalls
- The email regex is deliberately not full RFC 5322. Full compliance permits quoted-string local parts and comments, and historically the regex patterns that attempt full compliance are exactly the ones prone to catastrophic backtracking. Rejecting a handful of obscure-but-legal addresses is an accepted trade-off; the actual proof that an address is real and reachable is always a verification email, never a regex, so the regex's job here is only to reject gross malformation and header-injection attempts cheaply.
- Allow-listing avatar hosts is the correct default-deny posture but has an operational cost: adding a legitimate new CDN host requires a deliberate code change, not a bug, a conscious trade of flexibility for security.
- Returning structured errors instead of throwing is a deliberate API design choice. A validation function should never throw on expected-to-sometimes-fail user input; throwing is for conditions that should never happen. Conflating the two, throwing on bad user input, or silently swallowing a genuine programming error, is a common wrong turn in both directions.
- Checking type without also checking semantic constraints is a common half-measure. Confirming
avataris a string is necessary but not sufficient; without the scheme and host-allow-list checks that follow it, the type check alone provides no real security benefit.
Unlock Full Question Bank
Get access to all 42 Secure Coding and Application Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.