Secure Coding and Application Security Questions
Writing and reviewing code that resists attack. Covers the OWASP Top Ten and common web vulnerabilities (XSS, SQL injection, CSRF), input validation, secure coding practices and security code review, static application security testing (SAST), API and HTTP security, database and frontend security, and mobile app security. The application-layer defense discipline for engineers building software.
Compare JWT-based stateless authentication with stateful session stores in large microservice environments. Discuss the token-revocation challenges stateless tokens introduce, how each approach handles horizontal scaling, and when you would choose one over the other.
Sample Answer
Direct answer
A JSON Web Token (JWT) is a signed, self-contained credential: the server that validates it does not need to look anything up, because the token itself carries the claims and a signature proving they have not been tampered with. A stateful session store keeps the actual session data server-side and hands the client only an opaque reference (a session ID); every request requires a lookup against that store. JWTs scale horizontally with almost no coordination between service instances, at the direct cost of revocation being hard; stateful sessions give you instant, precise revocation, at the cost of needing a shared, available data store behind every service instance. In a large microservice environment, the choice is really about where you want the operational complexity to live: in token lifecycle management (short-lived tokens plus refresh), or in a shared session infrastructure that must be fast and available on every request path.
Structured elaboration
How each approach actually works
| Property | Stateless JWT | Stateful session store |
|---|---|---|
| Where session data lives | Encoded in the token itself, signed (not encrypted by default) | Server-side (Redis, a database, an in-memory store), client holds only an opaque ID |
| What a service needs to validate a request | The signing key (or public key for asymmetric signing) and the token itself, no external call | A round trip to the shared session store |
| Revocation | Hard: the token is valid until it expires, regardless of server-side state, unless you add an external revocation mechanism | Easy: delete the row/key in the store and the very next request fails |
| Horizontal scaling | Trivial: any service instance can validate any token independently, no shared state required | Requires a shared store reachable from every instance, or sticky sessions pinning a client to one instance (which itself complicates load balancing and failover) |
| Payload size / overhead | Token travels with every request, grows with the number of claims embedded | Session ID is small and fixed size; the actual data never leaves the server |
| Failure mode if the store is unavailable | Not applicable, validation needs no store | Every session-dependent request fails until the store recovers |
The token-revocation challenge, in detail
A signed JWT is valid proof of its claims for as long as its signature checks out and it has not expired, full stop; there is no built-in way to make the server "change its mind" about a token it already issued. This creates real problems: a user who changes their password or is deactivated for policy violations should lose access immediately, not whenever their existing token happens to expire; a stolen token is usable for its full remaining lifetime by whoever stole it. The practical mitigations all trade away some of the statelessness that made JWTs attractive in the first place:
- Short-lived access tokens. Keep the JWT's own lifetime very short (minutes), which bounds the damage window without needing a lookup on every request.
- Opaque, server-tracked refresh tokens. Pair the short-lived JWT with a longer-lived refresh token that is itself stateful (tracked and revocable server-side). This is the most common production pattern: the frequently-used access token stays stateless and fast, while the rarely-used refresh step is where you get a real revocation point.
- A revocation list or introspection endpoint for high-privilege tokens. For actions that matter enough to justify the cost, check the token's identifier against a fast, shared deny-list (or call an OAuth2 token-introspection endpoint) before honoring it, which reintroduces a lookup but only where the risk justifies it.
None of these make JWTs as instantly revocable as a stateful session; they narrow the exposure window instead.
Horizontal scaling, in detail
Stateless JWT validation needs nothing but the signing key, so adding a new service instance behind a load balancer requires zero coordination: any instance can validate any request from any user without talking to any other instance or a shared store. Stateful sessions need every instance to see the same session data, which means either a shared, low-latency, highly available session store (adding a new dependency and a new potential bottleneck and single point of failure to the request path) or sticky sessions (routing a given user consistently to the same backend instance, which complicates rolling deploys, instance failure handling, and even load distribution). In a microservice environment specifically, this difference is amplified: if every one of a dozen internal services needs to validate the caller's identity, a shared session store becomes a dependency of all twelve, while stateless JWTs let each service validate independently.
When to choose one over the other
- Choose stateless JWTs when the system is large, has many independently scaling services, session data does not need instant revocation for most flows, and you can tolerate (and are willing to build for) the short-lived-token-plus-refresh pattern. This is the common default for API-first and microservice architectures precisely because it avoids a shared-store dependency on the hot path.
- Choose stateful sessions when instant revocation genuinely matters for the primary flow (a banking application's active session, an admin console), the system does not need to scale to many independent services validating the same identity, or you already operate a fast, reliable shared store (many teams already run Redis for other reasons, which lowers the marginal cost of using it for sessions too).
- Choose the hybrid (the realistic answer for most large systems): short-lived stateless JWTs for routine, high-volume request authorization, backed by a stateful, revocable refresh token and a lightweight revocation check reserved for genuinely high-privilege actions. This gets most of the scaling benefit of statelessness while keeping a real revocation point where it matters.
Worked example
A microservice platform with an API gateway and fifteen backend services, evaluating both approaches for how a request's identity gets validated at each service:
Pure stateful sessions: every one of the fifteen services calls a shared Redis-backed session store on every request to resolve the session ID into user identity and claims. Adding a sixteenth service means it too now depends on that Redis cluster being available and fast; a Redis outage takes down authorization across the entire platform, not just one service.
Pure stateless JWTs: the API gateway validates the JWT's signature once and forwards a small set of claims downstream (or forwards the token itself for each service to validate independently). No service depends on a shared store to authorize a request. The cost shows up if a user is deactivated mid-session: their existing token remains valid, potentially for the token's full lifetime, at every one of the fifteen services, unless a revocation check is added somewhere.
Hybrid (the actual recommendation here): access tokens are JWTs with a 10-minute lifetime, validated statelessly by each service, no shared store on the hot path. A refresh token, stored server-side and revocable, is used to mint a new access token every ten minutes; deactivating a user immediately blocks their next refresh, and any already-issued access token expires naturally within ten minutes regardless. For the one or two genuinely high-privilege actions on the platform (issuing a refund, changing an account's payment method), an additional fast revocation-list check is added specifically to those endpoints, accepting a small latency cost only where the risk justifies it.
Trade-offs and pitfalls
- Treating "stateless" as strictly better because it "scales." Statelessness solves a specific problem (shared-store dependency on every request) at a specific cost (revocation difficulty); a system that genuinely needs instant revocation on its primary flow has traded away something it actually needed for a scaling property it may not have been bottlenecked on in the first place.
- Long-lived JWTs "for convenience." Extending a JWT's lifetime to reduce how often the client has to refresh directly extends the exposure window of a stolen token and the delay before a deactivated user actually loses access; this is the single most common way JWT-based systems end up with a worse security posture than the stateful sessions they replaced.
- Forgetting that a refresh token is itself a stateful session, just a slower-moving one. Teams sometimes believe they have "gone stateless" while still operating a server-side refresh-token store; that is fine and often the right design, but it means the shared-store dependency and the revocation capability did not disappear, they just moved to a lower-traffic part of the flow.
- Embedding too much in the JWT payload. Every claim added to a JWT is sent, unencrypted (base64-encoded, not encrypted, unless you additionally use an encrypted JWT format), with every single request; embedding a full user profile or a large permissions list bloats every request and, since it is only signed and not encrypted by default, exposes that data to anything that can see the token, including client-side JavaScript and browser storage.
- Assuming sticky sessions "solve" stateful scaling. Sticky sessions avoid the shared-store dependency but introduce their own failure modes: an instance restart or rolling deploy drops that instance's sessions, and load balancing becomes uneven if some users generate much more traffic than others.
A JSON-based API echoes a user-supplied parameter into a JavaScript client response inside a single-page application. Describe how you would test for DOM-based XSS: what instrumentation you would use (browser devtools, a Burp extension, automated DOM-sink scanning), a safe proof-of-concept payload skeleton, and how you would triage and classify the finding.
Sample Answer
Direct answer: Testing for DOM-based XSS means tracing the path from an attacker-controllable "source" (a URL parameter, location.hash, document.referrer, a postMessage payload) through the application's JavaScript to a dangerous "sink" (innerHTML, document.write, eval, setAttribute('onclick', ...)) and confirming the browser actually executes injected script when that path is exercised, since the server-side response here is only the messenger, not the vulnerable code.
Instrumentation and approach:
- Browser devtools first: open the page, search the rendered DOM and the page's JavaScript sources for the API in the response (here, the SPA reads a JSON API response and writes a field into the page). Set a DOM breakpoint ("break on subtree modification") on the container element, or use the Sources panel to step through the render function and see exactly which sink receives the value.
- Automated DOM-sink scanning: tools like a Burp Suite extension (e.g. DOM Invader) or a headless-browser crawler that hooks
innerHTML/document.write/evalat the JS engine level and reports every reachable source-to-sink flow, which is far faster than manually reading minified bundles. - Manual confirmation with a benign marker payload: before trying anything that executes, submit a distinctive string like
XSSPROBE12345as the parameter and confirm via devtools exactly where and how it lands in the DOM (as text, as an attribute, inside a script block). This tells you which encoding context you're up against before you craft a working payload.
Worked example / safe PoC skeleton. If the API response field is written via element.innerHTML = data.username, a proof-of-concept payload is <img src=x onerror=alert(document.domain)> submitted as the username - innerHTML parses this as real markup, the broken src triggers onerror, and alert(document.domain) fires, which is safe (no data exfiltration) but conclusively proves script execution. For a production engagement, swap alert(document.domain) for a non-destructive callback to a controlled logging endpoint you own, so you get a timestamped hit log rather than a visible popup that could alarm real users if the payload somehow reaches one.
Triage and classification once confirmed: record the exact source (which parameter/API field), the exact sink (which DOM write), and whether any encoding or sanitizer sits between them (many teams have a sanitizer library available but simply didn't call it on this particular field). Classify severity by what an attacker gains: session-cookie theft is high severity even without HttpOnly bypass tricks, since most SPAs also keep sensitive data in memory/localStorage that JavaScript CAN read directly.
Trade-offs and pitfalls: minified/bundled production JavaScript makes manual source-reading slow; work from source maps if available, or from the unminified dev build. A payload that works in Chrome devtools may be blocked by the site's CSP in production, so validate against the actual deployed headers, not just the sink's existence, before reporting severity. And a sink alone is not a finding: if the value is hard-coded or only ever set from a trusted internal source, there's no reachable "source," so always confirm the taint actually originates from something the attacker can control.
A microservice accepts arbitrary URLs to fetch thumbnails and was abused to perform SSRF calls to internal metadata endpoints. Propose a layered mitigation plan covering input validation and URL canonicalization, an allowlist of permitted destinations, egress filtering and network segmentation, proxying requests through a vetted fetch service, and runtime detection for anomalous outbound requests.
Sample Answer
Direct answer: For a microservice that was abused for SSRF via an arbitrary-URL thumbnail fetch, the fix needs to be layered: no single control here is sufficient alone, because each addresses a different way the previous layer could fail or be bypassed.
Structured elaboration, layer by layer:
- Input validation and URL canonicalization. Before anything else, parse the supplied URL properly (never with a regex) and reject anything that isn't
http/https, reject URLs with embedded credentials (http://user:pass@host), and canonicalize to catch equivalent-but-differently-encoded forms of the same dangerous target (decimal/octal/hex IP representations of127.0.0.1, IPv6 loopback forms,0.0.0.0). - An explicit allowlist of permitted destinations. Rather than trying to block every dangerous target, only fetch from domains the business actually needs (known image-hosting CDNs the feature is meant to support), resolved and IP-pinned at request time to defeat DNS rebinding - re-check the resolved IP right before connecting, not just at initial validation, since an attacker-controlled DNS record can change its answer between the two.
- Egress filtering and network segmentation. Even if the allowlist has a gap, network-level policy (a firewall rule, a cloud security group, or a service mesh egress policy) that blocks the thumbnail service from reaching internal-only IP ranges and the cloud metadata address provides an independent layer that doesn't depend on the application code being perfectly correct.
- Proxying requests through a vetted fetch service. Rather than letting every service that needs to fetch external URLs implement its own validation (and inevitably implement it inconsistently), centralize outbound URL-fetching behind one hardened internal service that owns the allowlisting, IP-pinning, redirect-following rules, and egress restrictions once, correctly, and every other service calls it instead of doing its own
requests.get(user_url). - Runtime detection. Even with the above, monitor for anomalous outbound request patterns from this service (requests to internal IP ranges, requests to the metadata address specifically, an unusual spike in failed/timed-out fetches that might indicate probing) as a safety net that catches both bypasses and future regressions.
Worked example. The vetted fetch service receives {url: "http://169.254.169.254/latest/meta-data/..."}, resolves the hostname, finds it's in the disallowed CIDR range for cloud metadata addresses, and rejects it before ever making the outbound connection - centralizing this logic means every consuming service (thumbnailing, link-preview generation, webhook delivery) gets this protection automatically rather than each team needing to remember to implement it correctly.
Trade-offs and pitfalls: a centralized fetch service becomes a new operational dependency (a single point of failure and a place where a misconfiguration affects every consumer at once), and it needs to correctly follow (or explicitly refuse) redirects - a service that validates only the initial URL but blindly follows a redirect to an internal address has done all this validation for nothing.
Explain the concept of memory safety and the common memory-related vulnerabilities: buffer overflow, use-after-free, integer overflow, and format-string bugs. For each, give a concise example of the coding pattern that causes it and how modern languages or tooling help prevent it.
Sample Answer
Direct answer
Memory safety means a program never reads or writes memory outside what it currently owns: it never touches memory after that memory has been freed, never writes past the end of a buffer it allocated, and never lets an arithmetic mistake turn into a wrong size or index used later. Four classic violations: buffer overflow (writing past a buffer's allocated bounds), use-after-free (using a pointer after the memory it refers to has been deallocated), integer overflow (an arithmetic result wraps or truncates and is then used as a size or index), and format-string bugs (attacker-controlled data used as the format specifier itself, not just as a value being formatted). Languages with bounds-checked collections, compile-time ownership tracking, or garbage collection eliminate whole classes of these by construction; C and C++ need coding discipline plus tooling (sanitizers, compiler warnings treated as errors) to catch what the language itself will not.
Structured elaboration
Buffer overflow (CWE-120 / CWE-787)
Writing (or reading) past the bounds of a fixed-size buffer.
char buf[8];
strcpy(buf, user_input); // no bound on user_input's length; overruns buf past 7 usable chars + NUL
Prevention: bounds-checked APIs that take an explicit size (snprintf rather than sprintf), and, more fundamentally, languages whose container and string types check every access at the point of use, Rust's Vec/slices panic instead of corrupting memory on an out-of-bounds index, and Python, Java, and Go strings and lists are bounds-checked by the runtime. Compiler and runtime mitigations (stack canaries, AddressSanitizer during testing) catch what slips through, but they are a safety net, not a substitute for bounds-safe code.
Use-after-free (CWE-416)
Continuing to use a pointer after the memory it points to has been freed; that memory may already have been reused for something else, so the program reads or writes attacker-influenceable data through a stale reference.
free(ptr);
/* ... other code runs, possibly reallocating this memory for something else ... */
printf("%s", ptr->name); // ptr's memory may no longer hold what this code expects
Prevention: garbage-collected languages (Python, Java, Go, JavaScript) eliminate this class entirely, memory is never freed while a reachable reference to it still exists. Rust's ownership and borrow checker rejects this at compile time, using a value after it has been moved or dropped is a compile error, not a runtime bug. In C++, smart pointers (unique_ptr, shared_ptr) free automatically and make an explicit free-then-use much harder to write by accident; in plain C, discipline (set pointers to NULL immediately after freeing) plus dynamic tooling (AddressSanitizer, Valgrind) catches what discipline alone misses.
Integer overflow (CWE-190)
An arithmetic operation produces a result outside the range its type can represent, so it wraps or truncates, which is dangerous specifically when that wrapped value is then used to size a buffer or bound a loop.
uint32_t count = attacker_controlled_value;
uint32_t size = count * sizeof(item_t); // wraps if the true product exceeds UINT32_MAX
char *buf = malloc(size); // allocates a much SMALLER buffer than intended
/* code that then writes `count` items into buf overflows it */
Prevention: languages with checked or arbitrary-precision arithmetic make the failure loud instead of silent, Python integers never silently wrap, and Rust panics on overflow in debug builds and offers explicit checked_add/checked_mul for code that must handle the boundary deliberately. In C/C++, an explicit range check before the multiplication, a safe-arithmetic helper, or a sanitizer/compiler flag that turns overflow into a detectable fault during testing (undefined-behavior sanitizer (UBSan), or -ftrapv) closes the gap the language leaves open.
Format-string bugs (CWE-134)
Attacker-controlled data passed as the format argument itself, rather than as a value being formatted, so format specifiers embedded in the attacker's input get interpreted by the runtime, letting them read arbitrary stack memory (%x) or, in the worst case, write to memory (%n).
printf(user_input); // vulnerable: user_input is interpreted AS the format string
printf("%s", user_input); // safe: user_input is a VALUE being formatted, not the format itself
Prevention: this is close to a C/C++-specific bug class, since it requires a printf-style variadic function that interprets its first argument as a format specifier. Modern languages' formatting facilities (Python f-strings and .format(), Java's String.format with typed arguments, Rust's format! macro) do not let a runtime string value be interpreted as a format specifier the way C's printf family does, so the bug class largely does not translate. In C/C++, the fix is a coding-standard rule enforced automatically: -Wformat-security flags a non-literal format-string argument, and treating that warning as a build error in continuous integration (CI) catches it before it ships.
Worked example
The integer-overflow case, compiled and run, because its arithmetic is exact and reproducible (unlike a buffer overflow or use-after-free, whose concrete effect depends on unspecified memory layout, unsigned-integer wraparound in C is well-defined behavior per the standard, so it can be demonstrated precisely without invoking undefined behavior):
uint32_t count = 268435457u; // attacker-controlled "how many items"
uint32_t itemsize = 16u; // sizeof(item_t)
uint64_t correct_size = (uint64_t)count * (uint64_t)itemsize; // computed in 64-bit: no wrap
uint32_t wrapped_size = count * itemsize; // computed in 32-bit: wraps
void *buf = malloc(wrapped_size);
Output:
count = 268435457
itemsize = 16
mathematically correct size (64-bit): 4294967312 bytes (4.295 GB)
actual 32-bit unsigned multiplication result: 16 bytes
malloc(wrapped_size) returned a buffer of 16 bytes (non-NULL: yes)
caller still believes 268435457 items fit; item #2 onward would write past this buffer
The true product of count * itemsize is about 4.295 GB, but computed as a 32-bit unsigned multiplication it wraps to exactly 16 bytes, so malloc hands back a 16-byte buffer while the surrounding code still believes 268,435,457 items fit in it. Writing the second item already runs past the buffer, this is exactly how CWE-190 (integer overflow) becomes CWE-122 (heap buffer overflow) in real vulnerabilities: the size computation, not the write itself, is where the bug actually lives.
Trade-offs and pitfalls
- These four are not equally likely across languages. Format-string bugs and manual buffer-size arithmetic are close to C/C++-exclusive; a team working entirely in a memory-safe, garbage-collected language should understand these classes conceptually (they show up in dependencies written in C, and in interview questions) without needing to defend against most of them in their own code.
- A sanitizer or bounds check caught in testing is a safety net, not proof of safety. These are dynamic checks; a code path a test suite does not exercise gets none of that protection at runtime.
- Signed integer overflow in C is undefined behavior, not a wraparound guarantee, unlike the unsigned case demonstrated above; do not assume a signed overflow will wrap predictably, a compiler is free to optimize on the assumption it never happens.
- The common wrong turn is treating memory safety as a "systems programming problem that does not apply to me." The size-computation mistake in the worked example is a plain arithmetic bug that could be written in any language; what differs by language is only whether the runtime catches it loudly (an exception) or lets it silently corrupt memory.
Design a safe sandbox to execute untrusted Python scripts uploaded by users (for example, user-provided data-preprocessing logic). Cover runtime isolation options, resource limits (CPU, memory, time), filesystem and network restrictions, and how you would prevent sandbox escapes.
Sample Answer
Direct answer
A safe sandbox for untrusted Python is a layered defense, not one mechanism: kernel/process-level isolation as the outer boundary, hard resource limits (CPU, memory, wall-clock time) enforced from outside the untrusted process rather than trusted to code running inside it, a locked-down filesystem, and default-deny network egress. Sandbox-escape prevention comes from minimizing what the isolation layer itself exposes, never from trying to make the Python language safe to run untrusted code in directly, because Python has no supported in-language sandboxing mechanism, and every attempt to build one by blocking dangerous builtins has a long history of being bypassed through object introspection.
Structured elaboration
Runtime isolation options
- Bare subprocess (weakest): a separate OS process, but it shares the host kernel entirely, so any unrestricted syscall or kernel vulnerability reaches the host directly. Only appropriate as one layer among several, never the whole solution.
- Container (namespaces, cgroups, seccomp): isolates the process's view of the filesystem, network, and process IDs (PIDs), but still shares the host kernel, so a vulnerability in the kernel's namespace or cgroup implementation is a full escape. Hardened further with a restrictive seccomp syscall allow-list and no added Linux capabilities.
- gVisor (a userspace reimplementation of the kernel interface intercepting syscalls): the untrusted process's syscalls are handled by a sandboxed userspace layer rather than the real host kernel directly, so most kernel vulnerabilities never reach the host at all. Real throughput cost from syscall-interception overhead in exchange for that extra isolation.
- Firecracker microVM (or similar hardware-virtualized VMs): the strongest practical isolation available; the untrusted code runs in its own minimal VM with its own kernel, so even a kernel exploit inside the sandbox stays inside that VM. Higher startup latency and resource overhead than a container, which matters for a system spinning up many short-lived executions per second.
Choose based on threat model and scale: a low-stakes internal tool might accept a hardened container; a public "run arbitrary code" feature (this is exactly the shape of a code-execution sandbox for interview or coding practice, genuinely arbitrary and potentially hostile input) should default to gVisor- or microVM-class isolation.
Resource limits: CPU, memory, time
- CPU: cap CPU-seconds consumed, via POSIX
RLIMIT_CPUor the container runtime's CPU quota/cgroup limit, so a busy loop or a cryptomining payload gets killed rather than starving co-located workloads. An in-process Python-level timer is not sufficient on its own, since the untrusted code could disable or out-race it; the limit must be enforced from outside the untrusted process. - Memory: cap via the container or VM's cgroup memory limit, enforced uniformly by the OS or hypervisor, not via an in-process mechanism the untrusted code itself runs inside of and could interfere with. This matters concretely: Python's own
resource.RLIMIT_ASis not reliably enforceable on every platform (confirmed below: it is rejected outright on macOS/Darwin), so an in-process rlimit call is a fragile, non-portable memory control; cgroup-based limits at the container/VM layer are the robust mechanism. - Wall-clock time: an outer supervisor (the orchestrator that spawned the sandboxed process) enforces a hard timeout independent of what is happening inside the sandbox, and kills and reaps the process, and anything it spawned, if the timeout is exceeded. This is the backstop even with CPU and memory limits already in place, since a process can be CPU-idle yet still need a hard ceiling on total run time.
Filesystem restrictions
A read-only root filesystem for everything except one designated scratch directory, mounted fresh and empty per execution and destroyed afterward, so nothing persists between runs and nothing outside that directory can be written to even if the untrusted code tries. No bind-mounts of host paths, and the orchestrator's own credentials, configuration, and source code live entirely outside the sandbox's mount namespace, not merely "not referenced" by convention.
Network restrictions
Default-deny egress: no outbound network access unless the specific use case genuinely requires it. A code-execution sandbox for practice/interview-style problems virtually never needs network access, so the correct default is none at all, not an allow-list to start from. If any egress is genuinely required, enforce an explicit allow-list of destinations at the network layer itself, not trusted to in-Python discipline about which hosts the code should or should not call, closing off the SSRF-style path of the sandboxed code reaching internal services or a cloud metadata endpoint.
Preventing sandbox escapes
Do not attempt to sandbox Python in Python: blocking import os, restricting __builtins__, or building an exec()-based "safe eval" is a defense with a long, consistent history of being bypassed through object introspection. This is precisely the object-graph-walk technique (__class__ / __mro__ / __subclasses__) that turns a server-side template-injection vulnerability into remote code execution: if an attacker can reach any Python object at all, they can typically walk that object graph to something dangerous. The correct trust boundary sits outside the language entirely, isolation at the process/kernel/VM level, so it does not matter what the Python code is able to do to itself, because it cannot reach anything outside the sandbox regardless. Inside that isolation layer, minimize the syscall surface further (a seccomp allow-list, dropped Linux capabilities) so even a full compromise of the sandboxed process has fewer paths to escalate through a kernel bug, and patch the isolation layer itself, the container runtime, gVisor, or the hypervisor, on a real cadence, since escapes are most often found in the isolation mechanism's own implementation, not in a cleverer attacker payload.
flowchart TD
A[Untrusted Python script] --> B[Process isolation: separate subprocess, no shared memory]
B --> C[Resource limits: CPU time via RLIMIT_CPU, wall-clock timeout]
C --> D[Filesystem: read-only rootfs, scratch tmpfs only, no host mounts]
D --> E[Network: default-deny egress, explicit allowlist only]
E --> F[Kernel-level containment: gVisor or a Firecracker microVM, seccomp syscall filter]
F --> G[cgroup memory and PID limits enforced by the container runtime]
Worked example
Two containment layers verified against a real runaway child process, plus a direct check of the memory-limit caveat above:
# Layer 0: RLIMIT_AS (memory) is not enforceable on this platform
resource.setrlimit(resource.RLIMIT_AS, (50 * 1024 * 1024, 50 * 1024 * 1024))
# -> ValueError: current limit exceeds maximum limit (rejected before the child even starts)
# Layer 1: wall-clock timeout kills an infinite loop
subprocess.run([PYTHON, "-c", "while True:\n pass\n"], timeout=1.0)
# -> subprocess.TimeoutExpired raised; process killed
# Layer 2: RLIMIT_CPU kills a CPU-bound computation via SIGXCPU
proc = subprocess.run(
[PYTHON, "-c", cpu_burn_code],
preexec_fn=lambda: resource.setrlimit(resource.RLIMIT_CPU, (1, 1)),
timeout=10,
)
# -> proc.returncode == -24 (SIGXCPU), confirmed via signal.SIGXCPU == 24
Output: RLIMIT_AS is rejected outright on this platform, exactly the fragility named above; the timeout kills the infinite loop within its pinned 1.0s budget; the CPU-bound computation is killed by SIGXCPU after its pinned 1 CPU-second limit, returncode -24. This is real evidence for the claim in the elaboration above: in-process rlimits are not a portable memory-containment mechanism, which is exactly why production sandboxes enforce memory via cgroups on the container/VM rather than a call the untrusted code's own process makes about itself.
Trade-offs and pitfalls
- Stronger isolation costs real startup latency and resource overhead. A system optimizing for many short executions per second needs to budget for this, not add it as an afterthought; production code-execution platforms commonly pre-warm a pool of containers or microVMs specifically to hide that latency from the end user.
- "It runs in a Docker container" is not sufficient isolation by itself. A container with no further hardening still shares the host kernel, and without separately restricted filesystem and network access, the sandboxed process can reach far more than intended.
- Forgetting to limit the number of processes or threads the sandboxed code can spawn is a distinct gap from per-process CPU/memory limits; a fork bomb exhausts the host through process count, not through any single process's resource usage, so a cgroup PID limit is a separate, necessary control.
- In-process rlimits are not a portable safety net. The demonstrated
RLIMIT_ASrejection on macOS is a concrete instance of a broader trap: a design validated only in one environment (commonly Linux-based continuous integration) can silently ship with a memory control that does nothing at all on a different deployment platform.
Unlock Full Question Bank
Get access to all Secure Coding and Application Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.