Secure Coding and Application Security Questions
Writing and reviewing code that resists attack. Covers the OWASP Top Ten and common web vulnerabilities (XSS, SQL injection, CSRF), input validation, secure coding practices and security code review, static application security testing (SAST), API and HTTP security, database and frontend security, and mobile app security. The application-layer defense discipline for engineers building software.
A microservice accepts arbitrary URLs to fetch thumbnails and was abused to perform SSRF calls to internal metadata endpoints. Propose a layered mitigation plan covering input validation and URL canonicalization, an allowlist of permitted destinations, egress filtering and network segmentation, proxying requests through a vetted fetch service, and runtime detection for anomalous outbound requests.
Sample Answer
Direct answer: For a microservice that was abused for SSRF via an arbitrary-URL thumbnail fetch, the fix needs to be layered: no single control here is sufficient alone, because each addresses a different way the previous layer could fail or be bypassed.
Structured elaboration, layer by layer:
- Input validation and URL canonicalization. Before anything else, parse the supplied URL properly (never with a regex) and reject anything that isn't
http/https, reject URLs with embedded credentials (http://user:pass@host), and canonicalize to catch equivalent-but-differently-encoded forms of the same dangerous target (decimal/octal/hex IP representations of127.0.0.1, IPv6 loopback forms,0.0.0.0). - An explicit allowlist of permitted destinations. Rather than trying to block every dangerous target, only fetch from domains the business actually needs (known image-hosting CDNs the feature is meant to support), resolved and IP-pinned at request time to defeat DNS rebinding - re-check the resolved IP right before connecting, not just at initial validation, since an attacker-controlled DNS record can change its answer between the two.
- Egress filtering and network segmentation. Even if the allowlist has a gap, network-level policy (a firewall rule, a cloud security group, or a service mesh egress policy) that blocks the thumbnail service from reaching internal-only IP ranges and the cloud metadata address provides an independent layer that doesn't depend on the application code being perfectly correct.
- Proxying requests through a vetted fetch service. Rather than letting every service that needs to fetch external URLs implement its own validation (and inevitably implement it inconsistently), centralize outbound URL-fetching behind one hardened internal service that owns the allowlisting, IP-pinning, redirect-following rules, and egress restrictions once, correctly, and every other service calls it instead of doing its own
requests.get(user_url). - Runtime detection. Even with the above, monitor for anomalous outbound request patterns from this service (requests to internal IP ranges, requests to the metadata address specifically, an unusual spike in failed/timed-out fetches that might indicate probing) as a safety net that catches both bypasses and future regressions.
Worked example. The vetted fetch service receives {url: "http://169.254.169.254/latest/meta-data/..."}, resolves the hostname, finds it's in the disallowed CIDR range for cloud metadata addresses, and rejects it before ever making the outbound connection - centralizing this logic means every consuming service (thumbnailing, link-preview generation, webhook delivery) gets this protection automatically rather than each team needing to remember to implement it correctly.
Trade-offs and pitfalls: a centralized fetch service becomes a new operational dependency (a single point of failure and a place where a misconfiguration affects every consumer at once), and it needs to correctly follow (or explicitly refuse) redirects - a service that validates only the initial URL but blindly follows a redirect to an internal address has done all this validation for nothing.
For a web application using cookie-based session tokens, describe the attack surface for CSRF, session fixation, and XSS. For each attack, describe at least two mitigation layers (application-level and infrastructure-level) you would implement, and explain the limitations or trade-offs of each control.
Sample Answer
Direct answer: For a web app using cookie-based session tokens, CSRF, session fixation, and XSS share the same underlying trust the browser places in cookies, but each needs its own layered defense; no single control covers all three. CSRF specifically works because the browser automatically attaches a user's session cookie to any request sent to a site, even one triggered by a different, attacker-controlled page - so a forged request looks legitimate to the server unless something explicitly checks for that.
Structured elaboration, attack surface and layered mitigations per attack:
| Attack | Application-level control | Infrastructure-level control | Limitation of relying on either alone |
|---|---|---|---|
| CSRF | Anti-CSRF token validated on every state-changing request | SameSite=Lax or Strict cookie attribute | Tokens alone can be forgotten on a new endpoint; SameSite alone doesn't protect against same-site attacks (a malicious subdomain) or older browsers that ignore the attribute |
| Session fixation | Regenerate the session ID on login/privilege change, invalidating the old one | None meaningful at the infrastructure layer; this is purely an application session-management concern | If you rotate the ID but don't invalidate the old session server-side, the old (attacker-known) session may still work until it naturally expires |
| XSS | Context-aware output encoding at every render point | Content-Security-Policy blocking inline scripts, HttpOnly on the session cookie so JavaScript can't read it even if XSS occurs | HttpOnly stops cookie theft via document.cookie but does NOT stop the attacker from making authenticated requests AS the victim from within the page itself (the browser still attaches the cookie) |
Worked example connecting the three. Consider a session-fixation attack as the entry point: an attacker gets a victim to log in under a session ID the attacker planted (via a crafted link). Now the attacker's known session ID is authenticated. From there, even without stealing anything, the attacker could ride that session to perform CSRF-protected actions if the CSRF token itself was also generated at session-creation time and thus also known to the attacker (a common mistake: binding the CSRF token to the pre-fixation session rather than regenerating it alongside the session ID). This is why session regeneration on login has to reset BOTH the session identifier and any session-bound CSRF token together, not just one of them.
Trade-offs and pitfalls: SameSite=Strict is the strongest CSRF mitigation but breaks legitimate cross-site navigation into the app (a user clicking a link from an email lands unauthenticated, since the cookie isn't sent on that top-level navigation); Lax is the common practical default, which still sends the cookie on top-level GET navigations but not on cross-site POSTs or embedded requests, covering the actual CSRF attack shape while preserving normal linking. HttpOnly and CSP are complementary, not redundant: HttpOnly protects the cookie's confidentiality even under successful XSS, while CSP tries to prevent the XSS from executing at all; a defense-in-depth posture wants both, since either can fail independently (a CSP misconfiguration, or a cookie set without HttpOnly by an older piece of legacy code).
Describe advanced WAF evasion techniques you might test during a penetration test (for example: mixed encoding, parameter pollution, polyglot payloads, chunked-request tricks). Pick one technique, explain how you would craft a proof of concept to demonstrate the bypass, and list the mitigations that should be applied at both the WAF and application levels.
Sample Answer
Direct answer
Web application firewall (WAF) evasion during a penetration test almost always exploits a gap between how the WAF parses a request and how the real application (or an intermediate proxy) eventually parses the same bytes; four common families are mixed/double encoding, HTTP parameter pollution, polyglot payloads, and chunked-request tricks. The strongest proof of concept picks one family, shows the specific parsing disagreement it exploits with a reproducible artifact, and pairs that with mitigations at both the WAF and the application layer, because a WAF-only fix (tightening one signature) does not remove the underlying application weakness the payload was trying to reach in the first place.
Structured elaboration
The four technique families
- Mixed/double encoding: a WAF signature is written against the decoded form of a payload, but only decodes a request once; an attacker encodes the payload one extra time, so the WAF's single decode still leaves it looking benign while the backend, which decodes twice (often because a proxy layer and an application layer each do their own decode step, unaware the other already did), reconstructs the dangerous value.
- HTTP parameter pollution: the same parameter name is submitted multiple times with different values (
?id=1&id=1' OR '1'='1); a WAF that inspects only the first or only the last occurrence can miss whichever one the application's own parameter-parsing convention actually uses. - Polyglot payloads: a single string is crafted to be syntactically valid, and dangerous, in more than one context at once (for example, a value that is inert as plain text but becomes an executable expression once it lands inside a specific templating or expression-language context); a generic WAF signature tuned for one interpreter's syntax can miss a payload shaped for a different one.
- Chunked-request tricks: an attacker splits a payload across multiple HTTP chunk boundaries (using
Transfer-Encoding: chunked); a WAF that inspects each chunk independently, rather than reassembling the full body first, can miss a signature that only becomes visible once the chunks are concatenated.
Deep dive: crafting a proof of concept for mixed/double encoding
This is the technique with the clearest, most mechanically reproducible bypass, so it is the one worked in full below. The core idea: a naive WAF filter that URL-decodes a request body exactly once before pattern-matching it can be bypassed by encoding the payload twice, if the backend (through a proxy layer, a framework layer, or both) ends up decoding it twice before use.
import re
import urllib.parse
SQLI_PATTERN = re.compile(r"(?i)\bor\b.{0,10}=.{0,10}|--|;|\bunion\b|\bselect\b")
def naive_waf_check(raw_body: str) -> bool:
"""Simulates a WAF that URL-decodes a request body exactly ONCE before
pattern-matching it against a SQL-injection signature. Returns True if
the request is BLOCKED."""
once_decoded = urllib.parse.unquote(raw_body)
return bool(SQLI_PATTERN.search(once_decoded))
def backend_effective_value(raw_body: str) -> str:
"""Simulates the backend, which (as many real frameworks and reverse
proxies do for legacy reasons) decodes the value TWICE before using it
in a query: once by the web server / framework layer, once by an
application-level helper that assumes the framework layer didn't
already do it."""
return urllib.parse.unquote(urllib.parse.unquote(raw_body))
def demo(raw_body: str, label: str) -> None:
blocked = naive_waf_check(raw_body)
effective = backend_effective_value(raw_body)
would_inject = bool(SQLI_PATTERN.search(effective))
print(f"--- {label} ---")
print(f"raw body sent: {raw_body}")
print(f"WAF single-decode view: {urllib.parse.unquote(raw_body)}")
print(f"WAF verdict: {'BLOCKED' if blocked else 'PASSED'}")
print(f"backend double-decoded: {effective}")
print(f"backend would execute injection: {would_inject}")
print()
if __name__ == "__main__":
plain_payload = "username=admin' OR '1'='1"
demo(plain_payload, "Plain payload (no encoding)")
single_encoded = urllib.parse.quote(plain_payload)
demo(single_encoded, "Single URL-encoded payload")
double_encoded = urllib.parse.quote(urllib.parse.quote(plain_payload))
demo(double_encoded, "Double URL-encoded payload (the evasion)")
Running this against the plain payload, a single-encoded payload, and a double-encoded payload produces:
--- Plain payload (no encoding) ---
raw body sent: username=admin' OR '1'='1
WAF single-decode view: username=admin' OR '1'='1
WAF verdict: BLOCKED
backend double-decoded: username=admin' OR '1'='1
backend would execute injection: True
--- Single URL-encoded payload ---
raw body sent: username%3Dadmin%27%20OR%20%271%27%3D%271
WAF single-decode view: username=admin' OR '1'='1
WAF verdict: BLOCKED
backend double-decoded: username=admin' OR '1'='1
backend would execute injection: True
--- Double URL-encoded payload (the evasion) ---
raw body sent: username%253Dadmin%2527%2520OR%2520%25271%2527%253D%25271
WAF single-decode view: username%3Dadmin%27%20OR%20%271%27%3D%271
WAF verdict: PASSED
backend double-decoded: username=admin' OR '1'='1
backend would execute injection: True
The plain and single-encoded payloads both get blocked, because the WAF's one decode pass reveals the pattern either way. The double-encoded payload passes the WAF (its single decode only strips one layer, leaving %3D, %27, and so on still present, so the signature does not match) while the backend's two decode passes reconstruct the exact same dangerous string the WAF was supposed to be catching. In a real engagement this simulation stands in for the actual next step, sending the crafted double-encoded body against the live target and confirming the response indicates the query executed (a successful login bypass, a different error signature, or a timing difference), which is environment-specific and not something to fabricate a result for here.
Mitigations
At the WAF level:
- Canonicalize/decode fully (repeatedly, to a fixed point, not a fixed count) before pattern matching, so a filter cannot be starved by an extra encoding layer.
- Reassemble chunked and fragmented request bodies fully before inspection rather than evaluating each chunk independently.
- Prefer a positive-security model (allow-list expected characters/shapes per field) over a purely negative, signature-based model for high-value endpoints, since a positive model does not depend on anticipating every encoding variant of a bad pattern.
At the application level:
- Use parameterized queries or an object-relational mapping (ORM) layer so that even a payload that reaches the query-construction code cannot alter the query's structure, regardless of what encoding it arrived in; this is the control that makes the WAF bypass harmless even when it succeeds.
- Decode exactly once, at a single well-defined layer, and make every other layer consume the already-decoded value; the double-decode weakness in this proof of concept exists specifically because two layers each assumed the other had not already decoded the input.
- Reject requests with unexpected or excessive encoding depth outright (a legitimate client essentially never needs to double-encode a form field).
Trade-offs and pitfalls
- Reporting a WAF bypass without the application-level root cause. A WAF signature update closes this specific encoding variant but not the next one; the report should make clear that the durable fix is parameterization, and the WAF fix is a compensating control, not a substitute for it.
- Assuming a bypass proven against a simulated or lab WAF generalizes to the client's actual product and configuration. Real WAF products vary widely in decode depth and reassembly behavior; a proof of concept like the one above demonstrates the class of weakness and gives the client something concrete to test their own configuration against, not a verified claim about their specific deployment.
- Chasing every encoding variant individually. Attackers can compose these four families (double-encode a chunked, polluted parameter, for example); testing and fixing one variant at a time is a losing game compared to canonicalizing input fully and relying on parameterization as the backstop.
You must produce a concise risk assessment for a new serverless payment-processing function that retrieves secrets from a store, calls external payment gateways, and writes results to a database. Produce: (A) a short textual data-flow summary, (B) the top five concrete vulnerability risks ranked by severity, and (C) prioritized, code-level mitigations that fit a compliance-minded but agile team.
Sample Answer
Direct answer
A risk assessment for this function should mirror how it actually moves data: retrieving a secret, calling an external payment gateway, and writing a result to a database are the three trust-boundary crossings that generate essentially all of the real risk here. The top five risks map onto those crossings plus the function's own execution environment, ranked by what a successful compromise would actually cost (payment-data exposure, fraud, non-compliance). Every mitigation is scoped to ship inside a normal sprint, not as a separate compliance initiative bolted on afterward, since a compliance-minded but agile team needs controls it can actually merge and deploy on its regular cadence.
Structured elaboration
(A) Data-flow summary
Trigger event arrives, the function starts (cold or warm), it retrieves the payment-gateway credential from a secrets store, validates and normalizes the incoming payment request, calls the external payment gateway to charge or authorize, receives the gateway's response, writes the transaction result to a database, and returns a response to the caller. Logging is a cross-cutting concern threaded through every one of these stages, not a separate stage of its own, and is exactly where sensitive data most often leaks even when the "main" data path is otherwise sound.
flowchart LR
U[Client] -->|HTTPS request| GW[API Gateway]
GW --> FN[Payment Function]
FN -->|fetch secret| VAULT[Secrets Store]
FN -->|charge request| PG[External Payment Gateway]
FN -->|write result| DB[(Database)]
PG -->|webhook callback| FN
(B) Top five risks, ranked by severity
- Sensitive data exposure through logging or error handling. Payment card data, tokens, or the retrieved secret itself ending up in logs, error messages, or monitoring output at any stage above. Ranked highest because it is simultaneously high-likelihood, extremely easy to introduce by accident (logging a full request or response object while debugging), and high-impact, direct payment-data exposure and a direct compliance violation, since the Payment Card Industry Data Security Standard (PCI DSS) explicitly prohibits logging full card numbers and sensitive authentication data.
- Insecure outbound call construction, an SSRF-shaped risk. If the payment-gateway endpoint or any parameter of the outbound call is built from anything influenced by the incoming request rather than a fixed, trusted configuration value, an attacker could redirect that call toward an internal endpoint. Serverless functions commonly have broad outbound network reachability by default, which is exactly what makes this risk class land here specifically.
- Secrets over-exposure through the function's execution identity. If the function's execution role can read more than the one payment-gateway credential it actually needs, for example every secret in the store rather than just its own, a compromise of this one function (through any of the other risks here) cascades into every other secret, not only this integration's own.
- Insufficient validation of the gateway's response before persisting it. Trusting fields in the response, amount, status, currency, without validating them against what was actually requested opens a path where a compromised or spoofed intermediary could cause the function to record a transaction result that does not match reality, a data-integrity risk with direct financial consequences.
- Overly broad database write access, or unparameterized query construction. The function's database credential having write access beyond the one transactions table it needs repeats the blast-radius problem from risk 3 at the database layer, and building the write query by concatenating any request or gateway-response field into a string, rather than using parameterized queries, is the same SQL injection class, untrusted data concatenated directly into a query string instead of bound as a parameter, applied here specifically.
(C) Prioritized code-level mitigations
- P1 (ships first, cheapest, highest impact): a shared, redacting logging wrapper. Scrub known-sensitive field names before anything is logged, applied at one shared utility so individual call sites cannot forget it. Directly closes risk 1, is shippable inside a single sprint, and happens to also satisfy PCI DSS's logging-restriction requirement without being framed as a separate compliance project.
- P2: pin the gateway endpoint to a fixed configuration constant, never derived from request data, and, where the runtime allows it, restrict outbound network reachability at the platform level to only the gateway's known address. Two complementary layers, code-level and platform-level, closing risk 2 without requiring a large architectural change.
- P3: scope the function's execution role to read access on exactly the one secret it needs, enforced by an automated policy-lint check in continuous integration (CI) rather than manual review alone, so scope creep is caught before merge instead of after an incident. Closes risk 3.
- P4: validate the gateway's response against an explicit schema (expected fields, types, an amount matching what was requested within tolerance, a bounded status enum) before writing anything to the database, rejecting or flagging a mismatch rather than trusting and persisting it. Closes risk 4.
- P5: scope the database credential to insert-only access on the transactions table specifically, and use parameterized queries or an ORM exclusively for the write. Closes risk 5.
The mitigations are ordered to match the severity ranking in (B): the highest-severity risk gets the cheapest, fastest-shipping fix first, and each subsequent mitigation closes the next risk down the list.
Worked example
The P1 mitigation (the redacting log wrapper), concretely, since it is the highest-priority fix:
SENSITIVE_KEYS = {"card_number", "cvv", "api_key", "secret", "authorization"}
def redact(payload):
if isinstance(payload, dict):
return {k: ("***REDACTED***" if k.lower() in SENSITIVE_KEYS else redact(v))
for k, v in payload.items()}
if isinstance(payload, list):
return [redact(v) for v in payload]
return payload
def safe_log(event_name, payload):
return f"{event_name}: {json.dumps(redact(payload))}"
Run against a realistic payment request:
naive_log output: payment_attempt: {"amount": 4999, "currency": "USD", "card_number": "4111111111111111",
"cvv": "123", "gateway_auth": {"api_key": "sk_live_abc123XYZ"}}
safe_log output: payment_attempt: {"amount": 4999, "currency": "USD", "card_number": "***REDACTED***",
"cvv": "***REDACTED***", "gateway_auth": {"api_key": "***REDACTED***"}}
The naive log call genuinely contains the raw card number and the live API key, confirming risk 1 is real and not merely theoretical; the redacting wrapper removes both while preserving non-sensitive fields (amount, currency) that a developer still needs for debugging, and works recursively, so a sensitive key nested inside gateway_auth is caught exactly the same as a top-level one.
Trade-offs and pitfalls
- A redaction wrapper is only as good as its key list. A field named something the list does not anticipate (a typo, a new field added later, a nested structure with an unexpected key name) slips through; pair the deny-list approach with periodic review of what is actually flowing through logs, not a one-time list written at rollout and never revisited.
- Least-privilege secrets and database scoping (P3, P5) are cheap to state and easy to under-invest in, since a broader role often "just works" during initial development and the gap only shows up during an incident; enforcing the scope check in CI, not just at initial setup, is what keeps it from drifting back open as the function evolves.
- Response validation (P4) has a real design cost: too strict a schema breaks on a legitimate gateway API change, too loose a schema does not actually catch a spoofed response. Version the schema deliberately alongside the gateway integration rather than treating it as a one-time check.
- This risk assessment intentionally does not repeat a design-time STRIDE/trust-boundary methodology exercise, that belongs to a sibling discipline; it stays at the concrete, implementation-level risk-and-mitigation altitude the question actually asks for.
Architect a secure API gateway for an enterprise that centralizes protection against injection, broken authentication/authorization, SSRF, and protocol abuse. Describe the components involved (authentication, authorization, WAF, mutual TLS, rate limiting, token introspection, egress controls, SSO protections), how the policies are enforced, how you would instrument detection, and trade-offs such as latency and operational complexity.
Sample Answer
Direct answer
A secure Application Programming Interface (API) gateway centralizes the security controls that would otherwise be duplicated (and inconsistently implemented) across every backend service: authentication, authorization, injection and protocol defense, Server-Side Request Forgery (SSRF)/egress control, and Single Sign-On (SSO) protections, all enforced at one well-instrumented choke point in front of a fleet of services that individually trust the gateway rather than the open internet. The design has to hold two things in tension: pushing enforcement to one place makes it consistent and auditable, but it also makes the gateway a single point of both failure and latency, so the architecture needs to be highly available and fast on the hot path while still being the place every security decision and every detection signal converges.
Structured elaboration
Component architecture.
flowchart LR
Client -->|"1: TLS handshake"| GW[API Gateway]
GW -->|"2: verify token"| AuthN[AuthN service<br/>OIDC/SSO IdP]
GW -->|"3: check scopes"| AuthZ[AuthZ / policy engine]
GW -->|"4: inspect request"| WAF[WAF layer]
GW -->|"5: rate check"| RL[Rate limiter]
GW -->|"6: emit signal"| Detect[Detection / SIEM pipeline]
GW -->|"7: forward, mTLS"| Backend1[Backend service A]
GW -->|"7: forward, mTLS"| Backend2[Backend service B]
Backend1 -->|"egress request"| EgressCtl[Egress control /<br/>allow-listed destinations only]
Authentication. Terminate authentication at the gateway, not in each backend, so there is exactly one place that validates tokens and exactly one place a token-validation bug can exist. For interactive users, the gateway participates in an SSO flow (OpenID Connect (OIDC) or SAML) against a central identity provider, exchanging the SSO session for a short-lived, gateway-issued access token that backends actually see. For service-to-service and third-party callers, validate a JSON Web Token (JWT) or opaque token via introspection against the issuing authorization server (OAuth 2.0 token introspection, RFC 7662) rather than trusting a locally cached public key indefinitely, so a revoked token stops working immediately instead of only once it naturally expires.
Authorization. Authentication answers "who is this," authorization answers "what are they allowed to call," and these need to be separate, composable decisions. A centralized policy engine (attribute-based, evaluating caller identity, requested route, and request context together) lets the gateway make a coarse-grained allow/deny decision before the request ever reaches a backend, while fine-grained, resource-level authorization (can this specific user see this specific record) still belongs in the backend service, which is the only place that actually knows the resource's ownership. The gateway's job is to cut off the large class of requests that should never reach a backend at all (wrong scope, wrong audience, expired token), not to replace the backend's own authorization logic.
Web Application Firewall (WAF). Sits in the request path to catch the OWASP-Top-Ten-shaped payloads (SQL (Structured Query Language) injection patterns, script-injection payloads, path traversal sequences, known exploit signatures for the frameworks in use) before they reach application code, as a defense-in-depth layer, never as a substitute for parameterized queries and output encoding in the backend itself. Tune it in detection-only mode first against real production traffic to characterize false positives before flipping to blocking mode, because a WAF that blocks legitimate traffic on day one erodes trust in the whole control and invites teams to request exceptions that quietly widen the hole.
Mutual TLS (mTLS). Two distinct mTLS relationships exist in this design and they serve different purposes: gateway-to-backend mTLS establishes that traffic reaching a backend really came through the gateway (backends can then refuse any connection that does not present the gateway's client certificate, closing off direct-to-backend bypass), while client-to-gateway mTLS (where the caller is a service or partner rather than a browser user) provides strong caller authentication independent of, and in addition to, the token-based authentication above.
Rate limiting. Apply it at multiple granularities simultaneously: per-caller-identity (the primary control, since it survives the caller rotating IPs), per-route (protecting expensive endpoints specifically, like search or export), and a coarse per-source-IP limit as a backstop against unauthenticated abuse before a caller identity is even established. Rate limiting is also a security control, not just a cost control: it is what turns a credential-stuffing or brute-force attempt from "instant" into "slow enough to detect and block."
Token introspection. Beyond initial validation, route sensitive operations through live introspection against the authorization server rather than relying solely on a cached JWT's embedded expiry, specifically because token revocation (a compromised session being killed, a user being deprovisioned) needs to take effect immediately, and a purely local, stateless JWT validation cannot express "this specific token was just revoked" without either a short token lifetime plus refresh (acceptable staleness window) or introspection (immediate, at the cost of a network round-trip per request).
Egress controls. The gateway is the natural place to also enforce outbound rules for any backend that itself makes server-side requests to caller-influenced URLs (an SSRF vector): centralizing an allow-list of legitimate outbound destinations here means one policy update closes the hole for every backend, instead of relying on each service team to have implemented its own allow-list correctly.
SSO protections. Beyond the authentication flow itself, this means the gateway (or the identity provider (IdP) it delegates to) enforces: strict redirect-URI allow-listing on the OAuth/OIDC flow (an open redirect here is a full account-takeover primitive, not a cosmetic bug), state/nonce validation to prevent cross-site request forgery (CSRF) and replay against the SSO callback, and short-lived session tokens with refresh rotation so a leaked session token has a bounded window of usefulness. Because SSO centralizes identity, a flaw in this specific flow compromises every downstream service simultaneously, which is exactly why it deserves explicit design attention rather than being treated as "just OAuth, handled by the library."
Detection instrumentation. Every one of the layers above should emit a structured event on both allow and deny decisions (not just denials; a stream of "everything is fine" telemetry is what lets you notice when it suddenly stops), feeding a Security Information and Event Management (SIEM) pipeline that correlates: repeated authentication failures for one identity (credential stuffing), authorization denials clustering on one route (probing for a missing check), WAF signature matches, and rate-limit trips. Because the gateway sees 100% of external traffic, it is the single richest source of this signal in the whole architecture, and instrumentation here should be treated as a first-class design requirement, not an afterthought bolted on after the routing logic is done.
Policy enforcement mechanics. Represent authentication and authorization requirements as declarative, versioned policy (per route: required scopes, rate limits, WAF ruleset, mTLS requirement) rather than as imperative code scattered through gateway plugins, so that a policy change is reviewable in a pull request and consistently applied, and so that a new backend service is secure by default the moment it is registered with the gateway rather than requiring every team to independently remember every control.
Worked example
A concrete route: POST /api/v1/payments/refund. The gateway's policy for this route declares: scope=payments:refund, rate_limit=10/min per identity, mtls_required=true for the calling service, waf_ruleset=strict. A request arrives with a valid SSO-derived JWT, but for a caller whose token has scope=payments:read only. The gateway's authorization check denies the request with a 403 before it reaches the payments backend at all, and emits a structured denial event tagged with the caller identity, the requested scope, and the granted scope. If ten of these denials arrive from the same caller identity within a minute, the SIEM correlation rule for "scope-probing" fires and pages the on-call security engineer, who can see from the single gateway log stream exactly which route and which identity, without needing to correlate logs across the payments service, the auth service, and the network layer separately. This is the concrete payoff of centralization: one denial is noise, ten correlated denials from one identity against one sensitive route is a signal, and the gateway is the only place positioned to see that pattern in real time.
Trade-offs and pitfalls
| Trade-off | Cost | Why it is usually still worth it |
|---|---|---|
| Latency | Every layer (authentication, authorization, WAF inspection, rate-limit check) adds hops before the request reaches the backend | Run authentication/authorization/rate-limit checks in-memory or against a local cache with async revalidation rather than a synchronous round-trip per layer per request, and only pay the full introspection round-trip cost for sensitive routes, not every request |
| Operational complexity | The gateway becomes a large, stateful, high-blast-radius component that a small team now has to run at very high availability, since every request depends on it | Treat gateway configuration with the same rigor as application code: versioned policy, staged rollout, automated rollback, and a documented bypass procedure for the gateway's own outage that does not simply disable security controls fleet-wide |
| Single point of failure | A gateway outage takes down every backend behind it, even backends that were themselves healthy | Design for graceful degradation per control (e.g., fail closed on authentication, but define explicitly whether WAF inspection fails open or closed under gateway resource pressure) rather than an undifferentiated "gateway is down, everything is down" |
| False confidence in backend teams | Backend engineers can start assuming "the gateway handles security" and skip resource-level authorization or input validation in their own service | Make explicit in the platform's contract with service teams that the gateway handles coarse-grained, cross-cutting controls only; fine-grained authorization and defense-in-depth input handling remain each backend's own responsibility, and this needs to be a stated architectural principle, not an assumption |
| WAF false positives | Overly aggressive rules block legitimate traffic (a customer's business data that happens to contain a string resembling a SQL keyword) | Stage new rules in detection-only mode against real traffic before blocking, and give backend teams a fast, auditable exception path so they are not tempted to work around the gateway entirely |
The single biggest pitfall in this design is architectural: building "one big gateway that does everything" without separating the concerns of authentication/authorization (identity-plane), WAF/rate-limiting (traffic-plane), and egress/SSRF control (network-plane) into independently scalable, independently failable components. A monolithic gateway that couples all three tends to fail all three together under load, exactly when the security controls matter most.
Unlock Full Question Bank
Get access to all Secure Coding and Application Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.