Threat Modeling and Attack Surface Analysis Questions
Systematically identifying how a system can be attacked and where its exposure lies. Covers structured methodologies (STRIDE, PASTA, DREAD, OCTAVE, attack trees), enumerating and reducing attack surface, mapping trust boundaries and data flows via DFDs, profiling likely threat actors, and prioritizing identified threats by likelihood and impact during design. Includes applying this methodology to specific architectural substrates (cloud-native and serverless, microservices, ML/AI systems, IoT, CI/CD pipelines, cryptographic subsystems) and operationalizing it as a recurring program (SDLC integration, governance, tooling, KPIs). The proactive 'think like an attacker before you build' discipline: distinct from live penetration testing (the adversarial validation of a built system), from runtime detection/monitoring (recognizing an attack already in progress), and from implementing the resulting security controls (a separate design-and-build discipline).
Explain the STRIDE threat modeling categories (Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege). For each category, provide a concise example vulnerability or attack path in the context of a web-based e-commerce application (authentication, payments, product catalog) and explain why it maps to that STRIDE category.
Sample Answer
Direct answer
STRIDE is a mnemonic for six threat categories, each corresponding to the violation of a security property: Spoofing (violates authentication), Tampering (violates integrity), Repudiation (violates non-repudiation), Information Disclosure (violates confidentiality), Denial of Service (violates availability), and Elevation of Privilege (violates authorization). Its value is that it forces you to ask all six questions at every trust-boundary crossing rather than relying on whatever threats happen to come to mind.
Structured elaboration
For a web-based e-commerce application (auth, payments, product catalog):
| STRIDE category | violates | e-commerce example | why it maps here |
|---|---|---|---|
| Spoofing | Authentication | An attacker replays a stolen session cookie to impersonate a logged-in customer's checkout session | The attacker is pretending to BE someone they are not; that is the defining test for Spoofing |
| Tampering | Integrity | A request to /api/cart/update has its unit_price field modified client-side before submission, bypassing server-side price validation | Data is being altered in a way the system did not intend; integrity, not confidentiality, is the property broken |
| Repudiation | Non-repudiation | A customer disputes a charge and the payment service has no signed, immutable log tying that specific customer session to that specific transaction | The system cannot PROVE who did what, which is exactly what non-repudiation logging exists to prevent |
| Information Disclosure | Confidentiality | The product catalog API returns full internal SKU cost and supplier fields in its JSON response, visible to any authenticated user via browser dev tools | Data intended to stay internal is exposed to an unintended audience |
| Denial of Service | Availability | An attacker scripts thousands of cart-abandonment requests against the checkout service, exhausting the database connection pool for legitimate shoppers | The system becomes unavailable to legitimate users; this is the availability axis specifically |
| Elevation of Privilege | Authorization | A regular customer account discovers that the admin order-refund endpoint only checks for a valid session token, not an admin role claim, and can call it directly | The attacker already has SOME access (a valid session) and uses it to gain MORE access than authorized; that upgrade-in-privilege is the defining test |
The classification test that keeps these six from blurring together: ask which single security PROPERTY is being violated (who you are, what you're allowed to do, whether data can be altered, read, proven, or kept available), not what TECHNIQUE the attacker used (an attacker might use SQL injection to achieve either Tampering or Information Disclosure depending on what the injected query does -- the technique doesn't determine the category, the violated property does).
Worked example
Trace one flow end to end through all six lenses: the checkout POST request. An attacker could (S) forge the session identity via a predictable session token, (T) alter the cart total in the request body, (R) complete a purchase then deny having done so if no audit trail ties their authenticated session to the order, (I) intercept the response and see another customer's saved card metadata due to an IDOR-style object reference bug, (D) flood the endpoint to exhaust checkout capacity during a sale, and (E) invoke an internal-only discount-application parameter that a normal customer request should never carry. Walking every STRIDE category against a SINGLE data flow like this, rather than against the application as a whole, is what makes the technique produce a complete rather than an ad hoc threat list.
Trade-offs and pitfalls
STRIDE tells you WHAT KIND of thing can go wrong per component or data flow; it does not tell you HOW LIKELY or HOW SEVERE each threat is -- that is DREAD's or a risk matrix's job, applied afterward. A common mistake is running STRIDE once against the whole application instead of once per component/data-flow-crossing-a-trust-boundary in the DFD; the latter is what actually produces a tractable, complete threat list rather than a handful of threats someone happened to think of.
For a multi-region active-active microservices platform using service mesh and automated CI/CD with cross-region data replication, produce a threat model identifying high-impact threats (misconfiguration, pipeline compromises, secrets leakage, replication divergence) and propose architecture and operational mitigations to preserve availability and security during region failures or CI/CD rollback scenarios.
Sample Answer
Direct answer
The four named threats each trace to a different root cause, a configuration drift, a compromised delivery pipeline, an exposed credential, and data that disagrees with itself across regions, so each gets its own architecture control (what is built) and operational practice (what is regularly exercised), rather than one blanket "add more security" response. The design has to hold up under two specific stress scenarios named in the question, a region failure and a CI/CD rollback, and the interesting risk is not either scenario alone but what happens when they overlap: a rollback initiated in the middle of a region failure is exactly where availability pressure and security shortcuts are most likely to collide.
Structured elaboration
flowchart TB
CI[CI/CD: build,\nsign, SBOM] --> GATEA{Region A gate:\ntwo-person approval}
CI --> GATEB{Region B gate:\ntwo-person approval}
GATEA --> A[Region A:\nmesh plus services]
GATEB --> B[Region B:\nmesh plus services]
GLB[Global load balancer] --> A
GLB --> B
A <-->|cross-region\nreplication| B
A -.->|lag and\nchecksum signal| DIVERGE{Divergence above\nthreshold?}
B -.->|lag and\nchecksum signal| DIVERGE
DIVERGE -- yes --> THROTTLE[Throttle writes,\nroute reads to\nknown-good region]
A -- region failure --> GLB
ROLLBACK[Rollback:\nsame signed artifact,\nsame gates] --> GATEA
ROLLBACK --> GATEB
Misconfiguration (mesh policies, load-balancer weights, DNS TTLs causing traffic blackholes or split-brain). Architecture: require every mesh policy and routing change to pass through policy-as-code validation (an automated check, akin to Open Policy Agent's Gatekeeper for Kubernetes admission control) before it can apply to any region, and roll changes out region by region rather than to every region simultaneously, so a bad policy is caught in the first region before it reaches the rest. Operational: run scheduled cross-region failover drills that specifically exercise Domain Name System (DNS) time-to-live (TTL) behavior and client reconnection under a real, timed failover, not only a synthetic health-check pass, since the health check passing and the client actually reconnecting cleanly are not the same thing.
CI/CD pipeline compromises (an attacker injects a malicious image or configuration, or promotes a bad build to every region at once). Architecture: sign every build artifact and verify the signature before deploy, keep pipeline runner environments hardened and isolated, and require a separate promotion gate per region rather than one global "promote everywhere" action, so a single compromised promotion cannot reach every region in one step. Operational: require two-person approval for any cross-region promotion, rotate pipeline credentials on a fixed cadence, and periodically red-team the pipeline itself as a target, not only the application it deploys.
Secrets leakage (pipeline, mesh sidecar, or replication credentials exposed). Architecture: issue short-lived credentials through a workload-identity mechanism rather than embedding long-lived static secrets in configuration or images, and ensure secrets are never written to pipeline logs by construction (redaction at the logging layer, not relying on developers to remember). Operational: run periodic automated secret scanning across repositories, built images, and log output, with a fast, rehearsed rotation runbook for anything a scan finds.
Replication divergence (a network partition or asymmetric replication produces conflicting writes). Architecture: choose the consistency model deliberately, per data domain, rather than one blanket choice for the whole platform. For data where a conflicting write is unacceptable (billing state, for example), route writes to a single designated primary region for that domain even in an otherwise active-active design, accepting a small availability cost during that region's own outage in exchange for correctness. For data where eventual consistency with defined merge semantics is acceptable, use a conflict-resolution strategy suited to the data shape, for example a conflict-free replicated data type where the structure allows automatic, order-independent merging, or a change-data-capture stream with explicit conflict-resolution rules where it does not. Operational: continuously monitor replication lag and data checksums between regions, and when divergence crosses a defined, per-domain threshold, automatically throttle writes to the affected region or route reads to a known-good region rather than serving data that may already be inconsistent.
How the design holds up under a region failure. When a region fails, the global load balancer shifts traffic to healthy regions using the same health-aware routing the drills above exercise, but availability alone is not the goal, security has to survive the failover too: the newly primary region must still enforce the same authorization and mesh-identity checks as before (nothing about failover should implicitly grant broader trust), and because secrets are already replicated through the workload-identity mechanism rather than stored only in the failed region, the surviving region can authenticate and authorize normally without an emergency, weaker fallback path being invented under pressure.
How the design holds up under a CI/CD rollback. A rollback is not exempt from the same controls a forward deploy uses: it should redeploy a previously signed, already-reviewed artifact through the same per-region promotion gates, not a separate "emergency, skip review" path, because an attacker who can convince an on-call engineer to trigger an unreviewed emergency deploy has found a way around every pipeline control described above. Database or schema changes tied to the original deploy need to be reversible, using an expand-contract pattern (adding new fields or tables without removing old ones until the rollback window has safely closed) so that rolling back the application code does not leave it running against a schema it can no longer read correctly.
Worked example
Trace the compound case the design most needs to survive: a network partition takes Region A offline while an engineer is mid-rollback of yesterday's bad deploy. The global load balancer begins shifting Region A's traffic to Region B using the health-aware routing exercised in drills; because replication-divergence monitoring was already watching lag between the two regions before the partition, it catches the resulting increase in lag as Region A drops out of sync and automatically throttles writes that would otherwise land only in the now-unreachable region, preventing a burst of writes that could never actually replicate. At the same time, the rollback in Region B goes through the same signed-artifact, two-person-approval gate a forward deploy would use, specifically because the region failure is already an unusually stressful moment where a shortcut would be most tempting, and that is exactly when the pipeline controls matter most, not a moment to informally suspend them. Because the original deploy used an expand-contract schema pattern, Region B's rollback to the previous application version runs cleanly against the still-present old and new schema fields without a data-compatibility break. When Region A recovers, its replication catches back up against Region B's now-authoritative state, and only once the divergence monitor reports the two regions back within the normal threshold does traffic resume being served from Region A again, rather than resuming immediately and risking a second round of conflicting writes.
Trade-offs and pitfalls
Applying strong, single-region-primary consistency to every data domain would defeat the purpose of an active-active design in the first place, since a partition would then force an explicit choice between availability and consistency for data that did not need that trade-off; the point of choosing the consistency model per domain is that only the data which genuinely cannot tolerate a conflicting write pays that cost. Two-person approval and per-region promotion gates add real friction at exactly the moment speed feels most urgent, during an incident; the design needs a pre-approved emergency path for rolling back to an artifact that was already reviewed once (fast, but still gated) rather than a separate "break glass, skip everything" escape hatch, since a standing bypass of the pipeline controls is itself a long-term vulnerability, not just a convenience. A common pitfall is testing failover drills only against a clean, planned scenario and never against a deploy or rollback happening at the same time; the worked example's compound case is exactly the kind of overlap that isolated drills miss, and it is where the four named threats can compound each other's actual impact rather than staying independent. Finally, tuning the replication-divergence threshold too aggressively generates alert fatigue during ordinary, benign eventual-consistency windows that were never actually a problem, so the threshold needs real tuning against each data domain's own tolerance rather than one number applied uniformly across very different kinds of data.
You discover a high-severity vulnerability that can be remediated only by disabling a widely used feature for 48 hours, which would reduce expected revenue by approximately 10%. As an SRE lead, how do you present the options to executive stakeholders, recommend a course of action, set measurable metrics to evaluate risk reduction, and plan communications and rollback? Explain how you would make the tradeoff clear and obtain buy-in.
Sample Answer
Direct answer
The core move in this scenario is refusing to present the choice as "security versus revenue," which is how it will sound if stated flatly, and instead presenting it as a bounded, time-boxed trade with a measurable exit condition: disable the feature for the shortest defensible window, define upfront exactly what "risk reduced enough to re-enable" looks like, and pair the revenue cost with a concrete estimate of what the alternative, an unpatched high-severity vulnerability sitting live, actually risks. Executive stakeholders should see two or three real options, not one recommendation presented as the only path, a clear recommendation with its reasoning, the specific metrics that will tell everyone the trade paid off, and a communications and rollback plan that removes the fear of "and then what" from the room.
Structured elaboration
Presenting options to executive stakeholders
Lay out genuine alternatives, not a single option dressed up as a choice:
- Disable the feature for 48 hours while the vulnerability is patched, accepting the roughly 10% revenue impact for that window.
- Leave the feature live with compensating controls (aggressive rate limiting, enhanced monitoring, restricting the feature to a subset of lower-risk traffic) while patching in the background, accepting a smaller but nonzero window of continued exposure in exchange for avoiding the full revenue hit.
- A partial disable, turning off only the specific code path that carries the vulnerability if the feature is separable, which may reduce both the revenue impact and the remediation timeline versus a full shutdown.
Each option should be presented with its own revenue impact, remediation timeline, and residual exposure, so the executive audience is choosing between three concretely described trades, not between "the security team's ask" and "doing nothing."
Recommended course of action
Recommend the option whose exposure-reduction-per-revenue-dollar is clearest and most defensible, which in most cases with a genuinely high-severity vulnerability is the partial disable if the feature is separable (option 3), since it captures most of the risk reduction of a full shutdown at a fraction of the revenue cost, or the full 48-hour disable (option 1) if the vulnerable code path cannot be cleanly isolated, since a compensating-controls-only approach (option 2) leaves a real, live vulnerability reachable by a sufficiently motivated attacker for the full remediation window, which is difficult to defend after the fact if it is exploited during that window.
Measurable metrics to evaluate risk reduction
- Exposure window closed: the number of hours the vulnerable path was reachable by an attacker, before versus after the action taken, which is the most direct measure of what the trade actually bought.
- Patch verification: a specific, named test confirming the vulnerability is closed (a reproduction of the original finding against the patched system, showing it no longer succeeds) before re-enabling, rather than re-enabling on a patch-deployed timestamp alone.
- Post-re-enable monitoring: a defined observation period (for example, the first 24-48 hours after re-enabling) with alerting specifically targeted at the previously vulnerable path, so a failed or incomplete patch is caught quickly rather than discovered later.
- Each of these is a concrete, checkable condition, not a vague "we'll monitor it," which is what makes the metrics section answer the executive's real underlying question: how will we know this worked.
Communications and rollback plan
- Before disabling: a short, plain-language notice to affected users or customers explaining the change and expected duration, framed around reliability or security rather than exposing internal vulnerability detail, since a vague or absent notice tends to generate more support burden than an honest, brief one.
- During the window: a status page or equivalent update if the remediation timeline shifts, so stakeholders are not surprised by a 48-hour estimate quietly becoming 72.
- Rollback plan: if the patch is not ready within the committed window, decide in advance, not in the moment, whether to extend the disable, ship a partial mitigation, or accept a defined smaller residual risk to re-enable early; having this decision pre-agreed with the executive stakeholders removes the pressure to make it under time stress on hour 47.
- After re-enabling: a brief closure communication confirming resolution, which closes the loop for the same audience that received the original notice and reduces the chance the incident resurfaces as a trust question later.
Making the trade-off clear and obtaining buy-in
Translate both sides of the trade into the same terms the executives already use. The revenue side is already in those terms (the roughly 10% figure). The security side needs the same treatment: state plainly what a successful exploit of this vulnerability would let an attacker do (for example, unauthorized access to customer data, or the ability to take actions as another user), and frame the 48-hour cost as bounded and quantifiable against an unpatched vulnerability's cost, which is unbounded and open-ended for as long as it stays live. Buy-in tends to follow once the comparison is stated as "a known, bounded cost now" versus "an unknown, open-ended cost for as long as we wait," rather than as an abstract severity rating the room has no intuitive way to weigh against a concrete revenue number.
Worked example
Say the vulnerability, if exploited, would let an attacker read other users' account data, a concrete, statable consequence. Presented to executives: "Option 1, full disable, costs an estimated 10% of revenue for 48 hours, roughly $X based on typical daily revenue for this feature, and fully closes the exposure. Option 2, controls only, costs under 1% of revenue but leaves the account-data-read vulnerability live for the same 48-hour patch window, exploitable by anyone who finds it during that time. Option 3, if the vulnerable path is separable from the rest of the feature, likely costs 2-4% of revenue and closes the exposure as fully as Option 1." With those three trades stated in comparable terms, the recommendation (option 3 if separable, otherwise option 1) becomes a specific, arguable claim rather than an appeal to authority, and the metrics section gives the room a concrete way to confirm afterward that the chosen option actually delivered the closed exposure it promised: the exposure-window and patch-verification metrics either show the vulnerability closed on schedule, or they do not, and either way the room has a fact to react to rather than a reassurance to take on faith.
Trade-offs and pitfalls
- The most common wrong turn is presenting only the recommended option, which forces executives to either rubber-stamp a decision they had no part in shaping or push back with no alternative on the table; presenting genuine options, even if one is clearly better, is what earns buy-in rather than compliance.
- Quantifying the revenue cost precisely while leaving the security cost as a qualitative severity label stacks the comparison unfairly and tends to produce a decision that under-weights the security side simply because it is the side with the harder-to-state number; the worked example's approach, stating what a successful exploit would concretely let an attacker do, is what closes that gap without fabricating a false-precision dollar figure for the security side.
- Skipping the pre-agreed rollback decision and improvising if the timeline slips is a common and costly mistake: a decision made under pressure at hour 47, with executives who were not part of the original trade-off conversation now being pulled in urgently, tends to produce worse outcomes than the same decision made calmly in advance.
- Re-enabling on a "the patch shipped" timestamp rather than a verified test result risks reopening the exposure if the patch is incomplete, which is exactly why patch verification is listed as its own metric rather than folded into "the timeline was met."
Summarize the Process for Attack Simulation and Threat Analysis (PASTA) methodology: list its stages and briefly describe the objective of each stage. Explain in what situations PASTA is more appropriate than a simpler framework like STRIDE.
Sample Answer
Direct answer
PASTA (Process for Attack Simulation and Threat Analysis) is a seven-stage, risk-centric methodology that starts from business objectives and works down to concrete attack simulation, in contrast to STRIDE's more mechanical per-component category sweep. PASTA is the better choice when the audience and stakes require an explicit business-risk narrative (funding decisions, compliance justification, executive buy-in), not just a technical threat list.
Structured elaboration
The seven stages, each with its objective:
- Define Objectives - capture business objectives and compliance requirements so later technical findings can be traced back to business impact.
- Define Technical Scope - enumerate the architecture, technologies, and dependencies in scope (the technical surface the rest of the process operates on).
- Application Decomposition - build the DFD-style model of components, data flows, and trust boundaries (this is where PASTA and STRIDE's inputs overlap).
- Threat Analysis - gather threat intelligence relevant to the decomposed architecture (industry-specific threat actors, known campaign patterns).
- Vulnerability and Weakness Analysis - map identified threats to actual vulnerabilities and weaknesses in the decomposed architecture (correlate stage 4 against stage 3).
- Attack Modeling - simulate plausible attack scenarios (attack trees or kill chains) that would exploit the identified vulnerabilities.
- Risk and Impact Analysis - quantify business impact and residual risk for each simulated attack, and prioritize remediation by business-risk-adjusted severity, not just technical severity.
When PASTA beats STRIDE: STRIDE is fast, mechanical, and per-component; it is the right tool when a team needs to sweep a specific system quickly and technical stakeholders will consume the output directly. PASTA is heavier (typically multi-day for a real system, requiring cross-functional participation from stage 1) but earns that cost when the audience includes non-technical business stakeholders who need to see risk in impact terms rather than category names, when the system is genuinely business-critical and regulatory-scrutinized (payments, healthcare, financial services) where a compliance-ready risk narrative is itself a deliverable, or when the goal is prioritizing a remediation budget across many findings, which requires the business-impact quantification PASTA builds in from stage 1.
Worked example
Consider threat-modeling a new payments API. A STRIDE pass on the API's DFD produces a solid technical finding quickly: the token-refresh endpoint is vulnerable to Spoofing via a predictable refresh token. A PASTA pass on the same system, run because leadership needs to decide whether to delay launch, produces a fuller chain: objectives (PCI DSS scope, Q3 launch date), technical scope (the API plus its 3 upstream dependencies), decomposition (the same DFD), threat analysis (payments APIs are a known target for credential-stuffing campaigns per current threat intel), vulnerability analysis (the predictable refresh token maps directly to that threat), attack modeling (a simulated credential-stuffing-to-refresh-token-prediction chain), and risk analysis (estimated fraud exposure if exploited, weighed against a 2-week launch delay to fix it). Same underlying technical finding, but PASTA's extra stages produce the artifact leadership can actually act on.
Trade-offs and pitfalls
PASTA's depth is also its cost: running full PASTA on every minor feature change is not sustainable, and teams that try tend to abandon it after one exhausting cycle. A defensible pattern is STRIDE for routine per-feature reviews and PASTA reserved for new business-critical systems, major re-architectures, or when a compliance or executive audience genuinely needs the business-risk narrative. Stopping at stage 3 (decomposition) and calling it PASTA produces a STRIDE-shaped DFD with none of PASTA's actual differentiator, for twice the process overhead.
Perform a threat model for a flow where an external partner uploads files via SFTP to a staging bucket, an Airflow DAG triggers a Spark job in Kubernetes to transform data, and the results populate a multi-tenant analytics database. Identify top attack surfaces, mitigations for supply-chain and insider threats, and detection controls to add at ingestion, processing, and serving layers.
Sample Answer
Direct answer
This pipeline has three layers with meaningfully different trust levels: an ingestion layer where an external partner is, by definition, untrusted; a processing layer that is internally trusted but pulls in third-party code and container images, which is where a supply-chain compromise would land; and a serving layer that must keep tenants isolated from each other, which is where an insider or a processing-layer compromise could otherwise read or write across tenant boundaries. The top attack surfaces are the SFTP upload path itself (the only point a genuinely external, unauthenticated-relative-to-your-org actor touches the system), the Spark job's dependency and container-image supply chain, and the multi-tenant data boundary at the database layer, and each needs its own mitigation and detection strategy rather than one blanket control.
Structured elaboration
Pipeline layout and trust boundaries:
flowchart LR
subgraph Ingestion["Ingestion layer: untrusted external partner"]
Partner[External Partner]
SFTP[SFTP Server]
Stage[(Staging Bucket)]
end
subgraph Processing["Processing layer: internal compute plus third-party deps"]
Airflow[Airflow DAG Scheduler]
K8s[Kubernetes Cluster]
Spark[Spark Job]
Deps[(Third-party Packages and Images)]
end
subgraph Serving["Serving layer: multi-tenant"]
DB[(Multi-tenant Analytics DB)]
end
Partner -->|uploads file| SFTP
SFTP --> Stage
Stage -->|triggers| Airflow
Airflow -->|schedules| K8s
K8s -->|runs| Spark
Deps -.->|pulled at build or run time| Spark
Spark -->|writes tenant-scoped rows| DB
Top attack surfaces.
- The SFTP upload itself. This is the one place a truly external actor interacts with the system directly. Threats include a compromised partner credential uploading malicious content, path traversal or filename manipulation in the uploaded file's metadata, and the uploaded file itself carrying a payload designed to exploit whatever parses it downstream (a crafted CSV or Parquet file targeting a parser vulnerability, for example).
- The Spark job's supply chain. Airflow scheduling Spark on Kubernetes typically means pulling a container image and a set of language-runtime dependencies (Python or Java/Scala packages) at build or deploy time. A compromised or typosquatted dependency, or a compromised base image, executes with whatever privileges the Spark job's Kubernetes service account holds, which in a poorly scoped cluster can be broad.
- The multi-tenant boundary at the database. Because a single Spark job's output lands in a shared analytics database serving multiple tenants, any bug in how the job scopes its writes, or any actor (insider or compromised process) with broader-than-necessary write access, can read or write across tenant lines.
Mitigations for supply-chain threats. Pin dependency versions and container base images to specific, verified digests rather than mutable tags, so a compromised upstream package cannot silently change what gets pulled into a future build. Scan images and dependency trees for known vulnerabilities and unexpected changes before deployment, and prefer a maintained internal mirror or artifact repository over pulling directly from public registries at run time, both to reduce exposure to a public registry being compromised or going unavailable and to create a control point for scanning. Scope the Spark job's Kubernetes service account to only the permissions the job actually needs (specific namespace, specific storage paths, no cluster-admin), so a compromised dependency executing inside the job cannot pivot broadly across the cluster.
Mitigations for insider threats. The core control is least-privilege access scoped per layer: an engineer who can modify the Airflow DAG definition should not automatically also hold broad write access to the serving database, and access to the staging bucket, the DAG source, and the production analytics database should be reviewed and logged separately rather than treated as one undifferentiated "data platform" permission set. Require code review and a deployment pipeline (not direct manual edits) for changes to the DAG and Spark job logic, since that is the layer where an insider could quietly alter transformation logic to leak or corrupt tenant data without an obvious signature. For the serving layer specifically, enforce tenant isolation at the database level itself (row-level security or per-tenant schema/database separation, not only application-layer filtering), so a compromised or malicious insider with query access still cannot casually read across tenants without an explicit, loggable privilege escalation.
Detection controls, one per layer.
- Ingestion: monitor SFTP authentication patterns (unusual source IP, upload volume, or timing for a given partner credential) and validate uploaded file structure and content type against an expected schema before it ever reaches the staging bucket that triggers processing, rejecting and alerting on anything that does not conform rather than passing it downstream for the pipeline to fail on later.
- Processing: monitor for unexpected outbound network connections from the Spark job's pods (a compromised dependency attempting to exfiltrate data or reach a command-and-control endpoint would typically need to make a network call the legitimate job never makes) and alert on any deployment of a container image or dependency version that does not match the pinned, scanned manifest.
- Serving: monitor for any query or write against the analytics database that crosses a tenant boundary (a query without the expected tenant-scoping predicate, or a write touching more than one tenant's partition), which should be rare-to-never in legitimate traffic and therefore a high-signal, low-false-positive detection.
Worked example
Trace one concrete failure through the diagram: a partner's SFTP credential is phished, and the attacker uploads a file crafted so that a benign-looking column actually contains a payload targeting a known deserialization weakness in an outdated version of a Spark dependency the job has not updated in over a year. Ingestion-layer detection (schema and content-type validation on upload) has a chance to catch this if the payload violates the expected schema, but a sufficiently well-crafted payload disguised as valid data might pass. If it reaches the Spark job, the outdated, unpinned dependency executes the payload, and because the Kubernetes service account for that job was scoped down to only the specific namespace and storage paths it needs (the supply-chain mitigation above), the compromised process cannot pivot to other workloads or escalate within the cluster, containing the blast radius to that one job's resources. The processing-layer detection control (unexpected outbound connections from the job's pods) is what actually surfaces the incident: the compromised process attempting to exfiltrate data to an external endpoint generates a network connection the legitimate job never makes, triggering an alert. Be precise about what that alert does and does not buy, because this is where threat models usually overclaim: egress monitoring is a detection control, not a preventive one, so on its own it tells responders the job is compromised while the job keeps running. Whether tainted output reaches the multi-tenant database depends on a separate decision taken in advance, namely whether an egress alert on a pipeline pod is wired to fail the job and quarantine its output automatically, or only to page a human. For a pipeline that writes across a shared tenant boundary the automated kill is usually worth its occasional false-positive job failure, since a rerun is cheap and cross-tenant contamination is not, but that is a design choice this pipeline has to make explicitly rather than a property detection gives you for free.
Trade-offs and pitfalls
The most common mistake is treating this as a single trust zone and applying uniform controls everywhere, when the whole value of separating attack surfaces by layer is that the ingestion layer needs input-validation-style controls, the processing layer needs supply-chain and least-privilege controls, and the serving layer needs data-isolation controls, and a control designed for one layer is often the wrong shape for another (schema validation does nothing for a compromised dependency, and dependency pinning does nothing for a malicious upload). A second pitfall is relying on application-layer tenant filtering alone at the serving layer; application code has bugs, and a defense-in-depth posture puts the real isolation guarantee in the database layer itself (row-level security or physical separation) so that an application-layer bug is a defect, not a tenant-isolation breach. A third is under-scoping detection to only the ingestion layer, since it is the most obviously external-facing point; the worked example above shows the ingestion control can plausibly miss a well-crafted payload, and it is the processing-layer network-egress detection that actually catches the incident, which is exactly why detection needs to exist at every layer rather than concentrating it at the perimeter.
Unlock Full Question Bank
Get access to all Threat Modeling and Attack Surface Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.