Threat Modeling and Attack Surface Analysis Questions
Systematically identifying how a system can be attacked and where its exposure lies. Covers structured methodologies (STRIDE, PASTA, DREAD, OCTAVE, attack trees), enumerating and reducing attack surface, mapping trust boundaries and data flows via DFDs, profiling likely threat actors, and prioritizing identified threats by likelihood and impact during design. Includes applying this methodology to specific architectural substrates (cloud-native and serverless, microservices, ML/AI systems, IoT, CI/CD pipelines, cryptographic subsystems) and operationalizing it as a recurring program (SDLC integration, governance, tooling, KPIs). The proactive 'think like an attacker before you build' discipline: distinct from live penetration testing (the adversarial validation of a built system), from runtime detection/monitoring (recognizing an attack already in progress), and from implementing the resulting security controls (a separate design-and-build discipline).
Design a threat model for a serverless data processing pipeline (API Gateway -> Lambda -> Kinesis -> analytics). Identify threats including event injection, function impersonation, excessive privileges, insecure dependencies, and cold-start related timing risks. Recommend mitigations across IAM, VPC placement, input validation, dependency management, observability, and secure deployment patterns.
Sample Answer
Direct answer
Model the pipeline (API Gateway into Lambda into Kinesis into downstream analytics) as a chain of trust boundaries where every hop independently verifies what the previous hop handed it, rather than inheriting trust from an earlier check. The five named threats each trace back to a different root cause, an unvalidated event, a spoofed or unauthorized invocation, an over-broad permission grant, a vulnerable packaged dependency, and a cold-start-related timing effect, so each gets addressed primarily through one or two of six mitigation categories (Identity and Access Management, or IAM; network placement; input validation; dependency management; observability; and secure deployment patterns) rather than one blanket fix covering everything.
Structured elaboration
flowchart LR
CLIENT[External client] -->|1: request| APIGW[API Gateway:\nschema validation,\nrate limiting]
APIGW -->|2: validated event| LAMBDA[Lambda:\nscoped IAM role,\ninput re-validation]
LAMBDA -->|3: PutRecord\nrestricted to this role| KINESIS[(Kinesis stream)]
OTHER[Other internal service\nno PutRecord permission] -.->|blocked by IAM| KINESIS
KINESIS -->|4: consume| ANALYTICS[Analytics Lambda:\nverifies producer marker,\nnarrow downstream scope]
ANALYTICS -->|5: scoped call only| DOWNSTREAM[Downstream service]
subgraph TB1[Trust boundary: public internet to platform]
CLIENT
end
subgraph TB2[Trust boundary: platform-internal, still\nper-function least privilege]
APIGW
LAMBDA
KINESIS
OTHER
ANALYTICS
DOWNSTREAM
end
Event injection or malformed payloads. Addressed primarily through input validation: enforce strict schema validation at API Gateway before a request ever reaches Lambda, rejecting unexpected fields, wrong types, or oversized payloads with a fail-closed default. For events written directly to Kinesis by internal producers rather than through the public API, add a signed or keyed marker that only an approved producer can generate, so a consumer can distinguish a genuinely produced event from one written by a compromised internal service, which schema validation alone cannot do.
Function impersonation. Addressed primarily through IAM: give each function a narrowly scoped execution role rather than a shared one, restrict who may invoke or update a given function to specific, named principals rather than any authenticated caller, and prefer short-lived, automatically rotated credentials over long-lived keys. Pair this with observability that flags an invocation or update coming from an unexpected principal, since that is exactly the signal that would indicate impersonation succeeded despite the IAM controls.
Excessive privileges. Also primarily an IAM concern, but distinct from impersonation: scope each function's execution role to the specific resources it actually touches, a specific Kinesis stream's Amazon Resource Name (ARN), a specific storage prefix, a specific secret, rather than account-wide or wildcard access, and keep the role that deploys or updates a function separate from the role the function uses at runtime. A function that only needs to read one stream and write to one downstream service should be structurally unable to reach anything else, so that a bug or a successful exploit inside the function's own code cannot be leveraged into a broader compromise.
Insecure dependencies. Addressed through dependency management: generate a software bill of materials (SBOM) at build time covering both direct and transitive dependencies, run software composition analysis (SCA) scanning against a maintained vulnerability feed before a package is allowed to ship, pin dependency versions rather than floating on the latest release, and keep the runtime's dependency footprint minimal since a smaller set of packages is a smaller set of things that can go wrong. Combine this with secure deployment patterns: immutable, versioned deployment artifacts (no editing a deployed function in place) and signed builds, so a dependency that passed the scan cannot be silently swapped for something else afterward.
Cold-start related timing risks. This threat actually covers two distinct risks worth separating. The first is a timing side channel: if a security-sensitive check (verifying a signature or comparing a secret) behaves differently on a cold start than on a warm invocation, an attacker who can distinguish cold from warm responses might infer something about the check's internal state; the mitigation is using constant-time comparison for any security-sensitive check and avoiding branching on cold-versus-warm state in code paths that handle secrets. The second is availability: an attacker who triggers a large burst of concurrent, unique invocations can force many simultaneous cold starts, degrading latency or throughput for legitimate callers even without exploiting anything directly. The mitigation here is a mix of reserved concurrency (capping how many concurrent instances a function can consume, containing the blast radius on shared infrastructure) and, for functions where consistent low latency is itself a security requirement, provisioned concurrency (keeping a set number of execution environments warm in advance), plus rate limiting at API Gateway so an invocation flood is throttled before it ever reaches Lambda.
Network placement. Functions that need to reach resources with no public endpoint (an internal database, a private service) should run inside a private subnet, connecting to AWS services like Kinesis through an interface Virtual Private Cloud (VPC) endpoint (powered by AWS PrivateLink) so that traffic never traverses the public internet. Functions that only call public AWS service endpoints can stay outside a VPC entirely, avoiding network overhead that provides no security benefit for a function with nothing private to reach; VPC placement is a targeted control for functions that actually need it, not a default applied uniformly regardless of what a given function talks to.
Observability, as its own design-time requirement. Beyond the impersonation-specific signal above, the pipeline as a whole needs tracing and metrics that let an operator reconstruct what happened across all three hops, API Gateway, Lambda, and Kinesis, for a given request or event: a propagated request identifier carried through every hop, per-function invocation and error-rate metrics, and Kinesis-level metrics such as iterator age (how far a consumer has fallen behind, which is itself a signal of a possible tampering or throttling issue upstream). This is design-time work in the same sense as the rest of this answer: specifying which signals the pipeline must emit so that any of the five threats above leaves a visible trace, not building the alerting and response system that consumes those signals, which is a separate discipline.
Secure deployment patterns, beyond the dependency-specific case above. The same discipline applies to the pipeline's own configuration and infrastructure, not only to application dependencies: manage infrastructure as code with a required review step, roll out changes to the pipeline (a new IAM policy, a new event-source mapping) through a staged or canary process rather than applying a change to the whole pipeline at once, and pair every deployment with an automated rollback trigger tied to the observability signals above, so a bad change is caught and reverted from its own telemetry rather than from a manual report.
Worked example
Suppose the analytics Lambda, on certain event types, calls a downstream administrative service to update account settings, a legitimate but sensitive part of the pipeline. An attacker who has compromised a low-privilege internal service, not the analytics Lambda itself, attempts to write a crafted event directly onto the Kinesis stream that looks like a legitimate "apply this admin action" event, trying to trigger that sensitive downstream call without going through the public API at all. In the intended configuration, only the specific producer Lambda's execution role has PutRecord permission on this stream; the compromised low-privilege service lacks that specific IAM grant, so the write attempt is denied at the IAM layer before the crafted event ever reaches the stream, let alone the analytics Lambda. Now consider a worse case: the attacker has instead compromised a service that legitimately does have PutRecord access to this stream. IAM alone no longer stops the write, but the analytics Lambda's input validation checks every consumed record for the signed producer marker described above; a record written by an unapproved source, even one with valid write access to the stream itself, fails that check and is discarded rather than acted on. If that check were somehow also defeated, the analytics Lambda's own execution role is scoped narrowly to the one downstream administrative action it actually needs, not a broad administrative capability, so even a fully successful forged event has a bounded blast radius rather than an open-ended one. Three independent mechanisms drawn from two of the six mitigation categories, IAM (both the stream-write restriction and the analytics function's narrowly scoped execution role) and input validation (the producer-marker check), each reduce the same attack path on their own, which is the point: no single layer is asked to be the only thing standing between a compromised internal service and a sensitive downstream action.
Trade-offs and pitfalls
VPC placement has a real, if now much smaller, cost: Lambda functions in a VPC historically incurred a significant cold-start penalty from provisioning a network interface per invocation, and while AWS's 2019 Hyperplane networking change meaningfully reduced that overhead by pre-provisioning shared network interfaces instead of one per function, VPC placement is still not free, so it should be applied to the functions that genuinely need to reach private resources rather than as a default for every function in the pipeline. Provisioned concurrency reduces cold-start-driven timing and availability risk but incurs an ongoing cost for capacity that sits idle between invocations, so it is worth reserving for the specific functions where cold-start behavior is actually security-relevant or latency-critical, not applied blanket across the whole pipeline. A common pitfall in practice is that teams grant a broad execution role early "to get it working" during initial development and never tighten it once the function is stable; the fix is a build-time policy check (policy-as-code review that rejects IAM statements broader than the pipeline's own declared usage) rather than relying on someone remembering to revisit permissions later. Finally, a subtle but costly mistake is securing the public API Gateway entry point thoroughly while forgetting that Kinesis is itself a separate write surface: if any internal service can write to the stream without IAM restricting who may do so, the API Gateway's careful validation is bypassed entirely by construction, since the stream, not the gateway, is what the downstream consumer actually trusts.
Perform a detailed threat model for a multi-tenant cloud data warehouse used by regulated customers. Focus on tenant isolation, side-channel risks, data exfiltration, privileged access, query logs, and metadata leakage. Recommend architectural mitigations (encryption per tenant, query sandboxing, workload isolation) and controls to demonstrate isolation to auditors.
Sample Answer
Direct answer
A multi-tenant cloud data warehouse for regulated customers needs a threat model built around one question repeated for every layer of the stack: can tenant A ever see, infer, or affect tenant B's data or performance? The six areas named in the question (tenant isolation, side-channel risks, data exfiltration, privileged access, query logs, metadata leakage) all reduce to variations of that question, and each needs both an architectural mitigation and a way to prove the isolation holds to an auditor who won't take "trust us" as an answer.
Structured elaboration
Tenant isolation failures
- Threat: a bug in row-level security, a missing tenant-ID filter in a query path, or a shared connection pool that leaks context between tenants lets one tenant's query return another tenant's rows.
- Mitigation: enforce tenant scoping at the lowest practical layer, not just in application code. Options in increasing strength and cost: row-level security policies enforced by the database engine itself (so even a buggy application query cannot bypass it), per-tenant schemas or databases, or fully separate compute clusters for the highest-sensitivity tenants. Never rely solely on application-layer
WHERE tenant_id = ?filters as the only control, since a single missed filter in one code path is a full isolation failure.
Side-channel risks, including noisy-neighbor effects
- Threat: tenants sharing physical compute (CPU cache, memory bus, disk I/O, or query-planner statistics) can infer information about each other's workload through timing, resource contention, or query-plan behavior, even with zero direct data access. The specific noisy-neighbor case is a tenant's heavy query load degrading or altering the observable performance of another tenant's queries, which itself is a low-bandwidth side channel (an attacker can sometimes infer when a competitor tenant runs large batch jobs, for example) as well as a plain availability problem.
- Mitigation: workload isolation through dedicated virtual clusters, VPC-level or compute-cgroup separation, and resource quotas per tenant so one tenant cannot exhaust shared capacity; for the highest-risk tenants, dedicated physical or virtual hosts rather than shared multi-tenant compute; query cost limits and admission control so a single tenant's query cannot starve the shared pool even accidentally.
Data exfiltration
- Threat: exfiltration via query results (a tenant, or an attacker who compromised a tenant's credentials, runs broad export queries), via user-defined functions (UDFs) that reach out to the network, or via a compromised internal service account with warehouse-wide access.
- Mitigation: sandbox UDF execution with no outbound network access by default; apply data loss prevention (DLP) scanning and rate limits on bulk export operations; require justification or approval workflows for large exports; restrict service accounts to the minimum tenant scope they actually need rather than warehouse-wide access as a default.
Privileged access
- Threat: database administrators, cloud platform administrators, or support staff with elevated access can read raw tenant data outside of any tenant-facing control, which regulated customers specifically ask about.
- Mitigation: just-in-time (JIT) privilege elevation instead of standing admin access, mandatory multi-factor authentication and approval for elevation, full session recording for privileged sessions, and separation of duties so no single administrator can both grant themselves access and use it unaudited.
Query logs
- Threat: query logs, which typically have broader read access than the production data itself (since they're often shipped to a general-purpose logging or observability platform), can contain literal tenant data if queries embed values directly, or can reveal query patterns that leak business information across tenants if logs aren't tenant-partitioned.
- Mitigation: redact or parameterize logged queries so literal values don't appear in plaintext logs; partition log storage and access by tenant, mirroring the data isolation model rather than treating logs as a separate, less-protected system; apply the same encryption and access controls to logs as to the underlying data.
Metadata leakage
- Threat: even without touching row data, metadata (table names, schema structure, row counts, query timing) can reveal a tenant's business activity to anyone with broader metadata access, and cross-tenant metadata stores are an easy place to under-protect because they don't feel like "the data" to engineers building the system.
- Mitigation: partition metadata by tenant with the same rigor as data, avoid global metadata views that span tenants unless explicitly required for platform operations, and treat metadata access grants as seriously as data access grants in the access review process.
Architectural mitigations, tied together
- Encryption per tenant: unique, KMS-backed data encryption keys per tenant (envelope encryption), so a key compromise or misconfiguration is scoped to one tenant rather than the whole warehouse.
- Query sandboxing: isolate UDF and ad hoc query execution in sealed environments with no unnecessary network egress and static analysis of submitted code where feasible.
- Workload isolation: dedicated compute paths for regulated or high-sensitivity tenants, resource quotas for everyone else, so noisy-neighbor effects are bounded even when full physical separation isn't cost-justified for every tenant.
Worked example
Trace how these controls combine for one concrete scenario: a support engineer needs to debug a slow query for tenant A. Without the controls above, that engineer might have standing warehouse-wide read access and pull raw rows from tenant A's tables directly, which is both a privileged-access risk and, if the query touches tenant B's shared execution plan cache, a potential metadata leak. With the controls above: the engineer requests JIT access scoped specifically to tenant A's schema, the request requires approval and is time-boxed, the session is recorded, and the query the engineer runs is logged with values redacted and stored in tenant A's own log partition. Nothing in that workflow required trusting the individual engineer's judgment; the controls make the isolation hold even for a well-intentioned support engineer, which is the property an auditor is actually testing for.
Trade-offs and pitfalls
Per-tenant encryption keys and dedicated compute cost real money and operational complexity: key rotation, backup, and restore workflows all get harder when every tenant has its own key material, and this cost should be stated plainly to leadership rather than presented as free. A common pitfall is protecting the primary data store carefully while leaving logs and metadata as an afterthought; both are named explicitly in this question precisely because they're the parts of the system engineers tend to under-protect, and an auditor evaluating "isolation" for a regulated customer will ask about them specifically. To demonstrate isolation to auditors concretely, bring: architecture diagrams showing per-tenant keys and workload boundaries, documented key lifecycle and rotation policy, access review records and JIT elevation logs, results from periodic side-channel and penetration testing, and a mapping of these controls to the relevant compliance framework (SOC 2 or ISO 27001 controls, for example) the customer expects. A model that only produces a risk list without this auditor-facing evidence trail has not actually answered the question's "controls to demonstrate isolation" requirement.
For a multi-region active-active microservices platform using service mesh and automated CI/CD with cross-region data replication, produce a threat model identifying high-impact threats (misconfiguration, pipeline compromises, secrets leakage, replication divergence) and propose architecture and operational mitigations to preserve availability and security during region failures or CI/CD rollback scenarios.
Sample Answer
Direct answer
The four named threats each trace to a different root cause, a configuration drift, a compromised delivery pipeline, an exposed credential, and data that disagrees with itself across regions, so each gets its own architecture control (what is built) and operational practice (what is regularly exercised), rather than one blanket "add more security" response. The design has to hold up under two specific stress scenarios named in the question, a region failure and a CI/CD rollback, and the interesting risk is not either scenario alone but what happens when they overlap: a rollback initiated in the middle of a region failure is exactly where availability pressure and security shortcuts are most likely to collide.
Structured elaboration
flowchart TB
CI[CI/CD: build,\nsign, SBOM] --> GATEA{Region A gate:\ntwo-person approval}
CI --> GATEB{Region B gate:\ntwo-person approval}
GATEA --> A[Region A:\nmesh plus services]
GATEB --> B[Region B:\nmesh plus services]
GLB[Global load balancer] --> A
GLB --> B
A <-->|cross-region\nreplication| B
A -.->|lag and\nchecksum signal| DIVERGE{Divergence above\nthreshold?}
B -.->|lag and\nchecksum signal| DIVERGE
DIVERGE -- yes --> THROTTLE[Throttle writes,\nroute reads to\nknown-good region]
A -- region failure --> GLB
ROLLBACK[Rollback:\nsame signed artifact,\nsame gates] --> GATEA
ROLLBACK --> GATEB
Misconfiguration (mesh policies, load-balancer weights, DNS TTLs causing traffic blackholes or split-brain). Architecture: require every mesh policy and routing change to pass through policy-as-code validation (an automated check, akin to Open Policy Agent's Gatekeeper for Kubernetes admission control) before it can apply to any region, and roll changes out region by region rather than to every region simultaneously, so a bad policy is caught in the first region before it reaches the rest. Operational: run scheduled cross-region failover drills that specifically exercise Domain Name System (DNS) time-to-live (TTL) behavior and client reconnection under a real, timed failover, not only a synthetic health-check pass, since the health check passing and the client actually reconnecting cleanly are not the same thing.
CI/CD pipeline compromises (an attacker injects a malicious image or configuration, or promotes a bad build to every region at once). Architecture: sign every build artifact and verify the signature before deploy, keep pipeline runner environments hardened and isolated, and require a separate promotion gate per region rather than one global "promote everywhere" action, so a single compromised promotion cannot reach every region in one step. Operational: require two-person approval for any cross-region promotion, rotate pipeline credentials on a fixed cadence, and periodically red-team the pipeline itself as a target, not only the application it deploys.
Secrets leakage (pipeline, mesh sidecar, or replication credentials exposed). Architecture: issue short-lived credentials through a workload-identity mechanism rather than embedding long-lived static secrets in configuration or images, and ensure secrets are never written to pipeline logs by construction (redaction at the logging layer, not relying on developers to remember). Operational: run periodic automated secret scanning across repositories, built images, and log output, with a fast, rehearsed rotation runbook for anything a scan finds.
Replication divergence (a network partition or asymmetric replication produces conflicting writes). Architecture: choose the consistency model deliberately, per data domain, rather than one blanket choice for the whole platform. For data where a conflicting write is unacceptable (billing state, for example), route writes to a single designated primary region for that domain even in an otherwise active-active design, accepting a small availability cost during that region's own outage in exchange for correctness. For data where eventual consistency with defined merge semantics is acceptable, use a conflict-resolution strategy suited to the data shape, for example a conflict-free replicated data type where the structure allows automatic, order-independent merging, or a change-data-capture stream with explicit conflict-resolution rules where it does not. Operational: continuously monitor replication lag and data checksums between regions, and when divergence crosses a defined, per-domain threshold, automatically throttle writes to the affected region or route reads to a known-good region rather than serving data that may already be inconsistent.
How the design holds up under a region failure. When a region fails, the global load balancer shifts traffic to healthy regions using the same health-aware routing the drills above exercise, but availability alone is not the goal, security has to survive the failover too: the newly primary region must still enforce the same authorization and mesh-identity checks as before (nothing about failover should implicitly grant broader trust), and because secrets are already replicated through the workload-identity mechanism rather than stored only in the failed region, the surviving region can authenticate and authorize normally without an emergency, weaker fallback path being invented under pressure.
How the design holds up under a CI/CD rollback. A rollback is not exempt from the same controls a forward deploy uses: it should redeploy a previously signed, already-reviewed artifact through the same per-region promotion gates, not a separate "emergency, skip review" path, because an attacker who can convince an on-call engineer to trigger an unreviewed emergency deploy has found a way around every pipeline control described above. Database or schema changes tied to the original deploy need to be reversible, using an expand-contract pattern (adding new fields or tables without removing old ones until the rollback window has safely closed) so that rolling back the application code does not leave it running against a schema it can no longer read correctly.
Worked example
Trace the compound case the design most needs to survive: a network partition takes Region A offline while an engineer is mid-rollback of yesterday's bad deploy. The global load balancer begins shifting Region A's traffic to Region B using the health-aware routing exercised in drills; because replication-divergence monitoring was already watching lag between the two regions before the partition, it catches the resulting increase in lag as Region A drops out of sync and automatically throttles writes that would otherwise land only in the now-unreachable region, preventing a burst of writes that could never actually replicate. At the same time, the rollback in Region B goes through the same signed-artifact, two-person-approval gate a forward deploy would use, specifically because the region failure is already an unusually stressful moment where a shortcut would be most tempting, and that is exactly when the pipeline controls matter most, not a moment to informally suspend them. Because the original deploy used an expand-contract schema pattern, Region B's rollback to the previous application version runs cleanly against the still-present old and new schema fields without a data-compatibility break. When Region A recovers, its replication catches back up against Region B's now-authoritative state, and only once the divergence monitor reports the two regions back within the normal threshold does traffic resume being served from Region A again, rather than resuming immediately and risking a second round of conflicting writes.
Trade-offs and pitfalls
Applying strong, single-region-primary consistency to every data domain would defeat the purpose of an active-active design in the first place, since a partition would then force an explicit choice between availability and consistency for data that did not need that trade-off; the point of choosing the consistency model per domain is that only the data which genuinely cannot tolerate a conflicting write pays that cost. Two-person approval and per-region promotion gates add real friction at exactly the moment speed feels most urgent, during an incident; the design needs a pre-approved emergency path for rolling back to an artifact that was already reviewed once (fast, but still gated) rather than a separate "break glass, skip everything" escape hatch, since a standing bypass of the pipeline controls is itself a long-term vulnerability, not just a convenience. A common pitfall is testing failover drills only against a clean, planned scenario and never against a deploy or rollback happening at the same time; the worked example's compound case is exactly the kind of overlap that isolated drills miss, and it is where the four named threats can compound each other's actual impact rather than staying independent. Finally, tuning the replication-divergence threshold too aggressively generates alert fatigue during ordinary, benign eventual-consistency windows that were never actually a problem, so the threshold needs real tuning against each data domain's own tolerance rather than one number applied uniformly across very different kinds of data.
You are onboarding a new SaaS tenant: describe how you would enumerate assets and the attack surface for their single-tenant web app deployed in AWS. Include cloud resources (compute, storage, IAM), developers' workstations, CI/CD pipelines, third-party integrations, mobile clients, and customers' browsers in your enumeration and explain how the attack surface expands with each asset class.
Sample Answer
Direct answer
Enumerate assets by walking outward from the tenant's data, not just by listing cloud resources: start with the AWS cloud resources actually holding and processing the tenant's data (compute, storage, IAM), then work through everyone and everything with a path to influence that data before it ever reaches the tenant, developer workstations, the CI/CD pipeline, and third-party integrations, and finally the client-side surfaces the tenant's own users interact with, mobile clients and browsers. Each asset class you add is not just "one more item on a list"; it is a genuinely new category of threat the earlier categories did not have, which is the part of this exercise that actually matters for onboarding, not the enumeration itself.
Structured elaboration
Cloud resources: compute, storage, IAM
- Compute (the application servers, containers, or serverless functions running the tenant's workload): each compute resource's own vulnerabilities (unpatched software, exposed management ports, over-permissive network access) are attack surface, and at onboarding time this is also where the boundary between "this tenant's compute" and "any other tenant's" needs to be confirmed, even for a single-tenant deployment, since the surrounding AWS account may host other tenants' infrastructure too.
- Storage (databases, object storage buckets holding uploads or backups): each storage resource is attack surface both for direct access (a misconfigured bucket policy allowing broader access than intended) and for what it reveals if the compute layer above it is compromised (what does a compromised application server's storage access actually permit).
- IAM (AWS Identity and Access Management: the roles and policies granting compute and storage access): this is attack surface in its own right, separate from compute and storage, because an over-broad IAM policy expands what an attacker gains even without needing to find a NEW vulnerability; each role attached to this tenant's resources is effectively part of the attack surface, since a compromise anywhere with that role attached inherits everything the role permits.
Developers' workstations
This asset class expands the attack surface in a way the cloud-resource layer alone does not capture: a developer's laptop, if it holds credentials, source code, or SSH access to any part of this tenant's infrastructure, becomes a path to that infrastructure that never touches the AWS account's own perimeter at all. A phished developer, or a compromised personal device used for work, can bypass every cloud-side control simply by using credentials the developer already legitimately holds.
CI/CD pipeline
The build-and-deploy pipeline expands attack surface differently again: it is the one component with a LEGITIMATE, automated path to modify what actually runs in production, which means compromising the pipeline (a poisoned dependency, a compromised build step) can result in attacker-controlled code being deployed without ever needing to compromise a running production system directly. This asset class is attack surface specifically because of what it is TRUSTED to do, not because of any inherent vulnerability in the pipeline software itself.
Third-party integrations
Every third-party service this tenant's application calls or is called by (webhooks, a payment processor, an analytics SDK, a single sign-on provider) extends the attack surface beyond anything this organization directly controls: a vulnerability or compromise in the third party itself, or in the credentials used to authenticate to it, becomes this tenant's exposure even though the vulnerable code lives outside this AWS account entirely. Onboarding needs an explicit inventory of every such integration, since "we didn't know that integration existed" is a common gap once a tenant's application has accumulated integrations over time.
Mobile clients
If the tenant's users interact via a mobile app, the app itself (and the device it runs on, outside this organization's control entirely) is attack surface: reverse-engineering the app can reveal embedded secrets or API contracts an attacker uses to call backend services directly, bypassing intended client-side logic, and a compromised or rooted user device changes the trust assumptions the backend can safely make about requests claiming to originate from the legitimate app.
Customers' browsers
The furthest-out asset class, and the one most outside this organization's direct control: the customer's own browser, running client-side code this organization shipped (JavaScript) inside an environment (the browser, on the customer's device) this organization does not own. This is attack surface both for direct client-side vulnerabilities (a cross-site-scripting flaw in the shipped code) and because the browser session itself, once authenticated, is a target an attacker can pursue through the customer's device rather than through any part of the infrastructure at all.
Worked example
A concrete illustration of how the attack surface actually widens as each class is added, using a single tenant's document-upload feature as the running example:
- Compute + storage + IAM only: the attack surface is "can someone reach the upload-processing service or the storage bucket directly, and what does the service's IAM role actually permit." A misconfigured bucket policy or an over-broad IAM role are the concrete risks at this layer alone.
- + Developer workstations: now ALSO "can someone phish a developer who has direct AWS console or SSH access to this compute/storage," a path that bypasses every control in step 1 entirely, since it does not go through the application or its AWS-side configuration at all.
- + CI/CD pipeline: now ALSO "can someone get a malicious change into the upload-processing service's next deployment," which, unlike step 2, does not even require compromising a specific person, only the automated pipeline that legitimately pushes code to the exact compute resource from step 1.
- + Third-party integrations: if the upload feature calls a third-party virus-scanning API, now ALSO "can that third party, or the credential used to call it, be compromised or abused to affect what this tenant's upload pipeline does with a file," an exposure that exists even if steps 1-3 are all perfectly secured.
- + Mobile client: if uploads can also happen from a mobile app, now ALSO "can the app be reverse-engineered to reveal the upload API's contract or an embedded credential, letting an attacker call the upload endpoint directly with a crafted request the app's own UI would never construct."
- + Customer browser: if uploads also happen via a web interface, now ALSO "can a cross-site-scripting flaw in the web upload page be used to act as the authenticated customer," an exposure that exists purely in code running on the customer's own device, nowhere in this organization's infrastructure at all.
Six asset classes, six genuinely distinct new categories of "how could this one feature be attacked," none of which subsumes any of the others; that non-overlap is exactly the reason onboarding enumeration needs to walk through all of them explicitly rather than assuming a thorough cloud-resource review alone has covered the attack surface.
Trade-offs and pitfalls
- Stopping the enumeration at cloud resources is the most common shortcut, because compute/storage/IAM is the part most directly under this organization's own control and easiest to scan with automated tooling; the worked example shows each of the other five classes represents a genuinely independent path an attacker can take, not a lower-priority variant of the same risk.
- Treating "single-tenant" as meaning the account itself has no isolation concerns is a subtle but real mistake. Even a single-tenant deployment often shares underlying AWS-account-level resources (a shared VPC, shared IAM boundaries with other workloads in the same account) with other things this organization runs, so the cloud-resource enumeration should not assume tenant isolation is automatically total just because this specific tenant does not share application-level infrastructure with another tenant.
- Third-party integrations accumulate silently over time, and an onboarding-time inventory goes stale unless it is re-checked; a webhook or SDK added six months after onboarding, with nobody updating the original attack-surface enumeration, is a common way this specific asset class's inventory drifts out of date faster than the others.
- Enumerating an asset class is not the same as having assessed it. Listing "mobile client" as an asset class is the first step; it still needs its own actual review (has the app been checked for embedded secrets, does the backend validate requests independent of trusting the app's own client-side logic) before the enumeration translates into an actual reduction in risk.
How does threat modeling change for serverless and cloud-native architectures compared to traditional VM-based designs? Identify unique attack surfaces (e.g., functions, event sources, IAM roles, managed services), and list recommended mitigation patterns and observability practices.
Sample Answer
Direct answer
Threat modeling a serverless or cloud-native system does not change the methodology, STRIDE and trust-boundary mapping still apply, but it changes what counts as an asset and where the boundaries actually sit. A traditional virtual machine (VM)-based design has a small number of coarse-grained trust boundaries (network perimeter, host operating system, application process); a serverless design has many more, finer-grained boundaries, because every function invocation, every event source, and every managed service call is itself a boundary crossing with its own identity and permission set. The four attack surfaces that specifically change are functions (short-lived, individually invokable units of compute), event sources (the triggers that invoke a function, each an entry point an attacker can target), Identity and Access Management (IAM) roles (the permission model, since there is no host to compromise, the permission boundary becomes the primary target), and managed services (databases, queues, storage, each with its own exposed configuration surface instead of being hidden behind an application server you control).
Structured elaboration
Why the boundaries move, not just multiply
In a VM-based design, an attacker who wants to reach a database typically has to compromise the network perimeter, then the host, then the application process, then use the application's own database credentials, a small number of sequential hops. In a serverless design, many of those hops are replaced by direct service-to-service calls authorized purely by IAM policy: a function invoked by an event has no host to compromise at all, so the permission boundary (does this function's role actually need to read this table) becomes the primary line of defense instead of one layer among several. This is the single biggest mental shift: the attack surface moves from "can I get code running here" to "what can already-authorized code reach."
The four attack surfaces, in detail
- Functions: individually invokable, ephemeral units of compute. Because each function typically has its own IAM role, an over-permissioned function (one granted broader access than its actual job needs) is a standing risk even if it is never directly compromised, since any bug in it (for example, unsanitized input passed to a downstream call) inherits the full blast radius of its role.
- Event sources: the triggers that invoke a function (an HTTP API gateway, a message queue, a storage-upload notification, a scheduled timer). Each is a distinct entry point with its own authentication model; a storage-upload trigger, for instance, fires on any object landing in a bucket, so if the bucket accepts uploads from untrusted parties, the function is effectively processing attacker-controlled input by design, not by accident.
- IAM roles: the permission model attached to each function and resource. Because there is no host-level compromise step to slow an attacker down, an overly broad role (wildcard permissions, or permissions scoped to a whole resource type rather than a specific resource) is directly exploitable the moment the function it's attached to has any other weakness.
- Managed services: databases, queues, object storage, and similar services that used to sit behind an application server you fully controlled now expose their own configuration surface directly (bucket policies, queue access policies, database network rules). A misconfiguration here is not hidden behind an application layer; it is the entire perimeter for that resource.
Trust boundaries: what changes when a boundary crosses from on-premises to a public cloud tenant
A useful way to see the shift is to threat-model the same component before and after a move from an on-premises data center to a public cloud tenant, since that move crosses several trust boundaries that used to not exist:
- Spoofing: on-premises, identity is often network-location-based (trusted because it's on the internal network); in a cloud tenant, identity must be explicit (IAM role, service identity, signed request) because network location no longer implies trust.
- Tampering: on-premises, data in transit between internal hosts is sometimes unencrypted under an assumption of a trusted network; crossing into a shared-tenancy cloud environment removes that assumption, so encryption in transit between services becomes mandatory rather than optional.
- Repudiation: on-premises logging is often host-based and can be tampered with by anyone who compromises the host; cloud-native logging (a managed, centralized audit trail) is harder for a compromised function to suppress, since the function itself typically has no permission to modify the log service's records.
- Information disclosure: a managed storage or database service is reachable from outside the traditional network perimeter by design (that is how it is managed), so a misconfigured access policy is directly internet-reachable in a way an on-premises database behind a firewall was not.
- Denial of service: on-premises capacity is fixed and a flood is visibly resource-exhausting; a serverless system auto-scales, so a flood instead becomes a cost-exhaustion and rate-limit problem (an attacker driving invocation counts, and therefore cost, rather than crashing a fixed pool of servers).
- Elevation of privilege: on-premises, privilege escalation often means compromising a host to gain broader network access; in a cloud tenant, it means a function's IAM role being usable to reach further than the function's actual job requires, since the role itself is the privilege boundary.
Recommended mitigation patterns
- Least-privilege IAM per function, scoped to specific resources rather than resource types, so a single function's compromise or misuse has the smallest possible blast radius.
- Explicit input validation at every event source, treating every trigger (API gateway request, queue message, storage-upload event) as untrusted input regardless of where it appears to originate, since the event source is the new perimeter.
- Resource-level policies on every managed service (bucket policies, queue access policies, database network rules) reviewed as part of the same threat model as the code, not as a separate infrastructure concern owned by a different team.
- Rate limiting and budget alerts to convert the denial-of-service and cost-exhaustion risk from an open-ended liability into a bounded one.
- Short-lived, scoped credentials (temporary security tokens rather than long-lived keys) wherever a function needs to call another service, so a leaked credential has a short useful life for an attacker.
Observability practices
- Centralized, tamper-resistant logging across every function and managed service, since there is no single host to install a traditional log agent on; the log aggregation has to be architected in from the start, not bolted on.
- Per-function invocation and permission-usage monitoring, specifically watching for a function exercising permissions it holds but has never used before, which is a strong signal of misuse of an over-permissioned role.
- Distributed tracing across event-driven chains, since a single user action can now trigger a chain of functions across multiple event sources, and a security-relevant anomaly (an unexpected fan-out, an unexpected downstream call) is only visible if the whole chain is traceable, not just each function in isolation.
Worked example
Consider a serverless image-processing pipeline: users upload images to object storage, a storage-upload event triggers a function that processes the file, and the processed result is written to a second storage location. Walking the four attack surfaces: the event source is the object-storage upload trigger, which fires on any file landing in the bucket, so if the upload endpoint is public-facing, the function must treat every uploaded file as untrusted, including its file type, size, and embedded metadata, not just its declared content-type. The function itself needs an IAM role scoped to read from the input bucket and write to the output bucket only, not broad storage access; if the processing library it uses has a known parsing vulnerability, a maliciously crafted image is the delivery mechanism, and the function's narrow role is what limits what that vulnerability can actually reach. The IAM role is the concrete mitigation surface: a role scoped to two specific buckets, rather than "storage:*", means that even a fully compromised function cannot read unrelated data in the account. The managed service boundary is the storage service's own bucket policy: it must reject uploads from outside the expected source and cap object size, since an attacker who can upload arbitrarily large or numerous files can drive both processing cost and, if the function scales without a concurrency limit, a real resource-exhaustion condition. Observability closes the loop: per-invocation monitoring on this function would catch it suddenly attempting to read from a bucket outside its normal two, which is the signal that either the role is misconfigured or the function's logic has been abused in a way the code review missed.
Trade-offs and pitfalls
- The most common wrong turn is threat-modeling a serverless system with a VM-based mental model, focusing on network perimeter and host hardening when there is no host, and missing that the real perimeter has moved to IAM policy and event-source validation.
- Over-permissioning IAM roles "to avoid breaking things during development" and never tightening them later is the single highest-leverage mistake, precisely because the absence of a host-compromise step means the role is the only thing standing between a function's weakness and a wide blast radius.
- Treating managed-service configuration (bucket policies, queue policies) as an infrastructure team's problem separate from the application threat model misses that these configurations are now part of the application's actual security boundary, not background plumbing.
- Under-investing in observability because "serverless has no servers to monitor" is a real trap: the lack of a host does not reduce the need for visibility, it changes what needs visibility, from host-level logs to per-invocation, per-permission, and cross-function trace data, and skipping that investment leaves the denial-of-service and privilege-misuse patterns above effectively invisible until real damage is done.
Unlock Full Question Bank
Get access to all 8 Threat Modeling and Attack Surface Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.