Cloud Security Architecture Questions
Designing and reasoning about the security posture of cloud and hybrid infrastructure: the shared responsibility model, network segmentation and boundary design, multi-account and multi-region security architecture, workload identity as an architectural choice, threat modeling a cloud architecture, cloud-specific attack vectors and mitigations, defense-in-depth control selection, secure cloud deployment patterns, and continuous cloud risk assessment and posture. IAM policy authoring, role/trust-policy mechanics, and secrets/credential lifecycle belong to identity-and-access-management; logging-pipeline design and SIEM/detection-rule engineering belong to security-monitoring-and-detection; encryption-key-management mechanics (KMS/CMK/BYOK) belong to data-protection-and-encryption; compliance-framework mapping (SOC2, PCI-DSS, HIPAA, GDPR) belongs to compliance-frameworks-and-certification-standards. This topic keeps identity, logging, or encryption content only when it is one ingredient inside a genuinely multi-control cloud-hardening question, not as a standalone ask.
Design an enterprise-scale Cloud Security Posture Management (CSPM) approach for dozens of cloud accounts and multiple regions. Cover drift detection, prioritized alerting, automated remediation workflows, integration with ticketing systems, suppression of false positives, onboarding process for new accounts, and metrics to measure policy coverage over time.
Sample Answer
Direct answer
An enterprise-scale Cloud Security Posture Management (CSPM) approach for dozens of accounts across multiple regions has to be designed as a pipeline, not a dashboard: drift detection feeds prioritized alerting, prioritized alerting feeds either automated remediation or a ticket, and the whole system needs a deliberate, low-friction path for onboarding a new account and suppressing a genuine false positive, or the program degrades into noise nobody trusts within a few months.
Structured elaboration
Drift detection. Continuous, not periodic, evaluation of every resource across every account against a shared policy set (a cloud-native tool such as Security Hub/Config, or a third-party CSPM platform aggregating findings centrally); "continuous" means minutes-to-hours between a misconfiguration existing and being detected, not a weekly batch scan, since the detection gap is exactly the window an attacker or an internet-wide scanner can exploit.
Prioritized alerting. Findings are scored by a combination of exploitability (internet-reachable right now) and business impact (does the resource hold sensitive data, is it in a regulated account), using account and resource tagging as the input to that scoring, not a flat severity list; a public bucket in a sandbox account and a public bucket in a production account holding customer data are the same finding type and very different priorities.
Automated remediation workflows. A narrow, explicitly-reviewed list of finding types with an unambiguous, safe fix (re-enabling Block Public Access on a bucket, closing an unrestricted security-group rule) are auto-remediated without waiting for a human; everything else routes to a human, because auto-remediating a finding whose "correct" fix depends on context (a database that might legitimately need broader access for a specific integration) risks causing an outage worse than the finding itself.
Integration with ticketing systems. Every finding that is not auto-remediated becomes a ticket in the organization's existing tracker, automatically assigned to the owning team via resource tagging, with a service-level agreement (SLA) tied to its severity score; a finding that only lives in the CSPM tool's own dashboard is a finding nobody is accountable for closing.
Suppression of false positives. A reviewed, time-boxed suppression mechanism (not a permanent, silent exception) lets a team mark a specific finding as an accepted, understood configuration, with an expiration date forcing periodic re-review; suppression without expiration is how a CSPM program's own dashboard becomes systematically less trustworthy over time, since suppressed findings accumulate and nobody re-checks whether the original justification still holds.
Onboarding process for new accounts. A new account is enrolled into the CSPM program automatically at creation (via the organization's account-vending process, not a manual step someone has to remember), inheriting the same policy baseline every existing account has, with a defined grace period before its findings count against the organization's coverage metrics, giving the new account's team time to remediate an inherited baseline gap before being penalized for someone else's earlier decisions.
Metrics to measure policy coverage over time. Percentage of accounts fully onboarded and reporting (the denominator matters as much as the numerator: an account not yet onboarded is invisible risk, not zero risk), mean time to remediate by severity, count of active (non-expired) suppressions as a signal of program health rather than success, and trend of finding volume per account over time (a declining trend indicates the program is working; a flat or rising trend across a mature program indicates either genuinely new risk or a detection or prioritization gap worth investigating).
Worked example
An organization with 60 AWS accounts across 4 regions rolls out this design. A new account for a recently-acquired subsidiary is created through the existing account-vending process, automatically inheriting Config rules and CSPM policy enrollment at creation, with a 30-day grace period before its findings affect the organization-wide coverage metric. Within the first week, the new account surfaces 40 findings inherited from the subsidiary's prior, less-mature security baseline; of these, 12 match the narrow auto-remediation list (public-access-block gaps, unrestricted security-group rules) and are fixed automatically within minutes of detection, while the remaining 28 become tickets routed to the subsidiary's engineering team by resource tag, each with an SLA based on severity. Three findings are reviewed and suppressed with a 90-day expiration, since they reflect a legitimate, documented integration the acquiring organization's policy set had not previously accounted for; those three will resurface for re-review when the suppression expires, rather than silently disappearing from the program's visibility.
Trade-offs and pitfalls
- A wide auto-remediation list is tempting at scale (dozens of accounts means dozens of times the manual triage burden) and is also the single fastest way to cause a self-inflicted outage. The list needs to stay narrow and reviewed, expanding only after a finding type has demonstrated, over real incidents, that its fix is genuinely unambiguous every time, not merely usually.
- Suppression without an expiration date is the most common way a mature CSPM program's dashboard becomes untrustworthy. A team under time pressure suppresses a finding "for now" with no forcing function to revisit it; six months later, the suppression is still active and nobody remembers why, and the program's coverage metric is quietly overstating its own accuracy.
- Onboarding automation is necessary but not sufficient if the grace period design is wrong. A grace period that is too short punishes a newly-onboarded team for a baseline they inherited and had no time to fix; one that is too long lets a genuinely risky new account sit outside the coverage metric's accountability for longer than is defensible. The right length depends on the organization's realistic remediation velocity, not a fixed industry number.
- Metrics that only count closed findings can be gamed by an account that simply suppresses everything rather than fixing it. Pairing the remediation-velocity metric with the active-suppression-count metric, as in the worked example, is what keeps the incentive pointed at actually fixing findings rather than making the dashboard look clean.
Design a secure VPC architecture in AWS for a three-tier web application (public load balancers, application layer, private database). Describe subnet placement across AZs, route tables, NAT gateways, security groups, bastion/jump host strategy, and where to place private endpoints and logging collectors. Consider both availability and security.
Sample Answer
Direct answer
A secure three-tier VPC (Virtual Private Cloud) design for a public load balancer, an application layer, and a private database puts each tier in its own subnet type, replicated across at least two Availability Zones (AZs) for availability, with the database subnet having no route to the internet at all rather than merely being blocked by a security group, since the network topology itself, not just an access-control rule, should make the database unreachable from outside.
Structured elaboration
flowchart TB
Internet(["Internet"]) --> IGW["Internet gateway"]
IGW --> PubA["Public subnet AZ-a: ALB, NAT GW"]
IGW --> PubB["Public subnet AZ-b: ALB, NAT GW"]
PubA --> AppA["Private app subnet AZ-a"]
PubB --> AppB["Private app subnet AZ-b"]
AppA --> DbA["Private DB subnet AZ-a (isolated, no NAT route)"]
AppB --> DbB["Private DB subnet AZ-b (isolated, no NAT route)"]
DbA -.-> DbB
AppA -->|"VPC endpoint"| KMS[("KMS / Secrets Manager")]
AppB -->|"VPC endpoint"| KMS
Bastion["Bastion / SSM Session Manager"] -.->|"admin access, no inbound SSH from internet"| AppA
AppA --> Flow[("VPC flow logs to log-archive account")]
Subnet placement across AZs. Three subnet tiers (public, private-application, private-database), each replicated in at least two AZs (three, if the workload's availability requirement justifies the added cost), so a single AZ failure does not take down the whole application. Public subnets host only the Application Load Balancer (ALB) and NAT gateways, nothing else, since minimizing what actually sits in a public subnet minimizes what is directly internet-reachable even before any security-group rule is considered.
Route tables. The public subnets' route table sends 0.0.0.0/0 to the Internet Gateway. The private application subnets' route table sends 0.0.0.0/0 to the AZ-local NAT gateway (each AZ's application subnet uses its own AZ's NAT gateway, not a shared one, so a single NAT gateway failure does not take down every AZ's outbound path). The private database subnets' route table has no 0.0.0.0/0 route at all, to either the internet gateway or a NAT gateway, so outbound internet access from the database tier is structurally impossible regardless of any security-group misconfiguration, a route-table-level guarantee, not merely a rule that could be loosened.
NAT gateways. One NAT gateway per AZ (not one shared NAT gateway for the whole VPC), placed in the public subnet of each AZ, giving the application tier outbound internet access (for package updates, external API calls) while remaining unreachable from inbound internet traffic, and avoiding a cross-AZ dependency that a single shared NAT gateway would introduce.
Security groups. The ALB's security group permits inbound 443 from the internet. The application tier's security group permits inbound only from the ALB's security group (by reference, not by CIDR), on the application's specific port. The database's security group permits inbound only from the application tier's security group, on the database's specific port, and nothing else, not even from the bastion host directly.
Bastion/jump host strategy. Prefer a session-manager-based access pattern (such as AWS Systems Manager (SSM) Session Manager) over a traditional bastion host with an open inbound Secure Shell (SSH) port: it requires no inbound security-group rule at all (the connection is initiated outbound from the managed instance to the SSM service), and every session is logged centrally without needing separate bastion-host session-logging infrastructure. Where a traditional bastion is still required for a specific tooling reason, place it in its own small public or dedicated subnet, restrict inbound SSH to a narrow, known administrative CIDR (never 0.0.0.0/0), and require multi-factor authentication (MFA) for any session.
Private endpoints. Traffic from the application tier to cloud-native services it depends on (a secrets manager, a Key Management Service (KMS) key, an object storage bucket) routes through VPC endpoints rather than out through the NAT gateway to the public internet-facing version of those services; this keeps that traffic on the provider's private network backbone entirely and, for gateway-type endpoints such as the one for object storage, removes a real cost (NAT gateway data-processing charges) as well as a security benefit.
Logging collectors. VPC flow logs are enabled on every subnet and shipped to a centralized, separate logging destination (ideally a dedicated log-archive account), not just stored locally in the same account the traffic originated from, so an incident investigation has a trustworthy record independent of whatever happened to the workload account.
Worked example
A concrete Classless Inter-Domain Routing (CIDR) layout for a VPC sized 10.0.0.0/16 across two AZs: public subnets 10.0.0.0/24 (AZ-a) and 10.0.1.0/24 (AZ-b), hosting the ALB and each AZ's own NAT gateway; private application subnets 10.0.10.0/24 (AZ-a) and 10.0.11.0/24 (AZ-b), each routing 0.0.0.0/0 to its own AZ's NAT gateway; private database subnets 10.0.20.0/24 (AZ-a) and 10.0.21.0/24 (AZ-b), with a route table containing only the VPC's local route, no default route at all. The database security group permits inbound on its port only from the application tier's security group ID; the application tier's security group permits inbound only from the ALB's security group ID; and VPC flow logs from every subnet ship continuously to a dedicated log-archive account. An operator needing to inspect an application-tier instance connects via SSM Session Manager, which requires no inbound port open on that instance's security group at all, and the session is centrally logged in the same log-archive account the flow logs already ship to.
Trade-offs and pitfalls
- A shared, single NAT gateway across all AZs is a common cost-saving shortcut that creates an availability trade-off worth stating explicitly. It is cheaper (one NAT gateway instead of one per AZ) but makes every AZ's outbound path depend on one AZ's infrastructure, which contradicts the multi-AZ availability goal the rest of the design otherwise achieves; the per-AZ NAT gateway design above costs more but removes that single point of failure.
- A traditional bastion host, even a well-configured one, is a standing piece of internet-facing attack surface that a session-manager-based approach removes entirely. This is not merely a modernization preference; a bastion's open inbound port, however narrowly scoped by CIDR, is one more thing that can be misconfigured or targeted, and the session-manager approach's zero-inbound-port design is a genuinely stronger default, not just a more convenient one.
- Database subnets with no default route are easy to accidentally break during a later change if someone adds a NAT route "to fix connectivity" without understanding why it was deliberately absent; this control needs to be documented explicitly as intentional, not just configured and left unexplained, so a future engineer does not silently undo the isolation guarantee while trying to fix an unrelated problem.
- VPC endpoints reduce both cost and exposure but need to be added deliberately for every external service the application tier depends on; a design that adds endpoints for the obvious cases (object storage, secrets manager) but misses a less obvious dependency (a specific regional API the application calls) leaves that one dependency's traffic routing out through the NAT gateway to the public internet, a gap that a dependency audit at design time, not just at initial rollout, is needed to catch.
Given a multi-tenant SaaS built on Kubernetes with an RDS backend, run a concise threat modeling exercise: identify top assets, likely entry points (external and internal), three high-risk threat scenarios, and concrete mitigations at network, platform, and application layers. Include residual risk and monitoring recommendations.
Sample Answer
Direct answer
A multi-tenant Software as a Service (SaaS) platform on Kubernetes with a relational database (RDS) backend has one asset that dominates every other consideration: tenant isolation itself, because the single worst outcome in a multi-tenant system is not "data was stolen," it is "tenant A read tenant B's data," which changes what counts as a high-risk scenario compared to a single-tenant threat model.
Structured elaboration
Top assets. Tenant data (the RDS-backed customer records, at both the row and schema level depending on the isolation model chosen), the Kubernetes control plane and its Role-Based Access Control (RBAC) bindings, service credentials (database credentials, Kubernetes secrets, cloud identity and access management (IAM) roles), and the continuous integration/continuous deployment (CI/CD) pipeline that can push code into every tenant's environment at once.
Likely entry points.
| Type | Entry point |
|---|---|
| External | The public API or ingress, a misconfigured or overly permissive database endpoint, the authentication service itself (since a flaw there compromises every tenant simultaneously, not just one) |
| Internal | A compromised pod or its service account pivoting to another tenant's workload on the same cluster, a compromised CI/CD credential pushing malicious code to production, an over-privileged internal tool with cross-tenant database access |
Three high-risk threat scenarios and layered mitigations.
- Tenant-boundary bypass at the data layer. A single shared database with row-level tenant scoping enforced only in application code (not the database engine itself) means one query-construction bug anywhere in the codebase can leak another tenant's rows. Mitigations: enforce tenant isolation at the database layer itself (row-level security policies, or separate schemas/databases per tenant for the highest-sensitivity tenants) as a second, independent layer beneath the application-level check, so an application bug alone is not sufficient to breach isolation.
- Cross-tenant lateral movement inside the cluster. A compromised pod belonging to one tenant's workload (through a vulnerable dependency, for instance) attempts to reach another tenant's pod or database credentials on the same shared cluster. Mitigations: network-layer segmentation (Kubernetes NetworkPolicy scoped per tenant namespace, default-deny), and identity-layer segmentation (a distinct, narrowly-scoped IAM role or database credential per tenant namespace, so a compromised pod's own credentials cannot reach another tenant's data even if network segmentation were somehow bypassed).
- Compromised CI/CD credential deploying to every tenant at once. Because the pipeline can push to the full multi-tenant fleet in one action, a compromised deployment credential is a single point of failure with blast radius across every tenant simultaneously. Mitigations: short-lived, OpenID Connect (OIDC)-federated deployment credentials (never long-lived static keys), mandatory code review and a signed-artifact requirement before any deployment, and a staged rollout (canary a subset of tenants before fleet-wide deployment) that limits how many tenants a single bad or malicious deployment reaches before detection.
Mitigations at network, platform, and application layers.
- Network layer: default-deny NetworkPolicy per tenant namespace, private database endpoint with no public reachability, a service mesh enforcing mutual TLS (mTLS) between services.
- Platform layer: per-tenant-namespace RBAC scoping (no cluster-wide role bound broader than the platform's own operators need), Pod Security Admission at the restricted level, admission-controller-enforced image signing.
- Application layer: tenant-scoping enforced at the data-access layer itself (an object-relational mapping (ORM) layer or query builder that cannot construct a query without an explicit tenant filter, rather than relying on every developer remembering to add one), and database-native row-level security as the second, independent layer described above.
Residual risk and monitoring recommendations. Even with all three layers implemented, a sufficiently sophisticated attacker who compromises a tenant-scoped credential with legitimate access to that tenant's own data retains the ability to exfiltrate that one tenant's data, since no isolation design prevents a credential from doing what it is legitimately scoped to do; this residual risk is addressed by monitoring, not architecture: per-tenant behavioral baselining on data-access volume and pattern (a specific tenant's typical read volume, flagged when a credential scoped to that tenant reads at an anomalous multiple of its baseline), and continuous verification that row-level security policies remain enabled and unmodified, since a database-layer control silently disabled during a maintenance operation removes the second independent layer without anyone necessarily noticing.
Worked example
A vulnerable dependency in one tenant's customization layer gives an attacker code execution inside that tenant's pod. Network-layer NetworkPolicy stops the attacker's attempt to reach another tenant's pod directly, since the default-deny policy only permits traffic to this tenant's own database credential path. The attacker instead attempts to widen the SQL query the compromised pod's application code issues, to read across tenant boundaries within the shared database; the database's own row-level security policy, evaluated independently of the application code that constructed the query, rejects the cross-tenant read regardless of what the (already-compromised) application layer attempted to construct. The attacker is limited to the one tenant's own data, which their compromised access already legitimately reaches, a residual risk the monitoring layer, not the architecture, is responsible for catching through anomalous access-volume detection.
Trade-offs and pitfalls
- Application-layer tenant scoping alone is the single most common point of failure in multi-tenant systems, and it is also the cheapest to implement, which is why teams frequently stop there. The worked example's second, independent database-layer check (row-level security) is what actually prevents a breach when, not if, an application-layer bug eventually occurs; treating application-layer scoping as sufficient on its own is the most consequential shortcut in this threat model.
- Per-tenant namespace isolation on a shared Kubernetes cluster is a real, ongoing operational cost as tenant count grows, since network policy, RBAC, and resource-quota configuration all scale roughly linearly with tenant count; a platform team needs to budget for this cost explicitly rather than discovering it as an unplanned burden once tenant count is already large.
- The residual-risk framing (a legitimately-scoped credential doing legitimate-looking things) is often the gap a threat model skips entirely, because it feels like "not a real vulnerability." It is exactly the scenario a real, patient attacker with any foothold inside one tenant's boundary ends up in, which is why monitoring for anomalous behavior from a valid credential, not just enumerating architectural controls, belongs in the threat model's conclusions.
- A staged CI/CD rollout limits blast radius but adds real deployment latency, and a team under delivery pressure may be tempted to skip the canary stage for an urgent fix; the canary stage's value is highest for exactly the deployments made under the most time pressure, which is also when it is most tempting to skip.
You discover a publicly accessible object storage bucket (e.g., S3/GCS) containing intermediary ETL outputs. Describe immediate remediation steps you would take to secure the bucket, and then list long-term measures to prevent recurrence, focusing on detection, automation, and process changes.
Sample Answer
Direct answer
Discovering a publicly accessible object storage bucket containing intermediary extract-transform-load (ETL) output data means treating the exposure window itself as the first thing to establish (how long has this been public, and was it actually accessed by anyone besides you), then closing the exposure immediately, and only after both of those, building the detection and process changes that stop this specific mistake from recurring silently again.
Structured elaboration
Immediate remediation, in order.
- Determine exposure duration and access history before changing anything, if the tooling to do so exists. Check the bucket's access logs (or, if not enabled, the cloud provider's data-event logging if it happens to be enabled account-wide) for the earliest evidence of the public setting and any actual read activity from outside the organization's own known identities; this determines whether the incident is "a misconfiguration that existed with no evidence of external access" or "a misconfiguration with confirmed external access," which changes the urgency and the notification obligations that follow.
- Remove the public exposure. Enable Block Public Access at the bucket level (and confirm the account-wide default is also enabled, since a bucket-level fix alone does not prevent the next bucket from repeating the same mistake), remove any explicit public-read grant on the bucket's policy or access control list (ACL).
- Rotate anything the exposed data could have compromised. If the ETL output data included any credential, connection string, or token, even as an intermediate artifact never intended to be sensitive on its own, rotate it; intermediary ETL output is easy to underestimate as "just processing data" when it can, in practice, contain exactly this kind of incidentally-sensitive content.
- Preserve evidence before any further remediation step that could overwrite it, a snapshot of the bucket's access logs and the object listing at the time of discovery, since the later detection and process-change work benefits from an accurate record of exactly what was exposed and for how long.
Long-term measures to prevent recurrence.
- Detection: enable a continuous, account-wide Cloud Security Posture Management (CSPM) check specifically for public storage exposure, rather than relying on incident discovery (as happened here) as the detection mechanism; the specific failure mode this incident represents, ETL intermediate output landing in a bucket that was never meant to be public, is exactly the class of drift a continuous check catches within minutes to hours rather than whenever someone happens to notice.
- Automation: enforce Block Public Access as an account-wide, Service Control Policy (SCP)-backed default that new buckets inherit automatically, so the default state for any newly-created bucket, including one an ETL pipeline provisions programmatically without a human directly configuring it, is private, requiring an explicit, reviewed exception to become public rather than an explicit action to become private.
- Process changes: require infrastructure-as-code (IaC) review for any new storage resource an ETL pipeline provisions, with a policy-as-code check specifically flagging a public-access setting before it ever reaches production, catching this exact mistake at review time rather than discovery time.
- Process changes: classify intermediary ETL output explicitly, not just final data products. A common root cause of this specific incident shape is that intermediate, "just processing" data is held to a lower security bar than a finished, customer-facing data product, even though it frequently contains the same underlying sensitive content in a rawer form; treating intermediate output with the same classification discipline as the final product closes the gap that made this bucket a lower-scrutiny target in the first place.
Worked example
The exposed bucket's access logs (enabled, fortunately, though only because of an unrelated organizational logging default) show the public-read setting has existed for 11 days, with three external read requests from an IP range not associated with the organization's own infrastructure or known partners. This confirms actual external access occurred, not merely theoretical exposure, changing the response from "close the gap and move on" to "close the gap, and separately investigate what those three external reads actually retrieved, since that determines whether a data-exposure notification obligation exists." Block Public Access is enabled at both the bucket and the account level within the hour. The ETL output is found to contain, among the intermediate processing data, a database connection string embedded in a debug-logging artifact the pipeline had written alongside its actual output; that credential is rotated immediately, independent of the broader investigation timeline, since a credential exposed for 11 days needs to be treated as potentially compromised regardless of whether the three confirmed external reads specifically retrieved it.
Trade-offs and pitfalls
- Rushing to fix the exposure before checking for evidence of access is an understandable but real mistake, since some remediation actions (deleting the bucket outright, for instance, rather than just changing its access setting) can destroy the very access-log evidence needed to determine whether this was a theoretical or an actual exposure. The correct order (check for evidence, then fix, preserving evidence throughout) matters specifically because it determines the incident's actual severity and legal exposure, not just its technical remediation.
- Treating intermediary ETL output as inherently lower-risk than a finished data product is the root-cause pattern behind this entire incident shape, and it is easy to reintroduce even after this specific bucket is fixed if the underlying classification discipline is not applied to every other intermediate-data location the same pipeline (or other pipelines) uses; fixing this one bucket without addressing the classification gap leaves the same mistake likely to recur in a different bucket.
- A CSPM check that only flags a bucket as "public" without distinguishing intentionally-public from accidentally-public content generates enough noise that a team may tune it down or ignore it over time, the same alert-fatigue risk present in any detection program; an explicit, reviewed allow-list of genuinely-intended-public buckets keeps this specific detection control credible and actionable.
- A policy-as-code gate on new IaC-provisioned storage resources does not, by itself, catch a resource an ETL pipeline creates dynamically at runtime rather than through a reviewed IaC deployment, a real gap if the pipeline's own code, not a Terraform module, is what provisions intermediate storage locations; the account-wide SCP-backed default (private unless explicitly and reviewedly made public) is the layer that catches this specific gap, which is exactly why both the IaC-review control and the account-wide default are both needed, not either alone.
Perform a threat modeling exercise for a given public web application that accepts file uploads and processes them in serverless functions. Use the STRIDE categories to identify top threats, then prioritize them by likelihood and impact and propose mitigations focusing on architectural changes a solutions architect should recommend.
Sample Answer
Direct answer
A public web application that accepts file uploads and processes them in serverless functions maps cleanly onto STRIDE (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege), and the highest-priority findings concentrate specifically on Tampering and Elevation of privilege, since an untrusted file is, by definition, attacker-controlled content reaching a processing function, which is exactly the shape of threat those two categories describe; the mitigations a solutions architect should recommend are architectural (network and identity boundaries), not code-level fixes the architecture review itself cannot verify.
Structured elaboration
Spoofing. An attacker impersonates a legitimate user to upload a file under someone else's identity, or spoofs the upload event itself to trigger processing without a genuine upload having occurred. Likelihood: medium (requires either a stolen credential or a flaw in the upload-authorization flow); impact: medium (primarily an attribution and audit-trail problem, unless combined with a Tampering finding). Mitigation: strong, short-lived upload-authorization tokens (a pre-signed URL scoped to one specific object key and a short expiration) rather than a broadly-reusable upload credential, and event-source validation in the processing function confirming the triggering event genuinely originated from the expected storage location, not an event a caller crafted directly.
Tampering. The uploaded file's content or metadata (filename, declared content type) is attacker-controlled and used unsafely by the processing function, the highest-priority finding in this threat model. Likelihood: high (this is the architecture's primary attacker-reachable surface); impact: high (can range from a processing function crash to remote code execution, depending on how the function parses the file). Mitigation: validate file type by content inspection, not by trusting the client-declared type or the filename's extension; never construct a file path, a shell command, or a downstream query using the filename or any other attacker-controlled metadata without strict validation first; and run the actual file-parsing logic in an isolated, minimally-privileged execution context.
Repudiation. Without sufficient logging, neither the platform nor the uploading user can later prove or disprove that a specific upload occurred, or what a processing function did with it. Likelihood: high if logging is not deliberately designed in; impact: low to medium on its own, but it compounds every other finding's investigability. Mitigation: log the upload event, the authenticated uploader's identity, and every stage of processing with enough detail to reconstruct what happened to a specific file, shipped to a centralized, tamper-resistant log destination.
Information disclosure. A processing function with an execution role broader than its actual function requires can read more data (other users' uploaded files, unrelated internal resources) than the specific upload it was invoked for; separately, an error message or a debug log accidentally including file content could leak sensitive data uploaded by one user to an operator with no legitimate need to see it. Likelihood: medium; impact: high, since this is a direct data-exposure path. Mitigation: per-function least-privilege execution roles scoped to only the specific object the triggering event names, and structured logging that explicitly excludes file content from log output.
Denial of service. A maliciously crafted or oversized file exhausts the processing function's memory, execution time, or downstream storage, or a flood of upload requests exhausts the platform's processing capacity. Likelihood: medium; impact: medium (availability, not data exposure). Mitigation: enforce a maximum file size before the file is even fully accepted, set function-level timeout and memory limits appropriate to legitimate file sizes, and rate-limit the upload endpoint itself.
Elevation of privilege. A compromised processing function (through a successfully exploited Tampering vulnerability) uses its own execution role to reach further than the immediate file it was invoked to process, the second-highest-priority finding, since it is the direct consequence of a successful Tampering attack turning into broader account access. Likelihood: medium (requires a prior successful Tampering exploit as the precondition); impact: high (turns a single-file compromise into a broader account compromise). Mitigation: the same per-function least-privilege role scoping named under Information disclosure, which is the single control doing the most work across both categories.
Prioritization by likelihood and impact
Tampering is prioritized highest (high likelihood, high impact, and the entry point every other high-impact finding in this model depends on). Elevation of privilege is prioritized second, specifically because it is what determines how bad a successful Tampering exploit actually becomes, the multiplier effect named in the elaboration above. Information disclosure and Denial of service follow as independently medium-to-high priority findings. Spoofing and Repudiation, while genuine findings, are lower standalone priority, since their impact is largely contingent on or compounds one of the other categories rather than being independently severe.
Architectural mitigations a solutions architect should recommend
Least-privilege, per-function execution roles (the single highest-leverage architectural control here, addressing both Information disclosure and Elevation of privilege at once); content-based file-type validation happening in a dedicated, isolated validation step before any business-logic processing touches the file; short-lived, narrowly-scoped upload authorization; centralized, tamper-resistant logging covering the full upload-to-processing lifecycle; and explicit size, timeout, and rate limits enforced at the platform edge, not left to the processing function's own default behavior.
Worked example
A file-sharing application's processing function extracts metadata from uploaded documents and stores the results in a database. An attacker uploads a file with a crafted filename containing a path-traversal sequence, exploiting the processing function's unsanitized use of that filename to write its output somewhere outside the intended location, a direct Tampering exploit. Because the function's execution role happens to be scoped broadly (shared across several processing functions "for simplicity"), the attacker's crafted output path lands in a location the function's role can write to, but that a properly-scoped, per-function role would not have permitted, turning a single Tampering finding into an Elevation-of-privilege finding as well. The architectural fix recommended is not a single patch to this one function's filename handling (a code-level fix outside this review's own scope, though also necessary), but the broader architectural correction: per-function role scoping across every processing function in the pipeline, so the next Tampering vulnerability discovered in a different function does not have the same broader-than-necessary blast radius this one did.
Trade-offs and pitfalls
- A threat model that stops at listing STRIDE categories independently, without tracing how a Tampering finding becomes an Elevation-of-privilege finding once it succeeds, misses the compounding relationship that actually determines real-world severity, exactly what the worked example demonstrates directly.
- A solutions architect's recommendations need to stay at the architectural level (role scoping, isolation boundaries, platform-level limits) rather than prescribing a specific code fix for a specific function, since the architectural review's own scope and expertise is the system's structure, not auditing every function's internal code; the worked example's filename-handling bug still needs a code fix, but the review's own deliverable is the broader per-function-role-scoping recommendation that limits the next such bug's impact too.
- Repudiation and Spoofing are genuinely lower standalone priority, and that ranking can be mistaken for "not worth fixing," when actually their value is specifically in supporting investigation of the higher-priority findings; without adequate logging (addressing Repudiation), an actual Tampering exploit in production is far harder to detect and investigate after the fact, even though Repudiation itself was ranked lower.
- A shared execution role "for simplicity" across multiple processing functions, as in the worked example, is a common, well-intentioned shortcut that directly converts what should be an isolated, single-function compromise into an account-wide risk; the cost of per-function role authoring is real but is precisely what the highest-priority finding in this model depends on to stay contained.
Unlock Full Question Bank
Get access to all Cloud Security Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.