Cloud Security Architecture Questions
Designing and reasoning about the security posture of cloud and hybrid infrastructure: the shared responsibility model, network segmentation and boundary design, multi-account and multi-region security architecture, workload identity as an architectural choice, threat modeling a cloud architecture, cloud-specific attack vectors and mitigations, defense-in-depth control selection, secure cloud deployment patterns, and continuous cloud risk assessment and posture. IAM policy authoring, role/trust-policy mechanics, and secrets/credential lifecycle belong to identity-and-access-management; logging-pipeline design and SIEM/detection-rule engineering belong to security-monitoring-and-detection; encryption-key-management mechanics (KMS/CMK/BYOK) belong to data-protection-and-encryption; compliance-framework mapping (SOC2, PCI-DSS, HIPAA, GDPR) belongs to compliance-frameworks-and-certification-standards. This topic keeps identity, logging, or encryption content only when it is one ingredient inside a genuinely multi-control cloud-hardening question, not as a standalone ask.
Perform a threat modeling exercise for a given public web application that accepts file uploads and processes them in serverless functions. Use the STRIDE categories to identify top threats, then prioritize them by likelihood and impact and propose mitigations focusing on architectural changes a solutions architect should recommend.
Sample Answer
Direct answer
A public web application that accepts file uploads and processes them in serverless functions maps cleanly onto STRIDE (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege), and the highest-priority findings concentrate specifically on Tampering and Elevation of privilege, since an untrusted file is, by definition, attacker-controlled content reaching a processing function, which is exactly the shape of threat those two categories describe; the mitigations a solutions architect should recommend are architectural (network and identity boundaries), not code-level fixes the architecture review itself cannot verify.
Structured elaboration
Spoofing. An attacker impersonates a legitimate user to upload a file under someone else's identity, or spoofs the upload event itself to trigger processing without a genuine upload having occurred. Likelihood: medium (requires either a stolen credential or a flaw in the upload-authorization flow); impact: medium (primarily an attribution and audit-trail problem, unless combined with a Tampering finding). Mitigation: strong, short-lived upload-authorization tokens (a pre-signed URL scoped to one specific object key and a short expiration) rather than a broadly-reusable upload credential, and event-source validation in the processing function confirming the triggering event genuinely originated from the expected storage location, not an event a caller crafted directly.
Tampering. The uploaded file's content or metadata (filename, declared content type) is attacker-controlled and used unsafely by the processing function, the highest-priority finding in this threat model. Likelihood: high (this is the architecture's primary attacker-reachable surface); impact: high (can range from a processing function crash to remote code execution, depending on how the function parses the file). Mitigation: validate file type by content inspection, not by trusting the client-declared type or the filename's extension; never construct a file path, a shell command, or a downstream query using the filename or any other attacker-controlled metadata without strict validation first; and run the actual file-parsing logic in an isolated, minimally-privileged execution context.
Repudiation. Without sufficient logging, neither the platform nor the uploading user can later prove or disprove that a specific upload occurred, or what a processing function did with it. Likelihood: high if logging is not deliberately designed in; impact: low to medium on its own, but it compounds every other finding's investigability. Mitigation: log the upload event, the authenticated uploader's identity, and every stage of processing with enough detail to reconstruct what happened to a specific file, shipped to a centralized, tamper-resistant log destination.
Information disclosure. A processing function with an execution role broader than its actual function requires can read more data (other users' uploaded files, unrelated internal resources) than the specific upload it was invoked for; separately, an error message or a debug log accidentally including file content could leak sensitive data uploaded by one user to an operator with no legitimate need to see it. Likelihood: medium; impact: high, since this is a direct data-exposure path. Mitigation: per-function least-privilege execution roles scoped to only the specific object the triggering event names, and structured logging that explicitly excludes file content from log output.
Denial of service. A maliciously crafted or oversized file exhausts the processing function's memory, execution time, or downstream storage, or a flood of upload requests exhausts the platform's processing capacity. Likelihood: medium; impact: medium (availability, not data exposure). Mitigation: enforce a maximum file size before the file is even fully accepted, set function-level timeout and memory limits appropriate to legitimate file sizes, and rate-limit the upload endpoint itself.
Elevation of privilege. A compromised processing function (through a successfully exploited Tampering vulnerability) uses its own execution role to reach further than the immediate file it was invoked to process, the second-highest-priority finding, since it is the direct consequence of a successful Tampering attack turning into broader account access. Likelihood: medium (requires a prior successful Tampering exploit as the precondition); impact: high (turns a single-file compromise into a broader account compromise). Mitigation: the same per-function least-privilege role scoping named under Information disclosure, which is the single control doing the most work across both categories.
Prioritization by likelihood and impact
Tampering is prioritized highest (high likelihood, high impact, and the entry point every other high-impact finding in this model depends on). Elevation of privilege is prioritized second, specifically because it is what determines how bad a successful Tampering exploit actually becomes, the multiplier effect named in the elaboration above. Information disclosure and Denial of service follow as independently medium-to-high priority findings. Spoofing and Repudiation, while genuine findings, are lower standalone priority, since their impact is largely contingent on or compounds one of the other categories rather than being independently severe.
Architectural mitigations a solutions architect should recommend
Least-privilege, per-function execution roles (the single highest-leverage architectural control here, addressing both Information disclosure and Elevation of privilege at once); content-based file-type validation happening in a dedicated, isolated validation step before any business-logic processing touches the file; short-lived, narrowly-scoped upload authorization; centralized, tamper-resistant logging covering the full upload-to-processing lifecycle; and explicit size, timeout, and rate limits enforced at the platform edge, not left to the processing function's own default behavior.
Worked example
A file-sharing application's processing function extracts metadata from uploaded documents and stores the results in a database. An attacker uploads a file with a crafted filename containing a path-traversal sequence, exploiting the processing function's unsanitized use of that filename to write its output somewhere outside the intended location, a direct Tampering exploit. Because the function's execution role happens to be scoped broadly (shared across several processing functions "for simplicity"), the attacker's crafted output path lands in a location the function's role can write to, but that a properly-scoped, per-function role would not have permitted, turning a single Tampering finding into an Elevation-of-privilege finding as well. The architectural fix recommended is not a single patch to this one function's filename handling (a code-level fix outside this review's own scope, though also necessary), but the broader architectural correction: per-function role scoping across every processing function in the pipeline, so the next Tampering vulnerability discovered in a different function does not have the same broader-than-necessary blast radius this one did.
Trade-offs and pitfalls
- A threat model that stops at listing STRIDE categories independently, without tracing how a Tampering finding becomes an Elevation-of-privilege finding once it succeeds, misses the compounding relationship that actually determines real-world severity, exactly what the worked example demonstrates directly.
- A solutions architect's recommendations need to stay at the architectural level (role scoping, isolation boundaries, platform-level limits) rather than prescribing a specific code fix for a specific function, since the architectural review's own scope and expertise is the system's structure, not auditing every function's internal code; the worked example's filename-handling bug still needs a code fix, but the review's own deliverable is the broader per-function-role-scoping recommendation that limits the next such bug's impact too.
- Repudiation and Spoofing are genuinely lower standalone priority, and that ranking can be mistaken for "not worth fixing," when actually their value is specifically in supporting investigation of the higher-priority findings; without adequate logging (addressing Repudiation), an actual Tampering exploit in production is far harder to detect and investigate after the fact, even though Repudiation itself was ranked lower.
- A shared execution role "for simplicity" across multiple processing functions, as in the worked example, is a common, well-intentioned shortcut that directly converts what should be an isolated, single-function compromise into an account-wide risk; the cost of per-function role authoring is real but is precisely what the highest-priority finding in this model depends on to stay contained.
Design a secure network segmentation strategy for a multi-account cloud environment that hosts public web front-ends, internal application services, and sensitive databases. Explain the roles and differences between security groups (or NSGs), network ACLs, cloud firewalls, and centralized WAF/proxy. Describe how you would use subnetting, route tables, transit gateways, and flow logs to prevent lateral movement and support incident investigations.
Sample Answer
Direct answer
A multi-account segmentation strategy for a public web front-end, internal application services, and sensitive databases needs two things working together: account-level separation (so a compromise cannot cross the account boundary through identity and access management (IAM) alone) and, within that, a consistent set of network-layer controls, security groups (or Network Security Groups (NSGs)), network access control lists (NACLs), a cloud firewall, and a centralized web application firewall (WAF) or proxy, each doing a genuinely different job so that lateral movement is stopped at multiple independent points and an incident investigation has the flow-log evidence to reconstruct exactly what happened.
Structured elaboration
Roles and differences between the four control types.
| Control | Scope | Statefulness | Primary role in this design |
|---|---|---|---|
| Security groups / NSGs | Per-instance or per-resource | Stateful | The fine-grained east-west control: which specific service may reach which other specific service, referenced by group ID rather than IP range |
| Network ACLs | Per-subnet | Stateless | The coarse, subnet-wide guardrail: broad allow/deny by Classless Inter-Domain Routing (CIDR) range at the network edge of each tier |
| Cloud firewall (a managed network firewall service, or an equivalent inspection appliance) | Per-VPC (Virtual Private Cloud) or centralized in a hub | Typically stateful, content-aware for some offerings | Deep inspection of traffic crossing account or region boundaries, and enforcement of organization-wide egress policy (blocking known-bad destinations, for instance) that individual account teams should not need to reimplement themselves |
| Centralized WAF/proxy | Fronting the public tier specifically | Application-layer aware | The only layer inspecting actual HTTP request content, catching an application-layer attack (injection, malformed payload) the other three structurally cannot see |
Subnetting, route tables, and transit gateways for lateral-movement prevention. Each trust tier, public web front-end, internal application services, sensitive databases, sits in its own subnet type, replicated within each workload account; the database tier's subnet has no default route to the internet at all, a routing-layer guarantee independent of any security-group configuration. Cross-account connectivity (an application-tier service in one account legitimately needing to reach a shared service in another account) routes through a transit gateway, which becomes the single, auditable chokepoint for all inter-account traffic, rather than direct account-to-account VPC peering relationships that would each need to be individually tracked and reviewed as the number of accounts grows.
Flow logs for lateral-movement prevention and incident investigation. Virtual Private Cloud (VPC) flow logs, enabled on every subnet across every account, capture connection-level metadata (source, destination, port, bytes, accept/reject) and ship continuously to a centralized, separate log-archive account that the originating accounts themselves have no delete access to. This serves two distinct purposes: as a near-real-time input to lateral-movement detection (an unexpected flow between two accounts that the transit gateway's routing should not have permitted, or an unusual volume between the application and database tiers), and as the forensic record an incident investigation depends on after the fact, one that remains trustworthy even if the account where the incident occurred is itself compromised.
Worked example
A compromised instance in the public web-tier account attempts to reach the sensitive-database account directly. Because the two accounts have no direct peering relationship, only a transit-gateway attachment each with its own explicit route table, the attempted connection has no path to the database account at all, it is rejected at the routing layer before any security group or NACL is even evaluated. The attacker's activity is nonetheless visible: the attempted connection (and its rejection) appears in the web-tier account's own flow logs, already streaming continuously to the centralized log-archive account, giving the incident-response team a record of the lateral-movement attempt independent of what the attacker does next inside the still-compromised web-tier account, including any attempt to disable that account's own local logging configuration.
Trade-offs and pitfalls
- Direct VPC peering between accounts, added ad hoc as specific integration needs arise, is the most common way this design's transit-gateway chokepoint benefit erodes over time. Each individual peering relationship might be reasonably justified on its own, but the accumulated effect is a set of undocumented, hard-to-audit direct paths that bypass the single auditable chokepoint the transit-gateway design was built around; new cross-account connectivity needs should route through the transit gateway by policy, not by convenience.
- The cloud firewall and the centralized WAF address different layers and are easy to conflate as redundant. The cloud firewall inspects and enforces policy on network-layer traffic crossing account or region boundaries; the WAF inspects application-layer request content at the public tier specifically. Treating one as covering the other's job leaves a real gap, an application-layer attack the cloud firewall cannot see, or an unauthorized cross-account network flow the WAF, sitting only at the public edge, never observes.
- Flow logs shipped to a centralized account only deliver their forensic value if that centralized account's own logs cannot be deleted or modified by the accounts that generated them, the same immutability principle that makes centralized logging trustworthy elsewhere in multi-account design; a log-archive account whose retention policy permits deletion by a sufficiently privileged principal in the source account undermines the worked example's core claim that the investigation remains possible even if the source account is compromised.
- A common wrong turn is treating account-level separation alone as sufficient and under-investing in the network-layer controls within each account, on the reasoning that "the account boundary already protects us." The account boundary limits IAM-based blast radius specifically; it does nothing to stop lateral movement between the application and database tiers within the same account if the subnet-level and instance-level controls inside that account are weak.
Explain the shared responsibility model in cloud computing. For each service model (IaaS, PaaS, SaaS) describe which security controls are typically the provider's responsibility and which are the customer's. Provide concrete examples (for example, EC2, RDS, and Gmail), describe a couple of common gray-area responsibilities, and explain how you would document responsibility boundaries for a new cloud service onboarding.
Sample Answer
Direct answer
Across Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS), the shared responsibility line moves in one direction only, toward the provider, as the service becomes more managed, but the customer's core responsibility, identity, access control, and what data goes into the service, never fully disappears at any point on that spectrum, which is exactly what makes the gray-area cases (not the clear-cut ones) the actual place organizations get this wrong.
Structured elaboration
IaaS, example: Amazon Elastic Compute Cloud (EC2). Provider responsibility: physical datacenters, the hypervisor, the host's own network infrastructure. Customer responsibility: guest operating system patching, network configuration (security groups, subnet placement), identity and access management (IAM) for who can access the instance, and everything running on top of the OS, since the customer chose and controls essentially the entire software stack above the hypervisor.
PaaS, example: Amazon Relational Database Service (RDS). Provider responsibility: the physical infrastructure, the hypervisor, and now also the database engine's own patching and the underlying operating system, since the customer never has direct OS-level access to a managed database instance. Customer responsibility: network access configuration (is the instance public, what security group scopes reachability), authentication method and credential management, encryption configuration, and the data itself, a narrower slice than IaaS but still real and still the customer's to get right.
SaaS, example: Gmail (as a representative business-email SaaS product). Provider responsibility: essentially the entire technical stack, the application itself, its infrastructure, its own security patching, with the customer having no direct infrastructure-level control at all. Customer responsibility: narrows to identity and access configuration (who has an account, multi-factor authentication (MFA) enforcement, single sign-on (SSO) integration), data governance (what data users choose to put into the product, sharing and retention settings the product exposes), and user behavior (a phished user credential is still the customer's problem to prevent and respond to, regardless of how secure the underlying SaaS platform itself is).
Common gray-area responsibilities
Encryption key management on a managed service. Whether the customer is responsible for key management depends entirely on whether they opted into a customer-managed key or accepted the provider's default key; this is nominally the customer's choice, but many organizations never make it deliberately, defaulting to whatever the console's default happens to be, which means the actual responsibility boundary in practice is often determined by an unexamined default rather than a considered decision.
Default configuration values. When a managed service is deployed with an insecure default (a database provisioned with a public endpoint enabled by default, for instance, in certain configurations), the provider technically offered the configuration option, but the customer is universally held accountable for the actual deployed state in every compliance framework and every real-world incident response; "the default was insecure" is not a defense, which is a gray area in principle but a settled question in practice, worth stating plainly because it is so commonly misunderstood.
Documenting responsibility boundaries for a new cloud service onboarding
Before onboarding any new managed service, produce a written responsibility matrix specific to that service (not a generic, one-size-fits-all shared-responsibility document), explicitly naming which of the customer's own controls apply (network configuration, IAM, encryption key choice, data classification) and confirming a specific, named owner for each; review the provider's own documented responsibility boundary for that specific service, since it varies by service even within one provider, rather than assuming it matches a different service the organization has already onboarded; and require this documented matrix as a gate in the onboarding process itself, not an afterthought produced once the service is already in production use.
Worked example
An organization onboarding Gmail as its business email platform for the first time documents its responsibility matrix explicitly: MFA enforcement and SSO integration are named, owned controls (assigned to the identity team), data-loss-prevention policy configuration for what leaves the organization via email is a named, owned control (assigned to the security team), and user security-awareness training for phishing resistance is a named, owned control (assigned to the security-awareness program), each with a specific owner rather than an implicit assumption that "Google handles security." Six months later, a user falls for a phishing email and enters their credentials on a fake login page; because MFA enforcement was a documented, owned control that had actually been implemented, the stolen password alone is insufficient for the attacker to access the account, the concrete payoff of having named and implemented that specific customer-side responsibility rather than assuming the SaaS provider's own security covered it.
Trade-offs and pitfalls
- "It's SaaS, the provider handles security" is the single most common and most consequential misunderstanding of this entire model, and it is exactly backwards about where the narrowing happens: the technical infrastructure responsibility narrows toward the provider as services become more managed, but the identity, access, and data-governance responsibility never narrows to zero, and SaaS is precisely where that narrower-but-real slice matters most, since it is often the only thing standing between a phished credential and an actual breach.
- The encryption-key-management gray area is a case where "the customer could have chosen differently" is technically true and practically misleading, since most organizations never make an active, considered choice about it; treating this as a settled customer responsibility without acknowledging how often it is an unexamined default is an honest gap worth naming rather than glossing over.
- A generic, one-size-fits-all shared-responsibility document produced once and reused for every new service onboarding misses that the actual boundary varies by specific service, even within the same provider; the worked example's Gmail-specific matrix would look meaningfully different from an EC2-specific or RDS-specific one, and treating them as interchangeable templates undermines the entire point of documenting the boundary explicitly.
- A responsibility matrix that names an owner on paper but is never actually implemented (MFA "assigned" to the identity team but never actually enforced, for instance) provides no real protection, only the appearance of it; the worked example's payoff depends specifically on the named control having been genuinely implemented, not merely documented as someone's responsibility.
You are reviewing an infrastructure-as-code repository (Terraform/CloudFormation) for a production cloud environment. List the most common high-risk misconfigurations you would look for across IAM, storage, networking, and compute. Also explain how you would automate detection of these misconfigurations in CI/CD before changes reach production.
Sample Answer
Direct answer
A production infrastructure-as-code (IaC) review should walk the same four surfaces every time: identity and access management (IAM), storage, networking, and compute, because that is the order in which a real breach typically chains (a broad credential finds an open door, and an open door leads to unencrypted data). Catching these at review time is necessary but not sufficient: the same checks need to run automatically on every pull request, not just when a human remembers to look.
Structured elaboration
| Category | High-risk misconfiguration to look for | Why it matters |
|---|---|---|
| IAM | Wildcard actions or resources ("Action": "*", "Resource": "*"); trust policies allowing AssumeRole from "Principal": "*"; missing multi-factor authentication (MFA) condition on privileged roles; static long-lived access keys checked into the repository | Turns any single compromised identity into an account-wide foothold |
| Storage | Bucket ACLs or policies allowing public read/write; Block Public Access disabled; missing default encryption; missing versioning on buckets holding data that must survive accidental deletion | Direct data exposure or irreversible data loss, often the visible symptom of a breach even when the entry point was elsewhere |
| Networking | Security groups with ingress from 0.0.0.0/0 on management ports (22, 3389); overly permissive network access control lists (NACLs); VPC (Virtual Private Cloud) flow logs disabled; resources placed in a public subnet with no clear reason | Widens the reachable surface for the earlier two categories to be exploited from the internet |
| Compute | Instances with unnecessary public IPs; hard-coded credentials in user-data or launch templates; instance metadata service left without http_tokens = "required" (IMDSv2 (Instance Metadata Service version 2) not enforced); containers configured to run as a privileged user | Gives an attacker who reaches a compute resource a path to the temporary IAM credentials or secrets on that host |
Worked example
The Terraform snippet below intentionally seeds one misconfiguration from each category; it is syntax-valid HashiCorp Configuration Language (HCL) against the AWS provider (validated with terraform validate), and each block is annotated with what a reviewer, or an automated check, should flag:
# IAM: wildcard action + wildcard resource
resource "aws_iam_policy" "too_wide" {
name = "app-policy"
policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = "*" # flag: no wildcard actions
Resource = "*" # flag: no wildcard resources
}]
})
}
# Storage: bucket has no Block Public Access resource attached
resource "aws_s3_bucket" "app_data" {
bucket = "example-app-data-bucket"
# flag: missing aws_s3_bucket_public_access_block, missing
# aws_s3_bucket_server_side_encryption_configuration
}
# Networking: SSH open to the entire internet
resource "aws_security_group" "app" {
name = "app-sg"
ingress {
from_port = 22
to_port = 22
protocol = "tcp"
cidr_blocks = ["0.0.0.0/0"] # flag: should be a bastion/VPN CIDR only
}
}
# Compute: IMDSv2 not enforced (http_tokens left at its default "optional")
resource "aws_instance" "app" {
ami = "ami-0123456789abcdef0"
instance_type = "t3.micro"
metadata_options {
http_endpoint = "enabled"
# flag: http_tokens = "required" is missing
}
}
Automating detection in CI/CD
- Static analysis on every pull request. A policy-as-code scanner (Checkov, tfsec, or Terrascan) runs against the raw HCL before a human ever reviews it, catching exactly the four patterns above by pattern-matching the resource configuration, not by executing anything.
- Plan-time policy gate. Beyond static patterns, evaluate the actual
terraform planJSON output with a policy engine (Open Policy Agent (OPA)/Conftest, or a managed equivalent such as HashiCorp Sentinel) so the check sees the fully resolved configuration, including values coming from variables or modules that a purely static scan might miss. - Secret scanning on the IaC repository itself. A tool such as gitleaks or truffleHog run in the same pipeline catches a credential accidentally committed into a
.tffile or aterraform.tfvars. - Post-deploy drift detection. A pipeline gate only catches what goes through the pipeline; a scheduled cloud-native check (AWS Config managed rules, or an equivalent cloud security posture management (CSPM) tool) catches the same four categories of misconfiguration when they are introduced through the console instead of through IaC.
Trade-offs and pitfalls
- A pipeline gate that only runs against production IaC misses the source of the problem. Misconfigurations are frequently written first in a development or staging module and then copied into production later; the same static and policy-as-code checks need to run against every environment's plan, not just the one that matters most.
- Over-strict gates get bypassed. If a policy-as-code check blocks a legitimate, reviewed exception (a genuinely public documentation bucket, for instance) with no override path, teams learn to route around the pipeline instead of fixing the finding; a documented, time-boxed exception mechanism keeps the gate credible.
- Static analysis alone cannot see values resolved at plan time from a module or a data source, which is why the plan-time policy gate is a separate, necessary layer rather than a duplicate of the static scan.
List and explain ten common cloud misconfigurations that frequently lead to breaches or data exposure across AWS, Azure, and GCP (for example: open storage buckets, overly permissive IAM policies, public database endpoints, default credentials). For each misconfiguration briefly state how you would detect it and the primary remediation step.
Sample Answer
Direct answer
Most cloud breaches trace back to a small, repeatable set of configuration mistakes rather than novel exploits: open storage, over-broad identity and access management (IAM), exposed management interfaces, and defaults left unchanged. The ten below are the ones that recur across AWS, Azure, and GCP specifically, generalized so the same mental checklist applies regardless of which provider is in front of you.
Structured elaboration
| # | Misconfiguration | Detection | Primary remediation |
|---|---|---|---|
| 1 | Publicly readable/writable object storage (S3 buckets, Azure Blob containers, GCS buckets) | Cloud security posture management (CSPM) inventory scan checking public ACLs (access control lists) and bucket/container policies | Enable Block Public Access (or the provider equivalent) as an account-wide default, not a per-bucket opt-in |
| 2 | Overly permissive IAM policies (wildcard actions/resources) | Policy analyzer or access-analysis tooling scanning for "*" in the action or resource fields of attached policies | Replace with a least-privilege policy scoped to actual usage, validated against real call history |
| 3 | Publicly reachable database endpoints (RDS/Cloud SQL/Azure SQL with a public IP and open security group) | Network configuration scan cross-referencing public IP assignment with the database's security group or firewall rule | Move the database to a private subnet with no public IP, and restrict access to the application tier's security group only |
| 4 | Default or unrotated credentials (default service account keys, unrotated root/admin keys) | Credential-age reporting (IAM credential report on AWS, equivalent identity audit on GCP/Azure) | Rotate immediately, then enforce a maximum credential age with automated rotation for anything long-lived |
| 5 | Management ports open to the internet (SSH/22, RDP (Remote Desktop Protocol)/3389 reachable from 0.0.0.0/0) | Security group/network security group audit for ingress rules with an unrestricted source CIDR on those ports | Restrict to a bastion host, VPN (Virtual Private Network) range, or a just-in-time access mechanism, never a direct internet-wide allow |
| 6 | Missing encryption at rest on storage or database resources | Configuration scan checking each resource's encryption setting against policy | Enable default encryption (provider-managed or customer-managed key) at the account or organization level so new resources inherit it |
| 7 | Overly permissive cross-account or cross-tenant trust relationships | IAM trust-policy scan for Principal values referencing an external account without a condition (such as an external ID) | Add an external ID condition and scope the trust to only the specific external principal that needs it |
| 8 | Logging or audit trail disabled or not centralized (CloudTrail/Activity Log/Cloud Audit Logs turned off or not shipped to a separate account) | Organization-level check confirming every account or subscription has logging enabled and forwarding to a dedicated log-archive destination | Enable organization-wide logging as a baseline requirement, enforced by policy, not left to each team to configure |
| 9 | Excessive network access between environments (no segmentation between development, staging, and production networks) | Network topology review confirming routing and security group boundaries actually separate environments, not just naming conventions | Enforce separate VPCs/virtual networks per environment with explicit, minimal peering rather than one flat network |
| 10 | Serverless functions or compute instances with broader IAM permissions than their event source or workload requires | Access-analysis tooling comparing a role's granted permissions against the resource's actual API call history | Scope down to the resources actually used, deployed behind a monitored canary period before full cutover |
Worked example
Applying items 1, 2, and 5 to one small environment: a startup's AWS account has a public documentation bucket (correctly public, item 1 does not apply here), an application role with dynamodb:* on Resource: "*" (item 2), and a bastion security group allowing SSH from 0.0.0.0/0 (item 5). Detection: the access-analysis tool flags the DynamoDB role because its actual CloudTrail history shows only GetItem and PutItem calls against one table, and the security group audit flags the bastion rule because its source CIDR is unrestricted. Remediation: the role is narrowed to dynamodb:GetItem/PutItem on that one table's ARN (Amazon Resource Name), and the bastion's security group is restricted to the company's VPN egress CIDR. Neither fix touches the intentionally-public documentation bucket, which is why detection needs to distinguish "public" from "wrongly public" rather than flagging every public resource identically.
Trade-offs and pitfalls
- Detection tooling is only as good as its baseline. A CSPM scan that flags every public bucket without a way to mark a bucket as intentionally public generates enough noise that a team starts ignoring its alerts, which is functionally the same as not scanning at all.
- Access-analysis-based least privilege depends on a representative sample of usage. A role scoped from 30 days of call history can miss a legitimate but infrequent code path (a monthly batch job, a quarterly report), so a scoped-down policy needs a monitored rollback window before being treated as final, not a one-shot cutover.
- Item 9 (environment segmentation) is the one most often skipped entirely, because a flat network is simpler to set up initially and the cost of the missing boundary is invisible until a development environment's weaker controls become the actual path into production.
- A single remediation step listed per item is a starting point, not a complete fix. Restricting a security group's source CIDR (item 5) reduces exposure but does not replace enforcing MFA (multi-factor authentication) or session logging on whatever the bastion ultimately grants access to; treat each remediation as the first of several controls, not the only one needed.
Unlock Full Question Bank
Get access to all Cloud Security Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.