Cloud Security Architecture Questions
Designing and reasoning about the security posture of cloud and hybrid infrastructure: the shared responsibility model, network segmentation and boundary design, multi-account and multi-region security architecture, workload identity as an architectural choice, threat modeling a cloud architecture, cloud-specific attack vectors and mitigations, defense-in-depth control selection, secure cloud deployment patterns, and continuous cloud risk assessment and posture. IAM policy authoring, role/trust-policy mechanics, and secrets/credential lifecycle belong to identity-and-access-management; logging-pipeline design and SIEM/detection-rule engineering belong to security-monitoring-and-detection; encryption-key-management mechanics (KMS/CMK/BYOK) belong to data-protection-and-encryption; compliance-framework mapping (SOC2, PCI-DSS, HIPAA, GDPR) belongs to compliance-frameworks-and-certification-standards. This topic keeps identity, logging, or encryption content only when it is one ingredient inside a genuinely multi-control cloud-hardening question, not as a standalone ask.
Explain the shared responsibility model in cloud computing. For each service model (IaaS, PaaS, SaaS) describe which security controls are typically the provider's responsibility and which are the customer's. Provide concrete examples (for example, EC2, RDS, and Gmail), describe a couple of common gray-area responsibilities, and explain how you would document responsibility boundaries for a new cloud service onboarding.
Sample Answer
Direct answer
Across Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS), the shared responsibility line moves in one direction only, toward the provider, as the service becomes more managed, but the customer's core responsibility, identity, access control, and what data goes into the service, never fully disappears at any point on that spectrum, which is exactly what makes the gray-area cases (not the clear-cut ones) the actual place organizations get this wrong.
Structured elaboration
IaaS, example: Amazon Elastic Compute Cloud (EC2). Provider responsibility: physical datacenters, the hypervisor, the host's own network infrastructure. Customer responsibility: guest operating system patching, network configuration (security groups, subnet placement), identity and access management (IAM) for who can access the instance, and everything running on top of the OS, since the customer chose and controls essentially the entire software stack above the hypervisor.
PaaS, example: Amazon Relational Database Service (RDS). Provider responsibility: the physical infrastructure, the hypervisor, and now also the database engine's own patching and the underlying operating system, since the customer never has direct OS-level access to a managed database instance. Customer responsibility: network access configuration (is the instance public, what security group scopes reachability), authentication method and credential management, encryption configuration, and the data itself, a narrower slice than IaaS but still real and still the customer's to get right.
SaaS, example: Gmail (as a representative business-email SaaS product). Provider responsibility: essentially the entire technical stack, the application itself, its infrastructure, its own security patching, with the customer having no direct infrastructure-level control at all. Customer responsibility: narrows to identity and access configuration (who has an account, multi-factor authentication (MFA) enforcement, single sign-on (SSO) integration), data governance (what data users choose to put into the product, sharing and retention settings the product exposes), and user behavior (a phished user credential is still the customer's problem to prevent and respond to, regardless of how secure the underlying SaaS platform itself is).
Common gray-area responsibilities
Encryption key management on a managed service. Whether the customer is responsible for key management depends entirely on whether they opted into a customer-managed key or accepted the provider's default key; this is nominally the customer's choice, but many organizations never make it deliberately, defaulting to whatever the console's default happens to be, which means the actual responsibility boundary in practice is often determined by an unexamined default rather than a considered decision.
Default configuration values. When a managed service is deployed with an insecure default (a database provisioned with a public endpoint enabled by default, for instance, in certain configurations), the provider technically offered the configuration option, but the customer is universally held accountable for the actual deployed state in every compliance framework and every real-world incident response; "the default was insecure" is not a defense, which is a gray area in principle but a settled question in practice, worth stating plainly because it is so commonly misunderstood.
Documenting responsibility boundaries for a new cloud service onboarding
Before onboarding any new managed service, produce a written responsibility matrix specific to that service (not a generic, one-size-fits-all shared-responsibility document), explicitly naming which of the customer's own controls apply (network configuration, IAM, encryption key choice, data classification) and confirming a specific, named owner for each; review the provider's own documented responsibility boundary for that specific service, since it varies by service even within one provider, rather than assuming it matches a different service the organization has already onboarded; and require this documented matrix as a gate in the onboarding process itself, not an afterthought produced once the service is already in production use.
Worked example
An organization onboarding Gmail as its business email platform for the first time documents its responsibility matrix explicitly: MFA enforcement and SSO integration are named, owned controls (assigned to the identity team), data-loss-prevention policy configuration for what leaves the organization via email is a named, owned control (assigned to the security team), and user security-awareness training for phishing resistance is a named, owned control (assigned to the security-awareness program), each with a specific owner rather than an implicit assumption that "Google handles security." Six months later, a user falls for a phishing email and enters their credentials on a fake login page; because MFA enforcement was a documented, owned control that had actually been implemented, the stolen password alone is insufficient for the attacker to access the account, the concrete payoff of having named and implemented that specific customer-side responsibility rather than assuming the SaaS provider's own security covered it.
Trade-offs and pitfalls
- "It's SaaS, the provider handles security" is the single most common and most consequential misunderstanding of this entire model, and it is exactly backwards about where the narrowing happens: the technical infrastructure responsibility narrows toward the provider as services become more managed, but the identity, access, and data-governance responsibility never narrows to zero, and SaaS is precisely where that narrower-but-real slice matters most, since it is often the only thing standing between a phished credential and an actual breach.
- The encryption-key-management gray area is a case where "the customer could have chosen differently" is technically true and practically misleading, since most organizations never make an active, considered choice about it; treating this as a settled customer responsibility without acknowledging how often it is an unexamined default is an honest gap worth naming rather than glossing over.
- A generic, one-size-fits-all shared-responsibility document produced once and reused for every new service onboarding misses that the actual boundary varies by specific service, even within the same provider; the worked example's Gmail-specific matrix would look meaningfully different from an EC2-specific or RDS-specific one, and treating them as interchangeable templates undermines the entire point of documenting the boundary explicitly.
- A responsibility matrix that names an owner on paper but is never actually implemented (MFA "assigned" to the identity team but never actually enforced, for instance) provides no real protection, only the appearance of it; the worked example's payoff depends specifically on the named control having been genuinely implemented, not merely documented as someone's responsibility.
Prepare a ransomware prevention and recovery design for cloud workloads and data. Cover immutable backups (WORM/Object Lock), backup account separation, cross-region replication strategy, backup encryption and KMS key separation, least-privilege for backup operators, automated testing of restores, and cost controls. Also describe detection signals that might indicate ransomware activity and the immediate playbook actions you'd take.
Sample Answer
Direct answer
A ransomware-resilient backup design assumes the attacker will eventually hold valid credentials in the production environment, and builds the recovery path so that credential is structurally incapable of reaching the backups: immutable storage (Write Once Read Many (WORM), enforced via Object Lock), a dedicated backup account the production credentials cannot touch, and least-privilege backup operators are the three controls that together mean "the attacker compromised production" does not also mean "the attacker can delete the recovery path."
Structured elaboration
flowchart TB
subgraph WA["Workload account"]
App["Application"] --> Prod[("Production data")]
end
subgraph BA["Dedicated backup account (separate from workload)"]
Vault1[("Backup vault, region A, Object Lock: Compliance mode")]
Vault2[("Backup vault, region B, cross-region copy")]
end
Prod -->|"one-way backup role, PutObject only"| Vault1
Vault1 -->|"scheduled cross-region copy"| Vault2
Restorer["Restore-test job (least-privilege, read-only)"] -->|"weekly automated restore drill"| Vault1
Attacker(["Compromised workload credential"]) -.->|"no delete/overwrite permission on vault"| Vault1
Immutable backups (WORM/Object Lock). Object Lock in Compliance mode (not Governance mode, which even an account root user can override) means no principal, including a compromised administrative credential, can delete or overwrite a locked backup object before its retention period expires; this is the single control that most directly defeats a ransomware actor's typical playbook of deleting backups before encrypting production data.
Backup account separation. The backup destination lives in a dedicated AWS account (or equivalent project/subscription on GCP/Azure) that the production workload's own credentials have no delete or administrative access to at all, only a narrow, one-way write path. This means a full compromise of the production account's identity and access management (IAM), including a compromised administrator credential in that account, still cannot reach into the backup account to remove the recovery path, since that permission was never granted in the first place, not merely restricted.
Cross-region replication strategy. Backups replicate to a second region on a scheduled basis, protecting against a region-level event independent of ransomware specifically (an infrastructure outage, a regional service disruption), and adding a second layer of recovery even in the unlikely case the primary backup vault's own region is somehow compromised.
Backup encryption and Key Management Service (KMS) key separation. Backup data is encrypted with a KMS key that lives in the backup account, not the production account; a compromised production credential cannot decrypt, and more importantly cannot request deletion of, a key it was never granted access to. Key separation matters as much as account separation, since a shared key would leave decrypt access reachable from production even if the storage itself were otherwise isolated.
Least-privilege for backup operators. The identity that writes new backups holds exactly PutObject (and equivalent) permission on the vault, nothing else, no delete, no ability to modify Object Lock settings, and no read access to unrelated data; the identity that runs restore tests holds read-only access scoped to the vault, never write or delete. Neither identity is broad enough to undo the immutability guarantee even if it were itself compromised.
Automated testing of restores. A scheduled job performs a genuine restore, not just a checksum verification, on a regular cadence (weekly, in the diagram above), because a backup that has never actually been restored is an unverified assumption, not a working recovery capability; ransomware-readiness reviews routinely find backups that technically exist but fail to restore when actually needed.
Cost controls. Immutable, cross-region, versioned backups accumulate storage cost indefinitely unless a lifecycle policy transitions older, still-locked backups to a cheaper storage tier after their compliance-critical recency window passes, and eventually expires versions once their retention period is satisfied; cost control has to be designed alongside immutability, not treated as a reason to shorten retention below what the recovery objective actually requires.
Detection signals for ransomware activity. A sudden spike in file-modification or encryption-pattern activity across a filesystem or object store; an unusual volume of PutObject calls overwriting existing keys in rapid succession; a spike in failed decryption or file-open errors reported by monitoring agents; and, specifically relevant to the backup design itself, any attempted delete or Object Lock modification call against the backup vault, which should never happen legitimately and is a near-certain indicator of an attack in progress given the least-privilege design above.
Immediate playbook actions. Isolate the affected workload (network-level quarantine, not deletion, to preserve forensic evidence) the moment ransomware activity is detected; confirm the backup vault's integrity and immutability status independently, since this design assumes it cannot have been touched, but confirming that assumption explicitly is still the first recovery-readiness check; identify the last known-clean restore point using the tested restore capability; and begin restoring to a clean, isolated environment rather than back into the still-compromised production account.
Worked example
A ransomware actor gains administrator-level credentials in a company's production AWS account through a phishing attack and begins encrypting data across the account's storage. Detection: the security team's monitoring flags an unusual spike in PutObject calls overwriting existing object keys across several buckets within a short window, well outside the account's normal write pattern. Immediate action: the affected instances and their network access are isolated. The team confirms, as designed, that the attacker's credentials, scoped entirely within the production account, have no path to the separate backup account at all; an attempted DeleteObject call against the backup vault (which the attacker's automated ransomware tooling did in fact attempt, per the design's detection signal) is denied outright by IAM before it ever reaches the Object Lock check, since the compromised credential was never granted delete permission on that account in the first place. Recovery proceeds from the most recent tested restore point, into a newly-provisioned, isolated environment, not back into the still-potentially-compromised original account.
Trade-offs and pitfalls
- Governance mode is a common, dangerous shortcut. It is easier to configure than Compliance mode because it permits an override by a sufficiently privileged principal, but that is exactly the property ransomware exploits if the attacker's compromised credential happens to hold that override permission; Compliance mode's inflexibility is the actual point of the control, not a limitation to work around.
- Backup account separation only holds if the separation is genuinely one-way. A backup account that grants the production account any administrative or delete-capable role back into itself (even one intended only for emergency operator use) reopens exactly the path this design exists to close; any emergency access needs to be a separate, tightly audited, out-of-band process, not a standing IAM relationship.
- Untested restores are the single most common gap discovered during an actual ransomware incident, not a hypothetical risk. A backup program with excellent immutability and account separation but no regular, automated restore testing can still fail at the moment of truth if the backup format has silently drifted from what the restore tooling expects; the weekly restore drill in the design above is not a nice-to-have, it is the control that validates every other control actually works end to end.
- Cost controls that shorten retention below the actual recovery objective undermine the whole design for the sake of a smaller storage bill. A retention window shortened to save cost, without re-evaluating whether it still covers a realistic "time to detect a slow, stealthy ransomware campaign," can leave the organization with only encrypted, already-compromised backups by the time the attack is finally noticed.
Explain the concept of 'hub-and-spoke' network topology in enterprise cloud networking. What are two security benefits and one potential single point of failure you must mitigate?
Sample Answer
Direct answer
Hub-and-spoke is a network topology where a central "hub" virtual network hosts shared services (a transit gateway or equivalent, centralized firewall/inspection, logging, VPN or Direct Connect termination) and every workload "spoke" network connects only to the hub, never directly to another spoke; the two clearest security benefits are centralized enforcement and reduced east-west exposure between workloads, and the one significant single point of failure to mitigate is the hub itself, since every spoke's connectivity, and every spoke's security enforcement, now depends on it.
Structured elaboration
Security benefit 1: centralized policy enforcement and inspection. Because every spoke's traffic to another spoke, or to the internet, or to an on-premises network passes through the hub, a single centralized firewall or inspection appliance in the hub can enforce one consistent policy for every workload, rather than each spoke team independently configuring and maintaining its own egress and inter-spoke rules, which in practice drift out of consistency over time.
Security benefit 2: default isolation between spokes. Spokes do not connect directly to each other by default in a hub-and-spoke design; two spokes can only reach each other by explicit route through the hub, which is itself a chokepoint the hub's own policy can restrict or deny. This means a compromise in one spoke does not automatically have a network path to another spoke, unless the hub's routing and policy explicitly permit it, a meaningfully smaller default blast radius than a flat, fully-meshed network where every workload can reach every other workload unless specifically blocked.
The single point of failure: the hub itself. Every spoke's connectivity to every other spoke, to the internet, and to on-premises infrastructure routes through the hub; an outage or a misconfiguration in the hub's transit gateway, its centralized firewall, or its own network path affects every spoke simultaneously, which is a materially larger blast radius for an availability failure than a flat topology would have (even though the security posture is generally better). The hub is also a concentrated target: a compromise of the hub's own centralized firewall or transit component potentially exposes the routing decisions for every spoke it serves.
Mitigating the hub's single-point-of-failure risk. Redundant hub infrastructure (a highly-available transit gateway configuration, redundant firewall appliances across multiple Availability Zones rather than a single instance) is the direct availability mitigation. For the security-concentration risk specifically, the hub's own administrative access needs tighter controls than any individual spoke (since compromising hub administration is a strictly worse outcome than compromising one spoke), and the hub's own configuration changes should go through a more deliberate change-control process than a typical spoke, given the blast radius of a mistake made there.
Worked example
An enterprise with 12 application spokes routes all inter-spoke and internet-bound traffic through a hub containing a redundant, multi-Availability-Zone firewall cluster and a highly-available transit gateway. A compromised workload in Spoke 3 attempts to reach a database in Spoke 7; because spokes do not connect directly, this traffic must route through the hub, where the centralized firewall's policy (permitting only the specific inter-spoke flows that have been explicitly approved) denies the connection, since Spoke 3 and Spoke 7 were never granted a route to each other. Separately, when one of the hub's two firewall appliances fails during a maintenance window, the redundant second appliance continues serving all 12 spokes without an availability interruption, which is the concrete payoff of the redundancy investment described above; without it, that same failure would have been a total outage for every spoke simultaneously, not just the one undergoing maintenance.
Trade-offs and pitfalls
- The security benefit and the availability risk come from the exact same architectural property, and treating them as separate concerns misses that they trade off against each other. The same centralization that makes policy enforcement consistent and blast radius small during a compromise is what makes an outage or a misconfiguration in the hub itself catastrophic across every spoke at once; a hub-and-spoke design is only a net improvement if the redundancy investment in the hub is genuinely made, not merely intended.
- A common wrong turn is under-investing in hub redundancy specifically because the hub "isn't a workload," so it does not get the same operational rigor (monitoring, on-call ownership, change review) that individual application spokes receive. The hub is infrastructure every spoke depends on, and it deserves at least the operational rigor of the most critical spoke it serves, not less.
- A hub-and-spoke design still allows a spoke-to-spoke path if the hub's policy explicitly permits it, and over time, as more legitimate integration needs accumulate, the hub's policy can gradually grow permissive enough to erode the default-isolation benefit. Periodic review of the hub's actual permitted spoke-to-spoke flows against what is still genuinely needed keeps this benefit real rather than nominal.
- A concentrated administrative-access risk at the hub is easy to overlook because the hub is infrastructure, not a customer-facing workload, and infrastructure access reviews sometimes receive less scrutiny than application access reviews. Given the hub's blast radius, its administrative access should receive at least as much scrutiny as the most sensitive spoke it serves, arguably more.
You are asked to perform a security review of a client's cloud migration plan. Provide a step-by-step assessment checklist covering identity and access, network architecture, data protection, logging and monitoring, compute/container hardening, automation/IaC, and third-party integrations. Explain how you'd present risks and prioritized remediation to business stakeholders.
Sample Answer
Direct answer
A cloud migration security review needs a checklist that walks the same seven domains every time, identity and access, network architecture, data protection, logging and monitoring, compute/container hardening, automation/infrastructure-as-code (IaC), and third-party integrations, because a migration plan that looks strong on the domains a team happened to focus on can still have a serious gap in one nobody thought to check, and presenting the findings to business stakeholders means translating each technical gap into a business-risk statement they can actually act on.
Structured elaboration
Step-by-step assessment checklist.
- Identity and access. Confirm the migration plan's identity model: multi-account or multi-project structure, least-privilege role design (not a broad "migration admin" role left in place after cutover), multi-factor authentication (MFA) enforcement, and a clear plan for retiring any temporary, migration-specific elevated access once the migration completes.
- Network architecture. Review the planned subnet tiering, whether the database tier will have genuine routing-layer isolation (not just a security group), and whether the migration introduces any temporary, wider-than-intended connectivity between the legacy and new environments that needs an explicit teardown plan.
- Data protection. Confirm encryption at rest and in transit for the migrated data, key management approach (a customer-managed key with its own access policy, not a default shared key by default), and whether the migration plan itself introduces a temporary exposure window (data staged somewhere less protected than its final destination during the transfer).
- Logging and monitoring. Confirm centralized, immutable logging is in place before the migration begins, not added afterward, since the migration itself is exactly the kind of high-change-volume period where an incident is more likely and audit trail matters most.
- Compute/container hardening. Review base-image hygiene, patching strategy, and (for a containerized workload) the image supply-chain and admission-control posture the migrated workload will run under, confirming the new environment meets at least the hardening bar of the environment being replaced, not a regression introduced by migration haste.
- Automation/IaC. Confirm the migration itself is executed through reviewed infrastructure-as-code with the same policy-as-code and static-analysis gates used for ongoing operations, rather than manual console configuration during a time-pressured cutover window, since manual migration steps are exactly where misconfigurations most often get introduced.
- Third-party integrations. Inventory every third-party service or partner integration the migrated workload depends on, confirming each one's access is scoped no more broadly than the legacy environment granted it, and flagging any integration whose access model does not translate cleanly to the new environment's identity structure.
Presenting risks and prioritized remediation to business stakeholders. Translate each technical finding into a business-risk statement (likelihood and impact in terms the business already tracks: regulatory exposure, customer trust, cost of a likely incident) rather than a raw technical severity score; group findings into "must fix before go-live," "fix within the first 30 days post-migration," and "longer-term hardening," since a business stakeholder needs to know what blocks the migration timeline versus what can proceed with a tracked follow-up plan, not just a flat list of findings.
Worked example
Reviewing a client's plan to migrate a customer-facing application from a legacy data center to the cloud, the checklist surfaces: the identity model correctly uses least-privilege per-service roles (step 1, no finding), but the network architecture plan places the database tier's subnet with a route table still carrying a temporary internet route "for the migration cutover, to be removed after" (step 2, a finding, since "temporary" routes are a common source of forgotten exposure); data protection is solid (step 3, no finding); logging is planned to be added two weeks after go-live rather than before (step 4, a significant finding, since the highest-risk period, the migration itself, would run without the audit trail needed to investigate anything that goes wrong during it); and a third-party payment integration's access model does not map cleanly to the new environment's role structure, currently planned to use a broader temporary role "until we figure out the right scoping" (step 7, a finding). These four findings (one clean network step aside) are presented to business stakeholders as: two must-fix-before-go-live items (the temporary database route and the payment integration's over-broad temporary role, both create real exposure during the highest-risk window) and one accelerate-the-timeline item (moving logging setup before, not after, go-live), framed around the specific business risk each one represents (an exposure window during the highest-scrutiny period of the migration, and a payment-integration access gap with direct compliance implications) rather than as a raw list of technical findings.
Trade-offs and pitfalls
- "Temporary" exceptions planned during a migration, the database route and the payment-integration role in the worked example, are the single most common source of a post-migration security gap, since the pressure of a cutover deadline creates a real incentive to accept a temporary shortcut "just to hit the date," and the planned teardown or proper-scoping step frequently does not happen on schedule once the migration itself is declared complete and attention moves elsewhere.
- Logging added after go-live rather than before is a subtle but consequential ordering mistake specifically because the migration window itself is a higher-than-normal-risk period, not a lower-risk one; a team that treats logging as a post-migration polish item has the priority backwards relative to when the audit trail is actually most likely to be needed.
- The must-fix-versus-follow-up categorization needs real technical judgment, not a mechanical severity score, since a technically "medium" severity finding occurring during the highest-risk migration window can matter more than a technically "high" severity finding in a lower-risk, more monitored steady-state environment; the worked example's categorization reflects this timing-aware judgment, not a rote severity mapping.
- A checklist walked in isolation, domain by domain, can miss an interaction between two domains that individually look fine; the "temporary" network route and the "temporary" payment-integration role in the worked example are individually explainable migration-timeline shortcuts, but together they represent a pattern (temporary exceptions accepted under deadline pressure) worth flagging to stakeholders as a process risk, not just as two separate technical findings.
Design a secure VPC architecture in AWS for a three-tier web application (public load balancers, application layer, private database). Describe subnet placement across AZs, route tables, NAT gateways, security groups, bastion/jump host strategy, and where to place private endpoints and logging collectors. Consider both availability and security.
Sample Answer
Direct answer
A secure three-tier VPC (Virtual Private Cloud) design for a public load balancer, an application layer, and a private database puts each tier in its own subnet type, replicated across at least two Availability Zones (AZs) for availability, with the database subnet having no route to the internet at all rather than merely being blocked by a security group, since the network topology itself, not just an access-control rule, should make the database unreachable from outside.
Structured elaboration
flowchart TB
Internet(["Internet"]) --> IGW["Internet gateway"]
IGW --> PubA["Public subnet AZ-a: ALB, NAT GW"]
IGW --> PubB["Public subnet AZ-b: ALB, NAT GW"]
PubA --> AppA["Private app subnet AZ-a"]
PubB --> AppB["Private app subnet AZ-b"]
AppA --> DbA["Private DB subnet AZ-a (isolated, no NAT route)"]
AppB --> DbB["Private DB subnet AZ-b (isolated, no NAT route)"]
DbA -.-> DbB
AppA -->|"VPC endpoint"| KMS[("KMS / Secrets Manager")]
AppB -->|"VPC endpoint"| KMS
Bastion["Bastion / SSM Session Manager"] -.->|"admin access, no inbound SSH from internet"| AppA
AppA --> Flow[("VPC flow logs to log-archive account")]
Subnet placement across AZs. Three subnet tiers (public, private-application, private-database), each replicated in at least two AZs (three, if the workload's availability requirement justifies the added cost), so a single AZ failure does not take down the whole application. Public subnets host only the Application Load Balancer (ALB) and NAT gateways, nothing else, since minimizing what actually sits in a public subnet minimizes what is directly internet-reachable even before any security-group rule is considered.
Route tables. The public subnets' route table sends 0.0.0.0/0 to the Internet Gateway. The private application subnets' route table sends 0.0.0.0/0 to the AZ-local NAT gateway (each AZ's application subnet uses its own AZ's NAT gateway, not a shared one, so a single NAT gateway failure does not take down every AZ's outbound path). The private database subnets' route table has no 0.0.0.0/0 route at all, to either the internet gateway or a NAT gateway, so outbound internet access from the database tier is structurally impossible regardless of any security-group misconfiguration, a route-table-level guarantee, not merely a rule that could be loosened.
NAT gateways. One NAT gateway per AZ (not one shared NAT gateway for the whole VPC), placed in the public subnet of each AZ, giving the application tier outbound internet access (for package updates, external API calls) while remaining unreachable from inbound internet traffic, and avoiding a cross-AZ dependency that a single shared NAT gateway would introduce.
Security groups. The ALB's security group permits inbound 443 from the internet. The application tier's security group permits inbound only from the ALB's security group (by reference, not by CIDR), on the application's specific port. The database's security group permits inbound only from the application tier's security group, on the database's specific port, and nothing else, not even from the bastion host directly.
Bastion/jump host strategy. Prefer a session-manager-based access pattern (such as AWS Systems Manager (SSM) Session Manager) over a traditional bastion host with an open inbound Secure Shell (SSH) port: it requires no inbound security-group rule at all (the connection is initiated outbound from the managed instance to the SSM service), and every session is logged centrally without needing separate bastion-host session-logging infrastructure. Where a traditional bastion is still required for a specific tooling reason, place it in its own small public or dedicated subnet, restrict inbound SSH to a narrow, known administrative CIDR (never 0.0.0.0/0), and require multi-factor authentication (MFA) for any session.
Private endpoints. Traffic from the application tier to cloud-native services it depends on (a secrets manager, a Key Management Service (KMS) key, an object storage bucket) routes through VPC endpoints rather than out through the NAT gateway to the public internet-facing version of those services; this keeps that traffic on the provider's private network backbone entirely and, for gateway-type endpoints such as the one for object storage, removes a real cost (NAT gateway data-processing charges) as well as a security benefit.
Logging collectors. VPC flow logs are enabled on every subnet and shipped to a centralized, separate logging destination (ideally a dedicated log-archive account), not just stored locally in the same account the traffic originated from, so an incident investigation has a trustworthy record independent of whatever happened to the workload account.
Worked example
A concrete Classless Inter-Domain Routing (CIDR) layout for a VPC sized 10.0.0.0/16 across two AZs: public subnets 10.0.0.0/24 (AZ-a) and 10.0.1.0/24 (AZ-b), hosting the ALB and each AZ's own NAT gateway; private application subnets 10.0.10.0/24 (AZ-a) and 10.0.11.0/24 (AZ-b), each routing 0.0.0.0/0 to its own AZ's NAT gateway; private database subnets 10.0.20.0/24 (AZ-a) and 10.0.21.0/24 (AZ-b), with a route table containing only the VPC's local route, no default route at all. The database security group permits inbound on its port only from the application tier's security group ID; the application tier's security group permits inbound only from the ALB's security group ID; and VPC flow logs from every subnet ship continuously to a dedicated log-archive account. An operator needing to inspect an application-tier instance connects via SSM Session Manager, which requires no inbound port open on that instance's security group at all, and the session is centrally logged in the same log-archive account the flow logs already ship to.
Trade-offs and pitfalls
- A shared, single NAT gateway across all AZs is a common cost-saving shortcut that creates an availability trade-off worth stating explicitly. It is cheaper (one NAT gateway instead of one per AZ) but makes every AZ's outbound path depend on one AZ's infrastructure, which contradicts the multi-AZ availability goal the rest of the design otherwise achieves; the per-AZ NAT gateway design above costs more but removes that single point of failure.
- A traditional bastion host, even a well-configured one, is a standing piece of internet-facing attack surface that a session-manager-based approach removes entirely. This is not merely a modernization preference; a bastion's open inbound port, however narrowly scoped by CIDR, is one more thing that can be misconfigured or targeted, and the session-manager approach's zero-inbound-port design is a genuinely stronger default, not just a more convenient one.
- Database subnets with no default route are easy to accidentally break during a later change if someone adds a NAT route "to fix connectivity" without understanding why it was deliberately absent; this control needs to be documented explicitly as intentional, not just configured and left unexplained, so a future engineer does not silently undo the isolation guarantee while trying to fix an unrelated problem.
- VPC endpoints reduce both cost and exposure but need to be added deliberately for every external service the application tier depends on; a design that adds endpoints for the obvious cases (object storage, secrets manager) but misses a less obvious dependency (a specific regional API the application calls) leaves that one dependency's traffic routing out through the NAT gateway to the public internet, a gap that a dependency audit at design time, not just at initial rollout, is needed to catch.
Unlock Full Question Bank
Get access to all Cloud Security Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.