AWS Core Services and Architecture Questions
Amazon Web Services' core service catalog and how the pieces compose into a working system: EC2, Lambda, S3, VPC, IAM, RDS, and the managed-service ecosystem. Covers service selection within AWS, common reference architectures, the AWS Well-Architected Framework pillars, and operational patterns specific to the platform. For provider-agnostic compute or storage trade-offs, see the cross-cloud entries.
At a high level, what are the main EC2 pricing models (On-Demand, Reserved, Savings Plans, Spot), and what workload characteristics make each a good fit?
Sample Answer
Direct answer
EC2 offers four main purchasing models, and they trade commitment for discount. On-Demand: pay by the hour or second with zero commitment, the most expensive per unit but the most flexible. Reserved Instances: commit to a 1- or 3-year term for a significant discount over On-Demand. Savings Plans: commit to a dollar-per-hour spend for 1 or 3 years, similar discount depth to Reserved Instances but more flexible about which instances the commitment applies to. Spot Instances: bid for AWS's spare capacity at a steep discount, with the risk that AWS can reclaim that capacity with short notice.
Structured elaboration
| Model | Commitment | Best fit | Trade-off |
|---|---|---|---|
| On-Demand | None | Unpredictable, bursty, or short-lived workloads: dev/test, proof-of-concept, spiky traffic | Highest per-hour cost, maximum flexibility |
| Reserved Instances (Standard or Convertible) | 1 or 3 years | Steady-state, predictable baseline load: always-on web servers, a database that runs continuously | Standard RIs give the deepest discount but are locked to a specific instance family/size/region; Convertible RIs allow changing the instance family or OS during the term for a somewhat smaller discount |
| Savings Plans (Compute or EC2 Instance) | 1 or 3 years, as a per-hour spend commitment | Organizations wanting predictable compute spend without locking to a specific instance configuration | Compute Savings Plans apply broadly, across instance families, regions, and even to Fargate and Lambda usage, which is more flexible but the deepest discount tier (EC2 Instance Savings Plans) is narrower, tied to a specific instance family in a specific region |
| Spot Instances | None, but interruptible | Fault-tolerant, stateless, or checkpoint-able workloads: batch processing, CI/CD runners, big-data jobs, ML training | Steepest discount available, but AWS can reclaim the instance with a short interruption notice, so the workload must tolerate that |
Why the discount ordering exists: On-Demand carries no commitment risk for AWS, so it costs the most. Reserved Instances and Savings Plans both trade a term commitment for a discount of similar depth; the real choice between them is flexibility (Savings Plans can shift across instance families and even services) versus the narrow but sometimes marginally deeper discount of a Standard Reserved Instance. Spot sits apart from that trade-off entirely: the discount reflects unused capacity risk, not a commitment, which is why it can be steeper than either of the commitment-based options.
Worked example
A fleet mixing Spot for cost, without losing availability, is the standard pattern: run an Auto Scaling group or Spot Fleet with mixed instance types across multiple AZs, using a capacity-optimized allocation strategy so the fleet favors instance pools AWS is less likely to reclaim. Diversifying across several instance types and AZs matters because a single instance type in a single AZ is a much easier target for a capacity reclaim than a fleet spread across many pools. On top of that, workloads should implement graceful shutdown handling for the two-minute Spot interruption notice delivered via instance metadata, checkpointing progress to S3 or a database so an interrupted job resumes rather than restarts from zero, and the fleet should keep a small baseline of On-Demand or Reserved capacity to absorb critical, latency-sensitive traffic that can't tolerate any interruption at all.
Trade-offs & pitfalls
The most common mistake is defaulting to On-Demand for a genuinely steady-state workload, teams under-invest in Reserved Instances or Savings Plans out of a (often outdated) fear of losing flexibility, when Savings Plans in particular were built specifically to remove that rigidity. The opposite mistake is putting a stateful, interruption-sensitive workload on Spot to chase the discount, if the workload can't checkpoint or gracefully hand off in-flight work, an interruption becomes a real outage, not just a cost optimization gone slightly wrong. And treating a Reserved Instance purchase as "set and forget" is a trap too, instance family needs shift over a 1- to 3-year term, which is exactly why Convertible RIs and Savings Plans exist as more flexible alternatives to a Standard RI locked to one configuration.
Compare EBS volumes and EC2 instance store for high-performance scratch storage. When is ephemeral instance-store NVMe preferable over provisioned gp3/io2 EBS volumes?
Sample Answer
EC2 instance store is physically attached Non-Volatile Memory express (NVMe) storage with the lowest possible latency and highest raw input/output operations per second (IOPS), but it is ephemeral: the data is gone the moment the instance stops, terminates, or the underlying hardware fails. Amazon Elastic Block Store (EBS), in the gp3 or io2 families, is network-attached, durable, persists independently of the instance lifecycle, and supports snapshots, but has to cross the network to reach the instance, which adds latency EBS cannot fully eliminate no matter how much performance is provisioned. Reach for instance store when the data is disposable and the workload needs raw local speed; reach for EBS when the data has to survive the instance.
Why the trade-off exists
- Instance store lives on the physical host, so there is no network hop between the instance and the disk; that is the entire source of its latency and IOPS advantage. The cost of that physical attachment is that the data cannot follow the instance if it stops, is replaced, or the host fails, and it is not covered by EBS snapshots.
- EBS is a separate, replicated storage service the instance attaches to over the network. That network hop is why EBS latency, even on the fastest provisioned-IOPS tiers, will not match local NVMe, but it is also why EBS data survives a stop and restart, supports point-in-time snapshots, and can be detached and reattached to a different instance.
When ephemeral instance-store NVMe is the right call
- Pure scratch space: a shuffle buffer, a temporary cache, or intermediate build or processing files that are cheap to regenerate if lost.
- The absolute lowest latency and highest throughput matters more than persistence, such as staging data locally right before a compute-heavy job processes it.
- Cost-sensitive, high-throughput workloads, where paying for durable, replicated EBS capacity for data you are going to discard anyway is wasted spend.
- Not every instance type includes instance store; it depends on the specific instance family, so this is a constraint to check before designing around it, not an option available everywhere.
When to prefer gp3 or io2 EBS instead
- Anything that has to survive a stop, restart, or instance replacement: databases, application state, or any artifact you cannot afford to regenerate.
- Anything that needs a snapshot for backup, disaster recovery, or cloning to another instance or AWS Region.
- Predictable, provisioned performance with a durability guarantee: gp3 for general steady performance, io2 when the workload needs the highest sustained IOPS with the strongest durability target.
Worked example
A batch analytics job on EC2 shuffles several hundred gigabytes of intermediate data during processing, then writes only the final aggregated result. The intermediate shuffle data goes on instance-store NVMe attached to the instance, since it needs the fastest possible local I/O and is fully disposable if the instance is replaced mid-job, the job would simply restart from its last checkpoint. The final aggregated result, and any checkpoint needed to resume the job, is written to a gp3 EBS volume (or Amazon S3) specifically because that data has to survive even if the instance disappears. Using EBS for the disposable shuffle data would add unnecessary network latency and cost for no durability benefit; using instance store for the final result would risk losing it on any instance interruption, including a Spot Instance reclaim.
Trade-offs and pitfalls
- The most common mistake is treating instance store as free durable storage because it is fast; it has no snapshot mechanism and no protection against instance interruption, which matters more on Spot Instances, where interruption is expected, not exceptional.
- Not every instance family offers instance store, and the ones that do vary in how much local storage they provide, so this decision is partly constrained by which instance types are already in use, not purely a preference.
- Striping multiple instance-store devices together increases throughput but also increases the blast radius of a single device or host failure; know that trade-off before applying it to anything that is not fully disposable.
- A hybrid pattern, using instance store for scratch and periodically checkpointing the durable result to EBS or S3, gets most of the speed benefit without the durability risk, and is the common real-world answer rather than picking one exclusively.
Propose a hub-and-spoke architecture for centralized security visibility across many AWS accounts using Security Hub, GuardDuty, and CloudTrail aggregation. What three services would you enable first on a brand-new AWS account to establish a baseline security posture, and why?
Sample Answer
A hub-and-spoke security visibility architecture designates one AWS account, the hub (often called the security or log archive account), as the delegated administrator for Amazon GuardDuty and AWS Security Hub, and as the destination for every other account's (the spokes') AWS CloudTrail logs, so one team gets a unified view without operating separately in every account. On a brand-new account, before anything else, the three services to enable first are CloudTrail (so an audit trail exists from minute one), GuardDuty (so threat detection starts baselining normal behavior from minute one), and IAM Access Analyzer (so accidental external access grants are caught immediately), because everything else in a security program depends on those three signals already flowing.
Architecture
flowchart TB
subgraph Org["AWS Organizations"]
Mgmt["Management account"]
end
subgraph Spokes["Spoke / member accounts"]
SpokeA["Account A workloads"]
SpokeB["Account B workloads"]
SpokeC["Account C workloads"]
end
subgraph Hub["Hub / security account"]
GD["GuardDuty\ndelegated administrator"]
SH["Security Hub\ndelegated administrator"]
S3["Central S3 log bucket\nSSE-KMS, Object Lock"]
EB["EventBridge"]
SIEM["SIEM / ticketing"]
end
Mgmt -- designates --> GD
Mgmt -- designates --> SH
SpokeA -- CloudTrail org trail --> S3
SpokeB -- CloudTrail org trail --> S3
SpokeC -- CloudTrail org trail --> S3
SpokeA -- findings --> GD
SpokeB -- findings --> GD
SpokeC -- findings --> GD
GD -- findings --> SH
S3 -- queried by --> SH
SH -- normalized findings --> EB
EB -- routes --> SIEM
Why hub-and-spoke, mechanically
- CloudTrail: enable a single AWS Organizations trail from the management account (or a delegated administrator), which automatically applies to every current and future member account and delivers all logs to one Amazon Simple Storage Service (S3) bucket in the hub account. This removes the "did someone forget to enable logging in the new account" failure mode entirely.
- GuardDuty: the management account designates one member account as the GuardDuty delegated administrator. That account can then auto-enable GuardDuty for every existing and future member account, and every member's findings roll up into the delegated administrator's view. GuardDuty is a regional service, so the delegated administrator must be the same account in every Region you operate in, not just one.
- Security Hub: designate a delegated administrator for Security Hub too, and use its central configuration feature to push standards and controls, for example the AWS Foundational Security Best Practices standard, to every account and organizational unit from one place, instead of configuring each account separately. Security Hub also ingests and normalizes GuardDuty findings alongside its own checks, so the hub ends up with one unified findings view rather than two dashboards.
- EventBridge and downstream routing: Security Hub findings can trigger Amazon EventBridge rules in the hub account to notify a security information and event management (SIEM) tool, open a ticket, or, for well-understood finding types, run an automated remediation through AWS Systems Manager Automation.
The first three services on a brand-new account, and why that order
- CloudTrail: with no audit trail, nothing else is investigatable after the fact. This has to exist before any other service produces something worth investigating.
- GuardDuty: a passive, low-effort detector with no infrastructure to run, that starts baselining normal network and API behavior immediately. The earlier it is on, the sooner its anomaly detection becomes reliable.
- IAM Access Analyzer: catches the most common real-world incident category, a storage bucket, IAM role, or encryption key policy that unintentionally grants access to an external account or the public, through static policy analysis rather than waiting to observe an actual attack.
Worked example
A company with a management account and fifteen workload accounts under one AWS Organization stands up a dedicated security account as the hub. Step one: enable an Organizations CloudTrail trail from the management account, delivering to a central S3 bucket in the security account, with S3 Object Lock in governance mode and server-side encryption using a customer-managed AWS Key Management Service (KMS) key that only the security account administers, so even a compromised workload account cannot delete its own audit history. Step two: designate the security account as the GuardDuty delegated administrator and set auto-enable to cover all accounts, so all fifteen existing accounts and any future account get GuardDuty on day one. Step three: designate the same security account as the Security Hub delegated administrator, turn on central configuration, and push the AWS Foundational Security Best Practices standard to every account from one configuration policy. Within a day, the security account's Security Hub console shows a single, deduplicated findings list spanning all fifteen accounts instead of fifteen separate consoles nobody has time to check individually.
Trade-offs and pitfalls
- Using one account as the delegated administrator for both GuardDuty and Security Hub is the common pattern, but it concentrates real blast radius in that account. Protect it with its own strict IAM boundary and multi-factor authentication, since it can see, and for some findings act, across the entire Organization.
- GuardDuty's regional nature is easy to get wrong: designating the delegated administrator in one Region does not cover another Region automatically, and AWS requires the same account to be the delegated administrator everywhere GuardDuty is enabled.
- Security Hub's central configuration model requires choosing, per account or organizational unit, whether it is centrally managed (only the delegated administrator can change its settings) or self-managed (the account manages its own). Mixing these without a clear policy leads to accounts drifting out of the baseline.
- A single central logging bucket is simple to reason about but becomes a single point of both cost and access-control complexity at scale. Partition by account and date, and separate read access for the security team from the write access granted to the log delivery role.
- Enabling every possible security service on day one sounds appealing but produces alert fatigue before anyone has tuned suppression rules. The three named above are the minimum viable baseline, not the complete target state.
How would you provide secure shell-like access to EC2 instances (and Kubernetes nodes) without exposing SSH to the internet or distributing SSH keys? Walk through using instance profiles and Systems Manager Session Manager, plus how you'd audit sessions.
Sample Answer
Direct answer
Route all shell access through AWS Systems Manager (SSM) Session Manager instead of SSH: give each instance, or Kubernetes node, an AWS Identity and Access Management (IAM) instance profile with just enough permission to act as an SSM managed node, close SSH in the security groups entirely, and grant human or CI identities ssm:StartSession scoped to specific instances and specific SSM documents. Session content is encrypted with a customer managed KMS key and can be streamed to CloudWatch Logs or S3 for audit, while the session API calls themselves, StartSession and TerminateSession, are already logged by CloudTrail as standard management events, no separate configuration needed for that part.
Structured elaboration
Building blocks
- Instance profile, not SSH keys. The instance needs the SSM Agent, preinstalled on most current AMIs, and an IAM role attached as an instance profile with just enough permission to register as a managed node. No key pair, no distributed private keys.
- Network path. Route SSM traffic through Virtual Private Cloud (VPC) interface endpoints, private connections into your own isolated AWS network, for the
ssm,ssmmessages, andec2messagesservices so sessions never need a public IP or NAT egress; a NAT path also works if you would rather keep egress simple, the endpoints are just the tighter-network option. - IAM authorization, scoped per document. Session Manager exposes different capabilities as different SSM documents: an interactive shell is the RunShell document, port forwarding is a separate pair of documents, and file transfer uses separate plugins. Your IAM policy grants
ssm:StartSessionagainst a Resource list of specific document Amazon Resource Names (ARNs), AWS's unique identifier strings for a resource, plus specific instance ARNs or a tag condition, so a user who is only granted the RunShell document ARN cannot start a port-forwarding session at all, there is no separate switch to remember, the capability simply is not in their policy.
{
"Effect": "Allow",
"Action": "ssm:StartSession",
"Resource": [
"arn:aws:ec2:us-east-1:111122223333:instance/i-0123456789abcdef0",
"arn:aws:ssm:us-east-1:111122223333:document/SSM-SessionManagerRunShell"
]
}
- Encryption and logging. In Session Manager preferences, turn on KMS encryption for session data, which grants the caller a data-key permission on the chosen key, and configure logging to CloudWatch Logs and S3, so session transcripts, what was typed and returned, are encrypted at rest and retained per your policy.
- Auditing, two distinct layers. CloudTrail records the control-plane calls, StartSession, TerminateSession, who, when, against which instance, as ordinary management events, logged by default with no extra CloudTrail configuration required. Separately, the session transcript itself, the actual command input and output, only exists if you turned on the CloudWatch Logs or S3 streaming in step 4; that is a distinct feature from CloudTrail and is what you would review to see what a user actually typed.
Worked example
An engineer needs to debug a production instance outside business hours. With this setup: they authenticate normally, their role's policy allows ssm:StartSession scoped to instances tagged for that environment and only the RunShell document, they run the start-session command, get a shell with no SSH port ever having been open, and every keystroke of that session streams, KMS-encrypted, to a CloudWatch Logs group with a retention policy your security team can search later. If that same engineer tried to start a port-forwarding session against the same instance, the call would fail with an IAM authorization error, because their policy's Resource list never included that document's ARN; port forwarding was never granted, not merely disabled by a switch.
Trade-offs and pitfalls
- This removes the bastion host's attack surface and the SSH key distribution problem entirely, but it makes your IAM hygiene the new perimeter: an overly broad
ssm:StartSessionpolicy, wildcarded documents or instances, reintroduces exactly the blast radius you were trying to avoid. - If the SSM Agent itself is unhealthy or the instance cannot reach its required endpoints, Session Manager access is unavailable; you need a documented emergency path, such as the EC2 serial console or a break-glass bastion kept cold, rather than discovering the gap during an actual incident.
- Session logging is opt-in per Session Manager preference, verify it is actually configured and encrypted rather than assuming it is; an unconfigured environment gives you the access-control benefit of SSM without the audit trail.
- IAM session policies can additionally restrict actions taken inside a session, but detecting and killing a session that looks compromised in real time still requires alerting on the logging stream or CloudTrail, SSM will not flag anomalous behavior for you on its own.
How do S3 Versioning and MFA Delete help protect against accidental deletes or overwrites? What operational and cost implications should you be aware of when enabling versioning on a large bucket?
Sample Answer
Direct answer
S3 Versioning keeps every write as a distinct object version instead of overwriting data in place, so an accidental overwrite just creates a new version alongside the old one, and an accidental delete adds a lightweight delete marker, which only hides the object from a normal listing or GET, rather than destroying it. Both are recoverable by fetching or restoring a prior version ID. MFA Delete adds a second, harder-to-automate layer on top: permanently deleting an object version, or changing the bucket's versioning state, requires a valid multi-factor authentication code from the account's root user, which stops both accidental scripted deletes and a compromised non-root credential from erasing history.
Structured elaboration
- Versioning mechanics: enabling versioning is effectively one-way at the bucket level; you can enable or suspend it, but you can never make it as if the bucket was never versioned, since versions created while it was enabled persist.
- MFA Delete specifics: it must be configured using the bucket owner's root account credentials, not just any AWS Identity and Access Management (IAM) admin, and once required, permanently deleting a version or disabling versioning/MFA Delete itself requires that MFA code at request time. That's why it's usually reserved for a bucket's highest-value data rather than applied everywhere.
- Operational cost of versioning: every write, including a rewrite of unchanged content, is billed as new, full-sized storage. A bucket with frequent overwrites can see storage cost grow substantially if nothing prunes old versions. The standard mitigation is a lifecycle rule that transitions or expires noncurrent versions after a chosen window, for example keeping noncurrent versions restorable for roughly 30 days and transitioning anything not needed sooner to a cheaper storage class before final expiry around 90 days, tuned to how much recovery time the team actually needs against storage spend.
- API and tooling overhead: version-aware operations, listing versions, deleting a specific version ID, restoring by copying an old version back as current, need version-ID-aware code. Naive scripts that assume "one object equals one thing" can behave unexpectedly on a versioned bucket, since a plain DELETE just adds a marker and doesn't free any storage.
Worked example
A bucket stores user-uploaded documents, and a bad deploy runs a cleanup script that deletes keys matching a stale prefix pattern, hitting 4,000 live objects by mistake. With versioning on, none of that data is actually gone: the 4,000 objects now show delete markers. Recovery is ListObjectVersions filtered to those keys, then removing the delete marker (or copying the prior version back as current) for each, which S3 Batch Operations can drive at scale from a manifest instead of a one-by-one script. Without versioning, the same script would have caused permanent, unrecoverable data loss.
Trade-offs & pitfalls
- Enabling versioning without a lifecycle rule to prune noncurrent versions is the most common mistake: storage cost creeps up silently, especially on buckets with high overwrite churn, until someone notices the bill.
- MFA Delete is a real operational tax: it must go through the root account, which is incompatible with fully automated, script-driven bulk cleanup unless that workflow can supply an MFA code. That's why teams usually restrict it to a small number of critical buckets.
- Versioning and MFA Delete protect against accidental overwrite, accidental delete, and many scripted mistakes, but neither is a substitute for cross-region or cross-account replication if the actual risk is bucket-level deletion or a compromised account with root-level access; that scenario needs mass-deletion controls beyond versioning alone.
- A delete marker still counts as an object version and, depending on lifecycle configuration, can itself accumulate cost or clutter if never cleaned up.
Unlock Full Question Bank
Get access to all AWS Core Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.