AWS Core Services and Architecture Questions
Amazon Web Services' core service catalog and how the pieces compose into a working system: EC2, Lambda, S3, VPC, IAM, RDS, and the managed-service ecosystem. Covers service selection within AWS, common reference architectures, the AWS Well-Architected Framework pillars, and operational patterns specific to the platform. For provider-agnostic compute or storage trade-offs, see the cross-cloud entries.
Design an EventBridge-driven automated response: when a GuardDuty finding fires (say, an EC2 instance making suspicious outbound connections), how would you wire EventBridge rules to trigger containment via Lambda or Systems Manager, and what least-privilege considerations apply to the automation itself?
Sample Answer
Direct answer
Wire an EventBridge rule to match GuardDuty findings by type and severity, routing matched events to a Lambda "triage" function and to Security Hub for tracking. The triage Lambda enriches the finding (instance tags, attached AWS Identity and Access Management (IAM) role, Virtual Private Cloud (VPC) context, your own isolated network in AWS) and invokes a Systems Manager (SSM) Automation document to contain the instance, typically by attaching a quarantine security group that blocks egress while preserving SSM connectivity, rather than terminating it outright, so evidence isn't destroyed. Every role in that chain (EventBridge target, the triage Lambda, the SSM Automation execution role) gets the minimum permissions to do exactly its one step, nothing broader, since the automation itself is now a privileged actor that needs to be as trustworthy as any human responder.
Structured elaboration
flowchart LR
GD["GuardDuty finding
(suspicious outbound traffic)"] --> EB["EventBridge rule
matches finding type/severity"]
EB --> TR["Triage Lambda
enrich: tags, IAM role, VPC, recent CloudTrail"]
TR --> SH["Security Hub
finding updated, workflow status set"]
TR --> SSM["SSM Automation document
runs on the instance via SSM Agent"]
SSM --> QSG["Attach quarantine security group
(block egress, allow SSM only)"]
SSM --> SNAP["EBS snapshot +
CloudTrail/VPC Flow Log export to S3"]
QSG --> NOTIFY["SNS to on-call / ticket"]
SNAP --> NOTIFY
NOTIFY --> APPROVE{"Analyst
approval"}
APPROVE -->|approve remediation| CLEAN["Rotate credentials,
rebuild instance, restore SG"]
APPROVE -->|false positive| REVERT["Remove quarantine SG,
close finding"]
Detection and routing: the EventBridge rule's event pattern filters on the GuardDuty finding's type and severity fields so only findings above your response threshold trigger automation; lower-severity findings can route to Security Hub for visibility without triggering containment.
Containment: attaching a quarantine security group (deny all egress except to SSM endpoints) is generally preferred over immediately terminating the instance, because termination destroys the evidence you need for the next step. The same pattern applies whether the finding is outbound network activity or, in a related variant, an EC2 instance showing suspicious PutObject patterns against S3 (a data-exfiltration signature): the response is still containment-first, not termination-first.
Forensic data collection: capture an Elastic Block Store (EBS) snapshot of the instance's volumes before any remediation, and export the artifacts that let an analyst reconstruct what happened: CloudTrail (API-level activity, including who/what made the suspicious calls), S3 access logs (for the exfiltration-to-S3 variant, showing which objects were read or written and from where), and VPC Flow Logs (network-level record of the suspicious connections GuardDuty flagged). Store all of it in a locked-down forensic S3 bucket with versioning and access logging of its own.
Credential handling: if the finding suggests the instance's IAM role or an associated credential may be compromised (either variant: outbound C2-style traffic or S3 exfiltration), revoke or rotate the exposed credentials as an explicit containment step, not just a follow-up cleanup task; a still-valid credential is an ongoing exposure even after the instance itself is quarantined.
Cross-account scoping: in a multi-account setup, the automation should be able to act on the affected account's resources without requiring standing, account-wide privileges. Use a dedicated automation/security-tooling account whose Lambda and SSM roles assume narrowly-scoped, temporary cross-account roles into the affected account, limited to the specific containment actions (attach SG, run SSM document, create snapshot) and denied everything else, including account-wide administrative actions.
Least-privilege for the automation itself: separate the Lambda execution role (read GuardDuty/EC2/SSM metadata, start SSM Automation executions, write to the forensic S3 prefix, update Security Hub) from the SSM Automation role (act only on the specific, tag- or ARN-scoped target instance, where an ARN, Amazon Resource Name, is AWS's unique identifier for a specific resource). Neither role should hold broad iam:*, ec2:*, or s3:* permissions; scope by resource ARN and, where possible, by condition keys tied to the finding's specific instance.
Reporting and prevention (mapped to Well-Architected pillars): close the loop with a written incident report and feed findings back into architecture, framed against four Well-Architected concerns:
- Detection: was the GuardDuty finding type/severity threshold right, or did this near-miss a lower bucket that wouldn't have triggered automation?
- Least privilege: did the compromised instance's IAM role have more access than the workload needed, and should its policy be tightened?
- Automation: did the playbook run cleanly end-to-end, or were there manual steps that should be automated for the next incident?
- Governance: are the account boundaries, tagging, and approval gates around this automation documented and enforced, so this response pattern is repeatable and auditable, not tribal knowledge?
Worked example
GuardDuty fires Backdoor:EC2/C&CActivity.B (suspicious outbound to a known command-and-control IP) for instance i-0123456789. EventBridge routes it to the triage Lambda, which tags the finding as IN_PROGRESS in Security Hub and invokes the SSM Automation document. The document attaches the quarantine security group, takes an EBS snapshot, exports the instance's last hour of CloudTrail and VPC Flow Log entries to the forensic bucket, and revokes the instance profile's temporary credentials by detaching and reissuing under a new role. An analyst is paged via SNS, reviews the captured evidence, confirms it's a genuine compromise (not a false positive from a legitimate but newly-added external integration), and approves the remediation branch: the instance is rebuilt from a known-good AMI rather than "cleaned in place," and the incident report goes into the next architecture review under the four pillars above.
Trade-offs and pitfalls
- Fully automated, unapproved termination is tempting for speed but destroys evidence and risks taking down a false positive in production; a human-approval gate before destructive remediation (not before containment) is the right balance for most environments.
- If the automation's own IAM role is over-privileged, an attacker who compromises the automation path (rather than the original instance) inherits broad account access, which is a worse outcome than the original finding; this is why the automation's roles get the same least-privilege scrutiny as the workload they're protecting.
- Snapshotting and log export take time and add cost; prioritize capturing the volatile, hard-to-recover evidence (memory-adjacent artifacts, recent flow logs) before slower steps, and don't let evidence collection delay containment (attaching the quarantine SG) which is the step that actually stops ongoing damage.
- Test the playbook against synthetic findings in a non-production account; a badly-scoped SSM document run for the first time against a real incident is a poor place to discover it also quarantines the bastion host used to investigate it.
What's the relationship between CloudWatch Metrics, CloudWatch Logs, and CloudWatch Alarms? When would you prefer a metric over a log-based metric, and how would you wire an alarm to notify an on-call person?
Sample Answer
Direct answer
The three work together as a pipeline. CloudWatch Metrics store numeric time series (CPU utilization, request latency, error counts). CloudWatch Logs store raw event text (application log lines, stack traces, audit events). CloudWatch Alarms watch a metric, either a native one or one derived from logs, against a threshold over a time window, and fire an action when it's breached. Metrics are what you alert on; logs are what you read to understand why.
Structured elaboration
When to prefer a native metric over a log-based metric
- Prefer a native metric (emitted directly by the service, or by your app via
PutMetricData) when you control the producer and can emit a number directly. It's cheaper, lower-latency, and easier to aggregate and dashboard. - Reach for a log-based metric (via a metric filter) when you can't change the producer, or you're deriving a new signal from logs you already collect, for example counting
ERRORlines in an existing application log without touching the app's code.
How a metric filter works: it scans incoming log events in a log group for a pattern (plain text, JSON field match, or a filter expression), and on a match increments a CloudWatch metric, a count or an extracted numeric value, with optional dimensions. The metric it produces behaves exactly like a native one: it can back a dashboard or an alarm.
What to actually monitor, by compute type
| Compute type | Core metrics to watch | Note |
|---|---|---|
| EC2 | CPUUtilization, StatusCheckFailed (system and instance), NetworkIn/NetworkOut | Memory and disk usage are not published by default; you need the CloudWatch agent installed on the instance to get those |
| Lambda | Invocations, Errors, Duration, Throttles, ConcurrentExecutions | Throttles and concurrency are the signals that catch you being rate-limited before errors show up |
Wiring an alarm to notify an on-call person: a CloudWatch Alarm evaluates a metric over a configured period and number of evaluation periods; when it transitions from OK to ALARM, it invokes an action, most commonly publishing to an SNS (Simple Notification Service) topic. Subscribers to that topic (email, SMS, a chat-ops integration, or a paging tool subscribed via its own endpoint) receive the notification. The same pattern extends to remediation: an SNS topic can also trigger a Lambda function that takes automated action, not just a page.
Worked example
A concrete alarm-plus-automation workflow for a CPU spike on an EC2 Auto Scaling group: define an alarm on CPUUtilization (statistic: average, period: several minutes, evaluation periods: more than one consecutive breach to avoid reacting to a single noisy sample). On breach, the alarm has two independent actions: (1) publish to an SNS topic that pages on-call, and (2) trigger the Auto Scaling group's scale-out policy so capacity increases automatically. The alarm doesn't choose between paging a human and automating a fix, it does both, because scaling out addresses the immediate load while the page lets a human confirm nothing else is wrong.
Trade-offs & pitfalls
A single noisy data point should not page anyone; setting evaluation periods and datapoints-to-alarm too low is the most common cause of alert fatigue. High-cardinality metric filters (for example, incrementing a metric per unique user ID pulled from a log line) can quietly explode both cost and the number of distinct metrics CloudWatch has to track, so log-based metrics should extract counts or aggregates, not raw identifiers, as dimensions. Finally, treating logs as your primary alerting source when a native metric already exists adds latency (log ingestion and filter evaluation both take time) and cost for no benefit; log-based metrics are a fallback for gaps, not the default.
What are the five (now six, including Sustainability) pillars of the AWS Well-Architected Framework? For each, give one concrete AWS service or practice that demonstrates it in a real production system.
Sample Answer
Direct answer
The AWS Well-Architected Framework has six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability (added in 2021, so a candidate who learned "the five pillars" before then should update that count). Each pillar is a different lens for evaluating an architecture's design decisions, and in practice they trade off against each other rather than all being maximized at once.
Structured elaboration
| Pillar | Focus | Example AWS service or practice |
|---|---|---|
| Operational Excellence | Run and improve systems and processes to deliver business value | A CI/CD pipeline (for example CodePipeline plus infrastructure as code) enabling repeatable, reversible deployments instead of manual changes |
| Security | Protect data, systems, and assets | Least-privilege IAM roles, encryption with AWS KMS, and CloudTrail for an immutable audit log |
| Reliability | Recover from failures and meet availability targets | Multi-AZ RDS with automated backups, and Route 53 (AWS's DNS service) health-check-based failover routing |
| Performance Efficiency | Use computing resources efficiently as demand changes | Auto Scaling matched to actual load, or a serverless compute choice (Lambda) where request patterns are spiky |
| Cost Optimization | Run workloads at the lowest cost that still meets requirements | Commitment discounts (Savings Plans or Reserved Instances) for steady-state load, and S3 lifecycle policies to move cold data to cheaper storage classes |
| Sustainability | Minimize the environmental impact of running workloads | Right-sizing so idle capacity isn't provisioned and run continuously, and shutting down non-production environments outside business hours |
Worked example
An alternate, practical framing worth knowing for an initial customer review: rather than trying to cover the entire framework's checklist in a first meeting, pick one concrete design decision per pillar and use it to anchor the conversation:
- Operational Excellence: "Do you deploy through a repeatable pipeline, or manually?"
- Security: "Are your IAM policies scoped per service, or does everything share one broad role?"
- Reliability: "Is your primary database Multi-AZ, or single-AZ?"
- Performance Efficiency: "Is compute autoscaled to load, or sized once at launch and left fixed?"
- Cost Optimization: "Are you using any commitment discounts, or is everything on-demand?"
- Sustainability: "Do non-production environments get shut down outside business hours, or run continuously?"
Each question is a single, concrete yes/no design decision rather than an abstract pillar description, which makes the framework tangible in a first conversation and gives the architect six specific follow-up threads to go deeper on afterward, rather than a full audit that wouldn't fit in one meeting.
Trade-offs and pitfalls
- Treating the pillars as independent, rather than in tension, is the most common mistake: maximizing Reliability (for example multi-region active-active) directly costs more, so Cost Optimization and Reliability pull in opposite directions, and the right balance depends on the workload's actual availability requirement, not on defaulting to the most reliable option everywhere.
- Naming only five pillars, or forgetting Sustainability specifically, is a common gap for anyone who learned the framework before it was added.
- Picking a token example per pillar without connecting it back to the actual workload's requirements turns the exercise into a checklist rather than a design discussion; the point of naming a concrete decision is to reveal a trade-off worth talking about, not just to demonstrate recall of the pillar names.
You have Lambda functions that need to query a relational database under high concurrency. How would you handle connection pooling, and what role does RDS Proxy play versus reusing connections across warm invocations?
Sample Answer
Direct answer
Put RDS Proxy between Lambda and the database rather than letting each Lambda execution environment open its own connection: Proxy pools and multiplexes a large number of client-facing connections onto a much smaller, stable set of physical database connections, absorbs the connection-storm problem that comes from Lambda's concurrency model, and manages AWS Identity and Access Management (IAM)-auth token rotation for you. Warm-invocation connection reuse, caching a client handle in the Lambda execution environment's global scope, still matters and is complementary, not a substitute: it avoids re-establishing the Proxy-facing connection on every invocation of an already-warm container, while Proxy is what keeps the database itself from seeing thousands of physical connections when Lambda scales out concurrently.
Structured elaboration
The problem RDS Proxy solves
A relational database has a hard ceiling on max_connections, in the low thousands at most, set by instance memory. Lambda can scale to hundreds or thousands of concurrent execution environments in seconds, and if each one opens its own direct connection to the database, you exhaust max_connections almost immediately under a burst, regardless of how efficient the query itself is. RDS Proxy sits in front of the database and:
- Pools and multiplexes: many Lambda-side logical connections share a much smaller pool of physical connections held open to the database, since most connections spend most of their time idle between queries.
- Handles auth centrally: works with IAM database authentication so Lambda does not need database credentials embedded in its code or environment; Proxy manages short-lived credential exchange with the database on the app's behalf.
- Smooths failover: for Aurora, Proxy keeps client-facing connections open across a failover event instead of every client immediately seeing a connection error, reconnecting transparently on the Proxy side.
Configuration knobs that matter
- MaxConnectionsPercent: the percentage of the target database's
max_connectionsthis Proxy, or a specific target group, is allowed to use, leaving headroom for other consumers of the same database. - MaxIdleConnectionsPercent: how many of those connections Proxy is willing to keep idle in the pool versus closing back down.
- ConnectionBorrowTimeout: how long a client waits for the Proxy to hand it a connection before failing; this is your effective backpressure signal when the pool is genuinely saturated.
- Session pinning: certain session-level operations, temp tables, session variables, some transaction patterns, force Proxy to pin a client to one specific physical connection for the rest of that session, which defeats multiplexing for that client. Know which patterns in your queries trigger pinning and avoid them where you can, since they reduce Proxy's actual pooling benefit.
Where warm-invocation reuse still helps
Within a single warm Lambda execution environment, cache the database client as a module-level variable rather than re-creating it on every invocation. This avoids repeating the TLS handshake and Proxy-side connection setup for every invocation of an already-warm container; it does not replace Proxy, because a cold start or a newly spun-up concurrent execution environment still needs a fresh connection, and that is exactly the storm Proxy exists to absorb.
What to do when the pool is genuinely exhausted
- Return a clear rejection with backoff guidance rather than letting the request hang until Lambda's own timeout.
- Buffer non-latency-sensitive writes through a queue and drain them into the database at a controlled rate instead of writing synchronously from every Lambda invocation.
- Keep queries short and transactions small; a long-held transaction ties up a pooled connection and reduces how many other invocations Proxy can serve from the same physical pool.
Worked example
A database sized for a max_connections value of 1,000 fronts a Lambda function that can burst to 2,000 concurrent invocations. Configure the Proxy's target group with MaxConnectionsPercent at 80, leaving 20% headroom for other clients such as an admin console or a batch job:
1000×0.80=800 connections available to this Proxy
With 2,000 concurrent Lambda invocations each needing a connection only for the brief duration of a query, Proxy multiplexes them across those 800 physical connections rather than needing 2,000 physical connections; invocations that arrive while all 800 are briefly in use wait up to the borrow timeout for one to free up, rather than the database itself ever seeing more than 800 connections.
Trade-offs and pitfalls
- RDS Proxy adds a small amount of latency per query, an extra network hop, and its own hourly cost; for very low-concurrency workloads, plain connection reuse in a warm Lambda might be enough and Proxy is unnecessary overhead.
- Session pinning is the most common way teams get less benefit from Proxy than they expected; if your ORM or query patterns rely heavily on session state, audit which patterns trigger pinning.
- Do not rely on Proxy alone to make an unbounded-concurrency Lambda safe for the database; borrow-timeout failures under sustained overload are still failures, just failures the database itself did not see. You still need to address the actual concurrency-to-capacity mismatch, reserved concurrency limits, queue-based buffering, or a bigger database.
- Provisioned Concurrency reduces cold starts, and therefore how often a fresh connection has to be established, but costs money for capacity held ready; it is a traffic-shaping choice independent of whether you are using Proxy, do not treat it as a pooling mechanism by itself.
Design a decoupled architecture where SNS fans out to multiple SQS queues with different consumers. Walk through retry behavior, dead-letter queues, and ordering considerations.
Sample Answer
Direct answer
Put an SNS topic in front of independent, per-consumer SQS queues. A single publish to the topic fans a copy of the message out to every subscribed queue automatically, and each queue gets its own retry behavior, visibility timeout, and dead-letter queue (DLQ), so a failure or backlog in one consumer never blocks or loses data for the others, because each queue is a separate durable buffer.
Structured elaboration
flowchart LR
P[Publisher Service] --> T((SNS Topic))
T --> Q1[SQS: Accounting Queue]
T --> Q2[SQS: Shipping Queue]
T --> Q3[SQS: Analytics Queue]
Q1 --> C1[Accounting Consumer]
Q2 --> C2[Shipping Consumer]
Q3 --> C3[Analytics Consumer]
Q1 -.maxReceiveCount exceeded.-> D1[DLQ: Accounting]
Q2 -.maxReceiveCount exceeded.-> D2[DLQ: Shipping]
Fan-out topology. One SNS topic, one SQS queue subscription per consumer (Accounting-queue, Shipping-queue, Analytics-queue for a PurchasePlaced event). Each consumer only ever reads from its own queue, so it never needs to know the others exist.
Retry and visibility timeout. When a consumer receives a message, SQS starts a visibility timeout during which the message is hidden from other receivers. If the consumer doesn't delete the message before that timeout expires (because it crashed, threw, or is still processing), the message becomes visible again and is redelivered, that's the retry mechanism, driven entirely by the consumer's own success or failure, not by any separate retry configuration.
Dead-letter queues. Configure a redrive policy with a maxReceiveCount on each queue pointing at a DLQ. After a message has been received that many times without being deleted, SQS routes it to the DLQ instead of retrying again, which is what actually stops a poison message from looping. Without a redrive policy configured, there's no automatic cap: a failing message keeps cycling in and out of visibility indefinitely.
Ordering. Standard SNS-to-SQS fan-out is at-least-once with best-effort ordering, which is fine for most of these consumers. If exactly one consumer (say Accounting, applying debits and credits in order per account) genuinely needs strict per-key ordering, that's a per-consumer requirement, not a topic-wide one. Since September 2023, an SNS FIFO topic supports delivery to both SQS FIFO and SQS Standard queues at the same time, so a mixed ordering requirement no longer forces a second topic: use one FIFO topic, subscribe Accounting with a FIFO queue (message-group ID set to the account ID) for strict per-account ordering, and subscribe Shipping and Analytics with ordinary Standard queues on that same topic for best-effort, lower-cost delivery. Every publish to a FIFO topic still needs a MessageGroupId, regardless of which subscriber types are attached, since ordering and deduplication are properties of the topic's publish contract, not of an individual subscription.
Worked example
A PurchasePlaced event is published once to the Standard SNS topic and fanned out to all three queues. The Accounting consumer throws an exception partway through processing one message. That message's visibility timeout expires, it becomes visible again, and Accounting's consumer receives it again (retry). Suppose the redrive policy on Accounting-queue is configured with maxReceiveCount: 5; after the fifth failed receive, SQS routes that message to DLQ: Accounting instead of retrying a sixth time, where it sits for manual inspection or a scripted replay. Meanwhile Shipping-queue and Analytics-queue each received their own independent copy of the same PurchasePlaced message and were entirely unaffected by Accounting's failure, because they're separate queues with separate state.
Trade-offs and pitfalls
- Assuming an SNS FIFO topic can only fan out to FIFO SQS queues is an outdated assumption; since September 2023 a single FIFO topic can mix FIFO and Standard queue subscribers, so a per-consumer ordering requirement discovered late only means adding a FIFO queue subscription for that one consumer, not standing up a second topic.
- Relying on retries without a DLQ means a poison message (one that will never succeed, such as one with malformed data) loops indefinitely, quietly consuming consumer capacity until someone notices.
- Both SNS delivery to a queue and SQS delivery to a consumer are at-least-once, not exactly-once (Standard queues); treating a message as guaranteed to be processed only once, without building idempotent handling in each consumer, risks duplicate side effects such as double-shipping an order.
Unlock Full Question Bank
Get access to all AWS Core Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.