AWS Core Services and Architecture Questions
Amazon Web Services' core service catalog and how the pieces compose into a working system: EC2, Lambda, S3, VPC, IAM, RDS, and the managed-service ecosystem. Covers service selection within AWS, common reference architectures, the AWS Well-Architected Framework pillars, and operational patterns specific to the platform. For provider-agnostic compute or storage trade-offs, see the cross-cloud entries.
You notice increased end-to-end request latency for a microservice. Walk through the diagnostic steps using CloudWatch metrics, ALB metrics, X-Ray, and logs: which would you check first, and what patterns tell you infrastructure versus application problem?
Sample Answer
Work from the outside in: start with Application Load Balancer (ALB) metrics since they show whether the problem sits between the client and your service or inside it, then use AWS X-Ray to find which downstream call within the request is actually slow, then drop into logs at that specific span to find the root cause. Amazon CloudWatch instance-level metrics answer one narrow question: is the underlying compute starved of resources. The pattern that separates infrastructure from application problems is whether the slowness correlates with resource saturation (CPU, memory, disk queue) or with a specific downstream call showing up consistently in X-Ray traces.
Diagnostic order and what each layer tells you
- ALB metrics first:
TargetResponseTime(time the backend took to respond),HTTPCode_Target_5XX_Count,HealthyHostCount, andRequestCountPerTarget(average request load per target in the target group; ALB has no client-facing queue-depth metric the way Classic Load Balancers do). RisingTargetResponseTimewith a steadyHealthyHostCountpoints at the application; a risingRequestCountPerTargetalongside a droppingHealthyHostCountpoints at capacity, the same targets absorbing more load because fewer of them are healthy. - CloudWatch instance and service metrics: CPU utilization, memory (if the CloudWatch agent publishes it), disk queue length, and status check failures. Sustained CPU above roughly 70 to 80 percent, a growing disk queue, or a failed status check is a real infrastructure signal; normal-looking infrastructure metrics next to a slow ALB response time is a strong sign the problem lives in application code or a downstream dependency, not the host.
- X-Ray traces: the service map and per-trace latency breakdown show which segment of the request is actually slow, an internal computation, a database call, or a third-party API. A long span on a downstream call, a slow SQL query, a saturated connection pool, is an application or dependency problem even though it shows up as backend latency at the ALB.
- Logs: once X-Ray points at a specific segment, application logs, database slow-query logs, and container standard output around that timestamp and trace ID confirm the actual cause, a stack trace, a lock wait, a garbage-collection pause, or a connection pool timeout.
Patterns that separate infrastructure from application
- Infrastructure: CPU, memory, or disk saturation across many hosts, failing status checks, or
RequestCountPerTargetrising alongside a drop inHealthyHostCount, all independent of what any individual trace shows. - Application: normal host-level metrics, but
TargetResponseTimeand X-Ray both point at a specific span (a database call, a downstream API, a lock), and logs at that timestamp show exceptions, long garbage-collection pauses, or connection pool exhaustion. - Mixed, and easy to misdiagnose: Auto Scaling Group cooldown delays, or a database hitting its own connection or throughput limit, look infrastructure-shaped on a dashboard but are actually caused by an upstream traffic pattern or an application-side connection leak.
Tying this to service-level objectives and alerting
Set alarm thresholds on the same metrics used above, tied to an actual service-level objective (SLO), for example 99.9 percent of requests under 500 milliseconds, rather than an arbitrary round number, so an alarm firing means the error budget (the allowed amount of failure or excess latency for the period) is actually at risk. Track the error budget's burn rate rather than paging on every threshold breach, since a two-minute blip that recovers on its own should not wake anyone up, while a sustained burn that will exhaust the monthly budget in a few hours should.
Worked example
Checkout latency creeps from a 99th-percentile of 200 milliseconds to 1.2 seconds over 20 minutes. ALB metrics show TargetResponseTime climbing while HealthyHostCount stays flat and RequestCountPerTarget stays flat too, ruling out a capacity problem. CloudWatch shows CPU at 25 percent across all targets, ruling out host saturation. X-Ray's service map shows the slow span is consistently a call to the inventory service, with its own latency climbing in the trace timeline. Logs on the inventory service around that window show a specific slow-query pattern: a full table scan on an unindexed column that only becomes slow once the table crossed a size threshold. That is a database and application problem end to end, not infrastructure, and the fix, adding an index, has nothing to do with scaling anything.
Trade-offs and pitfalls
- Jumping straight to logs before narrowing scope with ALB and CloudWatch metrics wastes time searching a haystack. The metrics exist specifically to tell you where to look before you start reading logs.
- X-Ray requires sampling and instrumentation to already be in place before the incident. If tracing is not enabled, or the sampling rate is too low to capture the slow requests, you lose the fastest path to the root cause and fall back to correlating logs by timestamp, which is slower and noisier.
- Alarming on raw metric thresholds instead of an SLO-tied burn rate produces either too many pages (threshold too tight) or missed real degradations (threshold too loose). Revisit thresholds against actual historical percentiles, not a guess.
- A downstream dependency showing up as "your" latency in X-Ray is still your incident to manage even though the root cause sits in a service you do not own. Know the escalation path to that team before you need it.
You have Lambda functions that need to query a relational database under high concurrency. How would you handle connection pooling, and what role does RDS Proxy play versus reusing connections across warm invocations?
Sample Answer
Direct answer
Put RDS Proxy between Lambda and the database rather than letting each Lambda execution environment open its own connection: Proxy pools and multiplexes a large number of client-facing connections onto a much smaller, stable set of physical database connections, absorbs the connection-storm problem that comes from Lambda's concurrency model, and manages AWS Identity and Access Management (IAM)-auth token rotation for you. Warm-invocation connection reuse, caching a client handle in the Lambda execution environment's global scope, still matters and is complementary, not a substitute: it avoids re-establishing the Proxy-facing connection on every invocation of an already-warm container, while Proxy is what keeps the database itself from seeing thousands of physical connections when Lambda scales out concurrently.
Structured elaboration
The problem RDS Proxy solves
A relational database has a hard ceiling on max_connections, in the low thousands at most, set by instance memory. Lambda can scale to hundreds or thousands of concurrent execution environments in seconds, and if each one opens its own direct connection to the database, you exhaust max_connections almost immediately under a burst, regardless of how efficient the query itself is. RDS Proxy sits in front of the database and:
- Pools and multiplexes: many Lambda-side logical connections share a much smaller pool of physical connections held open to the database, since most connections spend most of their time idle between queries.
- Handles auth centrally: works with IAM database authentication so Lambda does not need database credentials embedded in its code or environment; Proxy manages short-lived credential exchange with the database on the app's behalf.
- Smooths failover: for Aurora, Proxy keeps client-facing connections open across a failover event instead of every client immediately seeing a connection error, reconnecting transparently on the Proxy side.
Configuration knobs that matter
- MaxConnectionsPercent: the percentage of the target database's
max_connectionsthis Proxy, or a specific target group, is allowed to use, leaving headroom for other consumers of the same database. - MaxIdleConnectionsPercent: how many of those connections Proxy is willing to keep idle in the pool versus closing back down.
- ConnectionBorrowTimeout: how long a client waits for the Proxy to hand it a connection before failing; this is your effective backpressure signal when the pool is genuinely saturated.
- Session pinning: certain session-level operations, temp tables, session variables, some transaction patterns, force Proxy to pin a client to one specific physical connection for the rest of that session, which defeats multiplexing for that client. Know which patterns in your queries trigger pinning and avoid them where you can, since they reduce Proxy's actual pooling benefit.
Where warm-invocation reuse still helps
Within a single warm Lambda execution environment, cache the database client as a module-level variable rather than re-creating it on every invocation. This avoids repeating the TLS handshake and Proxy-side connection setup for every invocation of an already-warm container; it does not replace Proxy, because a cold start or a newly spun-up concurrent execution environment still needs a fresh connection, and that is exactly the storm Proxy exists to absorb.
What to do when the pool is genuinely exhausted
- Return a clear rejection with backoff guidance rather than letting the request hang until Lambda's own timeout.
- Buffer non-latency-sensitive writes through a queue and drain them into the database at a controlled rate instead of writing synchronously from every Lambda invocation.
- Keep queries short and transactions small; a long-held transaction ties up a pooled connection and reduces how many other invocations Proxy can serve from the same physical pool.
Worked example
A database sized for a max_connections value of 1,000 fronts a Lambda function that can burst to 2,000 concurrent invocations. Configure the Proxy's target group with MaxConnectionsPercent at 80, leaving 20% headroom for other clients such as an admin console or a batch job:
1000×0.80=800 connections available to this Proxy
With 2,000 concurrent Lambda invocations each needing a connection only for the brief duration of a query, Proxy multiplexes them across those 800 physical connections rather than needing 2,000 physical connections; invocations that arrive while all 800 are briefly in use wait up to the borrow timeout for one to free up, rather than the database itself ever seeing more than 800 connections.
Trade-offs and pitfalls
- RDS Proxy adds a small amount of latency per query, an extra network hop, and its own hourly cost; for very low-concurrency workloads, plain connection reuse in a warm Lambda might be enough and Proxy is unnecessary overhead.
- Session pinning is the most common way teams get less benefit from Proxy than they expected; if your ORM or query patterns rely heavily on session state, audit which patterns trigger pinning.
- Do not rely on Proxy alone to make an unbounded-concurrency Lambda safe for the database; borrow-timeout failures under sustained overload are still failures, just failures the database itself did not see. You still need to address the actual concurrency-to-capacity mismatch, reserved concurrency limits, queue-based buffering, or a bigger database.
- Provisioned Concurrency reduces cold starts, and therefore how often a fresh connection has to be established, but costs money for capacity held ready; it is a traffic-shaping choice independent of whether you are using Proxy, do not treat it as a pooling mechanism by itself.
Design a Step Functions workflow to orchestrate a multi-stage batch job (for example, ingest, transform, and load) that's triggered by an S3 upload or a schedule. How do you handle a single step failing partway through, retries, and notifying on final failure?
Sample Answer
Direct answer
Model each stage (ingest, transform, load) as its own Step Functions state, triggered by an EventBridge rule on an S3 PutObject event or a scheduled rule. Give each state a Retry block for transient errors (exponential backoff, a bounded number of attempts) and a Catch block that routes any state's failure to a shared failure-handler state, which records the failure and sends a notification, rather than letting the execution just fail silently.
Structured elaboration
flowchart TB
T1["EventBridge: S3 upload event
or schedule"] --> SF["Step Functions
state machine starts"]
SF --> ING["Ingest state
Retry: transient errors, backoff
Catch: routes to Failure handler"]
ING --> XFM["Transform state
Retry + Catch"]
XFM --> LOAD["Load state
Retry + Catch"]
LOAD --> DONE["Success:
notify via SNS"]
ING -.Catch.-> FAIL["Failure handler state"]
XFM -.Catch.-> FAIL
LOAD -.Catch.-> FAIL
FAIL --> SNS["SNS notification
with execution ARN and error"]
FAIL --> DLQ["Write failure record
for manual replay"]
- Trigger: an EventBridge rule matches
S3:ObjectCreatedfor the ingest bucket/prefix, or a scheduled rule (rate(...)/cron(...)) for a recurring batch, and starts a Step Functions execution with the triggering event as input. - Retry: each of ingest/transform/load gets a
Retryfield specifying which error types to retry (e.g., throttling or transient 5xx from the underlying service), a backoff rate, and a maximum attempt count, so a brief downstream blip is absorbed automatically instead of failing the whole run. - Catch (partial failure): each state's
Catchfield routes to a sharedFailureHandlerstate on any error that exhausts its retries, passing along the error info and which stage failed. The failure handler is a single place that publishes an SNS notification, including the execution's Amazon Resource Name (ARN), the unique identifier you use to look up that specific run, and error detail, and records the failure for later inspection, for example writing it to a dead-letter queue (DLQ), a holding area for failed items that need manual review or replay, rather than duplicating that logic in every state. - Idempotent steps: because a failed-then-retried execution may re-run an already-partially-completed step, make each stage idempotent where possible, writing intermediate output to a deterministic path so a retried "load" step overwrites rather than duplicates, or checking a completion marker before redoing work.
- Standard vs Express workflows: use a Standard workflow here: it's durable, gives you full execution history, and (unlike Express) supports a 14-day redrive window for a completed run. Keep the audit window realistic though: Step Functions retains a closed execution's history for 90 days by default (reducible to 30 days via a support request, but not extendable beyond 90), separate from and much shorter than the 1-year maximum execution duration a single run is allowed to take; if you need to audit a batch run's history longer than that, export it (e.g., to CloudWatch Logs or S3) rather than relying on the console. Express workflows trade the durable history off for higher throughput and lower cost, which fits high-volume, short-duration workloads better than a periodic batch job.
Worked example
A nightly batch: EventBridge's scheduled rule starts the state machine at a fixed time. Ingest reads the day's raw files from S3, transform reshapes and validates them, load writes the result into the destination store. Suppose transform fails on one execution because a downstream validation service returned a transient 503: the Retry block on the transform state retries it two or three times with backoff; if it keeps failing (a genuine data problem, not a blip), Catch routes to the failure handler, which notifies on-call via SNS with the execution ARN so they can open the specific failed run in the Step Functions console and see exactly which state and input caused it, rather than re-running the whole pipeline blind. Because ingest already succeeded and its output was already written to a stable S3 path, a corrected re-run can start from transform instead of re-ingesting from scratch, provided the pipeline is designed to check for and reuse already-completed stage output.
Trade-offs and pitfalls
- Don't let
Retrymask a real, non-transient failure: distinguish retryable error types (throttling, transient network errors) from ones that should fail fast (a malformed input file that will never succeed on retry), or you'll burn time and money retrying something that was never going to work. - A single shared failure-handler state keeps notification logic in one place, but make sure it still surfaces which stage failed and with what input; a generic "pipeline failed" alert without that context pushes the diagnosis work onto whoever gets paged.
- For genuinely large-scale fan-out (many thousands of items per run, well beyond what a normal
Mapstate's execution-history limits comfortably handle), Step Functions' Distributed Map mode runs each iteration as its own child workflow execution and can launch on the order of ten thousand concurrent iterations; that's more machinery than a typical three-stage batch pipeline needs, but it's the right tool once item counts or per-item state genuinely outgrow a standard Map state. - Test the failure path deliberately (inject a forced failure in staging) rather than assuming the
Catchblock works as intended; a silently-misconfigured catch (wrong error type match) fails the same way a missing one does, just later and more confusingly.
Your organization will provision hundreds of VPCs across accounts and regions. How would you design an IP address management strategy to avoid overlapping CIDRs and support automated provisioning of new VPCs at scale?
Sample Answer
Direct answer
Design a hierarchical, deterministic Classless Inter-Domain Routing (CIDR) allocation scheme (a large private block split predictably by region, then account, then Virtual Private Cloud, VPC), register every allocation in a central IP Address Management (IPAM) system so overlaps are structurally impossible, and automate allocation through the same infrastructure-as-code pipeline teams already use to provision VPCs, rather than letting anyone hand-pick a CIDR. AWS's own VPC IPAM service is the natural backbone for the registry and automation, but the hierarchy and guardrails matter more than which specific tool enforces them.
Design components
1. Hierarchical CIDR plan. Start from a large private range (for example, 10.0.0.0/8) and subdivide by a fixed number of bits at each level: region, then account, then VPC. Because each level takes a fixed number of bits, any team can compute their exact allocation from a simple formula instead of asking a human to hand out the next free block, and there's no way for two teams to land on the same range if they each follow the formula.
2. Central IPAM registry as the single source of truth. AWS VPC IPAM (or an equivalent central tool) tracks every allocation, its owner, its lifecycle state, and enforces that a new request can't overlap an existing one. This is what turns "we have a naming convention" into "overlap is actually impossible," since the registry, not convention, is what blocks a conflicting request.
3. Automated provisioning through infrastructure as code. A Terraform (or equivalent) module calls the IPAM registry to allocate a CIDR at VPC-creation time, uses the returned block to create the VPC and its subnets, and releases the allocation back to the pool through the same automation when the VPC is destroyed. This keeps allocation and infrastructure lifecycle in the same auditable pipeline instead of two systems that can drift apart.
4. Guardrails, not just tooling. Policy-as-code checks (AWS Config rules or a policy engine in the CI pipeline) reject any manually-created VPC that didn't go through the IPAM-integrated path, so the automation is the only supported way to get a CIDR, not just the convenient one.
Allocation hierarchy diagram
flowchart TD
Global["10.0.0.0/8\nglobal private range"]
Region["Region block: /12\n(4 bits for region)"]
Account["Account block: /20\n(8 bits for account)"]
VPC["VPC block: /22\n(2 bits for VPC)"]
Global -->|"split by region"| Region
Region -->|"split by account"| Account
Account -->|"split by VPC"| VPC
Worked example: sizing the hierarchy
Start from 10.0.0.0/8, which has 32−8=24 host bits.
Reserve 4 bits for region, giving a /12 block per region (8+4=12), with room for 24=16 regions.
Within each /12 region block (32−12=20 host bits available), reserve 8 bits for account, giving a /20 block per account:
accounts per region=220−12=28=256
Within each /20 account block (32−20=12 host bits available), reserve 2 bits for VPC, giving a /22 block per VPC (1,024 addresses, a common VPC size):
VPCs per account=222−20=22=4
Total addressable capacity under this scheme:
16 regions×256 accounts/region×4 VPCs/account=16,384 VPCs
That comfortably covers "hundreds of VPCs" today with two orders of magnitude of headroom, using a single, deterministic formula (region bits, then account bits, then VPC bits) that any automation can compute without querying a human.
Trade-offs and pitfalls
- Fixed bit-widths trade flexibility for predictability. A region or account that needs more than its allotted block requires either reclaiming unused space or accepting a discontiguous secondary allocation; that's the cost of a formula anyone can compute versus a system that hands out variable-size blocks on request.
- Overly generous per-VPC sizing wastes address space fast. A reflexive
/16per VPC (65,536 addresses) for workloads that need a few hundred would exhaust the hierarchy's headroom far sooner than the/22example above; size CIDRs to expected host counts, not habit. - Peering and Transit Gateway route-table limits are a downstream constraint: even a perfectly non-overlapping CIDR scheme can hit route-table entry limits if too many VPCs peer directly; a hub-and-spoke Transit Gateway topology scales further than full-mesh peering as VPC count grows.
- Reclamation discipline matters as much as allocation. Dev/test VPCs that are provisioned but never torn down slowly consume the address pool exactly like a memory leak; TTL-based automatic reclamation for non-production allocations keeps the registry honest.
- Migrating an existing overlapping VPC into the new scheme is a live-traffic re-IP, not a data-model exercise; it needs its own careful, and typically incremental, migration plan.
What's the relationship between CloudWatch Metrics, CloudWatch Logs, and CloudWatch Alarms? When would you prefer a metric over a log-based metric, and how would you wire an alarm to notify an on-call person?
Sample Answer
Direct answer
The three work together as a pipeline. CloudWatch Metrics store numeric time series (CPU utilization, request latency, error counts). CloudWatch Logs store raw event text (application log lines, stack traces, audit events). CloudWatch Alarms watch a metric, either a native one or one derived from logs, against a threshold over a time window, and fire an action when it's breached. Metrics are what you alert on; logs are what you read to understand why.
Structured elaboration
When to prefer a native metric over a log-based metric
- Prefer a native metric (emitted directly by the service, or by your app via
PutMetricData) when you control the producer and can emit a number directly. It's cheaper, lower-latency, and easier to aggregate and dashboard. - Reach for a log-based metric (via a metric filter) when you can't change the producer, or you're deriving a new signal from logs you already collect, for example counting
ERRORlines in an existing application log without touching the app's code.
How a metric filter works: it scans incoming log events in a log group for a pattern (plain text, JSON field match, or a filter expression), and on a match increments a CloudWatch metric, a count or an extracted numeric value, with optional dimensions. The metric it produces behaves exactly like a native one: it can back a dashboard or an alarm.
What to actually monitor, by compute type
| Compute type | Core metrics to watch | Note |
|---|---|---|
| EC2 | CPUUtilization, StatusCheckFailed (system and instance), NetworkIn/NetworkOut | Memory and disk usage are not published by default; you need the CloudWatch agent installed on the instance to get those |
| Lambda | Invocations, Errors, Duration, Throttles, ConcurrentExecutions | Throttles and concurrency are the signals that catch you being rate-limited before errors show up |
Wiring an alarm to notify an on-call person: a CloudWatch Alarm evaluates a metric over a configured period and number of evaluation periods; when it transitions from OK to ALARM, it invokes an action, most commonly publishing to an SNS (Simple Notification Service) topic. Subscribers to that topic (email, SMS, a chat-ops integration, or a paging tool subscribed via its own endpoint) receive the notification. The same pattern extends to remediation: an SNS topic can also trigger a Lambda function that takes automated action, not just a page.
Worked example
A concrete alarm-plus-automation workflow for a CPU spike on an EC2 Auto Scaling group: define an alarm on CPUUtilization (statistic: average, period: several minutes, evaluation periods: more than one consecutive breach to avoid reacting to a single noisy sample). On breach, the alarm has two independent actions: (1) publish to an SNS topic that pages on-call, and (2) trigger the Auto Scaling group's scale-out policy so capacity increases automatically. The alarm doesn't choose between paging a human and automating a fix, it does both, because scaling out addresses the immediate load while the page lets a human confirm nothing else is wrong.
Trade-offs & pitfalls
A single noisy data point should not page anyone; setting evaluation periods and datapoints-to-alarm too low is the most common cause of alert fatigue. High-cardinality metric filters (for example, incrementing a metric per unique user ID pulled from a log line) can quietly explode both cost and the number of distinct metrics CloudWatch has to track, so log-based metrics should extract counts or aggregates, not raw identifiers, as dimensions. Finally, treating logs as your primary alerting source when a native metric already exists adds latency (log ingestion and filter evaluation both take time) and cost for no benefit; log-based metrics are a fallback for gaps, not the default.
Unlock Full Question Bank
Get access to all AWS Core Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.