Serverless and Function-as-a-Service Architecture Questions
Building applications on managed, event-triggered compute: functions-as-a-service (AWS Lambda, Azure Functions, Google Cloud Functions, Cloudflare Workers) and serverless containers. Covers the invocation lifecycle and cold starts (init vs handler work, provisioned concurrency, packaging, layers and container images), statelessness and externalizing state, execution limits (timeouts, memory, payload and /tmp size), concurrency and scaling behavior (account limits, burst scaling, protecting downstream databases with connection proxies and throttling), event sources and trigger semantics (at-least-once delivery, retries, idempotent handlers, dead-letter handling), composing functions with managed services and workflow orchestrators (Step Functions and equivalents), and serverless-specific observability, security (per-function IAM, secrets) and pay-per-invocation cost modeling. Includes when serverless fits versus containers or VMs, and vendor lock-in trade-offs. General compute selection, generic messaging patterns, and general idempotency theory are covered by their own topics.
Design an observability architecture for a serverless platform: end-to-end distributed tracing from the API gateway through the function to whatever it calls downstream, plus service-level metrics and anomaly detection, all while keeping the observability bill under control. What sampling strategy would you use, where do traces and metrics get stored, and how do you correlate telemetry across an async hop to debug one specific failed request?
Sample Answer
Direct answer
Instrument every hop with OpenTelemetry (the open standard and SDKs for traces, metrics and logs), propagate one trace context through the API gateway, the functions and every message across async hops, and put a correlation ID in every log line. Keep costs bounded with low head sampling for successful traffic plus guaranteed capture of failures, cheap aggregated metrics for SLOs (service-level objectives) and anomaly detection, and logs with short hot retention and cheap archive. To debug one failed request, walk from the request ID the client saw, to the trace ID in its logs, to the trace, across the queue via its link to the consumer's span.
Requirements I am designing to
- 100 requests per second average at the gateway (259.2M requests per month), each producing a sync function call and an async message to a worker.
- p99 (the 99th-percentile latency: the value only the slowest 1% of requests exceed) latency SLO on the API; error-budget alerting (paging when the service has spent too much of the failure rate its SLO allows, rather than on every individual error; the full mechanism is below).
- Find any specific failed request within minutes, including failures that happen after the async hop.
- Observability spend should be a small fraction of compute spend, with arithmetic below.
Architecture
flowchart LR
C[Client] --> G[API gateway]
G --> F1[API function]
F1 -->|message plus trace context| Q[Queue]
Q --> F2[Worker function]
F2 --> D[Downstream API or DB]
F1 -.spans, EMF metrics, logs.-> T[Telemetry pipeline]
F2 -.spans, EMF metrics, logs.-> T
T --> TR[Trace store]
T --> M[Metrics store]
T --> L[Log store plus archive]
Components
- Tracing: the AWS Distro for OpenTelemetry Lambda layer (AWS's own supported build of the open-source OpenTelemetry instrumentation, attached to a function without changing its code) or native AWS X-Ray (AWS's distributed tracing service) on each function. Spans (timed records of one unit of work) cover the gateway, handler, each downstream call, and the enqueue.
- Metrics: SLO metrics (requests, errors, latency histograms) emitted as structured logs in the CloudWatch embedded metric format (EMF), which the platform turns into metrics without a synchronous API call inside the function; or a Prometheus-compatible managed store (Prometheus is a widely used open-source metrics system; a managed store that speaks its query language) if the organization standardizes on it.
- Logs: structured JSON with
request_id,trace_id,function_version, outcome. Kept in hot retention (the fast, queryable store, as opposed to cheap cold archive) for 14 days, then archived to object storage. - Anomaly detection: statistical bands (a normal range computed from a metric's own recent history, so an alert fires only once it moves outside that range) on a few high-value series (error rate, p99, throughput per route), plus multi-window burn-rate alerts: an SLO like 99.9% availability allows a fixed amount of failure over its measurement window, the error budget (here, 0.1% of requests over 30 days); a burn-rate alert pages when errors are consuming that budget faster than a sustainable rate, checked over both a short window (to catch a real incident fast) and a long window (so a brief blip does not page anyone) at once, which is why it is "multi-window" (for example, page when 2% of the monthly error budget is consumed in one hour). Anomaly detection runs on aggregated metrics, never on raw traces, which keeps it cheap.
Propagating context across the async hop
The trace context (a trace ID plus parent span ID, the ID of the span that caused this one, in the W3C (World Wide Web Consortium, the standards body) traceparent header or AWS's X-Amzn-Trace-Id) must ride inside the message, because there is no HTTP connection between producer and consumer.
- SQS (Amazon Simple Queue Service) to Lambda on AWS: SQS carries the X-Ray trace header in the reserved
AWSTraceHeadermessage system attribute, and a Lambda consumer picks it up automatically; the console shows the producer's trace linked to the consumer's. - Other brokers (EventBridge, Kafka (Apache Kafka, a distributed event-streaming log), SNS fan-out, one SNS message delivered to every subscriber at once): inject
traceparentinto message attributes or headers at publish and extract it in the consumer. - Batches: one consumer invocation may process 10 messages from 10 different traces. Model the consumer's work as its own span with span links (references from one span to other related spans that are not its direct parent or children, so one span can point back to many origins) to each message's producing span, rather than pretending it has a single parent.
- Redundant safety net (belt and braces): also put a business correlation ID (order ID, request ID) in the message body and in every log line, so the chain can be reconstructed even for unsampled traces.
Sampling strategy and the cost arithmetic
Prices: X-Ray records traces at $5.00 per million, and CloudWatch Logs standard ingestion is $0.50 per GB (US East, current list).
N=100×2,592,000=259,200,000 traces per month 100% sampling=259.2×5.00=$1,2965%=$64.801%=$12.96Decision: head-sample 5% of successful requests (the decision is made at the entry point and propagated, so a trace is complete or absent, never half-recorded), and capture 100% of errors and slow requests. Capturing all errors needs tail sampling: deciding after the trace finishes. A function cannot do that alone because it sees only its own spans, so either route spans through an OpenTelemetry Collector gateway (a standalone process that receives every function's spans and applies rules to them centrally, before they reach the trace store) running tail-sampling rules (errors, latency above the SLO threshold, specific tenants), or, if you stay on head sampling only, guarantee that every error log line carries the trace ID and full context, so unsampled failures are still debuggable from logs.
Logs: 1 KB per request is 259.2 GB per month, $129.60 of ingestion. Log at INFO for outcomes only, DEBUG sampled, errors in full.
Metrics cardinality is where observability bills explode. Each unique combination of metric name and dimension values is billed as a separate custom metric ($0.30 per metric per month for the first 10,000, $0.10 for the next tier). Adding a customer_id dimension across 10,000 customers for 3 metrics creates 30,000 metrics:
That is more than all tracing at 5%. Per-customer analysis belongs in logs or traces (queried on demand), not metric dimensions.
Where telemetry is stored
| Data | Store | Retention |
|---|---|---|
| Traces | X-Ray or an OpenTelemetry-compatible trace backend | 7 to 30 days; long-term value is low |
| SLO metrics | CloudWatch or managed Prometheus | 15 months; cheap because aggregated |
| Logs | Log service, hot | 14 days, then object-storage archive queried on demand |
Debugging one specific failed request
- Client reports failure and returns the
request_idfrom the error response (always return it). - Log query on
request_idfinds the API function's line and itstrace_id. - Open the trace: gateway span, handler span, enqueue span, status OK. The failure is after the hop.
- Follow the link from the enqueue span to the worker's span (automatic for SQS to Lambda; via span links for batches).
- The worker span shows a 5 s downstream timeout; its log lines, filtered by the same
trace_id, show the payload that triggered it; the message is in the dead-letter queue, ready to redrive (resend through the same processing path for another attempt) after the fix.
If the request was not in the 5% sample and tail sampling is not deployed, steps 3 and 4 fall back to logs by trace_id and correlation ID, which is why the log discipline is non-negotiable.
Pitfalls
- Head sampling only, then discovering the failure you need was in the 95% dropped.
- Losing trace context at the queue because the producer never injected it.
- High-cardinality metric dimensions.
- Synchronous telemetry exports inside the handler that add latency and billed duration to every request.
A Lambda-based function is CPU-bound and spends 60% of its cold-start time loading a large dependency into memory. Propose concrete optimizations (code changes, packaging, memory/CPU sizing, provisioned concurrency) to reduce cold-start overhead and improve throughput. For each optimization, state how you'd measure success and any side effects.
Sample Answer
Direct answer
Treat it as a measurement problem before an optimization problem. Split the cold start into its parts, then attack the 60% in order of cost: make the dependency cheaper to load (load less, load a faster format, load it once at module scope), package it so less has to be fetched and unpacked, give the function more memory, which on Lambda also means more CPU, and only then pay to hide what remains with provisioned concurrency. Because the function is CPU-bound, raising memory is unusually effective here: up to one full vCPU it can cut both init and handler time roughly in proportion with little change in cost per call.
Step 0: baseline, so every change is measurable
Record for at least a day (or a load test):
Init Durationfrom theREPORTlog line of cold invocations (p50 and p99), and the cold-start rate (cold invocations / all invocations).- Your own timers inside init, for example one around the dependency load and one around everything else, so you can confirm the "60%".
Durationp50/p99 for warm invocations (the handler), andMax Memory Usedfrom theREPORTline.- Cost per million invocations: requests plus GB-seconds (configured memory in GB multiplied by billed seconds).
Assume the baseline is: memory 1,024 MB, Init Duration 5.0 s of which the dependency load is 3.0 s (60%), warm handler 800 ms.
The optimizations
1. Code changes: make the load itself cheaper
- Load once, at module scope, never inside the handler, and never re-load per request.
- Load less. If only part of the dependency is needed (some tables, some model layers, one language's data), split it and load the needed part, or lazy-load the rarely used parts.
- Load a faster format. A lot of "loading" is really parsing: turning JSON, pickled objects (Python's built-in serialized-object format) or text into in-memory structures, which is CPU-bound. Pre-convert at build time into a format that needs little or no parsing (for example a binary layout you can memory-map, so pages, fixed-size chunks of the file, are read into memory on demand as they are touched instead of all up front).
- Measure: the dependency timer and
Init Durationp50/p99. Side effects: lazy loading moves the delay to the first request that touches the lazy part; a new format adds a build step that must stay in sync with the source data.
2. Packaging: fetch and unpack less
- Strip what is not needed (tests, docs, unused native binaries, compiled machine code bundled with the dependency rather than portable source files, and debug symbols, extra metadata compiled in for debuggers that is not needed at runtime) and check for duplicated libraries.
- If the dependency is large, ship it as a container image (up to 10 GB, versus 250 MB unzipped for zip packages including layers, separate zip archives of shared code that a function can attach at deploy time, counted toward the same 250 MB limit) built on an AWS base image, with the dependency in an early, rarely changing image layer (a cached, reusable slice of the image, a different concept from a Lambda layer; images are built by stacking layers, and unchanged ones do not need re-fetching). Lambda fetches container images in chunks as they are read and caches them, so what you read at init matters more than total size.
- Measure: package or image size,
Init Duration. Side effects: container images need a registry (a versioned storage service that Lambda pulls images from, for example Amazon ECR, Elastic Container Registry) and a build pipeline; a zip with layers is simpler if it fits.
3. Memory and CPU sizing
Lambda allocates CPU in proportion to memory; at 1,769 MB a function has the equivalent of one vCPU. At 1,024 MB it gets about 1,024 / 1,769 = 0.58 of one. For a single-threaded CPU-bound load, and assuming time scales inversely with CPU share (an ideal to verify, not a promise):
t1769=t1024×17691024| 1,024 MB | 1,769 MB | |
|---|---|---|
| Dependency load (init) | 3.0 s | 3.0 x 0.579 = 1.74 s |
| Warm handler | 800 ms | 463 ms |
| Handler GB-seconds per call | 1.0 GB x 0.800 s = 0.80 | 1.7275 GB x 0.463 s = 0.80 |
Cost per call is unchanged while latency drops by about 42%, and each environment completes about 1.7 times more requests per second, so the same traffic needs fewer concurrent environments (which means fewer cold starts). Above 1,769 MB extra vCPUs are added, but they help only if the code is multi-threaded; for single-threaded code you pay more for no speed.
- Measure: run the same test at several memory sizes. The open-source AWS Lambda Power Tuning tool automates this sweep and plots time against cost. Side effect: if the code turns out not to be CPU-bound, the extra memory is pure cost.
4. Provisioned concurrency: hide what is left
Keep N environments initialized so the remaining init never lands on a user request. With provisioned concurrency the Init phase runs when capacity is allocated, not when a user waits.
- Measure: cold-start rate on the provisioned alias (a named pointer, such as
live, to one published version of the function, which provisioned concurrency attaches to) should approach zero; watchProvisionedConcurrencyUtilization(how much of N is in use) andProvisionedConcurrencySpilloverInvocations(requests that overflowed N onto on-demand environments and may be cold). - Side effects: a standing charge (at 1,769 MB, one environment for a 30-day month is 1.7275 GB x 2,592,000 s x $0.0000041667 per GB-s = $18.66 in US East, N. Virginia); requires published versions and an alias; bursts beyond N still cold-start.
5. Alternatives worth one line each
- SnapStart (Java 11+, Python 3.12+, .NET 8+) snapshots the initialized environment, so a heavy dependency load is paid at publish time instead of per cold start. It cannot be combined with provisioned concurrency; choose one.
- Arm (Graviton) processors: Arm is the chip architecture, Graviton is AWS's own line of Arm-based processors, and they have a lower price per GB-second than x86 (the other common chip architecture, used by Intel and AMD processors); benchmark the CPU-bound path on both, since native libraries may perform differently.
Worked example: plan and expected outcome
| Change | Expected effect on init (from 5.0 s) | Running total | How to confirm |
|---|---|---|---|
| Pre-convert the dependency to a memory-mappable format | Dependency load drops well below 3.0 s; for illustration, say it falls by roughly half to about 1.5 s (measure it, don't assume) | ~3.5 s (1.5 s dependency + the other 2.0 s of the 5.0 s baseline, runtime start-up plus your own init code) | Dependency timer |
| 1,024 MB to 1,769 MB | Remaining CPU-bound init x 0.58 | ~3.5 x 0.58 \u2248 2.0 s | Init Duration p50/p99 |
| Provisioned concurrency sized to business-hours peak | Users stop seeing cold starts | 0 s felt by users during business hours | Cold-start rate, spillover metric |
Put together, that plan takes the baseline from 5.0 s to about 2.0 s of real Init Duration, a rough 60% cut, before provisioned concurrency hides whatever is left from users entirely; confirm every number in this row with your own timers; the 1.5 s and the combined 2.0 s are illustrative, not measured. Ship them one at a time. Changing everything at once makes it impossible to tell which change paid off, or which one caused a regression.
Pitfalls
- Optimizing the wrong 40%. Confirm the 60% claim with your own timer;
Init Durationalone doesn't say what inside init was slow. - Too close to the 10-second Init cap. On-demand Init is limited to 10 s; a slow load that overruns is retried inside the first invocation. Improvements that bring init well under 10 s remove that failure mode.
- Memory headroom. Loading a large dependency can double-count memory (the file bytes plus the parsed objects). Check
Max Memory Used; running out of memory kills the environment and forces another cold start.
Write language-agnostic pseudo-code for a wrapper that measures and emits two metrics per invocation: time spent in initialization (cold-start setup), and time spent in the handler itself. How does your wrapper tell a cold invocation apart from a warm one, and how would it push these metrics to a monitoring backend?
Sample Answer
Direct answer
Record a timestamp at the very first line of module-level code, record another when module-level setup finishes, and keep a module-level flag cold = true. Module-level code runs exactly once per execution environment, so the first invocation that sees cold == true is the cold one: it reports init duration (the two timestamps' difference) and flips the flag; every later invocation in that environment is warm. Around the handler, time from entry to exit in a finally block so failures are measured too. To push metrics, don't call a metrics API synchronously from the request path; write one structured log line per invocation in a format the backend turns into metrics (on AWS, CloudWatch, the account's monitoring and log service, reads log lines written in its Embedded Metric Format, EMF, and turns them into metrics on its own), or hand the data to a local agent that ships it asynchronously.
Pseudo-code (language-agnostic)
# ---- module scope: runs once per execution environment ----
INIT_START = monotonic_now()
COLD = true
... imports, clients, config, large dependency ...
INIT_END = monotonic_now()
function wrap(handler):
return function(event, context):
start = monotonic_now()
was_cold = COLD
COLD = false # every later call in this environment is warm
try:
return handler(event, context)
finally: # measured on success and on exception
metrics = { HandlerDuration: ms(monotonic_now() - start) }
if was_cold:
metrics.InitDuration = ms(INIT_END - INIT_START)
metrics.InitToFirstInvokeGap = ms(start - INIT_END)
emit(metrics, dimensions = { FunctionName, ColdStart: was_cold })
Use a monotonic clock (one that never jumps when the system clock is corrected) for durations, and wall-clock time only for the metric's timestamp.
Runnable implementation
The wrapper takes the clock as a parameter so the demo can use a fake, deterministic clock; in production you pass time.monotonic_ns. emit prints an EMF document: a JSON log line with an _aws block telling CloudWatch which fields are metrics and which are dimensions (the labels a metric is grouped by).
import json, os, time
class FakeClock:
"""Deterministic clock for the demo. In production pass time.monotonic_ns."""
def __init__(self):
self.ns = 0
def __call__(self):
return self.ns
def advance_ms(self, ms):
self.ns += int(ms * 1_000_000)
class ColdStartMetrics:
def __init__(self, namespace, function_name, clock=time.monotonic_ns,
wall_ms=lambda: int(time.time() * 1000), sink=print):
self.clock, self.wall_ms, self.sink = clock, wall_ms, sink
self.namespace, self.function_name = namespace, function_name
self.init_type = os.environ.get("AWS_LAMBDA_INITIALIZATION_TYPE", "on-demand")
self.init_start = clock() # first line of module scope
self.init_end = None
self.cold = True # module scope runs once per environment
def init_done(self):
self.init_end = self.clock()
def wrap(self, handler):
def wrapped(event, context):
start = self.clock()
was_cold, self.cold = self.cold, False
try:
return handler(event, context)
finally: # runs on success and on exceptions
handler_ms = (self.clock() - start) / 1e6
fields = {"HandlerDuration": round(handler_ms, 3)}
if was_cold:
fields["InitDuration"] = round((self.init_end - self.init_start) / 1e6, 3)
fields["InitToFirstInvokeGap"] = round((start - self.init_end) / 1e6, 3)
self.emit(fields, was_cold)
return wrapped
def emit(self, fields, was_cold):
doc = {
"_aws": {
"Timestamp": self.wall_ms(),
"CloudWatchMetrics": [{
"Namespace": self.namespace,
"Dimensions": [["FunctionName", "ColdStart"]],
"Metrics": [{"Name": k, "Unit": "Milliseconds"} for k in fields],
}],
},
"FunctionName": self.function_name,
"ColdStart": "true" if was_cold else "false",
"InitType": self.init_type, # logged property, not a dimension
**fields,
}
self.sink(json.dumps(doc, separators=(",", ":")))
# ---- demo driver: simulate one environment's life with a pinned clock ----
clock = FakeClock()
metrics = ColdStartMetrics("MyApp", "resize-images", clock=clock, wall_ms=lambda: 1767225600000)
clock.advance_ms(850) # module-scope work: imports, SDK clients, config load
metrics.init_done()
def handler(event, context):
clock.advance_ms(event["work_ms"])
if event.get("fail"):
raise RuntimeError("downstream timeout")
return {"ok": True}
h = metrics.wrap(handler)
clock.advance_ms(2) # time between init finishing and the first request
h({"work_ms": 120}, None) # cold
clock.advance_ms(5000) # environment frozen, then thawed
h({"work_ms": 95}, None) # warm
try:
h({"work_ms": 40, "fail": True}, None) # warm, handler raises
except RuntimeError as e:
print("handler error propagated:", e)
Output:
{"_aws":{"Timestamp":1767225600000,"CloudWatchMetrics":[{"Namespace":"MyApp","Dimensions":[["FunctionName","ColdStart"]],"Metrics":[{"Name":"HandlerDuration","Unit":"Milliseconds"},{"Name":"InitDuration","Unit":"Milliseconds"},{"Name":"InitToFirstInvokeGap","Unit":"Milliseconds"}]}]},"FunctionName":"resize-images","ColdStart":"true","InitType":"on-demand","HandlerDuration":120.0,"InitDuration":850.0,"InitToFirstInvokeGap":2.0}
{"_aws":{"Timestamp":1767225600000,"CloudWatchMetrics":[{"Namespace":"MyApp","Dimensions":[["FunctionName","ColdStart"]],"Metrics":[{"Name":"HandlerDuration","Unit":"Milliseconds"}]}]},"FunctionName":"resize-images","ColdStart":"false","InitType":"on-demand","HandlerDuration":95.0}
{"_aws":{"Timestamp":1767225600000,"CloudWatchMetrics":[{"Namespace":"MyApp","Dimensions":[["FunctionName","ColdStart"]],"Metrics":[{"Name":"HandlerDuration","Unit":"Milliseconds"}]}]},"FunctionName":"resize-images","ColdStart":"false","InitType":"on-demand","HandlerDuration":40.0}
handler error propagated: downstream timeout
The first line is the cold invocation (850 ms init, 2 ms gap before the first request, 120 ms handler). The next two are warm and carry only HandlerDuration. The third invocation raised an exception, and its metric was still emitted before the error propagated. Notice what the _aws.CloudWatchMetrics[0].Dimensions list is doing: [["FunctionName", "ColdStart"]] tells CloudWatch to turn those two fields into metric labels you can filter and group a graph by. InitType sits outside that list, so CloudWatch stores it as a plain property on the log entry, searchable in Logs Insights but not usable to split a metric.
How cold versus warm detection works, and where it breaks
- Why the flag works. The runtime imports your module once per environment. If the environment crashes or times out, Lambda resets it and re-runs module code, so the flag correctly resets to
true. - Provisioned concurrency and pre-initialization. Provisioned concurrency is a Lambda setting that pays to keep a chosen number of environments already initialized and sitting idle, ready before any request arrives, instead of only initializing one when a request shows up needing it. So an environment can be initialized long before the first request. The first request is then "first in this environment" but the user did not wait for init. That is why the wrapper also emits
InitToFirstInvokeGapand logsAWS_LAMBDA_INITIALIZATION_TYPE(on-demand, meaning an environment created on demand by an incoming request;provisioned-concurrency;snap-start, covered next; orlambda-managed-instances, an instance type where Lambda runs your function on capacity you provision and one execution environment can serve more than one concurrent invocation). A user-visible cold start isColdStart=truewith a gap near zero. - SnapStart. SnapStart is a Lambda feature that initializes one environment fully, freezes a snapshot of its memory and disk state once that init finishes, and boots every later environment for that function by restoring from that snapshot instead of re-running your module code. Because module code never runs again in a restored environment, the
COLDflag is lefttrueinside the snapshot itself (init ran once, before the snapshot was taken), so every environment that restores from it starts already believing it is cold, and reports the same staleInitDurationthe original environment measured, not its own. Worse,INIT_STARTandINIT_ENDwere read from a monotonic clock (a clock that only measures elapsed time going forward on the machine that reads it, never wall time) running on the environment that made the snapshot, so a reading taken before the snapshot cannot say anything about how long this restored environment took to come up; the clock's origin point is frozen into the snapshot along with everything else. Lambda gives you an escape hatch for this: an after-restore runtime hook, a callback the runtime invokes once a restore finishes and before the handler runs again, where you can reset the timers. Then take the real start-up cost from the platform instead of your own clock: for SnapStart functions, theREPORTlog line AWS writes after each invocation carries aRestore Durationfield (the time to thaw the snapshot and run your after-restore hook), and AWS defines the user-visible cold-start duration for these functions asRestore DurationplusDuration(the handler's own run time). - What the wrapper cannot see. Code download, sandbox start and runtime boot happen before your first line runs. The platform's own
Init Durationfield on theREPORTlog line covers the whole Init phase, and theDurationmetric excludes init, so compare the two: platform init minus your init is the part you cannot optimize from code. - Timeouts. When a function times out, the runtime is stopped before
finallyruns, so no metric is emitted. Detect those from the platform'sREPORTline (it carriesStatus: timeout) or theErrorsmetric.
Pushing to a monitoring backend
| Option | How | Trade-off |
|---|---|---|
| Structured log line (EMF on AWS) | Print the JSON; CloudWatch Logs extracts metrics asynchronously. | No network call in the request path and no extra API cost per call. Metrics appear with log-ingestion delay; delivery is at least once, so rare duplicates are possible. |
Synchronous metrics API call (for example CloudWatch PutMetricData) | Call the API before returning. | Adds a network round trip to every invocation's latency and bill, and a failed call can fail the request. Avoid in the hot path. |
| Local agent or extension (an OpenTelemetry collector, OpenTelemetry being the open standard for traces, metrics and logs, running as a Lambda extension; or a vendor extension) | Send to a local process that batches and flushes. | Keeps vendor neutrality and batching; costs extra init time and memory, and flush time at shutdown. |
Cardinality. Every distinct combination of dimension values becomes a separate billed metric. FunctionName and ColdStart (two values) are safe; putting request IDs or user IDs in dimensions creates millions of metrics. High-cardinality fields go in the log line as properties (like InitType here), not dimensions.
Complexity and edge cases
- Cost: O(1) time per invocation and one log line; memory O(1).
- Edge cases: exceptions (handled by
finally); timeouts (not visible to the wrapper, use platform logs); SnapStart restores (reset timers after restore); provisioned concurrency (use the gap and init type to separate "first in environment" from "user waited"); clock adjustments (monotonic clock only for durations).
What event-driven design patterns are commonly used in serverless architectures? Walk through a concrete example: a pipeline that needs to kick off a downstream job whenever a new file lands in object storage. What components and events would you wire together, and how would you make sure the trigger is reliable and doesn't fire the job twice for the same file?
Sample Answer
Direct answer
Serverless systems are usually built by wiring managed services together with events, not by writing servers that poll. The common patterns are:
- event notification (a service announces "something happened")
- queue-based load leveling (a queue sits between producer and consumer)
- publish/subscribe fan-out (one event, many independent consumers)
- orchestration (a workflow engine drives the steps)
- the claim-check pattern (pass a pointer to large data, not the data)
- dead-letter handling (failures go somewhere you can see them)
For "a file lands, so run a job" I would wire it like this:
- The bucket emits an Object Created event to an event router (Amazon EventBridge, a managed service that receives events from many sources and routes each one to whichever targets match its filtering rules).
- A rule filtered on the input prefix (the folder-like portion of the object key that events are matched against, such as
incoming/) sends it to a queue. - A small function starts a workflow (Step Functions Standard, one of two execution modes: Standard keeps a full execution history and can run for up to a year; Express, mentioned below, is cheaper and higher-throughput but keeps no long history, and suits short executions better) whose name is derived from the object's identity.
- The workflow runs the job.
Every hop is at-least-once (each event may be delivered more than once, though never fewer), so reliability comes from queues with dead-letter queues and alarms. "Not twice" comes from a deterministic execution name. A second start with the same name does not start a second job.
The patterns, briefly
| Pattern | What it is | Use it when |
|---|---|---|
| Event notification | A service emits "this happened" (object created, row changed) | Something should react to a change without the producer knowing who |
| Queue-based load leveling | A queue between producer and consumer | The consumer is slower, rate-limited or sometimes down |
| Pub/sub fan-out | One event delivered to many independent consumers (an Amazon SNS, Simple Notification Service, topic; EventBridge rule targets) | Several teams react to the same fact |
| Orchestration | A workflow engine calls the steps and holds the state | Multi-step jobs with retries, waits and branching |
| Choreography | Each service reacts to the previous one's events; no central controller | Loosely coupled, short chains |
| Claim check | Put the large payload in object storage; pass only its location | Payloads bigger than event or message limits |
| Dead-letter queue (DLQ) | Where messages go after repeated failures | Always, on every async hop |
The pipeline
flowchart LR
up[Upload to incoming/] --> bucket[(Object storage bucket)]
bucket -->|Object Created event| router[EventBridge rule: prefix incoming/]
router --> q[Queue]
q --> starter[Starter function]
starter -->|StartExecution, deterministic name| wf[Workflow: Step Functions Standard]
wf --> job[Batch or data-processing job]
wf --> done[Completion event]
q -.->|after 5 failures| dlq[DLQ + alarm]
- Why EventBridge rather than a direct bucket-to-function trigger. Rules can filter on prefix and suffix, send one event to several targets, and archive and replay events. A direct S3-to-Lambda notification also works, but it is an asynchronous invoke (Lambda queues the event internally and retries it on failure without the caller waiting for a response) with Lambda's own retry queue (a separate, less visible retry mechanism than a queue you configure and can alarm on yourself), and you get fewer controls.
- Why a queue before the starter. The queue gives a buffer, a retry policy you control, a DLQ, and a visible metric (age of oldest message) to alarm on.
- Why a workflow rather than running the job inside the function. A Lambda invocation lasts at most 15 minutes. A Standard workflow can wait on a job for far longer, and it records every step, so retries and failures are visible.
- Multipart uploads (large files uploaded to object storage in separate chunks that are only assembled into the final object once every part has arrived) emit one event, when the upload completes, so the job never sees a half-written file.
Making the trigger reliable
- Assume duplicates and delays. S3 documents its event notifications as delivered at least once, usually within seconds but sometimes a minute or longer. The event router and the queue are also at-least-once.
- Every async hop gets a DLQ and an alarm. Alarm on DLQ depth above 0, and on the age of the oldest message in the main queue.
- Reconcile. A scheduled sweep compares objects under
incoming/with the set of files the workflow has processed. It re-triggers anything missing. This catches the rare lost event, a misconfigured rule, or an outage longer than the retry windows. - Avoid recursion. If the job writes output back into the same bucket, write it under a different prefix (or bucket) than the rule listens on. Otherwise every output file triggers another job.
Making sure the job fires once per file
Derive the workflow's execution name from the object's identity: bucket + key + version id (enable versioning on the bucket first: with it on, every overwrite of the same key is kept as a separate, uniquely identified version instead of replacing the object in place, and it is that per-write version id, not the key alone, that changes on a genuine re-upload). For a Standard workflow, a second StartExecution (the API call that starts a workflow execution) with the same name and input returns the existing execution while it is running, and returns ExecutionAlreadyExists once it has closed. Either way, no second job starts. Names can be reused 90 days after an execution closes.
import hashlib
def execution_name(bucket, key, version_id):
# Same object version -> same name; a re-upload (new version) -> new name.
digest = hashlib.sha256(f"{bucket}/{key}#{version_id}".encode()).hexdigest()
return f"ingest-{digest[:40]}" # 47 chars, within the 80-char limit, no banned characters
class FakeStandardWorkflow:
"""Models documented StartExecution semantics for STANDARD workflows."""
def __init__(self):
self.executions = {} # name -> (input, status)
def start(self, name, payload):
if name in self.executions:
prev_input, status = self.executions[name]
if status == "RUNNING" and prev_input == payload:
return "same execution returned (idempotent)"
raise RuntimeError("ExecutionAlreadyExists")
self.executions[name] = (payload, "RUNNING")
return "started"
def handle(event, sfn):
d = event["detail"]
name = execution_name(d["bucket"]["name"], d["object"]["key"], d["object"]["version-id"])
payload = f'{d["bucket"]["name"]}/{d["object"]["key"]}'
try:
return name, sfn.start(name, payload)
except RuntimeError as e:
return name, f"{e}: already handled, acknowledge the message"
sfn = FakeStandardWorkflow()
evt = {"detail": {"bucket": {"name": "raw-landing"},
"object": {"key": "incoming/2026-09-27/orders.csv", "version-id": "v1"}}}
print(handle(evt, sfn)) # first delivery
print(handle(evt, sfn)[1]) # duplicate while running
sfn.executions[execution_name("raw-landing", "incoming/2026-09-27/orders.csv", "v1")] = (
"raw-landing/incoming/2026-09-27/orders.csv", "SUCCEEDED")
print(handle(evt, sfn)[1]) # duplicate after it finished
evt["detail"]["object"]["version-id"] = "v2"
print(handle(evt, sfn)[1]) # genuine re-upload
('ingest-f5772e2d0a15754e5cd838f6fdd74a26602f48ad', 'started')
same execution returned (idempotent)
ExecutionAlreadyExists: already handled, acknowledge the message
started
- Duplicate while the job is running: the same execution is returned.
- Duplicate after it finished:
ExecutionAlreadyExists, which the starter treats as "already handled" and acknowledges. - A genuine re-upload (new version id) gets a new name and correctly starts a new job.
Two edges to handle deliberately:
- A retry after a FAILED execution needs a new name (for example, append an attempt number after checking the old execution's status). Otherwise
ExecutionAlreadyExistsblocks the legitimate retry. - The job itself should also be idempotent. It writes its output to a deterministic path and overwrites, so a manual re-run can't double-count.
Pitfalls
- Hashing the event id instead of the object identity. Duplicates carry different event ids, so the dedupe never matches.
- Dropping the version id and deduping on key alone. A legitimate re-upload of
orders.csvwould be silently ignored. - No DLQ alarm. Failed files wait unseen until someone asks where the report is.
A deployment artifact you need to run in Lambda is 800MB, well past the deployment-package limit and too big to comfortably fit in /tmp. Compare architectures for serving it: mounting EFS, streaming it from S3 at cold start, packaging it as a container image, or moving it to a managed endpoint outside Lambda entirely. For each, discuss cold-start latency, throughput, cost, and operational complexity.
Sample Answer
Direct answer
For an 800 MB artifact that changes with the code, package the function as a container image: it lifts the size limit from 250 MB to 10 GB, needs no private network, and Lambda loads image contents on demand from a cache, so cold starts track what init actually reads. Choose Amazon EFS (Elastic File System, a shared network file system) when the artifact changes independently of the code or is shared by several functions. Avoid downloading from S3 at every cold start: the arithmetic below shows 800 MB alone takes over 10 seconds per cold start. Move to a managed always-on endpoint outside Lambda when traffic is steady enough that you pay for idle capacity anyway, when you need a GPU, or when no cold start is acceptable.
First, a correction to the premise: /tmp on Lambda is configurable from 512 MB up to 10,240 MB, so 800 MB can fit there (extra ephemeral storage is billed per GB-second). The real problems are getting 800 MB there on every cold start, and the memory needed if the artifact is then loaded into RAM.
The constraints that decide it
- Package limits. A zip deployment, including layers, may be at most 250 MB unzipped; a container image may be up to 10 GB uncompressed.
- Init time limit. On-demand Init (module-level set-up) gets 10 seconds; if it overruns, Lambda retries Init inside the first invocation, bounded by the function timeout.
- Network bandwidth. Lambda's documented network bandwidth per execution environment is 625 Mbps (megabits per second, 10^6 bits). Anything fetched over the network at cold start is bounded by it.
- Memory. If the artifact must sit in RAM, configured memory must hold it plus the runtime and working space, and you pay for that memory on every invocation.
Worked example: how long does 800 MB take to fetch?
Lambda's "MB" is 2^20 bytes, so 800 MB = 800 x 1,048,576 x 8 = 6,710,886,400 bits. At the 625 Mbps ceiling:
625×106 bits/s6,710,886,400 bits=10.74 sThat is the best case, before any decompression or parsing, and it already exceeds the 10-second Init cap. Every new environment pays it, so a burst that needs 50 new environments downloads 50 x 800 MB = 40,000 MB.
The four options compared
| Cold-start latency | Throughput once warm | Cost | Operational complexity | |
|---|---|---|---|---|
Stream from S3 into /tmp at cold start | Worst: at least 10.74 s of transfer per new environment (above), plus unpacking | Fine once loaded | Init time is billed; extra /tmp above 512 MB billed per GB-second; S3 request costs trivial | Low to build, but fragile: sits on the 10 s Init cap; each environment repeats the download |
| Mount Amazon EFS | Depends on read pattern. Reading all 800 MB eagerly is also a network read, so expect the same order as S3. Reading only the parts needed is much faster: memory-mapping (mapping the file so pages load into memory only as your code actually touches them, instead of reading the whole file upfront) or indexed access (seeking straight to the record you need instead of scanning) | Good for reads served from the environment's page cache, the operating system's own in-memory copy of recently read file blocks, which sticks around until the environment is recycled; first reads pay network latency | EFS storage plus throughput charges; the throughput mode you pick, bursting (earns credits while idle and spends them under load), elastic (scales automatically with actual usage) or provisioned (a fixed throughput you pay for whether you use it or not), must sustain N environments reading at once | Highest: function must join a VPC (Virtual Private Cloud, an isolated private network you control) with mount targets, the network endpoints EFS exposes inside a subnet so anything there can attach the file system, in each Availability Zone (an isolated data-center group inside the Region) it runs in, and security groups, the firewall rules attached to those mount targets, allowing NFS (Network File System, the protocol EFS speaks, port 2049); a separate process has to update the files |
| Container image (artifact baked in) | Good: Lambda splits images into chunks, deduplicates them (stores one copy of any chunk shared across images instead of a separate copy per image) and caches them, then fetches chunks as the code reads them (described in AWS's 2023 USENIX ATC paper, USENIX ATC being a peer-reviewed systems-research conference, titled "On-demand Container Loading in AWS Lambda"). Cold start depends on bytes read during init, not image size. Deploys and long-idle functions warm the cache again | Same as zip | Registry storage (Amazon ECR, Elastic Container Registry); no VPC or EFS cost; Init billed as with any image | Moderate: image build pipeline; every artifact change is a new image and deploy |
| Managed endpoint outside Lambda (a container service such as Amazon ECS on Fargate, or a model-hosting endpoint, a managed, always-running service built specifically to serve a loaded ML model over an API, such as Amazon SageMaker's real-time inference endpoints) | None for users once running; the artifact loads once per replica | Highest: long-lived processes, concurrency per replica, GPUs available | Pay per replica-hour, busy or idle | Moderate to high: scaling policy, health checks, deployments; Lambda (if kept) calls it over the network |
flowchart TD
A[800 MB artifact] --> B{Changes with code releases?}
B -- yes --> C{Traffic steady or GPU needed?}
B -- no, updated separately or shared --> E[EFS mount, read lazily]
C -- no, spiky --> D[Container image]
C -- yes --> F[Managed always-on endpoint]
Recommendation and what would change it
Default: container image. It is the only option that removes the size limit without adding a network dependency at start-up that you have to operate. Put the artifact in an early image layer that rarely changes and your code in the last layer, so ordinary code changes don't re-upload the artifact.
- Switch to EFS if the artifact updates on its own schedule (for example daily data refreshes) and rebuilding an image each time is unacceptable, or several functions share it. Structure access to read only what each request needs, and provision throughput for peak concurrency.
- Switch to a managed endpoint once sustained load makes always-on replicas cheaper than per-invocation billing, or when the artifact needs a GPU or must stay in memory without per-environment duplication.
- Use S3 download only for rare background jobs where a slow cold start does not matter, and then fetch during the invocation (subject to the function timeout) rather than squeezing it into the 10-second Init.
Pitfalls
- RAM, not just disk. If the 800 MB is fully deserialized, meaning converted from the file's on-disk bytes into in-memory objects your code works with directly, you may need well over 1 GB configured, paid on every invocation. Memory-mapping keeps resident memory closer to what is actually touched.
- EFS throughput exhaustion. Many environments starting at once can exhaust burst capacity, the reserve of extra throughput EFS's bursting mode grants based on file-system size and lets you spend down temporarily; the symptom is init times that stretch only during scale-outs.
- Inactive container functions. A container-image function not invoked for weeks becomes inactive; the next call fails while Lambda re-optimizes the image. Scheduled invocations or monitoring the function state avoids surprise.
Unlock Full Question Bank
Get access to all 24 Serverless and Function-as-a-Service Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.