Serverless and Function-as-a-Service Architecture Questions
Building applications on managed, event-triggered compute: functions-as-a-service (AWS Lambda, Azure Functions, Google Cloud Functions, Cloudflare Workers) and serverless containers. Covers the invocation lifecycle and cold starts (init vs handler work, provisioned concurrency, packaging, layers and container images), statelessness and externalizing state, execution limits (timeouts, memory, payload and /tmp size), concurrency and scaling behavior (account limits, burst scaling, protecting downstream databases with connection proxies and throttling), event sources and trigger semantics (at-least-once delivery, retries, idempotent handlers, dead-letter handling), composing functions with managed services and workflow orchestrators (Step Functions and equivalents), and serverless-specific observability, security (per-function IAM, secrets) and pay-per-invocation cost modeling. Includes when serverless fits versus containers or VMs, and vendor lock-in trade-offs. General compute selection, generic messaging patterns, and general idempotency theory are covered by their own topics.
A greenfield service needs to choose between serverless functions and containers managed by Kubernetes. What practical criteria would you use to decide: think about latency and cold starts, traffic pattern, cost model, operational overhead, vendor lock-in, and your team's existing skills. In what kind of scenario does one option clearly win?
Sample Answer
Direct answer
Default to serverless functions when traffic is low, bursty or unpredictable, the work is short and stateless, and the team has no Kubernetes platform already running. Default to Kubernetes-managed containers when traffic is steady and high, requests need consistently low latency, the work is long-running or needs special hardware, or the organization already operates a Kubernetes platform with people to run it. For a typical greenfield service with no existing platform, I would start serverless and write down the traffic level at which I would move.
Two terms first:
- Serverless functions (FaaS, functions as a service), for example AWS Lambda: you deploy a function; the provider starts copies on demand, scales them to zero when idle, and bills per request plus per millisecond of run time.
- Containers on Kubernetes: you package the service as a container image and Kubernetes (a container orchestrator: it runs your containers as pods, one or more containers scheduled and scaled together, on nodes, the physical or virtual machines that make up the cluster) keeps a chosen number of copies running on servers you pay for whether or not they are busy. On AWS the managed version is EKS (Amazon Elastic Kubernetes Service).
The criteria
| Criterion | Serverless functions | Kubernetes containers | Question to ask |
|---|---|---|---|
| Latency and cold starts | A cold start (booting a new function environment on first request or during scale-up) adds latency to some requests; mitigated by provisioned concurrency (paying to keep environments warm) | Pods are already running, so no per-request cold start; scaling up new pods and nodes is slower | Does the p99 (99th percentile) latency target tolerate an occasional slow request? |
| Traffic pattern | Scales to zero and up quickly; ideal for spiky or idle-most-of-the-day traffic | Needs capacity running for the baseline; autoscaling reacts more slowly | What is the ratio of peak to average traffic? |
| Cost model | Pay per request and per GB-second (memory allocated, in GB, times run time, in seconds; the unit Lambda's execution price is billed in); nothing when idle | Pay for provisioned capacity (servers kept running and billed 24x7, whether or not they are doing work) plus the cluster itself | What is the average utilization? |
| Operational overhead | No servers, OS patching or cluster upgrades | Cluster upgrades, node patching, autoscaler tuning, networking | Who will be on call for the platform? |
| Vendor lock-in | Handler signatures, event formats and triggers are provider-specific | Container images and Kubernetes manifests move between clouds fairly easily | How likely is a move, and what would it cost? |
| Team skills | Needs event-driven design (structuring the system around reacting to events rather than direct calls) skills and distributed tracing (following one request's path across multiple services to debug it) | Needs Kubernetes operations skills | What does the team already run well? |
Two hard limits also decide it on their own: AWS Lambda caps one invocation at 15 minutes and 10,240 MB of memory, and it has no GPUs. Work that exceeds those goes to containers.
Worked example: where the cost lines cross
Compare Lambda with one always-on container sized at 1 vCPU (one virtual CPU core, a unit of allocated compute) and 2 GB, using AWS us-east-1 list prices (Lambda $0.0000166667 per GB-second plus $0.20 per million requests; Fargate, AWS's serverless container runtime, $0.000011244 per vCPU-second and $0.000001235 per GB-second):
How to read the code before running it: LAMBDA_GBS/LAMBDA_REQ are Lambda's per-GB-second and per-request prices; FG_VCPU_S/FG_GB_S are Fargate's per-vCPU-second and per-GB-second prices; H = 730 is the average hours in a billing month, used to turn an hourly or per-second rate into a monthly one. per_req prices one 100 ms, 1 GB request: 0.1*1.0*LAMBDA_GBS is 100 ms (0.1 s) times 1 GB, i.e. 0.1 GB-seconds of execution, priced at LAMBDA_GBS, plus the flat LAMBDA_REQ per-request fee. The container's monthly price, task, sums its vCPU cost and its memory cost, each a rate times seconds in the month.
LAMBDA_GBS, LAMBDA_REQ = 0.0000166667, 0.20/1e6
FG_VCPU_S, FG_GB_S = 0.000011244, 0.000001235
H = 730
task = (1*FG_VCPU_S + 2*FG_GB_S) * 3600 * H
per_req = 0.1*1.0*LAMBDA_GBS + LAMBDA_REQ # 100 ms at 1 GB
print(f"1 vCPU/2 GB Fargate task 24x7: ${task:.2f}/month")
print(f"Lambda per request (100 ms, 1 GB): ${per_req:.8f}")
be = task/per_req
print(f"breakeven: {be/1e6:.1f}M requests/month = {be/(H*3600):.1f} req/s average")
print(f"EKS control plane: ${0.10*H:.2f}/month")
1 vCPU/2 GB Fargate task 24x7: $36.04/month
Lambda per request (100 ms, 1 GB): $0.00000187
breakeven: 19.3M requests/month = 7.3 req/s average
EKS control plane: $73.00/month
How to read it:
- Below about 7 requests per second on average, Lambda is cheaper than even one always-on container. A new internal API doing 50,000 requests a day (about 0.6 requests/s) costs roughly 1.5M x $0.00000187, under $3 a month, on Lambda.
- On Kubernetes you would also pay $73 a month for each EKS cluster's control plane (the managed pieces that run the cluster itself, scheduling and API server, separate from and in addition to the nodes that run your pods), run at least two copies for availability, and spend engineer time on the cluster.
- Above the breakeven, containers win on raw compute price, and the gap widens as steady traffic grows, because a busy container is cheaper per unit of work than the same work billed by the millisecond.
The example is deliberately simple: it ignores the free tier, assumes 100 ms per request, and assumes one container could serve that load. Rerun it with your own duration and memory.
Scenarios where one option clearly wins
- Serverless clearly wins: a webhook receiver, a scheduled nightly report, or an image thumbnailer triggered by uploads. Traffic is bursty or idle most of the day, each task takes seconds, there is no platform team, and scale-to-zero means paying almost nothing when idle.
- Kubernetes clearly wins: a high-traffic API steady at around 2,000 requests/s with a strict 20 ms p99 target, maintaining long-lived WebSocket connections (persistent two-way connections), at a company that already runs a Kubernetes platform. Utilization is high, cold starts are unacceptable, and the operational cost is already paid.
Trade-offs and pitfalls
- Counting only the compute bill. The larger cost of Kubernetes for a small team is people: upgrades, security patching, on-call. The larger hidden cost of serverless at scale is the per-millisecond premium on steady traffic.
- Assuming cold starts always matter. For asynchronous work (queues, file uploads) a few hundred extra milliseconds is invisible; they only matter for user-facing, latency-sensitive requests.
- Treating the decision as permanent. Packaging the function as a container image and keeping business logic out of the handler keeps a later move cheap.
- Ignoring the downstream database. Functions can scale to hundreds of copies quickly and exhaust a relational database's connection limit, so plan a connection proxy (a lightweight service, such as RDS Proxy, that pools and limits database connections on functions' behalf) or a concurrency cap.
A serverless function has intermittent cold-start latency spikes that are hurting your p99. Name three practical strategies to reduce cold-start latency, and for each one explain the mechanism by which it helps and its trade-offs in cost, complexity, or effectiveness.
Sample Answer
Direct answer
First confirm the spikes really are cold starts: compare p99 (the 99th-percentile latency, the value only 1% of requests exceed) for invocations whose logs carry an Init Duration against those that don't. If they are, three practical strategies, from cheapest to most expensive: (1) make initialization cheaper (smaller package, fewer and lazier imports, more memory for more CPU); (2) restore from a snapshot instead of initializing (AWS Lambda SnapStart); (3) pre-initialize environments so users never hit a cold one (provisioned concurrency, or minimum instances on other platforms). Do (1) always, then pick (2) or (3) based on runtime support and how strict the latency target is.
Why cold starts land exactly in p99
AWS reports cold starts are typically under 1% of invocations. p99 is the boundary of the slowest 1%. If 0.5% of requests are cold, the slowest 1% is half cold and half slow-warm, and p99 sits at the edge of the warm tail. If 1.5% are cold (low traffic, frequent deploys, bursts that force new environments), the p99 value itself is a cold request. A modest rise in cold-start rate can therefore move p99 from roughly "slow warm request" to "full cold start". This is why the fix list starts with cutting the cold-start cost, and why reducing how often they happen also counts.
Strategy 1: make initialization cheaper
Mechanism. Cold-start time is roughly code download + runtime start + your init code. Shrinking any term shrinks the spike.
- Import only what you use (a single SDK service client rather than the whole SDK); remove unused dependencies, test fixtures and docs from the package.
- Lazy-load heavy objects needed only on some code paths.
- Raise memory: Lambda allocates CPU in proportion to memory (1,769 MB equals one full vCPU, one virtual CPU core, a unit of allocated compute), so CPU-bound init (parsing, class loading, JIT warm-up, where JIT is just-in-time compilation of hot code) runs faster.
Trade-offs.
| Cost | Complexity | Effectiveness |
|---|---|---|
| Free to cheap. More memory costs more per millisecond, but if run time falls in proportion the cost per call is roughly unchanged. | Low; ordinary engineering. Lazy loading moves some delay to the first request that needs it. | Helps every cold start and every runtime. Bounded: you cannot optimize below runtime start-up plus unavoidable work. |
Strategy 2: snapshot restore (Lambda SnapStart)
Mechanism. When you publish a function version, Lambda runs Init once, takes a snapshot of the initialized environment's memory and disk (a Firecracker microVM snapshot, Firecracker being the lightweight virtual machine technology Lambda runs on), and caches it. New environments resume from the snapshot instead of re-running Init.
Trade-offs.
| Cost | Complexity | Effectiveness |
|---|---|---|
| No extra charge on Java. On Python and .NET you pay a snapshot caching charge per published version (minimum 3 hours) plus a restore charge each time an environment is restored. | Moderate. Supported only on Java 11+, Python 3.12+ and .NET 8+, on published versions (numbered, immutable snapshots of your code and configuration), not on $LATEST (the mutable pointer that always reflects whatever you deployed most recently). Not combinable with provisioned concurrency, Amazon EFS (Elastic File System, a network filesystem) mounts (attaching that shared filesystem to the environment), or more than 512 MB of /tmp. Anything unique created at Init (random seeds, IDs, credentials) is duplicated across every restored copy and must be regenerated after restore; network connections may need re-establishing. | Large for heavy-init runtimes (JVM, Java Virtual Machine, the runtime Java code executes in, start-up; big framework loads). AWS notes it works best at scale; rarely invoked functions benefit less. Restores are faster than a full Init but not free. |
Strategy 3: pre-initialized capacity (provisioned concurrency)
Mechanism. You tell Lambda to keep N environments initialized for a version or alias at all times. Requests up to N concurrent never see a cold start; above N, they spill over (run on ordinary on-demand environments instead of the pre-warmed ones) to normal on-demand environments (and can be cold). The equivalents elsewhere are minimum instances on Google Cloud Run and always-ready instances on Azure Functions.
Trade-offs, with numbers. Price in US East (N. Virginia): $0.0000041667 per GB-second for keeping capacity provisioned, and duration drops to $0.0000097222 per GB-second (from $0.0000166667) for work that runs on it.
- 10 environments at 1 GB for a 30-day month: 10 GB x 2,592,000 s = 25,920,000 GB-s x $0.0000041667 = $108.00 per month, even with no traffic.
- If those environments serve 5 million requests of 80 ms at 1 GB (400,000 GB-s), duration costs $3.89 instead of $6.67 on-demand. The $108 floor dominates unless utilization is high.
| Cost | Complexity | Effectiveness |
|---|---|---|
| A 24/7 floor you pay whether used or not. | Needs published versions and aliases (an alias is a named pointer such as live to one version); you must size N and ideally scale it on a schedule or on utilization (Application Auto Scaling, the AWS service that scales non-server resources, can change N on a schedule or on utilization). Watch ProvisionedConcurrencySpilloverInvocations to see when you are under-provisioned. | The strongest: cold starts disappear up to N. Does not help bursts above N. |
Worked example: choosing
A checkout API on Node.js with a 250 ms p50 (median latency: half of requests finish faster than this) and 1.8 s p99, 40 requests/s in business hours, 2 requests/s at night; cold-starting requests show ~1.5 s Init Duration.
- Trim the bundle and replace a whole-SDK import with single clients:
Init Durationfalls from ~1.5 s to, say, ~600 ms (measure it; don't assume). - SnapStart is not available for Node.js, so for the remaining spikes use provisioned concurrency. By Little's law (concurrency equals arrival rate times how long each request takes; here the 250 ms p50 stands in for that per-request time), 40 req/s x 0.25 s means about 10 requests are in flight at once, so provision ~12 (headroom) during business hours via a schedule and ~2 at night. Priced at the same $0.0000041667 per GB-second as above (1 GB, roughly 12 business hours and 12 night hours a day): 12 environments busy 12 h/day plus 2 environments the other 12 h/day, over a 30-day month, is about $75.60/month, cheaper than holding a flat 10 environments 24/7 ($108, from the pricing above) precisely because it drops to 2 overnight.
- Re-check p99 and the spillover metric after a week.
Pitfalls
- "Keep-warm" pings (a scheduled event calling the function every few minutes) keep one environment warm. A burst that needs 20 environments still cold-starts 19. Provisioned concurrency is the supported version of this idea.
- Blaming cold starts for everything. Slow downstream calls, network set-up for functions attached to a VPC (Virtual Private Cloud, a private network), or large payloads also show up in p99. Split latency by cold versus warm before paying for a fix.
- Deploy frequency matters. Every deployment replaces all environments, so ten deploys a day means ten waves of cold starts.
Design a serverless, event-driven pipeline that ingests telemetry at 100k events/second, does lightweight enrichment and aggregation, and writes results to analytic storage. Cover event ingestion and buffering, processing concurrency, idempotency, error handling, storage choice, observability, and the cost profile of your design.
Sample Answer
Direct answer
At a sustained 100,000 events per second I would build this as Kinesis Data Streams (provisioned shards) feeding a Lambda consumer that enriches and pre-aggregates in batches, with a per-shard checkpoint committed atomically alongside the aggregates, plus a Firehose copy of the raw stream into S3 for replay. The two decisions that matter most are both about batching: producers pack many events into each stream record, and Lambda is invoked once per batch of about 1,600 events rather than once per event. Invoking once per event would cost about $52,560 a month in request fees alone. The batched design runs at about $10,900 a month, and the raw archive (not the compute) is the biggest line item.
Quick glossary for the rest of the answer:
- Serverless / FaaS (functions as a service): you upload a function, the cloud runs as many copies as traffic needs and bills per invocation and per millisecond. AWS Lambda is the example here.
- Kinesis Data Streams (KDS): AWS's managed, ordered, replayable log (similar in spirit to Kafka, Apache Kafka: the most common self-hosted alternative doing the same ordered-log job). It is split into shards; each shard accepts up to 1 MB/s or 1,000 records/s of writes.
- Event source mapping (ESM): the Lambda-managed poller that reads a shard and invokes your function with a batch of records.
- At-least-once delivery: every record is delivered, but some may be delivered twice (after a retry). The pipeline must tolerate duplicates.
- Idempotent: processing the same input twice leaves the same result as processing it once.
- Amazon Data Firehose: a managed service that buffers a stream and writes large files to S3. S3 is object storage; Parquet is a columnar file format; Athena is a serverless SQL engine that queries files in S3.
- DynamoDB: AWS's managed key-value database.
Requirements I am designing to
| Assumption | Value | Why it matters |
|---|---|---|
| Throughput | 100,000 events/s sustained | Drives shard count and cost |
| Event size | about 1 KB JSON | 100 MB/s, 262,800 GB per 730-hour month |
| Freshness | aggregates queryable within about 3-5 minutes | Allows 2 s batching windows and a short grace period |
| Aggregation | per-minute counts and sums by (device group, metric) | Keeps aggregate state small |
| Correctness | aggregates must not double count | Delivery is at-least-once, so this needs design |
Architecture
flowchart LR
P[Devices and gateways] -->|packed 24 KB records| K[Kinesis stream 125 shards]
K --> L[Lambda consumer: enrich and pre-aggregate]
L -->|one transaction per batch| D[(DynamoDB: per-shard minute aggregates and checkpoint)]
L -.->|poison records| Q[SQS on-failure queue]
D --> R[Rollup job every minute]
R --> S3A[(S3 aggregates as Parquet, queried by Athena)]
K --> F[Firehose]
F --> S3R[(S3 raw archive for replay)]
1. Ingestion and buffering
- Shard count. 100 MB/s divided by 1 MB/s per shard is 100 shards by bytes. With 24 events packed per record, the record rate is 100,000 / 24 = 4,167 records/s, far below the 100,000 records/s limit of 100 shards. Bytes bind, so I provision 125 shards (25% headroom for skew and bursts).
- Why pack events. Unpacked, 100,000 records/s would need 100 shards just for the record-count limit, and every downstream fee charged per record (Kinesis PUT payload units, which bill each record in 25 KB chunks; Firehose's 5 KB billing increment) would be charged 100,000 times a second. Packing is done by the gateway or by the Kinesis Producer Library's (KPL, an AWS client library that batches and compresses records before sending them) aggregation feature.
- Partition key. Device ID, hashed, so load spreads evenly across shards. A partition key concentrated on a few tenants creates a hot shard (one shard at its 1 MB/s ceiling while others idle), which shows up as write throttling.
- Provisioned vs on-demand. On-demand mode (no shard planning, billed per GB) costs about $42,048/month at this volume versus about $1,522 provisioned (arithmetic below). For flat, known 100k/s load, provisioned wins by more than 20x. I would flip to on-demand only for a new stream whose traffic I cannot yet predict.
- Buffering. The stream itself is the buffer: default retention is 24 hours, so a consumer outage of a few hours loses nothing; the backlog is replayed when the consumer recovers.
2. Processing concurrency
- The Kinesis ESM runs one concurrent batch per shard by default.
ParallelizationFactor(1 to 10) raises that, but it breaks per-shard ordering (ordering is then only per partition key), and my deduplication below relies on per-shard order. So I keep the factor at 1 and scale by adding shards instead. - Maximum concurrency is therefore 125 execution environments, well inside Lambda's default regional quota of 1,000 concurrent executions. I set reserved concurrency (a per-function concurrency allotment that is both a guaranteed floor and a hard ceiling) of about 150 on this function so it can never starve other functions, and so nothing else can starve it.
- Batching settings:
BatchSize10,000 (the maximum),MaximumBatchingWindowInSeconds2. Each shard receives 100,000 / 125 = 800 events/s, so each 2 s batch carries about 1,600 events (about 67 packed records, about 1.6 MB, safely under the 6 MB invocation payload limit even after base64 encoding inflates it by a third). - That gives 125 / 2 = 62.5 invocations/s. If a batch takes about 120 ms, average concurrency is 62.5 x 0.12 = 7.5 environments. Cold starts (the extra latency when a new execution environment boots) are irrelevant here: environments stay warm under constant load, and a few hundred ms of extra delay on a 2 s batch changes nothing.
- Enrichment (joining each event to device metadata) uses a reference table loaded into memory at initialization and refreshed every few minutes, never a per-event database call. 100,000 lookups/s against a database would itself be a larger system than the pipeline.
3. Idempotency: exactly-once aggregates on at-least-once delivery
Kinesis plus Lambda is at-least-once. The ESM retries a whole batch if the function errors or times out, so if the function wrote aggregates and then timed out before returning, the retry would add the same events again. For aggregates (counters), duplicates silently inflate numbers.
The fix: store a per-shard checkpoint (the highest sequence number already counted) in the same atomic write as the aggregate increments. Each Kinesis record has a sequence number that increases within a shard. On every invocation the function:
- Reads the checkpoint for its shard (cached in memory, re-read if the conditional write in step 3, a database write that only takes effect if the value has not changed since it was last read, fails).
- Drops records whose sequence number is at or below the checkpoint (already counted).
- Commits, in one DynamoDB
TransactWriteItemscall (an all-or-nothing multi-item write): increment the (shard, minute) aggregate items, and set the checkpoint to the batch's last sequence number, with a condition that the checkpoint still equals the value just read.
Either both the counts and the checkpoint land or neither does, so a retry of any subset of the batch finds its records already covered and adds nothing.
A small simulation proves it. It models 4 shards of 5,000 records, with a 10% chance each attempt crashes before writing and a 10% chance it crashes after writing, and halves the batch after each failure (a simplified model of BisectBatchOnFunctionError):
import random
from collections import Counter
def make_stream(num_shards=4, per_shard=5000, seed=7):
rng = random.Random(seed)
shards = {}
for s in range(num_shards):
recs = []
for seq in range(1, per_shard + 1):
minute = (seq - 1) * 60 // per_shard # events spread over 60 minutes
group = rng.choice(["sensor-a", "sensor-b", "sensor-c"])
recs.append((seq, minute, group))
shards[s] = recs
return shards
class Store:
"""Stands in for the database: aggregate counters plus a per-shard checkpoint.
commit() is atomic, like one TransactWriteItems call."""
def __init__(self):
self.agg = Counter()
self.hwm = {}
def commit(self, shard, deltas, new_hwm):
self.agg.update(deltas)
if new_hwm is not None:
self.hwm[shard] = new_hwm
def handler(store, shard, batch, use_checkpoint):
last_done = store.hwm.get(shard, 0) if use_checkpoint else 0
fresh = [r for r in batch if r[0] > last_done]
deltas = Counter((shard, minute, group) for _, minute, group in fresh)
new_hwm = max(last_done, batch[-1][0]) if use_checkpoint else None # never move backwards
store.commit(shard, deltas, new_hwm)
def run(use_checkpoint, batch_size=500, p_fail_before=0.10, p_fail_after=0.10, seed=42):
shards = make_stream()
rng = random.Random(seed)
store = Store()
for shard, recs in shards.items():
pos, size = 0, batch_size
while pos < len(recs):
batch = recs[pos:pos + size]
roll = rng.random()
if roll < p_fail_before: # crash before the write: nothing committed
size = max(1, len(batch) // 2) # simplified bisect-on-error
continue
handler(store, shard, batch, use_checkpoint)
if roll < p_fail_before + p_fail_after: # write committed, then timeout: batch retried
size = max(1, len(batch) // 2)
continue
pos += len(batch)
size = batch_size
truth = Counter((s, m, g) for s, recs in shards.items() for _, m, g in recs)
return truth, store.agg
for mode in (False, True):
truth, got = run(use_checkpoint=mode)
label = "with per-shard checkpoint" if mode else "naive (no checkpoint) "
print(f"{label}: true events={sum(truth.values())}, counted={sum(got.values())}, "
f"every (shard, minute, group) cell exact={truth == got}")
Output:
naive (no checkpoint) : true events=20000, counted=22250, every (shard, minute, group) cell exact=False
with per-shard checkpoint: true events=20000, counted=20000, every (shard, minute, group) cell exact=True
The naive version over-counts by 11.25%. The max(...) on the checkpoint line is load-bearing: my first version set the checkpoint to the retried half-batch's last sequence number, which moved it backwards after a bisect and still double counted (21,125). A checkpoint must only ever move forward.
Two things this does not cover, stated so nobody assumes otherwise: duplicates created by the producer (a device retrying a PUT gets a new sequence number) need an event ID and a dedupe step, and the raw archive in S3 is at-least-once, so queries over raw events should deduplicate by event ID.
4. Error handling
| Failure | Handling |
|---|---|
| Transient (throttle, timeout, database conflict) | Function raises; ESM retries. MaximumRetryAttempts 5, MaximumRecordAgeInSeconds 3,600 so a stuck shard cannot block for the whole 24 h retention |
| One malformed record | Validate and skip it inside the handler, writing it to an error log with its sequence number. It must not fail the batch |
| A record that crashes the code | BisectBatchOnFunctionError splits the batch to isolate it; ReportBatchItemFailures (partial batch response) lets the function say "everything before sequence X succeeded" so good records are not reprocessed |
| Retries exhausted | ESM on-failure destination: an SQS (Amazon Simple Queue Service) queue that receives the batch's shard and sequence range (metadata, not the records), from which an operator replays the records out of the stream |
| Late events | The minute rollup runs at minute + 3 and again at minute + 10, overwriting that minute's output, so late arrivals within 10 minutes are counted |
The DLQ (dead-letter queue) idea here is the on-failure destination: a place failed work goes so it is not lost and does not block the shard.
5. Storage choice
- Aggregates: DynamoDB holds the live per-shard, per-minute partial counts (small, write-heavy, needs atomic conditional writes). A rollup (a job that combines many small partial records into one summarized one) Lambda run every minute by EventBridge Scheduler (AWS's managed cron) sums the 125 shard items for a closed minute and writes one Parquet object per minute to S3 with a deterministic key such as
aggregates/dt=2026-09-27/hh=14/mm=05.parquet. Rewriting the same key is naturally idempotent. Analysts query with Athena. A small daily compaction merges per-minute files into hourly ones to avoid a small-file problem (many tiny files make each file's fixed per-file read overhead dominate, slowing every downstream query). - Raw events: Firehose reads directly from the stream and writes compressed files to S3, so enrichment or aggregation logic can be replayed after a bug fix.
- What would change it: if dashboards need sub-second freshness or high-concurrency interactive queries, I would write aggregates to a real-time OLAP (online analytical processing) database such as ClickHouse or Apache Druid instead of S3 plus Athena.
6. Observability
- Consumer lag: Lambda's
IteratorAgemetric (how old the last record in each batch was when the batch was sent to the function). This is the single most important alarm: rising iterator age means processing is falling behind. Alarm if it exceeds 60 s for 5 minutes. - Stream health: Kinesis
WriteProvisionedThroughputExceeded(a hot or under-provisioned shard) andGetRecords.IteratorAgeMilliseconds. - Function health:
Errors,Throttles,Duration(p99, meaning the 99th percentile),ConcurrentExecutions. - Business correctness: a reconciliation job (one that independently checks two derived numbers against each other and flags any mismatch) comparing hourly raw-archive event counts with aggregated counts, alarming on a difference above 0.1%. This catches silent double counting or silent drops that no infrastructure metric will show.
- Structured logs with shard ID and sequence range per batch, so any aggregate can be traced to the records behind it.
7. Cost profile
List prices for us-east-1 as published on the AWS pricing pages; the per-event CPU cost is an assumption to be measured in a load test. DynamoDB is billed in write request units (WRU): under on-demand pricing a normal write costs 1 WRU per KB, and (as the pitfall below explains) a transactional write costs double that per KB:
import math
# us-east-1 list prices (verify on the pricing pages before relying on them)
LAMBDA_GBS = 0.0000166667 # $ per GB-second, x86, first tier
LAMBDA_REQ = 0.20 / 1e6 # $ per request
KDS_SHARD_HR = 0.015 # provisioned shard-hour
KDS_PUT_UNIT = 0.014 / 1e6 # per 25 KB PUT payload unit
KDS_OD_IN, KDS_OD_OUT = 0.08, 0.04 # on-demand $/GB ingested, retrieved
FH_GB = 0.029 # Firehose ingestion $/GB, billed in 5 KB increments
DDB_WRU = 0.625 / 1e6 # on-demand write request unit
SEC_MONTH = 730 * 3600 # AWS bills a month as 730 hours
events_s, event_kb = 100_000, 1.0
events_per_record = 24 # producers pack ~24 KB records
shards = 125 # 100 MB/s needs 100 shards at 1 MB/s each, plus 25% headroom
records_s = events_s / events_per_record
mb_s = events_s * event_kb / 1000
gb_month = mb_s / 1000 * SEC_MONTH
kds_prov = shards * 730 * KDS_SHARD_HR + records_s * SEC_MONTH * KDS_PUT_UNIT
kds_od = gb_month * (KDS_OD_IN + 2 * KDS_OD_OUT) # two consumers: Lambda and Firehose
window_s, mem_gb = 2, 1.0
inv_s = shards / window_s
events_per_inv = events_s / inv_s
duration_s = 0.040 + events_per_inv * 0.00005 # ASSUMED: 40 ms fixed + 0.05 ms per event
lam = inv_s * SEC_MONTH * (LAMBDA_REQ + duration_s * mem_gb * LAMBDA_GBS)
per_event_invoke_requests = events_s * SEC_MONTH * LAMBDA_REQ
wru_per_inv = 2 * 1 + 2 * 4 * 1.05 # txn: 1 KB checkpoint + ~1.05 x 4 KB aggregate item, 2 WRU per KB
ddb = inv_s * SEC_MONTH * wru_per_inv * DDB_WRU
fh_billed_kb = math.ceil(events_per_record * event_kb / 5) * 5
fh = records_s * fh_billed_kb / 1e6 * SEC_MONTH * FH_GB
print(f"records/s={records_s:.0f} data={mb_s:.0f} MB/s GB/month={gb_month:,.0f}")
print(f"Kinesis provisioned={kds_prov:,.0f} on-demand={kds_od:,.0f}")
print(f"Lambda: {inv_s:.1f} inv/s, {events_per_inv:.0f} events/inv, {duration_s*1000:.0f} ms, "
f"avg concurrency={inv_s*duration_s:.1f}, cost={lam:,.0f}")
print(f"Lambda request fee alone if invoked once per event={per_event_invoke_requests:,.0f}")
print(f"DynamoDB aggregates={ddb:,.0f}")
print(f"Firehose raw archive (billed {fh_billed_kb} KB/record)={fh:,.0f}")
print(f"TOTAL (provisioned Kinesis)={kds_prov + lam + ddb + fh:,.0f} $/month")
Output (dollars per month):
records/s=4167 data=100 MB/s GB/month=262,800
Kinesis provisioned=1,522 on-demand=42,048
Lambda: 62.5 inv/s, 1600 events/inv, 120 ms, avg concurrency=7.5, cost=361
Lambda request fee alone if invoked once per event=52,560
DynamoDB aggregates=1,068
Firehose raw archive (billed 25 KB/record)=7,939
TOTAL (provisioned Kinesis)=10,890 $/month
The ~4 KB aggregate-item size is an assumption for this example: it is enough room for counts and sums across roughly a few dozen (device group, metric) key combinations packed into one item; size it against your own key cardinality. The ~1.05 multiplier on that item accounts for the occasional batch that straddles a minute boundary and touches two minute items. S3 storage and Athena query charges are excluded because they depend on retention and query volume.
What the numbers say:
- Compute is the cheap part ($361). Lambda cost is dominated by how you batch, not by how fast the code is.
- The raw archive is 73% of the bill. Its value (replay, ad hoc analysis) should be confirmed with the data consumers. If a 7-day replay window is enough, extending the stream's retention instead of archiving everything could remove most of it.
- Firehose bills each record rounded up to 5 KB, so 1 KB unpacked records would be billed at 5x their size. Packing records to just under a 5 KB multiple matters.
Trade-offs and pitfalls
- Is serverless right at this volume? The load is flat and high, which is where serverless is weakest on unit price. It still wins here because the managed pieces (stream, poller, retries, scaling to 125 parallel consumers, no servers to patch) cost less in engineering time than running and patching your own Kafka and stream-processing cluster, while the Lambda line itself is only a few hundred dollars. If enrichment grew heavy, say 2 ms of CPU per event instead of 0.05 ms, two things break. First, one consumer per shard can then process at most 1 / 0.002 = 500 events/s, but each shard receives 800, so the stream falls steadily behind unless you add shards: 100,000 events/s / 500 events/s per shard needs at least 200 shards by that arithmetic (up from 125), before headroom. Second, compute becomes 100,000 x 0.002 = 200 GB-seconds every second at 1 GB, which is 200 x 2,628,000 s/month x $0.0000166667/GB-s, about $8,760 a month in Lambda duration alone. At that point a long-running container consumer (for example Apache Flink, an open-source stream-processing engine that keeps one continuously running process instead of Lambda's start-stop-per-batch model, on a managed service) is the better call, because steady, CPU-heavy work is exactly where per-millisecond billing loses.
- Per-event invocation is the classic mistake: 262.8 billion requests a month.
- ParallelizationFactor as the first scaling lever silently breaks the per-shard ordering the checkpoint relies on. Scale with shards.
- Aggregating in the function's memory across invocations is unsafe: environments are recycled at any time and two environments can serve the same shard over time. State lives in the database or in Lambda's managed tumbling-window state (a built-in feature that carries a small amount of state between consecutive batches from the same shard, for windowed aggregation, which AWS documents as at-least-once and capped at 1 MB per shard, so it does not remove the double-count problem).
- Transactions are not free: each item written in a DynamoDB transaction costs 2 write request units (WRU) per KB instead of 1. That is why there is one transaction per batch, not per event.
You're being asked to recommend whether to adopt a cloud provider's proprietary managed serverless inference offering: autoscaling and monitoring are built in and deployment is simpler, but it increases vendor lock-in. Draft the key points of a decision memo: what business and technical trade-offs would you weigh, how would you quantify the cost and velocity gains, and what would you propose to mitigate lock-in if you go ahead with it?
Sample Answer
Direct answer
My memo would recommend adopting the managed serverless inference offering with a deliberate exit path (inference is running a trained model against new input to get a prediction, as opposed to training it; "serverless" means the vendor runs and autoscales that serving layer for you), provided the numbers show a payback within a few months and the lock-in is bounded by keeping the model, the serving contract and the infrastructure definitions portable. Lock-in is not a reason to refuse on its own; it is a cost to be priced (what it would take to leave) and compared with what the managed service saves every month.
Memo structure
- Decision requested: adopt the managed serverless inference service for production model serving, yes or no, with conditions.
- Context: current state (self-managed serving on our own cluster, who operates it, incidents, deploy lead time).
- Options: (A) stay self-managed, (B) adopt managed with portability guardrails, (C) adopt managed fully, using every proprietary feature.
- Trade-offs, quantified.
- Risks and lock-in mitigation.
- Recommendation, conditions and the triggers that would reverse it.
Business and technical trade-offs to weigh
| Factor | For adopting | Against adopting |
|---|---|---|
| Operations | Autoscaling, patching, monitoring built in; frees platform engineers | Less control over scaling behaviour, instance types, tail latency (how slow the slowest requests are, not just the typical one) |
| Velocity | Faster path from trained model to endpoint | Constrained by the vendor's supported frameworks and container contract |
| Cost | No paying for idle capacity with serverless scaling; fewer people-hours | Higher unit price per compute-hour; vendor pricing power over time |
| Reliability | Vendor's SLA (service-level agreement) and multi-zone deployment | Shared-fate with the vendor's regional incidents; cold starts (extra delay while an idle endpoint spins back up) on scale-from-zero (the endpoint scaling down to nothing between requests to avoid paying for idle capacity) |
| Security and compliance | Vendor-managed patching, integrated identity and encryption | Data residency, audit requirements, model artifacts in vendor storage |
| Strategy | Focus engineering on models, not plumbing | Harder to go multi-cloud or negotiate; switching cost grows with every proprietary feature used |
Quantifying cost and velocity gains
Every input below is an assumption for illustration; the memo would replace each with our measured figures.
- Current infrastructure spend: $18,000 per month.
- Self-managed operations: 1.5 engineers at a loaded cost (salary plus benefits, payroll taxes and overhead, the true cost of employing someone, not just their salary) of $200,000 per year, so $25,000 per month.
- Managed service unit-price premium: 30% on infrastructure, so $23,400 per month.
- Residual operations after adoption: 0.3 engineers, so $5,000 per month.
Sensitivity: the saving comes from people, not infrastructure. The infrastructure premium could rise to (25,000 minus 5,000) / 18,000 = 111% before it erased the operations saving. That break-even premium is the single most useful number in the memo, because it tells the reader how wrong the price assumption can be before the decision flips.
Velocity: measure lead time from "model approved" to "serving production traffic" and deploys per month, before and after a pilot. If lead time drops from 5 working days to 1, each update reaches production 4 working days sooner; across 4 model updates a month that is 4 x 4 = 16 working days a month of earlier value delivery; tie it to a business metric the model moves (conversion, fraud caught) rather than claiming a dollar figure you cannot defend.
Lock-in, priced as exit cost: estimate the engineering work to move to another platform. With portability guardrails in place, suppose 12 engineer-weeks: 12 / 52 × $200,000 = $46,154. At $14,600 per month of savings, the exit cost is repaid in about 3.2 months. Without guardrails (proprietary SDK calls in application code, vendor-specific model packaging, a vendor feature store, the vendor's managed store of precomputed model inputs, tightly coupled to its own pipeline), the exit estimate might triple, which is precisely why option B beats option C.
Mitigating lock-in if we go ahead
- Portable model artifacts: store weights in open formats (ONNX, the Open Neural Network Exchange format, or safetensors, a fast, safe file format for storing model weights) in our own bucket, not only in the vendor's model registry (the vendor's own catalog and version store for trained models).
- Standard serving container: use a bring-your-own-container option (deploying your own packaged container image instead of the vendor's pre-built one) with a standard HTTP inference contract, so the same image runs on Kubernetes (an open-source system for running and scaling containers across a cluster of machines).
- Thin client boundary: application code calls an internal
predict()interface; only one adapter module knows the vendor SDK. - Infrastructure-as-code (declaring cloud resources in version-controlled files that a tool applies, instead of clicking them into existence, so the setup is readable and reproducible elsewhere) for endpoints, scaling policies and permissions.
- Vendor-neutral telemetry: OpenTelemetry (the open standard for traces and metrics) for traces and metrics, so dashboards and alerts survive a move.
- Exit drill: once or twice a year, deploy one production model to the fallback platform and send it shadow traffic (a copy of real production requests, so the fallback sees real load without its answers being served to users). This turns the exit estimate from a guess into a measurement.
- Commercial terms: negotiate price protection and data-export terms at signing, when leverage is highest.
Recommendation and reversal triggers
Adopt option B. Reverse or re-evaluate if: the infrastructure premium measured in the pilot exceeds about 80% (well inside the 111% break-even, leaving margin), p99 latency (the response time only the slowest 1% of requests exceed) or cold starts violate the product SLO (service-level objective, the target we hold ourselves to, such as "p99 under 300 ms"), a compliance requirement cannot be met, or the exit drill's measured cost grows beyond a year of savings.
Pitfalls
- Comparing unit prices only and ignoring people costs, which dominate at small scale.
- Treating lock-in as binary rather than as an exit cost that grows with each proprietary feature.
- Claiming velocity gains without a before-and-after measurement.
- Adopting every proprietary add-on on day one because it is convenient, which silently multiplies the exit cost.
Create a migration plan for moving a monolithic, Kubernetes-based serving stack (model servers, feature caches, batch jobs) to a serverless-first architecture. Cover your assessment criteria for what moves where (FaaS, serverless containers, or a managed service), the cutover strategy, the rollback plan, how you'd load-test before cutover, and what organizational or process changes the new model requires.
Sample Answer
Direct answer
I would not move the stack wholesale. I would classify every workload against a small set of criteria, move the ones serverless fits, leave the ones it does not, and migrate one workload at a time behind a traffic switch that makes rollback a configuration change. Model servers usually go to serverless containers or a managed inference service (not plain FaaS), feature caches go to a managed cache, and batch jobs go to an orchestrated workflow. The plan fails more often on organization (ownership, on-call, cost accountability) than on technology, so that is a workstream of its own.
1. Assessment criteria: what moves where
Score each workload on these questions:
| Criterion | Points toward FaaS | Points toward serverless containers | Points toward a managed service or staying put |
|---|---|---|---|
| Run time per unit of work | Seconds; always under 15 min (Lambda's hard limit) | Minutes to hours; long-lived HTTP | Continuous streaming, days-long jobs |
| Hardware | CPU, up to 10 GB memory | CPU or GPU (on platforms that offer it) | Specialized accelerators, big-memory nodes |
| Traffic shape | Spiky, with long idle periods | Bursty but with long requests or big images | Flat and high utilization (always-on is cheaper) |
| State | Stateless; state externalized | Stateless per instance | Stateful in memory (cache, session, model sharded across nodes) |
| Artifact size | Under 250 MB zip or 10 GB image | Any image size | N/A |
| Latency | Tolerates occasional cold start or can pay for pre-warming | Same, with minimum instances | Hard real-time tails: even the slowest request must finish inside a strict, non-negotiable deadline |
Applied to this stack:
| Component | Destination | Why |
|---|---|---|
| Model servers (GPU) | Managed inference endpoint, or serverless GPU containers if traffic is bursty | FaaS has no GPU; the choice between the two is utilization (always-on wins when busy most of the time) |
| Model servers (small CPU models) | FaaS or serverless containers | Short requests, stateless, bursty |
| Feature caches (Redis in the cluster) | Managed cache service | A cache is state; functions cannot host it. A managed service removes the operational work without pretending it is stateless |
| Batch jobs (feature backfills, retraining triggers, scoring) | Workflow orchestrator (AWS Step Functions, a managed state machine service) driving functions for short steps and container batch tasks for long ones | Fan-out (splitting one job into many parallel sub-tasks), retries and checkpointing (the orchestrator recording how far a long job has progressed, so a restart resumes from there instead of the beginning) come from the orchestrator; steps longer than 15 minutes cannot be functions |
| Anything with flat, near-100% utilization | Stay on Kubernetes | Serverless charges a premium for elasticity you would not use |
2. Cutover strategy
Use the strangler-fig pattern: put a routing layer in front, move one workload at a time, and let the old system shrink.
flowchart LR
A[Assess and score] --> B[Build serverless version]
B --> C[Shadow traffic]
C --> D[Canary 5 percent]
D --> E[Ramp 25 to 100 percent]
E --> F[Soak 2 weeks]
F --> G[Decommission K8s workload]
D -.rollback.-> H[Weight back to K8s]
E -.rollback.-> H
(K8s above is the common abbreviation for Kubernetes, used to keep the diagram's node labels short.)
- Order: start with a low-risk, stateless, bursty workload (a batch scoring job or a small CPU model) to build the paved road (the standard, supported path other teams then follow: the defaults and tooling that make the safe way the easy way); do GPU serving and anything on the checkout path last.
- Shadow first: mirror live requests to the new path, compare outputs, discard results.
- Canary and ramp: a canary release sends a small slice of real traffic to the new version first, so a bug is caught while it can only hurt a few users. Weighted routing at the gateway or load balancer (5%, 25%, 50%, 100%), each step held through a full daily traffic cycle.
- Data: the feature cache is derived data, so warm the managed cache from the source of truth before the canary rather than dual-writing (writing every update to both the old and new stores in parallel to keep them in sync) forever.
3. Rollback plan
- Keep the Kubernetes deployment running at reduced replicas during the ramp and for a two-week soak (running at full production traffic without incident, long enough to catch issues that only show up over days, such as a slow memory leak or a cache that gradually fills up) after 100%. Rollback is flipping the routing weight, measured in minutes.
- Define rollback triggers in advance: SLO burn rate (how fast you are consuming your error budget, the allowed amount of SLO violation, relative to the rate that would exhaust it before the period ends) above threshold, error rate above baseline, prediction mismatch above tolerance, cost per request above budget.
- Make it possible: no one-way schema or data-format changes during the migration window; if the batch pipeline changes its output format, the old consumers must still read it.
- Rehearse: execute one rollback during the canary on purpose, so the first real one is not the first one.
4. Load-testing before cutover
Serverless fails differently under load, so test the serverless-specific limits, not just throughput:
- Concurrency quota. The default Lambda quota is 1,000 concurrent executions per Region, shared by all functions in the account. Size it with Little's law (concurrency equals arrival rate multiplied by time per request): 800 requests per second at 250 ms average is 800 × 0.25 = 200 concurrent; check the peak, not the average, and request increases ahead of time.
- Scaling rate. Each function can add at most 1,000 execution environments every 10 seconds. A step change larger than that is throttled with HTTP 429 (Too Many Requests) until it catches up.
- Downstream protection. 200 concurrent environments means up to 200 database connections; put a connection proxy in front (a small pooling layer, such as RDS Proxy, that sits between the many function environments and the database and multiplexes their connections into a much smaller, stable pool) or cap the function with reserved concurrency (a per-function ceiling on how many concurrent executions it may run, guaranteed and capped, so it cannot consume the whole account's quota or overwhelm a downstream dependency).
- Cold-start share at burst. Ramp from idle to peak and measure the fraction of requests that hit new environments.
- GPU capacity. For scale-to-zero GPU services, test acquiring GPUs at burst time in the target region.
- Cost under load. Record billed duration and memory during the test to project the monthly bill against the Kubernetes baseline.
5. Organizational and process changes
| Area | Change |
|---|---|
| Ownership | Teams own their functions end to end (code, permissions, alarms), instead of a platform team owning "the cluster" |
| Platform team's new job | A paved road: templates, IaC (infrastructure-as-code) modules, a standard logging and tracing library, default alarms, least-privilege role patterns |
| On-call | Runbooks (step-by-step written procedures an on-call engineer follows during an incident) for throttling, cold-start regressions, quota exhaustion, dead-letter queue growth; alarms per function |
| Cost accountability | Cost is now per request and per team; tag every function and review cost per 1,000 predictions monthly (FinOps: the practice of making cloud cost a shared, visible, actively managed engineering concern rather than a monthly surprise) |
| Quota management | Someone owns account-level limits, because one team's runaway function can starve everyone's concurrency |
| Skills | Event-driven design, idempotent handlers (safe to run twice), distributed tracing across async hops |
Pitfalls
- Moving the always-busy workloads. They get more expensive and gain nothing.
- Forcing GPU serving into FaaS or splitting a stateful cache across functions.
- Big-bang cutover, which makes rollback a second migration.
- Treating it as an infrastructure project. Without ownership and cost accountability, a hundred functions become a distributed mess no one owns.
Unlock Full Question Bank
Get access to all 21 Serverless and Function-as-a-Service Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.