Serverless and Function-as-a-Service Architecture Questions
Building applications on managed, event-triggered compute: functions-as-a-service (AWS Lambda, Azure Functions, Google Cloud Functions, Cloudflare Workers) and serverless containers. Covers the invocation lifecycle and cold starts (init vs handler work, provisioned concurrency, packaging, layers and container images), statelessness and externalizing state, execution limits (timeouts, memory, payload and /tmp size), concurrency and scaling behavior (account limits, burst scaling, protecting downstream databases with connection proxies and throttling), event sources and trigger semantics (at-least-once delivery, retries, idempotent handlers, dead-letter handling), composing functions with managed services and workflow orchestrators (Step Functions and equivalents), and serverless-specific observability, security (per-function IAM, secrets) and pay-per-invocation cost modeling. Includes when serverless fits versus containers or VMs, and vendor lock-in trade-offs. General compute selection, generic messaging patterns, and general idempotency theory are covered by their own topics.
Your team needs to pick a compute model for a platform with spiky, unpredictable traffic, and one concrete workload on the table is a CPU-bound, latency-sensitive job like image processing under strict SLAs. Compare serverless (FaaS) and managed Kubernetes across cost model (pay-per-use versus reserved), cold-start latency, burst concurrency, observability, vendor lock-in, and operational burden. Give a recommended phased roadmap, including migration considerations, for getting there.
Sample Answer
Direct answer
For a platform with spiky, unpredictable traffic, I would run the CPU-bound image-processing workload on serverless functions (FaaS) first, with a small amount of provisioned concurrency for the baseline, behind a queue wherever the SLA (service-level agreement) allows asynchronous processing. The reason is arithmetic: to meet a strict SLA through spikes on Kubernetes, you must keep peak capacity running or accept queueing while nodes boot, and at a 20x peak-to-baseline ratio that is roughly 7x the cost of the serverless option in the example below. I would plan the roadmap so the workload can move to managed Kubernetes when measured average utilization makes containers cheaper (about 48% on the prices below) or when a hard limit such as GPUs or 15-minute runtimes forces it.
Terms used throughout:
- FaaS (functions as a service): the provider runs your function on demand and bills per request and per GB-second (memory size times run time). AWS Lambda is the example.
- Managed Kubernetes: a cloud-run Kubernetes control plane (for example EKS, Amazon Elastic Kubernetes Service) that runs your containers as pods (one or more containers scheduled and scaled together) on servers (nodes) you pay for while they run; a pod autoscaler adds or removes pod copies to match load, and if no node has room, a cluster autoscaler adds or removes nodes themselves.
- Cold start: the delay when the platform creates a new execution environment (download code, start the runtime, run initialization) before it can serve a request.
- Provisioned concurrency: paying Lambda to keep a set number of environments initialized so those requests skip the cold start.
- p99: the 99th percentile latency, the time within which 99% of requests finish.
Assumptions for the worked example
| Input | Value |
|---|---|
| Work per image | 300 ms of CPU at 1 vCPU (assumed; measure it) |
| Baseline | 20 images/s |
| Spikes | 400 images/s, totalling 60 minutes a day, arriving unpredictably |
| SLA | p99 under 2 s per image for synchronous requests |
| Prices | AWS us-east-1 list prices: Lambda $0.0000166667 per GB-s, $0.20 per million requests; provisioned concurrency $0.0000041667 per GB-s allocated and $0.0000097222 per GB-s of duration; Fargate (AWS's serverless container runtime, billed per second for the vCPU and memory a container reserves) $0.000011244 per vCPU-second and $0.000001235 per GB-second |
Lambda allocates CPU in proportion to memory, reaching one full vCPU at 1,769 MB, so a CPU-bound function should be sized at 1,769 MB or more, not at the 128 MB default.
LAMBDA_GBS, LAMBDA_REQ = 0.0000166667, 0.20/1e6
PC_ALLOC, PC_DUR = 0.0000041667, 0.0000097222
FG_VCPU_S, FG_GB_S = 0.000011244, 0.000001235
SEC_MONTH = 730*3600
mem_gb = 1769/1024 # 1 vCPU equivalent
cpu_s = 0.300 # ASSUMED: 300 ms of CPU per image at 1 vCPU
baseline, spike, spike_min_per_day = 20, 400, 60 # img/s; spikes total 60 min/day
days = 730/24
imgs = (baseline*(24*60-spike_min_per_day)*60 + spike*spike_min_per_day*60)*days
avg_rate = imgs/SEC_MONTH
# Lambda on-demand only
lam = imgs*(cpu_s*mem_gb*LAMBDA_GBS + LAMBDA_REQ)
# Lambda with provisioned concurrency covering baseline (conc = rate*duration, +50%)
pc = int(round(baseline*cpu_s*1.5))
pc_alloc = pc*mem_gb*SEC_MONTH*PC_ALLOC
base_imgs = baseline*SEC_MONTH
pc_dur = base_imgs*cpu_s*mem_gb*PC_DUR
od_imgs = imgs-base_imgs
lam_pc = pc_alloc + pc_dur + od_imgs*cpu_s*mem_gb*LAMBDA_GBS + imgs*LAMBDA_REQ
print(f"images/month={imgs/1e6:.2f}M avg={avg_rate:.1f}/s peak concurrency={spike*cpu_s:.0f} vCPU")
print(f"Lambda on-demand: ${lam:,.0f}/month")
print(f"Lambda + {pc} provisioned envs for baseline: ${lam_pc:,.0f}/month")
def fargate(vcpus): return vcpus*(FG_VCPU_S+2*FG_GB_S)*SEC_MONTH
peak_vcpu = spike*cpu_s/0.7 # run at 70% target utilization
print(f"containers sized for peak ({peak_vcpu:.0f} vCPU always on): ${fargate(peak_vcpu):,.0f}/month")
avg_vcpu = avg_rate*cpu_s/0.7
print(f"containers if autoscaling were perfect ({avg_vcpu:.1f} vCPU avg): ${fargate(avg_vcpu):,.0f}/month")
lam_vcpu_s = mem_gb*LAMBDA_GBS
fg_vcpu_s = FG_VCPU_S+2*FG_GB_S
print(f"$ per busy vCPU-second: Lambda={lam_vcpu_s:.3e} Fargate(1vCPU,2GB)={fg_vcpu_s:.3e} ratio={lam_vcpu_s/fg_vcpu_s:.2f}")
print(f"containers win on compute price above {fg_vcpu_s/lam_vcpu_s:.0%} average utilization")
images/month=94.17M avg=35.8/s peak concurrency=120 vCPU
Lambda on-demand: $832/month
Lambda + 9 provisioned envs for baseline: $813/month
containers sized for peak (171 vCPU always on): $6,178/month
containers if autoscaling were perfect (15.4 vCPU avg): $553/month
$ per busy vCPU-second: Lambda=2.879e-05 Fargate(1vCPU,2GB)=1.371e-05 ratio=2.10
containers win on compute price above 48% average utilization
Reading those numbers before the takeaway: the provisioned count of 9 and the peak concurrency of 120 both come from the same idea (conc = rate x duration in the code). Keeping up with the 20 images/s baseline without queueing needs 20 x 0.3 s = 6 environments busy at once; add 50% headroom for the noise around that average and you provision 9 (pc). The 400 images/s spike needs 400 x 0.3 s = 120 concurrent environments to absorb instantly (peak concurrency=120 vCPU), and sizing containers for that peak at a 70% target utilization (headroom so a real traffic wobble does not immediately throttle) needs 120 / 0.7 ≈ 171 vCPU held always-on.
Provisioned concurrency has two prices, not one: a reservation rate ($0.0000041667/GB-s, PC_ALLOC) charged for every second an environment sits ready whether or not it is handling a request, and a lower execution rate ($0.0000097222/GB-s, PC_DUR, versus $0.0000166667 on-demand) charged only while it is actually running code, because you already paid to reserve it. The baseline's 52.56M images a month (base_imgs, 56% of the 94.17M total) run at that discounted execution rate, about $170 to hold 9 environments ready plus $265 while they are busy; every image above baseline (od_imgs, about 41.6M) still pays the full on-demand rate (about $359), plus the flat per-request fee on all 94.17M images (about $19). That totals the printed $813, a little under the $832 of running every image at the plain on-demand rate, because the discount on the baseline's share of images outweighs the cost of holding 9 environments ready around the clock.
Why 48%, specifically: Lambda's price per unit of actual work is fixed no matter how busy the account is, so its cost per busy vCPU-second is the constant $2.879e-05 printed above. A container's price is fixed per second it runs whether or not it has work, so its cost per unit of actual work falls as utilization rises, a container running at half utilization effectively pays double its sticker price for the work it does. Setting the container's cost per busy second at utilization U equal to Lambda's fixed rate gives the crossover:
U1.371×10−5=2.879×10−5⇒U=2.8791.371≈0.48Below about 48% utilization the container is paying for idle time Lambda never bills; above it, the container's busy-time price undercuts Lambda's fixed one.
Container cost is priced at Fargate (per-second container) rates for a like-for-like unit price; EC2 virtual-machine nodes under EKS would be somewhat cheaper per vCPU but add the $0.10/hour cluster fee and node management.
The key reading: the realistic Kubernetes number lies between $553 and $6,178, and where it lands depends on how much capacity you must hold warm to survive a spike you cannot predict. Scaling a Kubernetes deployment means the pod autoscaler adds pods, and if no node has room, the cluster autoscaler must launch new nodes, which takes time you must measure in your own environment. Unpredictable spikes of 20x therefore push you toward the $6,178 end. Serverless costs about $830 whichever way the spikes fall. Provisioned concurrency for the baseline is roughly cost-neutral here ($813 versus $832) because the baseline keeps those environments busy, and it removes cold starts from the steady part of the traffic.
Comparison across the six named dimensions
| Dimension | Serverless (FaaS) | Managed Kubernetes | Verdict for this workload |
|---|---|---|---|
| Cost model | Pay per use: per request plus per GB-second, zero when idle; about 2.1x the price of a busy container per vCPU-second | Reserved capacity: pay for nodes whether busy or not; cheaper per unit when utilization is high | FaaS while utilization is low and spiky; K8s above about 48% average utilization |
| Cold-start latency | New environments add latency, larger for big container images with native image libraries; provisioned concurrency removes it for the baseline | No per-request cold start, but new nodes take a long time to arrive during a spike | FaaS with provisioned baseline; spikes pay some cold starts, which the 2 s SLA must be tested against |
| Burst concurrency | Lambda scales each function by up to 1,000 new environments every 10 seconds, up to the account concurrency quota (default 1,000 per region, raisable) | Limited by spare node capacity and node-launch time | FaaS clearly: 120 concurrent at peak is far inside Lambda's limits |
| Observability | Per-invocation metrics and logs out of the box; distributed tracing needs instrumentation; no host to log into when debugging | Full control: Prometheus (an open-source metrics system) metrics, profilers, sidecars (helper containers running next to each service); you run the stack | K8s gives deeper profiling for CPU-bound tuning; FaaS is adequate with tracing added |
| Vendor lock-in | Triggers, event shapes, IAM (AWS Identity and Access Management) permissions and orchestration are provider-specific | Images and manifests are portable across clouds | K8s, but mitigable on FaaS by shipping the function as a container image with a thin handler |
| Operational burden | No nodes, patching or cluster upgrades | Cluster upgrades, node patching, autoscaler tuning, capacity planning | FaaS, unless a platform team already runs K8s |
Design specifics for the image workload on FaaS
- Pass references, not bytes. Clients upload to object storage (S3) with a pre-signed URL (a time-limited upload link that needs no credentials); the function receives the object key. Lambda's synchronous payload limit is 6 MB per request and response, too small for many images.
- Asynchronous by default. Where the product allows it, an upload event goes onto a queue (SQS, Amazon Simple Queue Service, AWS's managed message queue) that triggers the function; the queue absorbs spikes and a reserved-concurrency cap (a hard ceiling on how many concurrent environments this function may use, protecting shared account capacity and downstream systems from being overrun; distinct from provisioned concurrency, which keeps environments pre-warmed) protects anything downstream. Only truly interactive paths call synchronously.
- Deterministic output keys (for example
thumbnails/{image_id}/{size}.webp) so a retried event overwrites the same object instead of creating duplicates. - Memory sized for CPU: benchmark at 1,769 MB and above. Extra memory adds more vCPU, but only helps if the image library uses multiple threads.
- Container-image packaging with the native imaging library baked in; keep the image small because image size affects cold starts.
Phased roadmap
Assume the current state is image processing inside an existing monolith on virtual machines.
- Phase 0: measure (1-2 weeks). Instrument the current path: CPU time per image by size, image-size distribution, real peak-to-baseline ratio, and SLA misses. Replace my assumed 300 ms with the measured number and rerun the cost model.
- Phase 1: extract behind an interface. Put image processing behind an internal API and an upload event, with the business logic in a plain library that does not know whether it runs in Lambda or a container. This is the lock-in insurance.
- Phase 2: shadow, then canary, on FaaS. First shadow it: send a copy of production traffic to the function and compare outputs byte for byte and latency at p99, without serving its results to users. Then canary it: route 5%, 25%, 100% of real traffic, with a flag that routes back to the old path instantly. Load-test a synthetic 20x spike before 100%, including a cold-start-heavy run.
- Phase 3: harden. Provisioned concurrency sized to the measured baseline, reserved concurrency as a ceiling, a DLQ (dead-letter queue, where messages that fail repeatedly are parked for inspection) on the async path, alarms on throttles (invocations rejected because a concurrency limit was hit), errors, duration p99 and queue age (how long a message waits before being processed, a sign the consumer is falling behind).
- Phase 4: re-evaluate quarterly. Track average utilization (busy vCPU-seconds divided by provisioned vCPU-seconds had it run on containers). If it stays above roughly 50%, or the workload needs GPUs or more than 15 minutes per task, move it to managed Kubernetes. Because Phase 1 kept the logic in a portable container image, that move is a deployment change, not a rewrite.
Migration considerations: run old and new paths in parallel until output parity is proven; keep the old path deployable for rollback through at least one full traffic cycle (for example a month-end peak); make every processing step idempotent (running it twice has the same effect as running it once) so a replay during cut-over is harmless; and budget for the database or storage behind the function, which sees the same 20x spikes once the compute stops being the bottleneck.
Trade-offs and pitfalls
- Comparing FaaS with a perfectly autoscaled cluster. That comparison favors Kubernetes ($553 versus $832) but assumes capacity appears the instant a spike starts, which is the one thing unpredictable traffic denies you.
- Under-sizing Lambda memory for CPU-bound work. At low memory the function gets a fraction of a vCPU and runs proportionally slower, so it can cost about the same while blowing the SLA.
- Ignoring the 2.1x unit-price gap. It is why serverless is a phase for this workload's current traffic shape, not necessarily its permanent home.
- Choosing Kubernetes for portability alone when nobody on the team runs it. The operational burden is real on day one; lock-in cost is only paid if you actually move.
You're leading a cross-functional evaluation to choose between serverless functions and containerized microservices for a new component, for example a CPU-bound batch or ingestion workload. Build an evaluation framework: what criteria would you weight (cost, latency, cold starts, scalability, operational complexity), how would you weight them, and what's your plan for a short pilot or spike to gather real evidence before committing?
Sample Answer
Direct answer
I would run this as a weighted decision matrix whose weights are agreed before anyone scores the options, followed by a time-boxed two-week pilot that measures the handful of criteria where the options actually differ, with pass/fail thresholds written down in advance. The matrix does not make the decision on its own: its real job is to show which criteria swing the result, so the pilot spends its time measuring those and not the ones that do not matter. For a CPU-bound batch or ingestion workload, the swing factors are almost always cost per unit of work and operational complexity, not cold starts.
Terms:
- Serverless functions (FaaS, functions as a service), for example AWS Lambda: code runs on demand, scales automatically, billed per request and per millisecond; each invocation is capped (15 minutes on Lambda).
- Containerized microservices: the component runs as long-lived containers on an orchestrator such as Kubernetes or a managed container service, billed for the capacity you keep running.
- Cold start: extra latency when the platform creates a new function environment.
- Spike: a short, throwaway experiment built to answer one question with evidence.
Step 1: set up the evaluation (week 0)
Who is in the room and what each owns:
| Stakeholder | Owns the evidence for |
|---|---|
| Engineering lead for the component | Build effort, code shape, team fit |
| SRE / platform (site reliability engineering) | Operational complexity, on-call load, observability |
| Finance or FinOps (cloud cost management) | Cost model at current and 3x volume |
| Security | Isolation, secrets, identity per workload |
| Product | Latency and freshness requirements, what "good enough" means |
Decision rules agreed up front:
- Criteria and weights are fixed before scoring, and signed off by the group. Changing weights after seeing scores is how teams launder a preference into a number.
- Hard constraints are gates, not weights. For example: a single task that needs more than 15 minutes, more than 10,240 MB of memory, or a GPU rules out Lambda regardless of score.
- Each criterion has a 1-5 scale defined in measurable terms (for cost, for example: 5 = under $50 per million records processed, 1 = over $200; every score in between is a number, not a feeling), so two people scoring the same evidence get the same number.
Step 2: criteria and weights
For a CPU-bound batch ingestion workload (say, nightly and hourly files that must be parsed, validated and loaded):
| Criterion | Weight | Why this weight for this workload |
|---|---|---|
| Cost per unit of work | 30 | CPU-bound and recurring, so compute dominates the bill |
| Scalability (fan-out, running many instances of the work in parallel, for backlogs, catch-up after outages) | 20 | Batch backlogs arrive in bursts |
| Operational complexity | 20 | Determines on-call load for years |
| Team fit (skills, existing platform) | 15 | Execution risk |
| Latency | 10 | Batch has a freshness target (how old the newest loaded data is allowed to be, for example "loaded within 1 hour of arriving"), not a per-request target |
| Cold starts | 5 | Almost irrelevant for batch; a few hundred ms on a minutes-long job is noise |
The ordering is itself a finding worth saying out loud to stakeholders: cold starts get most of the attention in serverless debates and deserve the least weight here.
Step 3: score, then find the swing factors
Illustrative scores before any pilot (1-5; these are the team's priors, which the pilot will replace):
weights = {"cost": 30, "latency": 10, "cold_starts": 5, "scalability": 20,
"ops_complexity": 20, "team_fit": 15} # sums to 100
scores = { # 1-5, illustrative until the pilot replaces them
"serverless": {"cost": 3, "latency": 4, "cold_starts": 3, "scalability": 5,
"ops_complexity": 4, "team_fit": 3},
"containers": {"cost": 4, "latency": 4, "cold_starts": 5, "scalability": 3,
"ops_complexity": 3, "team_fit": 4},
}
def total(opt, w): return sum(w[c] * scores[opt][c] for c in w) / 100
for opt in scores:
print(f"{opt}: {total(opt, weights):.2f} / 5")
# sensitivity: shift weight from scalability to cost, 5 points at a time
for shift in range(0, 21, 5):
w = dict(weights, cost=weights["cost"] + shift, scalability=weights["scalability"] - shift)
s, c = total("serverless", w), total("containers", w)
print(f"cost={w['cost']} scalability={w['scalability']}: serverless={s:.2f} containers={c:.2f} "
f"-> {'serverless' if s > c else 'containers' if c > s else 'tie'}")
serverless: 3.70 / 5
containers: 3.65 / 5
cost=30 scalability=20: serverless=3.70 containers=3.65 -> serverless
cost=35 scalability=15: serverless=3.60 containers=3.70 -> containers
cost=40 scalability=10: serverless=3.50 containers=3.75 -> containers
cost=45 scalability=5: serverless=3.40 containers=3.80 -> containers
cost=50 scalability=0: serverless=3.30 containers=3.85 -> containers
What this tells the group: the totals are 0.05 apart, and moving just 5 weight points from scalability to cost flips the winner. A matrix this close has not decided anything. It has told us exactly what to measure: the real cost per million records on each option, and whether serverless's scalability advantage is needed (how big do backlogs actually get?). Latency and cold starts score the same or barely matter, so the pilot should spend no time on them.
Step 4: the pilot (two weeks, time-boxed)
Hypotheses with thresholds, written before the pilot starts:
| Hypothesis | Measurement | Pass threshold (example) |
|---|---|---|
| Cost per million records on each option | Replay one representative day of real input; record billed duration or container-hours | The option within 20% of the cheaper one scores 5 on cost |
| Backlog catch-up | Inject a 6-hour backlog; time to drain | Drains within the 1-hour freshness target |
| Operational complexity | Count steps to deploy, roll back, and diagnose one injected failure; have an engineer outside the pilot team do it | Diagnose the injected failure in under 30 minutes using only the dashboards |
| Hard limits | Largest input file's processing time and memory | Largest task fits in 15 minutes with 30% margin, or it is split |
Build both options as thin as possible, against the same code. Keep the parsing and loading logic in one library; wrap it in a Lambda handler (fanned out with a Step Functions Distributed Map, AWS's managed workflow service's mode for processing large datasets in S3 object storage in parallel, which runs up to 10,000 child executions, separate parallel runs of the same workflow step, one per item or batch of the dataset, concurrently by default) and in a container job (a Kubernetes Job, a Kubernetes workload that runs a container to completion instead of keeping it running indefinitely, or an equivalent managed container task such as an ECS/Fargate task). This makes the comparison about the platform, not about two different implementations.
Guard against pilot bias: replay the same input on both; run each at least three times; report the spread, not the best run; and have the platform team, not the component team, score operational complexity.
Outcome: a one-page decision record (a short written artifact stating what was decided, why, and what would change it) with the measured numbers substituted into the matrix, the resulting scores, the sensitivity check re-run, the decision, and the conditions under which it would be revisited. Utilization is the right kind of trigger because a container's cost is fixed whether or not it is busy, while the FaaS option costs about the same per unit of work regardless of utilization, so above some utilization the fixed-cost container option becomes the cheaper one; the specific crossover is illustrative here (for example, "revisit if volume triples or average utilization of the container option would exceed 50%") and should be replaced, the same way $X and $Y were above, once the pilot has measured each option's real cost per unit of work.
Trade-offs and pitfalls
- Scoring before weighting turns the matrix into a justification for whatever the loudest person already wanted.
- Averaging away a hard constraint. A 15-minute timeout cannot be offset by good scores elsewhere; treat it as a gate or redesign the work into smaller tasks.
- A pilot on toy data. CPU-bound work scales with real file sizes and real skew; a synthetic dataset of uniform small files will flatter serverless.
- Letting the pilot become the production system by default. Time-box it and throw away what does not win.
- Ignoring team fit because it is "soft". A cheaper option the team cannot operate at 3 a.m. is not cheaper.
Walk through what actually determines your bill for a workload running on a serverless platform, then use it: given a request rate, average execution duration, and memory size, estimate the monthly cost. Which levers would you tune first to bring the cost down without hurting your latency SLOs?
Sample Answer
Direct answer
On a FaaS (functions-as-a-service) platform such as AWS Lambda you pay for two things per function: a per-request fee and a compute fee measured in GB-seconds (memory allocated in GB multiplied by billed execution time in seconds). Memory is also the CPU dial, so memory and duration trade against each other. Around the function you also pay for the front door (API gateway), logs, and data transfer, and those often rival the function itself. The first levers are: cut billed duration (especially time spent waiting on I/O), right-size memory by measurement, move to the cheaper ARM architecture, and cut log volume, checking p99 (99th-percentile) latency against your SLO (service-level objective, the latency target you have committed to) after each change.
What determines the bill
Prices below are AWS Lambda, US East (N. Virginia), from the current public price list.
| Component | Unit price | Driven by |
|---|---|---|
| Requests | $0.20 per 1M requests | Invocation count |
| Compute, x86 | $0.0000166667 per GB-second (first tier) | Memory size multiplied by billed duration (rounded up to 1 ms) |
| Compute, ARM (Graviton, AWS's own ARM-based processor family) | $0.0000133334 per GB-second (first tier) | Same, about 20% cheaper per GB-second |
| Free tier | 1M requests and 400,000 GB-seconds per month | Applies once per account |
| Init (cold start: the one-time setup work, code download, runtime start-up, and your own init code, that only the first request on a new execution environment pays) time | Billed as duration | Since 1 August 2025, INIT is billed for on-demand zip functions (functions deployed as a plain code archive rather than a container image, the default and simplest packaging) on managed runtimes (AWS's built-in language runtimes, as opposed to a custom one you bring yourself) too |
| Provisioned concurrency (paying to keep a set number of environments already initialized and idle, so even the first request to each hits a warm environment) | $0.0000041667 per GB-second reserved, plus a lower duration rate | Only if you pre-warm environments |
Ephemeral /tmp beyond 512 MB | $0.0000000309 per GB-second | Rarely material |
Outside the function: an HTTP API on API Gateway is $1.00 per million requests for the first 300M, and log ingestion into CloudWatch Logs is $0.50 per GB at the standard rate.
The formula
GB-s=N×d×1024M Cost=106N−106×0.20+(GB-s−400,000)×pGB-swhere N is requests per month, d is average billed duration in seconds, M is memory in MB, and pGB-s is the per GB-second price.
Worked estimate
Inputs: 50 requests per second on average, 200 ms average duration, 512 MB memory, x86, 30-day month.
N=50×2,592,000=129,600,000 requests GB-s=129,600,000×0.2×0.5=12,960,000 Compute=(12,960,000−400,000)×0.0000166667=$209.33 Requests=(129.6−1)×0.20=$25.72Function total: about $235 per month. Add an HTTP API front door at 129.6M multiplied by $1.00 per million = $129.60, so the API gateway alone is more than half of the function bill. If each invocation writes 1 KB of logs, that is 129.6 GB, or about $65 of ingestion at $0.50 per GB.
This estimate correctly subtracts the free tier once, because the formula is exact billing math for one account. Do not let the free tier itself drive a design decision: at this example's 50 requests per second, 400,000 GB-s is about 3.1 percent of the 12,960,000 GB-s used, worth roughly $6.67 of compute at the x86 rate, not the well-under-half-a-percent rounding error it becomes once traffic reaches roughly 300 requests per second or higher, where the same fixed 400,000 GB-s free allowance is a shrinking share of a much larger GB-second total. Account for it once for precision at whatever scale you are estimating, and do not let its exact size change which lever you pull first.
Which levers to pull first, and their latency risk
| Lever | Effect in the example | Latency risk |
|---|---|---|
| Cut billed duration (reuse connections created outside the handler, so opening one at module scope during Init means later warm requests skip connection setup and pay only for the request itself; parallelize downstream calls; stop waiting synchronously on slow services) | 200 ms to 120 ms drops compute to (129.6M × 0.12 × 0.5 − 400,000) × 0.0000166667 = $122.93 | None; usually improves latency |
| Switch to ARM | Same GB-seconds at $0.0000133334 gives $167.47 (20% less compute) | Low, but native dependencies (compiled libraries, not pure interpreted code) must be rebuilt for the new CPU architecture, a compiled binary built for x86 will not run on ARM, and benchmarked |
| Right-size memory by measurement | If the work is CPU-bound, 1024 MB may halve duration to 100 ms: 129.6M × 0.1 × 1.0 GB-s costs the same $209.33 but p99 improves | Going down in memory on CPU-bound code raises duration and p99; going up can be free |
| Cut log volume (sample debug logs, drop full-event dumps) | 1 KB to 200 bytes per call saves about $52 of the $65 | None for latency; keep error logs complete |
| Replace the front door where possible (HTTP API, API Gateway's newer, cheaper, simpler product, instead of REST API, its older, more feature-rich and more expensive one, or a direct function URL, a dedicated HTTPS endpoint Lambda can generate for a function that bypasses API Gateway entirely, for internal callers) | Up to the $129.60 line | Check you are not dropping features such as caching or request validation |
| Batch for asynchronous sources (process 10 queue messages per invocation) | Divides request count and per-invocation overhead | Adds per-message latency, fine off the user path |
The ordering principle: duration first because it is almost always free latency-wise, architecture second, memory by measurement (a power-tuning sweep: run the same workload at several memory settings and plot cost and latency at each, for example with the open-source AWS Lambda Power Tuning tool, to find the cheapest setting that still meets your latency target), then the non-function lines.
Pitfalls
- Lowering memory to save money. Because CPU scales with memory (1,769 MB is one full vCPU, virtual CPU, one full CPU core's worth of compute, on Lambda), halving memory on CPU-bound code roughly doubles duration: no saving and a worse p99.
- Ignoring the non-function lines. Engineers optimize GB-seconds while the gateway, logs, NAT (network address translation) gateway data processing, or cross-region transfer dominate.
- Paying for idle waiting. A function that polls a job for 30 seconds is billed for 30 seconds of memory. Hand waiting to a workflow orchestrator or an event instead.
- Forgetting that the free tier is per account, so it is noise at production scale and should not be used to justify a design.
What does 'serverless' and Functions-as-a-Service (FaaS) mean in production? Compare FaaS to running the same workload as containers, covering lifecycle, typical billing model, and operational overhead. When is FaaS a good fit versus a poor one?
Sample Answer
Direct answer
"Serverless" means you give a cloud provider your code plus a rule for when to run it, and the provider owns the machines: provisioning, scaling, patching and idle capacity. FaaS (functions-as-a-service) is the most common form. You deploy one function (AWS Lambda, Azure Functions, Google Cloud Run functions, Cloudflare Workers), it runs when an event arrives (an HTTP request, a queue message, a file landing in storage), and you pay per request and per millisecond of running time, with zero cost while idle. Compared with running the same code as containers, FaaS trades control and steady-state cost efficiency for near-zero operations and pay-for-use billing. It fits spiky, event-driven, short tasks; it fits poorly for steady high-volume traffic, long jobs, and anything that needs a long-lived process.
A plain-language picture: containers are like leasing a car (you pay every month whether you drive or not, and you look after it); FaaS is like a taxi (you pay only for trips, someone else maintains it, but a lot of trips every day costs more than the lease, and sometimes you wait for one to arrive).
How FaaS and containers actually differ
| Dimension | FaaS (AWS Lambda as the example) | Containers (Kubernetes, Amazon ECS, Cloud Run services, Cloud Run's always-on container mode, distinct from the Cloud Run functions mentioned above) |
|---|---|---|
| Lifecycle | The platform creates an execution environment (a small isolated sandbox with your code loaded) when a request needs one. Creating it is the cold start. It is reused for later requests while "warm", frozen between requests, and destroyed whenever the platform decides (Lambda recycles environments every few hours even under steady traffic). One environment handles one request at a time. | You start long-lived processes that run until you stop them. You choose how many replicas run and when to add more (autoscaling rules). One process usually serves many requests concurrently. |
| Billing | Per request plus duration in GB-seconds (memory configured in GB multiplied by seconds of run time), measured to 1 ms. Lambda in US East (N. Virginia), x86: $0.20 per million requests and $0.0000166667 per GB-second. Idle costs nothing. | Per vCPU-hour (one virtual CPU core running for one hour) and GB-hour of whatever you provisioned, busy or idle. You usually need at least two replicas for availability, so there is a cost floor. |
| Operational overhead | No servers, OS patching, cluster upgrades or capacity planning. You still own: code, permissions (a least-privilege IAM role, where IAM is AWS Identity and Access Management, per function), timeouts, concurrency limits, retries, logs and metrics. | You own image patching, cluster and node management (less on managed options such as AWS Fargate, which runs containers without you managing servers), health checks, rollout strategy and scaling policy. |
| Hard limits | Lambda: 15-minute maximum run time, 128 MB to 10,240 MB memory, 6 MB request and 6 MB response payload for synchronous calls, 250 MB unzipped package (10 GB as a container image), 1,000 concurrent executions per Region by default (raisable). | Effectively your cluster's size. Jobs can run for days; processes can hold WebSocket connections open (a WebSocket is a persistent, two-way network connection that stays open between messages, unlike a normal request/response HTTP call). |
Worked example: where the cost lines cross
An HTTP API whose handler takes 100 ms on average with 512 MB (0.5 GB) of memory. A 30-day month has 2,592,000 seconds. Cost per month at a steady rate of r requests per second:
cost(r)=r×2,592,000×(106$0.20+0.1s×0.5GB×$0.0000166667)| Steady load | Requests / month | Request charge | GB-seconds | Duration charge | Total |
|---|---|---|---|---|---|
| 1 req/s | 2,592,000 | $0.52 | 129,600 | $2.16 | $2.68 |
| 50 req/s | 129,600,000 | $25.92 | 6,480,000 | $108.00 | $133.92 |
So every 1 request/second of sustained load costs about $2.68 a month on Lambda, and the bill scales linearly. Now suppose (an assumed figure; substitute your real quote) a pair of small always-on containers that can comfortably serve 50 req/s costs $30 a month. The break-even is $30 / $2.68 per req/s, about 11.2 req/s sustained. Below that, FaaS is cheaper and you also skip the operations work; well above it, the containers win on cost. (This ignores the Lambda free tier of 1 million requests and 400,000 GB-seconds a month, and extras common to both, such as the API gateway, the managed service that receives HTTP requests and routes them to your function, billed per request, and logging charges.)
Most real traffic is not flat: an internal tool busy 2 hours a day and idle 22 costs FaaS almost nothing during the idle hours, while the containers bill all 24. That idle ratio, more than the peak rate, is what usually decides it.
When FaaS is a good fit versus a poor one
Good fit
- Event glue: react to an uploaded file, a queue message, a database change stream, a schedule (a nightly report), or a database change stream (a feed of row-level insert, update and delete events a database emits as they happen).
- Spiky or unpredictable traffic with long idle periods, where paying for idle replicas is waste.
- Low-to-moderate volume APIs where a few hundred ms of occasional cold-start latency is acceptable.
- Small teams that would rather not run a cluster.
Poor fit
- Steady, high-volume traffic above the break-even: you pay a premium for elasticity you are not using.
- Work longer than 15 minutes, or that needs a long-lived process (WebSocket servers, streaming consumers holding state, in-memory caches you rely on).
- Strict tail-latency targets where a cold start in the p99 (the 99th-percentile latency: the value only 1% of requests exceed) breaks the SLO (service-level objective), unless you pay to keep environments warm.
- GPU work, or large in-memory datasets that every environment would load separately.
My default: start event-driven and low-traffic workloads on FaaS, and move a component to containers when its measured sustained load passes the break-even, it hits a hard limit, or cold starts break its latency target.
Trade-offs and pitfalls
- "Serverless" still has servers, you just do not manage them. You gain speed and lose knobs: no choosing the OS, limited control over networking, CPU tied to the memory setting.
- Scaling can hurt what sits behind you. 1,000 concurrent functions can mean 1,000 database connections, and most relational databases cap total connections in the low hundreds (each one holds memory and overhead on the database server), so that alone can take the database down before it ever drops a request for load reasons. Relational databases need a connection proxy (a lightweight service that sits between many function instances and the database, holding a small pool of real database connections open and multiplexing many callers onto them) or a concurrency cap on the function (a hard ceiling on how many copies run at once, so the connection count never exceeds what the database can hold).
- Retries are part of the model. Asynchronous and queue triggers deliver at least once (a message may arrive more than once), so handlers must be idempotent: running twice has the same effect as once.
- Hidden line items: API gateway fees, log ingestion, NAT gateway (network address translation gateway: the managed component that lets resources inside a private network reach the internet, billed per hour and per GB) traffic for functions inside a private network. Price the whole path, not just the function.
- Vendor lock-in comes mostly from the triggers, permissions and managed services around the function, not the function code. Keeping business logic in plain modules the handler calls keeps a later move to containers cheap.
You're being asked to recommend whether to adopt a cloud provider's proprietary managed serverless inference offering: autoscaling and monitoring are built in and deployment is simpler, but it increases vendor lock-in. Draft the key points of a decision memo: what business and technical trade-offs would you weigh, how would you quantify the cost and velocity gains, and what would you propose to mitigate lock-in if you go ahead with it?
Sample Answer
Direct answer
My memo would recommend adopting the managed serverless inference offering with a deliberate exit path (inference is running a trained model against new input to get a prediction, as opposed to training it; "serverless" means the vendor runs and autoscales that serving layer for you), provided the numbers show a payback within a few months and the lock-in is bounded by keeping the model, the serving contract and the infrastructure definitions portable. Lock-in is not a reason to refuse on its own; it is a cost to be priced (what it would take to leave) and compared with what the managed service saves every month.
Memo structure
- Decision requested: adopt the managed serverless inference service for production model serving, yes or no, with conditions.
- Context: current state (self-managed serving on our own cluster, who operates it, incidents, deploy lead time).
- Options: (A) stay self-managed, (B) adopt managed with portability guardrails, (C) adopt managed fully, using every proprietary feature.
- Trade-offs, quantified.
- Risks and lock-in mitigation.
- Recommendation, conditions and the triggers that would reverse it.
Business and technical trade-offs to weigh
| Factor | For adopting | Against adopting |
|---|---|---|
| Operations | Autoscaling, patching, monitoring built in; frees platform engineers | Less control over scaling behaviour, instance types, tail latency (how slow the slowest requests are, not just the typical one) |
| Velocity | Faster path from trained model to endpoint | Constrained by the vendor's supported frameworks and container contract |
| Cost | No paying for idle capacity with serverless scaling; fewer people-hours | Higher unit price per compute-hour; vendor pricing power over time |
| Reliability | Vendor's SLA (service-level agreement) and multi-zone deployment | Shared-fate with the vendor's regional incidents; cold starts (extra delay while an idle endpoint spins back up) on scale-from-zero (the endpoint scaling down to nothing between requests to avoid paying for idle capacity) |
| Security and compliance | Vendor-managed patching, integrated identity and encryption | Data residency, audit requirements, model artifacts in vendor storage |
| Strategy | Focus engineering on models, not plumbing | Harder to go multi-cloud or negotiate; switching cost grows with every proprietary feature used |
Quantifying cost and velocity gains
Every input below is an assumption for illustration; the memo would replace each with our measured figures.
- Current infrastructure spend: $18,000 per month.
- Self-managed operations: 1.5 engineers at a loaded cost (salary plus benefits, payroll taxes and overhead, the true cost of employing someone, not just their salary) of $200,000 per year, so $25,000 per month.
- Managed service unit-price premium: 30% on infrastructure, so $23,400 per month.
- Residual operations after adoption: 0.3 engineers, so $5,000 per month.
Sensitivity: the saving comes from people, not infrastructure. The infrastructure premium could rise to (25,000 minus 5,000) / 18,000 = 111% before it erased the operations saving. That break-even premium is the single most useful number in the memo, because it tells the reader how wrong the price assumption can be before the decision flips.
Velocity: measure lead time from "model approved" to "serving production traffic" and deploys per month, before and after a pilot. If lead time drops from 5 working days to 1, each update reaches production 4 working days sooner; across 4 model updates a month that is 4 x 4 = 16 working days a month of earlier value delivery; tie it to a business metric the model moves (conversion, fraud caught) rather than claiming a dollar figure you cannot defend.
Lock-in, priced as exit cost: estimate the engineering work to move to another platform. With portability guardrails in place, suppose 12 engineer-weeks: 12 / 52 × $200,000 = $46,154. At $14,600 per month of savings, the exit cost is repaid in about 3.2 months. Without guardrails (proprietary SDK calls in application code, vendor-specific model packaging, a vendor feature store, the vendor's managed store of precomputed model inputs, tightly coupled to its own pipeline), the exit estimate might triple, which is precisely why option B beats option C.
Mitigating lock-in if we go ahead
- Portable model artifacts: store weights in open formats (ONNX, the Open Neural Network Exchange format, or safetensors, a fast, safe file format for storing model weights) in our own bucket, not only in the vendor's model registry (the vendor's own catalog and version store for trained models).
- Standard serving container: use a bring-your-own-container option (deploying your own packaged container image instead of the vendor's pre-built one) with a standard HTTP inference contract, so the same image runs on Kubernetes (an open-source system for running and scaling containers across a cluster of machines).
- Thin client boundary: application code calls an internal
predict()interface; only one adapter module knows the vendor SDK. - Infrastructure-as-code (declaring cloud resources in version-controlled files that a tool applies, instead of clicking them into existence, so the setup is readable and reproducible elsewhere) for endpoints, scaling policies and permissions.
- Vendor-neutral telemetry: OpenTelemetry (the open standard for traces and metrics) for traces and metrics, so dashboards and alerts survive a move.
- Exit drill: once or twice a year, deploy one production model to the fallback platform and send it shadow traffic (a copy of real production requests, so the fallback sees real load without its answers being served to users). This turns the exit estimate from a guess into a measurement.
- Commercial terms: negotiate price protection and data-export terms at signing, when leverage is highest.
Recommendation and reversal triggers
Adopt option B. Reverse or re-evaluate if: the infrastructure premium measured in the pilot exceeds about 80% (well inside the 111% break-even, leaving margin), p99 latency (the response time only the slowest 1% of requests exceed) or cold starts violate the product SLO (service-level objective, the target we hold ourselves to, such as "p99 under 300 ms"), a compliance requirement cannot be met, or the exit drill's measured cost grows beyond a year of savings.
Pitfalls
- Comparing unit prices only and ignoring people costs, which dominate at small scale.
- Treating lock-in as binary rather than as an exit cost that grows with each proprietary feature.
- Claiming velocity gains without a before-and-after measurement.
- Adopting every proprietary add-on on day one because it is convenient, which silently multiplies the exit cost.
Unlock Full Question Bank
Get access to all 6 Serverless and Function-as-a-Service Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.