Distributed Systems and Microservices Testing Questions
Testing systems composed of many interacting services. Covers integration and end-to-end testing across service boundaries, handling eventual consistency and partial failure, and validating behavior in distributed, specialized architectures. Includes fault injection and testing at scale.
Design a deterministic fault-injection harness for a distributed workflow engine that executes multi-step transactions with compensating actions. Describe how you would inject faults at a precise, chosen step and timing, how you would capture a deterministic trace of what happened, and how you'd assert that the retry and compensation logic still produces a correct final state.
Sample Answer
Direct answer
Build the fault-injection harness around a hook keyed by step index and elapsed time within the workflow (so a fault can be triggered precisely at "step 3, 50 milliseconds after it starts" rather than randomly), capture a full, ordered trace of every step transition and compensation action the workflow takes, and assert that after the fault, the compensation logic drives the workflow to one of its explicitly-defined terminal states (fully completed, or fully and correctly rolled back), never a state that is neither.
Structured elaboration
- Precise, addressable injection points. Instrument the workflow engine (or wrap each step's execution) with a hook that the test can parameterize:
inject_fault(at_step=3, after_ms=50, fault_type="exception"). This lets the test target the exact moment a compensating-transaction pattern is most likely to have a bug: mid-step, after some side effect has already happened but before the step recorded its own completion. - Deterministic trace capture. Every step transition (started, succeeded, failed, compensating, compensated) gets logged with the workflow's instance ID and a monotonic sequence number, so a test can assert not just the FINAL state but the exact SEQUENCE of compensations that ran, in the right order, exactly once each.
- Assert convergence to a valid terminal state. For a saga-style workflow (payment charged, inventory reserved, shipping scheduled, each with its own compensating action), a fault at step 2 must drive the workflow to EITHER "all steps after 2 skipped, step 1's compensation ran, workflow is CANCELLED" or, if the design retries instead of compensating, "step 2 eventually succeeded on retry, workflow proceeded normally." There is no valid state where step 1 committed, step 2 failed, and no compensation and no retry ran; that stuck state is exactly what this harness exists to catch.
- Idempotency of compensation itself. Since the workflow engine may itself retry a failed compensation, assert that running a given compensating action twice (because of the engine's own retry-on-failure behavior) does not double-refund, double-restock, or otherwise apply its effect twice.
Worked example
class FaultSchedule:
def __init__(self, at_step, after_ms, fault_type="exception"):
self.at_step, self.after_ms, self.fault_type = at_step, after_ms, fault_type
def test_fault_at_step_2_triggers_compensation_of_step_1_only():
trace = []
engine = SagaEngine(
steps=[charge_payment, reserve_inventory, schedule_shipping],
compensations=[refund_payment, release_inventory, cancel_shipping],
on_transition=lambda evt: trace.append(evt),
)
engine.inject_fault(FaultSchedule(at_step=2, after_ms=10, fault_type="exception"))
engine.run(order_id="ord-99")
assert engine.final_state("ord-99") == "CANCELLED"
step_names = [e.step for e in trace if e.kind == "compensated"]
assert step_names == ["charge_payment"], (
f"expected only step 1's compensation to run, got: {step_names}"
)
# idempotency of the compensation itself: replay it and assert no double refund
balance_before = ledger.balance("acct-1")
engine.replay_compensation("ord-99", step="charge_payment")
assert ledger.balance("acct-1") == balance_before, "replaying a compensation must not refund twice"
Trade-offs and pitfalls
- Precise step-and-timing fault injection requires the workflow engine to expose a hookable extension point; for a fully managed third-party workflow service without such hooks, you are limited to injecting faults at the boundary of each step's own external call (still useful, just less precise about the exact millisecond).
- The most common bug this class of test finds is a compensation that assumes it will only ever run once; always include a deliberate double-invocation of the compensation path in the test suite, since the engine's own retry logic is a completely realistic way for that to happen in production.
- Asserting only the final state (
CANCELLED) without asserting the exact sequence of compensations can hide a bug where the RIGHT final state is reached by the WRONG path (for example, two compensations ran instead of one, coincidentally leaving the same final balance); capture and assert the full trace, not just the endpoint.
Design tests that fail based on SLO violations rather than basic unit assertions. For a distributed payment service with an SLO of 99.9% success rate and a p95 latency under 200 milliseconds, explain how you'd implement automated test suites that collect metrics during runs, evaluate the SLO with appropriate statistical significance, and report a meaningful failure with diagnostic artifacts suitable for CI gating.
Sample Answer
Direct answer
Gate the test on the actual service-level objective (SLO) computed FROM a batch of runs, not a single request: run enough requests to have statistical power, compute the observed success rate and the p95 latency across that batch, apply a statistical test appropriate for a proportion (success rate) with a defined confidence level, and fail the gate with the actual computed numbers and a link to raw per-request data attached, rather than a bare pass or fail.
Structured elaboration
- Why a single-request assertion is the wrong tool. An SLO like "99.9% success rate" is a statement about a POPULATION of requests, not any one request; a single request succeeding or failing tells you almost nothing about whether the SLO holds. The test therefore needs to generate a batch of requests (synthetic load, not necessarily huge, but large enough to say something statistically meaningful) and reason about the batch's aggregate behavior.
- Statistical significance for a success-rate SLO. Treat each request's success/failure as a Bernoulli trial and use a proportion confidence interval (a Wilson score interval is a reasonable, well-behaved choice for small-to-moderate sample sizes) around the OBSERVED success rate; only fail the gate if the interval's upper bound is still below the SLO target, which avoids flagging a fine system as failing purely due to a small sample size producing sampling noise near the boundary.
- Latency SLOs need a percentile, not a mean. Compute the p95 directly from the collected latency samples (do not average percentiles across batches; recompute from the raw pooled samples) and compare against the SLO's stated threshold.
- Reporting a meaningful failure. On a gate failure, report the exact observed success rate, its confidence interval, the observed p95, the sample size, and (critically) enough raw per-request data (which specific requests failed, their error codes/latencies) that a developer can start debugging immediately rather than re-running the whole batch first just to see what happened.
Worked example
import math
def wilson_lower_bound(successes, n, z=1.96):
if n == 0:
return 0.0
p_hat = successes / n
denom = 1 + z**2 / n
center = p_hat + z**2 / (2 * n)
margin = z * math.sqrt((p_hat * (1 - p_hat) + z**2 / (4 * n)) / n)
return (center - margin) / denom
def test_payment_service_meets_slo_under_synthetic_load():
n = 5000
results = [run_payment_request(sample_payment()) for _ in range(n)]
successes = sum(1 for r in results if r.status == "success")
latencies = sorted(r.latency_s for r in results)
observed_rate = successes / n
p95_latency = latencies[int(n * 0.95)]
# Fail only if we're confident (95%) the TRUE success rate is below the 99.9% SLO,
# not just because THIS batch happened to see a couple of failures. The sample size
# itself has to be large enough that a genuinely healthy service can actually clear
# this bound: at n=500, even ZERO observed failures gives a Wilson lower bound of only
# about 0.992, meaning the gate could never pass at all regardless of true reliability.
# At n=5000, a perfect run's lower bound is about 0.999, leaving real headroom to
# distinguish a healthy service from a genuinely regressed one.
lower_bound = wilson_lower_bound(successes, n)
assert lower_bound >= 0.995, (
f"SLO breach: observed success rate {observed_rate:.4f} over n={n} requests, "
f"95% confidence lower bound {lower_bound:.4f} is below the 99.9% target "
f"({n - successes} failures; see attached per-request error codes for detail)"
)
assert p95_latency < 0.200, f"p95 latency {p95_latency*1000:.0f}ms exceeded the 200ms SLO"
With n=5000 requests and, say, 10 observed failures (success rate 0.998), the Wilson lower bound comes out to about 0.9963, comfortably above 0.995, so the gate correctly PASSES: a handful of failures out of 5000 is still statistically consistent with a true rate at or above 99.9%. But with 25 failures out of 5000 (success rate 0.995), the lower bound drops to about 0.9926, below 0.995, so the gate correctly FAILS: at that failure rate the true success rate is very unlikely to actually be 99.9% or better. The choice of n matters as much as the statistical test itself: at n=500, even a perfect run (zero failures) only reaches a Wilson lower bound of about 0.992, which never clears the 0.995 cutoff, so that sample size would fail the gate unconditionally no matter how reliable the service actually is. Choosing n large enough that a genuinely healthy service has real headroom to pass is part of implementing this correctly, not an afterthought.
Trade-offs and pitfalls
- Too small a sample size cuts both ways, and the more dangerous direction is the less intuitive one: it's not just that a marginal regression might be missed for lack of statistical power, it's that if n is too small relative to how close the SLO target sits to 100%, a PERFECTLY healthy service may never be able to clear the lower-bound threshold at all (at n=500 above, zero failures still only reaches a 0.992 lower bound against a 0.995 cutoff, an unconditional failure). Choose n large enough that a genuinely healthy service's best-case lower bound has real headroom above the cutoff, not just enough requests to sound statistically serious.
- Averaging percentiles ACROSS batches (rather than recomputing from pooled raw samples) is a common statistical error that silently distorts the reported p95; always keep raw per-request latency samples and recompute percentiles directly from them.
- A gate that only reports "FAILED" without the observed rate, confidence bound, and per-request detail turns every failure into a re-investigation from scratch; the extra reporting effort pays for itself the first time someone has to debug a 2am CI failure.
Design a soak and stress test strategy for microservices communicating through an asynchronous message bus. Include how you'd generate realistic load patterns with ramp-up and ramp-down phases, a warmup requirement, resource monitoring, how you'd observe consumer lag, validate correctness under sustained load, and a teardown procedure that avoids resource leaks and gives reproducible results.
Sample Answer
Direct answer
Ramp load up and down gradually rather than stepping directly to peak, include an explicit warmup phase before measuring anything (so cold caches and JIT/connection-pool warmup don't contaminate results), continuously observe consumer lag as the primary signal of whether the message-bus pipeline is keeping up, validate correctness (not just throughput) at multiple points during the sustained-load window rather than only at the end, and make teardown an explicit, verified step that confirms no resources (consumer groups, topics, worker processes) were left behind.
Structured elaboration
- Load-pattern shape. A soak/stress test for a message-bus-connected system should ramp UP gradually (over minutes, not instantly) to let auto-scaling, connection pools, and caches reach a steady state realistically, sustain a target load for the actual soak duration (long enough to surface issues that only appear over time: slow memory leaks, gradually growing consumer lag, log/disk growth), then ramp DOWN gradually and confirm the system returns to a quiescent, healthy baseline rather than leaving stuck consumers or unprocessed backlog.
- Warmup as a distinct, excluded phase. Explicitly exclude the warmup window from your measured metrics (do not let a cold-start latency spike count against your soak-test's pass/fail criteria); measure only the steady-state window once warmup has completed.
- Resource monitoring. Track CPU, memory, and disk/IO on both the producer and consumer sides throughout the run, watching specifically for gradual upward trends (a slow memory leak, growing thread counts) that a short test would never surface but a multi-hour soak test is specifically designed to catch.
- Consumer lag as the central distributed-systems-specific signal. For a message-bus-connected system, consumer lag (how far behind the latest produced offset the consumer group currently is) is the most direct signal of whether the pipeline is keeping up with sustained load; alert on and assert against lag staying within a bounded range throughout the sustained-load window, not just at the very end.
- Validating correctness under sustained load, not just at the end. Sample and verify correctness (using the same deterministic-entity technique as a shorter load-correctness test) at multiple points DURING the soak, not only after teardown, since a correctness regression that only appears after hours of sustained load (a slow resource leak that eventually causes dropped messages, for example) would otherwise only be caught at the very end, well after it started, and possibly not caught at all if the final check happens to land in a moment of transient recovery.
- Teardown avoiding resource leaks. Explicitly verify, as part of teardown, that consumer groups are properly deregistered, worker processes are terminated, and any test-specific topics are deleted; a soak test that leaves orphaned consumer groups behind can itself become a source of the exact resource-leak problems it exists to detect in the system under test, in the test infrastructure instead.
Trade-offs and pitfalls
- Skipping the warmup exclusion is a common mistake that makes soak-test results noisy and hard to compare run-over-run, since early cold-start effects get mixed into what should be steady-state measurements.
- A soak test that only checks correctness at the very end can miss a regression that appeared and then transiently self-corrected mid-run (a garbage-collection pause causing temporary lag that later recovers, for instance); sampling correctness throughout the run, not just at teardown, closes this gap.
- Failing to verify teardown completeness (leaked consumer groups, orphaned topics) accumulates cruft across repeated nightly soak-test runs, eventually degrading the shared test environment's own reliability in a way that looks like, but isn't, a regression in the system under test.
Explain how different service-discovery mechanisms (DNS, a central registry, Kubernetes Service objects, and a service mesh) affect how you write tests for microservices. Describe how you would simulate or stub service-discovery behavior in local and CI environments, and what pitfalls, such as DNS caching or sidecar timing, can break tests.
Sample Answer
Direct answer
Stub service discovery at the same layer your application actually calls it (a DNS resolver override for DNS-based discovery, a fake registry client for a central-registry pattern, a local sidecar/proxy config for a mesh), matching the real mechanism as closely as possible, and specifically test the two most common pitfalls: DNS answers being cached longer than intended, and a sidecar or proxy not yet being ready when the application starts trying to route through it.
Structured elaboration
- DNS-based discovery. In local and CI environments, either run a lightweight local DNS server pre-loaded with the test's intended name-to-address mappings, or override resolution at the HTTP client level (many HTTP client libraries support a custom resolver or a hosts-file-style override) so the application code path exercises real DNS-style resolution logic, not a hardcoded URL. Test specifically for stale-cache behavior: change what a name resolves to mid-test and assert the application picks up the change within its configured TTL, not indefinitely cached from the first lookup.
- Central registry (e.g., a service-registry pattern). Run a real (lightweight, disposable) instance of the registry locally/in CI and register test service instances against it directly, rather than mocking the registry client, so the application's actual registry-client code (including its retry/caching behavior) is exercised.
- Kubernetes Service objects. In CI, a real (if minimal) cluster or a kind/minikube-style local cluster lets you exercise real
Service-based DNS resolution; where a full cluster is too heavy for a given test tier, fall back to the DNS-override approach above, accepting the fidelity trade-off. - Service mesh. Meshes typically route via a local sidecar proxy; test with the real sidecar running locally against a test configuration, specifically checking startup-ordering: does the application's outbound call fail or hang if it starts sending traffic before its own sidecar has finished initializing and loading its routing configuration.
- The two named pitfalls. DNS caching: many HTTP clients and even the OS resolver cache lookups more aggressively or for longer than the DNS record's own TTL suggests; test for this explicitly by changing an address mid-test and asserting the expected propagation delay, not assuming it. Sidecar timing: test the application's behavior specifically during the window right after its own process starts but before its sidecar is confirmed ready, since this window is exactly when service-discovery-related requests are most likely to fail in a way that's invisible once both are fully up and stable.
Trade-offs and pitfalls
- The closer your test's discovery mechanism matches the real one (a real lightweight registry instance, a real sidecar), the more genuine confidence you get, at the cost of test speed and setup complexity; DNS-level or client-level overrides are faster and lighter but can miss discovery-specific bugs that only manifest against the real mechanism.
- DNS-caching bugs are notoriously environment-dependent (they can behave differently between a developer's laptop, CI, and production depending on which resolver and which HTTP client library is in play); a passing test locally does not guarantee the same caching behavior in every environment, so treat this class of test as a signal, not a guarantee.
- Sidecar-readiness bugs are often timing-sensitive and can be hard to reproduce reliably in a fast CI environment where everything starts quickly; consider an explicit, deliberately-slowed-down startup-ordering test rather than relying on a naturally-occurring race to reproduce it.
You have an API that starts a background job and returns a job_id; clients poll /job/{id} for status. Design tests to validate behavior under timing variations: the job finishes quickly, the job takes a very long time, the job fails mid-work, a client sends duplicate start requests, and there is visibility lag before the status reflects reality. How would you simulate each condition and assert correctness?
Sample Answer
Direct answer
Test each timing variation as its own explicit scenario rather than hoping one "happy path" test exercises them all: fast completion, slow completion, mid-work failure, duplicate start requests, and read-after-write lag on the status endpoint each need their own fault-injected or time-controlled test, because each exercises a different code path in the job-lifecycle state machine.
Structured elaboration
An async-job-with-polling API is really a small state machine (SUBMITTED -> RUNNING -> SUCCEEDED / FAILED), and the interesting bugs live at the TRANSITIONS, not the steady states. Structure the test suite around the transitions:
- Fast completion. Submit, then immediately poll. The test must handle the (correct) possibility that the job already finished before the first poll returns; a test that assumes there is always at least one
RUNNINGobservation is itself buggy. - Slow / long-running job. Use a controllable delay in the job's own logic (a test double, not a real long sleep) so the test can poll multiple times and assert the status stays
RUNNINGwith a stablejob_id, then eventually transitions. - Mid-work failure. Inject a failure partway through the job's work and assert the status becomes
FAILEDwith an actionable error, and that (if applicable) any partial side effects are rolled back or marked so a client pollingFAILEDdoesn't misread partial output as complete. - Duplicate start requests. Submit the same logical job twice (same idempotency key, if the API supports one) and assert you get back the SAME
job_idrather than two independent jobs silently doing the work twice; if the API has no idempotency key, assert (and document) that duplicate submissions are the caller's responsibility, which is itself a finding worth surfacing rather than assuming. - Visibility lag. After the job actually finishes internally, assert that a poll immediately after may still show
RUNNINGfor the read-after-write consistency window, and that a bounded poll (with backoff) eventually showsSUCCEEDED; do not assume the status flips synchronously with completion just because it usually does in a fast test environment.
Worked example
import time
def poll_until_terminal(client, job_id, timeout_s=5.0, interval_s=0.05):
deadline = time.monotonic() + timeout_s
last = None
while time.monotonic() < deadline:
last = client.get_status(job_id)
if last in ("SUCCEEDED", "FAILED"):
return last
time.sleep(interval_s)
raise AssertionError(f"job {job_id} still {last!r} after {timeout_s}s")
def test_mid_work_failure_reports_failed_not_stuck():
client = JobClient(work_fn=raises_after_partial_progress)
job_id = client.start()
status = poll_until_terminal(client, job_id)
assert status == "FAILED"
detail = client.get_error_detail(job_id)
assert detail is not None, "a FAILED job must carry an actionable error detail, not a bare failure"
marked_as_partial = "partial" in detail.lower()
rolled_back = client.get_side_effects(job_id) == []
assert marked_as_partial or rolled_back, (
"a job that fails mid-work must either roll back its partial side effects, or explicitly "
"mark the error detail as partial, so a caller never mistakes partial output for a complete result"
)
def test_duplicate_start_returns_same_job_id():
client = JobClient(work_fn=lambda: None)
idempotency_key = "submit-42"
job_id_1 = client.start(idempotency_key=idempotency_key)
job_id_2 = client.start(idempotency_key=idempotency_key)
assert job_id_1 == job_id_2, "duplicate submission under the same key must not spawn a second job"
Trade-offs and pitfalls
- Do not use real
time.sleep(N)to simulate a "slow job"; inject a controllable delay or hook into the job's own execution so the test suite doesn't become slow AND flaky at the same time. - The duplicate-start test is only meaningful if the API design actually intends idempotent submission; if it doesn't, this test should assert the documented behavior (two independent jobs), not an aspirational one, or you will be "fixing" a correctly-behaving system.
- Visibility-lag testing is easy to skip because it rarely reproduces on a fast, quiet test environment; if the read path and write path for job status are backed by different stores or caches, add an explicit test that reads immediately after a status change and tolerates (and asserts you tolerate) a brief stale read, rather than assuming synchronous consistency because it happened to hold locally.
- The mid-work-failure assertion above is deliberately written as an explicit either/or (
marked_as_partial or rolled_back) rather than a compound boolean expression: an earlier, more compact phrasing of this same check used operator precedence in a way that made it pass regardless of whether rollback actually happened, as long as the error message text didn't happen to contain the word "partial" -- a silent, vacuous pass. Prefer named intermediate booleans over compactand/orchains in any assertion that encodes more than one condition.
Unlock Full Question Bank
Get access to all 45 Distributed Systems and Microservices Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.