Flaky Test Management and Test Reliability Questions
Detecting, isolating, and eliminating non-deterministic tests. Covers root-causing flakiness, quarantine and remediation systems, distinguishing product bugs from test bugs, and maintaining suite health over time. Emphasizes keeping automated suites trustworthy so failures mean something.
Design a system to detect groups of tests whose failures strongly correlate (for example because they share a flaky fixture or rely on the same external resource). Describe what data to collect, statistical or ML techniques to group correlated failures (for example: co-failure matrix, pointwise mutual information, clustering), and how to present actionable hypotheses to engineers for confirmation.
Sample Answer
Direct answer: Represent test failures as a co-occurrence matrix across CI runs, and use that structure both to statistically detect correlated groups and to generate a human-readable hypothesis (a shared fixture, service, or infra dependency) for engineers to confirm, since raw correlation alone doesn't explain WHY tests are grouped.
Structured elaboration
- Data to collect: per-test-run outcome (pass/fail) tagged with run ID, timestamp, CI node/worker, and any known shared-resource identifiers the test declares (which fixture, which external service it depends on, which database it touches), if that metadata is available. Even without explicit resource declarations, the outcome matrix alone (tests as rows, runs as columns, pass/fail as values) is enough to start.
- Co-failure matrix: for each pair of tests, compute how often they fail in the SAME run relative to how often each fails independently. A simple, defensible starting metric is the Jaccard-like co-failure rate: (runs where both failed) divided by (runs where at least one failed). High-correlation pairs are candidates for sharing a root cause.
- Pointwise mutual information (PMI): PMI corrects a weakness of raw co-failure rate, two tests that are BOTH individually very flaky can co-fail often just by chance, even with no shared cause. PMI compares the observed joint failure rate to what independence would predict, so two tests that are individually rare failures but almost always fail TOGETHER get flagged strongly, while two tests that are both just generally noisy do not get over-flagged.
- Clustering: once pairwise correlation (or PMI) scores exist for all test pairs, hierarchical or graph-based clustering (treating high-correlation pairs as edges) groups tests into candidate clusters that likely share a root cause, rather than requiring an engineer to inspect every pair manually.
- Conditional-independence / leader-test detection: a refinement worth naming explicitly: within a correlated cluster, check whether the correlation of a "follower" test to the rest of the cluster DISAPPEARS once you condition on a suspected "leader" test's outcome. If it does, the leader test is a strong candidate for being the actual root-cause signal (perhaps it directly exercises the shared broken resource) and the others are downstream symptoms, which helps engineers focus on the one test worth debugging deeply rather than all of them equally.
- Correlating with code and infra events: complement the pure failure-correlation signal with git metadata (did a specific commit's merge time align with when this cluster's correlation appeared), CI job history, and deployment windows; a cluster that correlates with a specific deploy timestamp is a much stronger, more actionable hypothesis ("this deploy broke a shared dependency") than correlation alone.
- Presenting actionable hypotheses: don't just output a list of correlated test IDs. Present the cluster together with its most likely SHARED CAUSE (inferred from any available resource metadata, or from the leader-test analysis, or from an aligned deploy/commit event), and let the engineer CONFIRM or REJECT the hypothesis rather than starting their investigation from a blank correlation number.
Worked example: Six tests across three unrelated feature areas start co-failing on the same set of CI runs, with a co-failure rate near 80% between all six pairs. Individually, none of the six has a high standalone failure rate, so this pattern would be missed by per-test flakiness scoring alone. PMI confirms the correlation is far above what chance would predict. Leader-test analysis shows that conditioning on one specific test (which directly calls a shared caching service) eliminates most of the residual correlation among the other five, and cross-referencing against deploy history shows the pattern started exactly at a caching-service redeploy timestamp. The presented hypothesis to engineers: "these 6 tests likely share a dependency on the caching service redeployed at [timestamp]; investigate that service first," rather than six separate flaky-test tickets.
Trade-offs & pitfalls: pure statistical correlation can produce clusters that are correlated by COINCIDENCE (especially with a modest number of runs and many tests, spurious pairwise correlations are expected by chance), so any correlation-detection system needs either a significance threshold that accounts for the number of pairs tested (a multiple-comparisons correction) or a minimum-run-count floor before a cluster is surfaced; presenting a low-confidence spurious cluster to engineers as an "actionable hypothesis" wastes their trust exactly the way an over-aggressive quarantine system does.
Write a small utility (describe input/output and algorithm) that generates deterministic test data for property-based or randomized tests using a seed. Requirements: allow seeding per test run, produce reproducible sequences across environments and languages (explain constraints), and include at least one strategy to generate unique but predictable identifiers.
Sample Answer
Direct answer: Derive every generator's random stream from a single top-level seed plus a NAMESPACE string via a cryptographic hash (not a language's built-in, often per-process-randomized string hash), so multiple independent generators within the same test run don't accidentally desynchronize each other, and the whole scheme reproduces identically across environments and, with care about hash choice, across languages.
Approach and code
import hashlib
import random
def seeded_rng(test_run_seed, namespace):
"""Derive a per-namespace deterministic RNG from one top-level seed, so
independent generators in the same run don't share (and desynchronize)
a single global stream, while staying fully reproducible."""
digest = hashlib.sha256(f"{test_run_seed}:{namespace}".encode()).hexdigest()
derived_seed = int(digest[:16], 16)
return random.Random(derived_seed)
def generate_unique_predictable_id(test_run_seed, namespace, index):
"""Unique within a run (via index), but fully reproducible given the
same seed -- unlike a real UUID4, which is unique but never reproducible."""
digest = hashlib.sha256(f"{test_run_seed}:{namespace}:{index}".encode()).hexdigest()
return f"test-{namespace}-{digest[:12]}"
def generate_test_records(test_run_seed, namespace, count, value_range=(1, 1000)):
rng = seeded_rng(test_run_seed, namespace)
return [
{"id": generate_unique_predictable_id(test_run_seed, namespace, i),
"value": rng.randint(*value_range)}
for i in range(count)
]
Reproducibility across environments and languages, and its real constraints: hashlib.sha256 is a STANDARDIZED algorithm with an identical, specified output for identical input across every language and platform, so the derived seed computation itself is portable. The constraint is what happens AFTER deriving the seed: random.Random's specific pseudo-random ALGORITHM (Mersenne Twister in CPython) is implementation-specific, so the same derived integer seed fed into Python's random.Random and, say, a Java java.util.Random will NOT produce the same sequence of values, since the two languages' PRNG algorithms differ. True cross-LANGUAGE reproducibility (not just cross-environment reproducibility within one language) requires either standardizing on a specific, publicly-specified PRNG algorithm implemented identically in every language involved (a real, nontrivial engineering commitment), or scoping the reproducibility guarantee explicitly to "same seed, same language" and treating cross-language reproduction as a separate, harder problem to solve only if genuinely needed.
Unique but predictable identifiers: generate_unique_predictable_id combines the seed, namespace, AND an index into the hash input, so every index produces a distinct, deterministic 'ID, while remaining fully reproducible: rerun with the same seed and the same index, and you get the identical ID, unlike a genuinely random UUID4 which is unique but can never be reproduced on a later run.
Verification (executed this session, python3): the code above originally had a real bug: .encode).hexdigest (twice) called neither method, str.encode and hashlib digest's .hexdigest were referenced as bound-method OBJECTS rather than actually invoked (missing ()), which raises TypeError: object supporting the buffer API required on the very first call, before any record can be generated at all. Fixed by adding the missing () to both call sites above. With the fix applied, I ran the four claimed adversarial cases for real: (1) the SAME seed and namespace called twice independently, confirming byte-identical output both times; (2) a DIFFERENT seed producing different output, confirming the seed genuinely drives the generation rather than being silently ignored; (3) the SAME top-level seed but a DIFFERENT namespace ("users" vs "orders"), confirming the two namespaces produce independent, non-colliding streams (zero ID overlap, different value sequences), verifying the namespace-isolation design goal specifically; (4) generating 200 IDs in one call and confirming all 200 are unique. All four passed on the corrected code:
PASS: identical seed produces byte-identical records across two independent calls
PASS: different seed produces different records
PASS: 'orders' and 'users' namespaces under the SAME top-level seed produce
independent, non-colliding streams (0 id overlap, different value sequences)
PASS: 200/200 generated IDs are unique within a single run
Trade-offs & pitfalls: relying on a language's DEFAULT string-hashing function (Python's built-in hash, for instance) instead of an explicit hashlib call is a real, easy mistake, hash for strings is deliberately randomized per-process by default in modern Python (a security hardening measure against hash-flooding attacks), so using it would silently break cross-run reproducibility despite the code otherwise looking correct; this is exactly the kind of bug that would pass a quick local smoke test (within one process, hash is CONSISTENT) and only reveal itself as broken reproducibility across separate CI runs (different processes), a genuinely subtle, easy-to-ship defect worth calling out explicitly rather than assuming any hash function works.
You encounter a Heisenbug that only appears under nightly CI load (high parallelism) and cannot be reproduced locally. Describe an end-to-end plan to collect artifacts and instrument the system to root-cause the bug: include low-overhead non-invasive traces to enable triage, more invasive instrumentation for reproductions, resource throttling and synthetic load generation, and strategies to create a minimal load-based reproduction for debugging.
Sample Answer
Direct answer: Since you cannot reproduce it locally, the plan has to work in layers of increasing invasiveness, always-on lightweight tracing that's cheap enough to run in every nightly build, targeted heavier instrumentation activated only when a failure is suspected, and finally a deliberately-constructed synthetic-load reproduction that lets you iterate locally instead of waiting for the next nightly occurrence.
Structured elaboration
- Low-overhead, non-invasive traces for baseline triage: enable lightweight, always-on tracing (request-correlation IDs, coarse-grained timing spans, sampled profiling) on every nightly run, cheap enough not to distort the very timing behavior you're trying to observe, but detailed enough that when a Heisenbug DOES occur, you have SOME evidence from that exact occurrence rather than nothing. This is the critical first layer, since without it, every occurrence is a wasted data point.
- More invasive instrumentation, activated conditionally: rather than running expensive instrumentation (fine-grained thread-state dumps, full memory snapshots) on every run, which itself risks changing timing enough to hide the bug (the classic Heisenbug trap, observing it changes it), activate it CONDITIONALLY: for example, on the first sign of trouble (a slow span, a retry), escalate to capturing a full diagnostic snapshot for that specific run only.
- Resource throttling and synthetic load generation to reproduce locally: since the bug only appears under nightly high-parallelism load, the most direct path to local reproducibility is recreating that load condition deliberately, artificially constrain local resources (CPU throttling, limited connection-pool size) and generate synthetic concurrent load (many parallel instances of the suspected test, or a load-generation tool hitting the same code path concurrently) rather than waiting to get lucky during an actual nightly run.
- Building a minimal load-based reproduction: once you can reproduce it locally under synthetic load, apply the same shrink-and-confirm discipline as the minimal-failing-example technique (remove unrelated load sources, narrow the concurrency pattern) until you have the SMALLEST synthetic load scenario that still reproduces the bug reliably enough to iterate on quickly, ideally in seconds or a few minutes per attempt, not requiring a full nightly run's worth of load every time.
- Sequencing across the plan: start passive (layer 1) to gather evidence from real occurrences without disturbing anything; escalate to layer 2 only once you have a working hypothesis worth actively instrumenting for; use layer 2's evidence to inform what synthetic load pattern (layer 3) to attempt; iterate on the minimal local reproduction once you have one, since it's now far cheaper to test a fix hypothesis locally than to wait for the next nightly occurrence.
Worked example: a service intermittently deadlocks only during the nightly suite's peak parallelism (roughly 200 concurrent test workers). Layer-1 tracing over several nights shows the deadlock consistently correlates with a specific connection-pool exhaustion pattern just before it occurs. Layer-2 instrumentation (activated specifically when pool utilization crosses 90%) captures a full thread-state dump on the next occurrence, confirming two specific code paths racing for the same pooled connection. Layer 3 recreates this locally by throttling the local connection pool to a small size (say, 5 connections) and generating 50 concurrent synthetic requests hitting the two racing code paths directly, reliably reproducing the deadlock in under 10 seconds locally, letting the team iterate on and verify a fix (a lock-ordering change) in minutes rather than waiting for tomorrow's nightly run.
Trade-offs & pitfalls: over-instrumenting from the start (jumping straight to layer 2 or 3 without the layer-1 evidence to target it) wastes effort on broad, low-signal instrumentation and risks the observer-effect problem, heavy tracing everywhere can itself change timing enough to make the bug stop reproducing, which is exactly backwards from what you want. The disciplined, staged approach above (cheap and broad first, expensive and targeted only once justified) is what avoids that trap.
Implement a retry-with-circuit-breaker wrapper for API tests in Python or JavaScript. The wrapper should: retry idempotent requests with exponential backoff up to N attempts, trip a circuit and stop further retries after M consecutive downstream failures, and expose metrics for retries and circuit state. Provide code and explain idempotency and metric reporting considerations.
Sample Answer
Direct answer: Layer two independent mechanisms, per-call exponential-backoff retry for ordinary transient failures, and a circuit breaker tracking CONSECUTIVE failures ACROSS calls that trips open and stops attempting entirely once the downstream service looks genuinely unhealthy, since retry alone doesn't protect against hammering an already-struggling downstream service, and a circuit breaker alone doesn't help with an isolated, one-off transient blip.
Approach and code
import random
import time
class CircuitOpenError(Exception):
pass
class RetryCircuitBreaker:
"""Retries idempotent requests with exponential backoff up to
max_retries; trips OPEN after trip_threshold CONSECUTIVE failures
(across calls, not just within one call's own retries), rejecting
further calls until a cooldown elapses."""
def __init__(self, max_retries=3, base_delay=0.01, trip_threshold=5,
cooldown_seconds=0.2, sleep_fn=time.sleep, clock_fn=time.monotonic):
self.max_retries = max_retries
self.base_delay = base_delay
self.trip_threshold = trip_threshold
self.cooldown_seconds = cooldown_seconds
self.sleep_fn = sleep_fn
self.clock_fn = clock_fn
self.consecutive_failures = 0
self.state = "closed"
self.opened_at = None
self.metrics = {"total_calls": 0, "total_retries": 0, "circuit_trips": 0, "rejected_calls": 0}
def _maybe_recover(self):
if self.state == "open" and (self.clock_fn() - self.opened_at) >= self.cooldown_seconds:
self.state = "half_open" # allow one trial call through
def call(self, fn, *args, is_idempotent=True, **kwargs):
if not is_idempotent:
raise ValueError("retry logic requires an idempotent operation")
self._maybe_recover()
if self.state == "open":
self.metrics["rejected_calls"] += 1
raise CircuitOpenError("circuit open; downstream considered unhealthy")
self.metrics["total_calls"] += 1
for attempt in range(self.max_retries + 1):
try:
result = fn(*args, **kwargs)
self.consecutive_failures = 0
self.state = "closed"
return result
except Exception as exc:
self.consecutive_failures += 1
if self.consecutive_failures >= self.trip_threshold:
self.state = "open"
self.opened_at = self.clock_fn()
self.metrics["circuit_trips"] += 1
raise CircuitOpenError(f"tripped after {self.consecutive_failures} failures") from exc
if attempt == self.max_retries:
raise
self.metrics["total_retries"] += 1
delay = self.base_delay * (2 ** attempt) + random.uniform(0, self.base_delay)
self.sleep_fn(delay)
Idempotency considerations: the wrapper explicitly REQUIRES is_idempotent=True (raising immediately otherwise), since retrying a NON-idempotent request (one that has a real, non-repeatable side effect, like charging a payment twice) risks a genuinely dangerous double-execution bug, far worse than the flakiness it's meant to fix; this makes the safety requirement an explicit, enforced PART of the API rather than a comment a caller could accidentally ignore.
Metric reporting considerations: total_calls vs total_retries separately lets you compute a real retry RATE (not just a count), circuit_trips tracks how often the downstream looked genuinely unhealthy (a strong, aggregate signal of downstream reliability worth its own alerting, separate from any single test's flakiness), and rejected_calls (calls refused outright while the circuit was open) is what distinguishes "we tried and failed" from "we didn't even try", an important distinction for correctly interpreting a burst of failures during an outage.
Verification (executed this session, python3): the code as originally drafted had two compounding bugs in the recovery path: self._maybe_recover was called without parentheses (so it never actually ran, a no-op statement), and self.opened_at = self.clock_fn was likewise missing parentheses (so it stored the CLOCK FUNCTION ITSELF, not a timestamp). Even after fixing the call site, calling _maybe_recover() on the unfixed body raised TypeError: unsupported operand type(s) for -: 'function' and 'function', because it tried to subtract two function objects. This means the claimed half-open cooldown recovery never worked: I reproduced this directly, tripping the circuit with an always-failing function then advancing a fake clock 2 seconds past a 1-second cooldown, the shipped code still raised CircuitOpenError on the next call instead of allowing a trial call through. After fixing both call sites (self._maybe_recover() and self.opened_at = self.clock_fn()), I reran all four adversarial cases for real: (1) a function failing twice then succeeding recovers via retry without tripping the circuit; (2) an always-failing function trips the circuit after trip_threshold consecutive failures and a subsequent call while open is rejected without invoking the downstream function; (3) using a fake, controllable clock, after the cooldown elapses the circuit allows a trial call through and closes on success; (4) a non-idempotent call is rejected immediately, before any attempt. All four now genuinely pass:
tripped at iteration 3: tripped after 4 failures, state=open
call after cooldown result: ok state: closed
non-idempotent rejected: retry logic requires an idempotent operation
Trade-offs & pitfalls: consecutive_failures resets to zero on ANY success, which is the correct behavior for detecting a genuinely unhealthy downstream (sustained failure), but means a downstream service ALTERNATING between failing and succeeding (a partial, flapping degradation) never trips the circuit at all, since it never accumulates trip_threshold CONSECUTIVE failures; a production version handling that pattern would need a complementary rate-based trigger (trip if failure rate over a recent window exceeds a threshold, regardless of whether failures were strictly consecutive) alongside the consecutive-failure trigger implemented here, worth naming explicitly as a real gap in this specific implementation rather than assuming consecutive-failure tracking alone covers every unhealthy-downstream pattern.
Explain key behavioral differences between mobile emulators/simulators and real devices that cause automated mobile tests to pass on emulators but fail on physical devices. List at least six differences (e.g., sensors, GPU, manufacturer OS customizations, WebView versions, network variability, hardware performance) and say how to mitigate them in a test strategy.
Sample Answer
Direct answer: Emulators approximate hardware in software and are internally consistent by construction, so tests pass reliably against that consistent approximation; real devices carry genuine hardware variability, real sensors, and manufacturer-specific OS customizations the emulator never models, so passing on an emulator establishes far less confidence than it appears to.
Structured elaboration
Six concrete differences, each with a mitigation:
- Sensors: emulators typically provide synthetic, perfectly-behaved sensor data (GPS, accelerometer) on demand; real devices have genuine sensor noise, latency, and permission-prompt timing that can affect app behavior. Mitigation: include a real-device test tier specifically for sensor-dependent flows, don't rely on emulator sensor simulation as sufficient coverage.
- GPU: emulators often use software rendering or a different GPU abstraction than the real device's actual GPU driver; rendering-timing-sensitive UI tests can pass reliably on emulator's consistent (if slower or different) rendering path and then flake on real hardware's actual GPU timing characteristics. Mitigation: for GPU-timing-sensitive assertions, prefer real-device testing or add generous, explicit tolerance for rendering completion rather than assuming emulator timing transfers.
- Manufacturer OS customizations: many Android manufacturers ship customized OS layers (different power-management/background-process-killing behavior, custom permission dialogs) that a stock emulator image doesn't replicate. Mitigation: maintain a real-device test matrix covering the manufacturer/OS-version combinations your actual user base concentrates in, informed by real usage analytics, not an arbitrary sample.
- WebView versions: an emulator's bundled WebView version can lag or differ from what's actually deployed on real devices in the field (WebView updates independently of the OS on many Android versions). Mitigation: explicitly pin and verify the WebView version in both emulator and real-device test environments, and treat a version mismatch as a known coverage gap rather than an unknown one.
- Network variability: emulators typically run on a stable, fast host-machine network path; real devices experience genuine cellular/WiFi variability (latency spikes, brief disconnects) that can expose real timeout and retry-handling bugs. Mitigation: use network-condition simulation tools (throttling, packet loss injection) in BOTH emulator and real-device testing rather than assuming the emulator's clean network is representative.
- Hardware performance: emulators often run on a powerful host machine and can be FASTER (or, under host contention, unpredictably slower) than the actual range of real devices in the field, especially lower-end devices. Mitigation: include a deliberately lower-spec real device (or a resource-throttled configuration) in the test matrix specifically to catch performance-dependent flakiness that a fast host-machine emulator would never surface.
How to mitigate them as a test strategy overall, not just per-difference: use emulators for the BULK of automated testing (fast, cheap, parallelizable, deterministic) as the first line of defense, and reserve a smaller, targeted real-device test tier (run less frequently, perhaps nightly rather than per-PR) specifically for the categories above where emulator behavior is known to diverge from reality. This tiered approach captures most of automation's speed and cost benefits from emulators while still catching the real-device-specific failure classes that would otherwise ship undetected.
Worked example: an app's push-notification handling passes reliably on emulator (which delivers a synthetic notification event instantly and consistently) but fails intermittently on real mid-range Android devices from a specific manufacturer, whose custom OS layer aggressively kills background processes to save battery, occasionally killing the app before the notification handler runs. This is invisible to emulator testing entirely (no such power-management behavior exists there) and was only caught by the manufacturer-specific real-device test tier, confirming why category 3 (manufacturer customizations) needs deliberate real-device coverage informed by actual field device distribution, not assumed away.
Trade-offs & pitfalls: a real-device test matrix is expensive to maintain and slower to run than emulator tests, so the temptation is to skip it or run it rarely; the mitigation is choosing WHICH real devices to include deliberately (informed by actual field usage data on device/manufacturer distribution) rather than either an arbitrary small sample or attempting comprehensive device coverage that isn't cost-effective.
Unlock Full Question Bank
Get access to all Flaky Test Management and Test Reliability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.