Flaky Test Management and Test Reliability Questions
Detecting, isolating, and eliminating non-deterministic tests. Covers root-causing flakiness, quarantine and remediation systems, distinguishing product bugs from test bugs, and maintaining suite health over time. Emphasizes keeping automated suites trustworthy so failures mean something.
You inherit a test suite of 2,000 automated tests with a 12% flakiness rate and high maintenance cost. Prepare an audit plan: how will you triage tests (metrics to collect), classify failure causes, prioritize which tests to fix, rewrite, or delete, and propose a remediation roadmap with expected outcomes over a six-month horizon.
Sample Answer
Direct answer: Triage by computing a per-test flakiness score AND a root-cause category from historical data first (cheap, mechanical, and comprehensive across all 2,000 tests), THEN apply the impact-weighted prioritization (which tests actually block releases or waste the most developer time) to decide fix order, and be honest that at a 12% aggregate rate, a meaningful fraction of the 2,000 tests likely need to be DELETED, not fixed, since not every test is worth the ongoing maintenance investment.
Structured elaboration
Metrics to collect for triage: per-test flakiness rate (with a minimum-run-count floor before trusting the number, per the statistical sub-area); a root-cause CATEGORY classification (from the standard taxonomy this topic uses throughout) inferred from failure-signature analysis where possible; and an IMPACT signal (how often this test's failures actually block a merge or a release, distinct from its raw failure rate, since a rarely-run, low-traffic test failing often is a smaller problem than a frequently-run, release-gating test failing occasionally).
Classifying failure causes: apply the failure-clustering/fingerprinting approach (from the correlation sub-area) across all 2,000 tests' historical failures to group them into the standard root-cause taxonomy at scale, mechanically, rather than manually investigating each of potentially hundreds of currently-flaky tests one at a time, which wouldn't be tractable for a suite this size within a reasonable timeframe.
Prioritizing fix vs rewrite vs delete:
- FIX: tests with high impact and a well-understood, addressable root cause (a known pattern like a fixed sleep, a shared-fixture issue) where the fix effort is modest relative to the test's ongoing value.
- REWRITE: tests with high impact but a root cause suggesting the test's fundamental DESIGN is the problem (deep architectural coupling to shared state, for example), where a targeted fix isn't sufficient and the test needs to be rebuilt against the isolation patterns covered throughout this topic.
- DELETE: tests with low impact (rarely block anything, redundant with other coverage) and/or a poor genuine-catch ratio (per the disable/quarantine-decision sub-area's retrospective analysis), where the ongoing maintenance cost isn't justified by the marginal coverage value, a real, honest category that a 12%-flaky, 2,000-test inherited suite likely needs applied to a meaningful fraction of tests, not just the smallest handful.
A six-month remediation roadmap with expected outcomes:
- Month 1: build/deploy the telemetry and mechanical classification pipeline across all 2,000 tests (the triage foundation); expected outcome: a complete, categorized inventory replacing the vague "12% flaky" headline number with a concrete breakdown by root cause and impact.
- Months 2-3: apply the impact-weighted prioritization and execute the highest-value FIX and DELETE decisions first (fixes because they're addressable and high-value; deletions because they're immediate, low-risk wins that reduce the ongoing burden with no further investigation needed); expected outcome: the aggregate flakiness rate drops meaningfully as the "cheap wins" (both fixes and deletions) land.
- Months 4-5: tackle the REWRITE category, the highest-effort tier, now that the easier wins are already banked and the team has more bandwidth and momentum; expected outcome: the remaining, harder-to-fix high-impact tests get addressed, not just deferred indefinitely.
- Month 6: stand up the ONGOING program (quarantine SLA, ownership, standing dashboard, per the organizational-process sub-area) so the suite doesn't simply DRIFT back to a similar state after this one-time cleanup effort concludes; expected outcome: a sustainable, monitored baseline rather than a one-time fix that erodes again over the following year.
Worked example: applying this process to the 2,000-test suite, mechanical classification reveals roughly 240 tests (12%) are currently flaky, of which retrospective impact analysis shows about 40 account for the large majority of actual release-blocking pain (the familiar concentration pattern seen throughout this topic's prioritization discussions), those 40 become the month 2-3 FIX focus; a further roughly 60 tests are found to have near-zero genuine-catch value and low impact, becoming month 2-3 DELETE candidates; the remaining flaky tests, lower-impact or requiring deeper rework, spread across the REWRITE tier (months 4-5) and the standing quarantine process (month 6 onward) rather than blocking the whole roadmap on addressing every single one within the six months.
Trade-offs & pitfalls: a plan that treats DELETE as an uncomfortable last resort rather than a legitimate, first-class outcome tends to under-use it, spending disproportionate fix effort on tests that genuinely aren't earning their keep; be willing to delete a meaningful fraction of a 2,000-test suite with a 12% flakiness rate and high maintenance cost, since that combination (age, size, and inherited status) is exactly the profile where a real fraction of accumulated tests have stopped being worth their ongoing cost, and treating deletion as taboo rather than a legitimate outcome of honest triage wastes remediation effort that could go toward higher-value fixes instead.
Explain key behavioral differences between mobile emulators/simulators and real devices that cause automated mobile tests to pass on emulators but fail on physical devices. List at least six differences (e.g., sensors, GPU, manufacturer OS customizations, WebView versions, network variability, hardware performance) and say how to mitigate them in a test strategy.
Sample Answer
Direct answer: Emulators approximate hardware in software and are internally consistent by construction, so tests pass reliably against that consistent approximation; real devices carry genuine hardware variability, real sensors, and manufacturer-specific OS customizations the emulator never models, so passing on an emulator establishes far less confidence than it appears to.
Structured elaboration
Six concrete differences, each with a mitigation:
- Sensors: emulators typically provide synthetic, perfectly-behaved sensor data (GPS, accelerometer) on demand; real devices have genuine sensor noise, latency, and permission-prompt timing that can affect app behavior. Mitigation: include a real-device test tier specifically for sensor-dependent flows, don't rely on emulator sensor simulation as sufficient coverage.
- GPU: emulators often use software rendering or a different GPU abstraction than the real device's actual GPU driver; rendering-timing-sensitive UI tests can pass reliably on emulator's consistent (if slower or different) rendering path and then flake on real hardware's actual GPU timing characteristics. Mitigation: for GPU-timing-sensitive assertions, prefer real-device testing or add generous, explicit tolerance for rendering completion rather than assuming emulator timing transfers.
- Manufacturer OS customizations: many Android manufacturers ship customized OS layers (different power-management/background-process-killing behavior, custom permission dialogs) that a stock emulator image doesn't replicate. Mitigation: maintain a real-device test matrix covering the manufacturer/OS-version combinations your actual user base concentrates in, informed by real usage analytics, not an arbitrary sample.
- WebView versions: an emulator's bundled WebView version can lag or differ from what's actually deployed on real devices in the field (WebView updates independently of the OS on many Android versions). Mitigation: explicitly pin and verify the WebView version in both emulator and real-device test environments, and treat a version mismatch as a known coverage gap rather than an unknown one.
- Network variability: emulators typically run on a stable, fast host-machine network path; real devices experience genuine cellular/WiFi variability (latency spikes, brief disconnects) that can expose real timeout and retry-handling bugs. Mitigation: use network-condition simulation tools (throttling, packet loss injection) in BOTH emulator and real-device testing rather than assuming the emulator's clean network is representative.
- Hardware performance: emulators often run on a powerful host machine and can be FASTER (or, under host contention, unpredictably slower) than the actual range of real devices in the field, especially lower-end devices. Mitigation: include a deliberately lower-spec real device (or a resource-throttled configuration) in the test matrix specifically to catch performance-dependent flakiness that a fast host-machine emulator would never surface.
How to mitigate them as a test strategy overall, not just per-difference: use emulators for the BULK of automated testing (fast, cheap, parallelizable, deterministic) as the first line of defense, and reserve a smaller, targeted real-device test tier (run less frequently, perhaps nightly rather than per-PR) specifically for the categories above where emulator behavior is known to diverge from reality. This tiered approach captures most of automation's speed and cost benefits from emulators while still catching the real-device-specific failure classes that would otherwise ship undetected.
Worked example: an app's push-notification handling passes reliably on emulator (which delivers a synthetic notification event instantly and consistently) but fails intermittently on real mid-range Android devices from a specific manufacturer, whose custom OS layer aggressively kills background processes to save battery, occasionally killing the app before the notification handler runs. This is invisible to emulator testing entirely (no such power-management behavior exists there) and was only caught by the manufacturer-specific real-device test tier, confirming why category 3 (manufacturer customizations) needs deliberate real-device coverage informed by actual field device distribution, not assumed away.
Trade-offs & pitfalls: a real-device test matrix is expensive to maintain and slower to run than emulator tests, so the temptation is to skip it or run it rarely; the mitigation is choosing WHICH real devices to include deliberately (informed by actual field usage data on device/manufacturer distribution) rather than either an arbitrary small sample or attempting comprehensive device coverage that isn't cost-effective.
In JavaScript using Playwright, design a helper function findStableLocator(page, selectors, timeoutMs=5000) that tries selectors in order and returns a Locator only if the element is present and stable (no layout change) for a short duration (e.g., 200ms). Describe the algorithm and boundary conditions; you do not need to supply full runnable code but show key pseudo-code steps and Playwright primitives you would use.
Sample Answer
Direct answer: Try each selector in the given priority order, but for each candidate that resolves to a visible element, additionally confirm its bounding box is UNCHANGED across a short polling window before accepting it, since "present" and "stable" are two different conditions and a layout still actively shifting is exactly what causes a subsequent click or read to hit the wrong location.
Algorithm
function findStableLocator(page, selectors, timeoutMs=5000):
deadline = now + timeoutMs
for selector in selectors: # try in priority order
locator = page.locator(selector)
while now < deadline:
if locator.count == 0:
break # this selector: not present, try next
if not locator.first.isVisible:
sleep(pollIntervalMs)
continue
boxAtStart = locator.first.boundingBox
sleep(stabilityWindowMs) # e.g. 200ms
boxAtEnd = locator.first.boundingBox
if boxAtStart is not None and boxAtStart == boxAtEnd:
return locator # present, visible, AND stable
# box changed (or vanished) during the window: layout still shifting, retry
# this selector never stabilized within the remaining time budget; fall through
raise TimeoutError(f"no selector in {selectors} became stable within {timeoutMs}ms")
Key Playwright primitives used: page.locator(selector) (lazy, re-queried each check, unlike a one-time querySelector), locator.count to check presence without throwing, locator.first.isVisible for the visibility gate, and locator.first.boundingBox twice, separated by a short sleep, as the STABILITY check specifically; Playwright's own built-in actionability checks (auto-waiting before click, for example) partially overlap with this but don't expose a reusable "wait for layout stability" primitive directly, which is why this helper constructs it explicitly from two bounding-box reads.
Boundary conditions:
- Multiple selectors, first one present but never stabilizing: the algorithm should not get stuck forever on selector 1 if it's present-but-perpetually-shifting; the outer
while now < deadlineloop is scoped per-selector inside the shared overall deadline, so time spent failing to stabilize on selector 1 eats into the SAME budget available for trying selector 2, meaning a pathological "present but never stable" first selector could exhaust the whole timeout before trying any fallback; a more robust version would allocate a PER-SELECTOR sub-budget (dividingtimeoutMsacross the candidate list) rather than one shared clock, worth calling out explicitly as a real design choice with a real trade-off. - Element present, visible, but bounding box is
None(can happen for certain detached-but-matched states): treated as "not yet stable," continues polling rather than crashing on a null comparison. - Zero selectors provided: should raise a clear
ValueErrorupfront rather than silently returningNoneor looping until timeout with nothing to check. stabilityWindowMsversuspollIntervalMs: these are deliberately separate parameters conceptually (this pseudocode collapses them for brevity), a longer stability window gives more confidence layout has truly settled but adds latency to every successful resolution, even the common case where the element was already stable; 200ms is a reasonable default balancing the two.
Trade-offs & pitfalls: the shared-deadline-across-selectors boundary condition above is a genuine, easy-to-miss bug class, a helper meant to provide FALLBACK resilience across several selectors can accidentally consume its entire timeout budget on the FIRST selector if that selector is present but pathologically unstable (for example, a continuously-animating loading skeleton that matches the first selector), never actually trying the more robust fallback selectors at all. Explicitly deciding (and documenting) whether the timeout is shared or per-selector is a design decision worth making deliberately, not accidentally, since the two choices produce meaningfully different behavior on exactly the adversarial case this helper exists to handle.
CI shows intermittent failures that pass locally. You suspect flaky tests due to timeouts and async operations. Describe a plan to identify sources of flakiness: how to reproduce locally (increased logging, deterministic time control), what tests to run repeatedly, and how to change tests to be deterministic or tolerant (mocking time, explicit synchronization). Explain trade-offs between flakiness fixes and test coverage.
Sample Answer
Direct answer: Treat "passes locally, fails intermittently in CI" as a diagnostic question with three candidate explanations (test-code timing, application-level async behavior, or environment difference), and design your reproduction to distinguish between them rather than guessing and patching the first plausible cause.
Structured elaboration
- Reproduce locally under CI-like conditions: running the test once locally rarely reproduces a CI-only intermittent failure, because CI runners are typically slower and noisier (shared CPU, network variance) than a developer machine. Instead: (a) run the SAME test many times in a loop locally (50 to 100 iterations) to establish a baseline local failure rate, even if it is near zero; (b) if it stays at zero locally, artificially degrade the local environment (CPU throttling, added network latency) to see if the failure reproduces under load; (c) increase logging verbosity around the async boundary you suspect, and run in CI itself with that extra logging to capture evidence from an actual failure rather than a hypothesized one.
- Deterministic time control: if the suspected timeout is timing-sensitive, replace real waits with a controllable clock (a time-mocking library, or an injected clock dependency) so you can deterministically simulate the "just barely too slow" window that CI's variance occasionally produces, rather than waiting for it to happen by chance.
- Which tests to run repeatedly: prioritize the specific failing test in isolation many times (to separate "this test is flaky" from "the suite has cross-test pollution"), and separately run the full suite in the SAME order and parallelism CI uses, since order- or parallelism-dependent flakiness will not show up running the test alone.
- Changing tests to be deterministic or tolerant: two different fixes address two different diagnoses. If the root cause is a genuine async race in test code (asserting before an operation completes), the fix is an explicit, condition-based wait instead of a fixed sleep. If the root cause is real environmental variance (CI is sometimes just slower), the fix is a more GENEROUS but still bounded timeout, paired with clear logging that distinguishes "genuinely stuck" from "was slow but finished."
- Deciding test vs application vs infra: if the same operation, called directly (bypassing the test framework) under a tight timing loop, occasionally takes longer than the test's timeout, the root cause is likely genuine async application latency, not a test bug; if the operation is always fast when isolated but the TEST still times out under full-suite load, the root cause is more likely resource contention from concurrent tests (CPU, DB connections) rather than the operation itself.
- Trade-offs between flakiness fixes and test coverage: the fastest way to make a flaky test "pass reliably" is to make its assertions weaker or its scope narrower, but that silently reduces the coverage it provides. A better trade-off is to keep the assertion strength and instead make the WAIT mechanism deterministic (explicit condition waiting) rather than trading away what the test actually verifies.
Worked example: a checkout test intermittently times out waiting for an order-confirmation API call in CI, never locally. Looping the isolated call 500 times locally under artificial CPU throttling reproduces occasional latency spikes above the test's 2-second timeout, roughly 1 in 150 iterations, confirming this is real async latency variance rather than a test-code bug. The fix: raise the wait to a bounded 5 seconds AND log the actual observed latency on every run (not just on timeout), so a future latency regression is visible as a trend in the logs rather than only as an intermittent CI failure.
Trade-offs & pitfalls: a common mistake is fixing the FIRST plausible cause found (usually "just increase the timeout") without confirming it against the reproduction evidence; this frequently just delays the next flake rather than resolving it. Another pitfall: adding verbose logging only in CI and not validating it locally first means you may ship logging that itself changes timing (e.g., synchronous log writes slowing the exact operation you're diagnosing), which can mask or shift the very race you are trying to observe.
Write a small utility (describe input/output and algorithm) that generates deterministic test data for property-based or randomized tests using a seed. Requirements: allow seeding per test run, produce reproducible sequences across environments and languages (explain constraints), and include at least one strategy to generate unique but predictable identifiers.
Sample Answer
Direct answer: Derive every generator's random stream from a single top-level seed plus a NAMESPACE string via a cryptographic hash (not a language's built-in, often per-process-randomized string hash), so multiple independent generators within the same test run don't accidentally desynchronize each other, and the whole scheme reproduces identically across environments and, with care about hash choice, across languages.
Approach and code
import hashlib
import random
def seeded_rng(test_run_seed, namespace):
"""Derive a per-namespace deterministic RNG from one top-level seed, so
independent generators in the same run don't share (and desynchronize)
a single global stream, while staying fully reproducible."""
digest = hashlib.sha256(f"{test_run_seed}:{namespace}".encode()).hexdigest()
derived_seed = int(digest[:16], 16)
return random.Random(derived_seed)
def generate_unique_predictable_id(test_run_seed, namespace, index):
"""Unique within a run (via index), but fully reproducible given the
same seed -- unlike a real UUID4, which is unique but never reproducible."""
digest = hashlib.sha256(f"{test_run_seed}:{namespace}:{index}".encode()).hexdigest()
return f"test-{namespace}-{digest[:12]}"
def generate_test_records(test_run_seed, namespace, count, value_range=(1, 1000)):
rng = seeded_rng(test_run_seed, namespace)
return [
{"id": generate_unique_predictable_id(test_run_seed, namespace, i),
"value": rng.randint(*value_range)}
for i in range(count)
]
Reproducibility across environments and languages, and its real constraints: hashlib.sha256 is a STANDARDIZED algorithm with an identical, specified output for identical input across every language and platform, so the derived seed computation itself is portable. The constraint is what happens AFTER deriving the seed: random.Random's specific pseudo-random ALGORITHM (Mersenne Twister in CPython) is implementation-specific, so the same derived integer seed fed into Python's random.Random and, say, a Java java.util.Random will NOT produce the same sequence of values, since the two languages' PRNG algorithms differ. True cross-LANGUAGE reproducibility (not just cross-environment reproducibility within one language) requires either standardizing on a specific, publicly-specified PRNG algorithm implemented identically in every language involved (a real, nontrivial engineering commitment), or scoping the reproducibility guarantee explicitly to "same seed, same language" and treating cross-language reproduction as a separate, harder problem to solve only if genuinely needed.
Unique but predictable identifiers: generate_unique_predictable_id combines the seed, namespace, AND an index into the hash input, so every index produces a distinct, deterministic 'ID, while remaining fully reproducible: rerun with the same seed and the same index, and you get the identical ID, unlike a genuinely random UUID4 which is unique but can never be reproduced on a later run.
Verification (executed this session, python3): the code above originally had a real bug: .encode).hexdigest (twice) called neither method, str.encode and hashlib digest's .hexdigest were referenced as bound-method OBJECTS rather than actually invoked (missing ()), which raises TypeError: object supporting the buffer API required on the very first call, before any record can be generated at all. Fixed by adding the missing () to both call sites above. With the fix applied, I ran the four claimed adversarial cases for real: (1) the SAME seed and namespace called twice independently, confirming byte-identical output both times; (2) a DIFFERENT seed producing different output, confirming the seed genuinely drives the generation rather than being silently ignored; (3) the SAME top-level seed but a DIFFERENT namespace ("users" vs "orders"), confirming the two namespaces produce independent, non-colliding streams (zero ID overlap, different value sequences), verifying the namespace-isolation design goal specifically; (4) generating 200 IDs in one call and confirming all 200 are unique. All four passed on the corrected code:
PASS: identical seed produces byte-identical records across two independent calls
PASS: different seed produces different records
PASS: 'orders' and 'users' namespaces under the SAME top-level seed produce
independent, non-colliding streams (0 id overlap, different value sequences)
PASS: 200/200 generated IDs are unique within a single run
Trade-offs & pitfalls: relying on a language's DEFAULT string-hashing function (Python's built-in hash, for instance) instead of an explicit hashlib call is a real, easy mistake, hash for strings is deliberately randomized per-process by default in modern Python (a security hardening measure against hash-flooding attacks), so using it would silently break cross-run reproducibility despite the code otherwise looking correct; this is exactly the kind of bug that would pass a quick local smoke test (within one process, hash is CONSISTENT) and only reveal itself as broken reproducibility across separate CI runs (different processes), a genuinely subtle, easy-to-ship defect worth calling out explicitly rather than assuming any hash function works.
Unlock Full Question Bank
Get access to all Flaky Test Management and Test Reliability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.