Programming for Test Automation Questions
General Java and Python language proficiency questions where the test-engineering framing is load-bearing: it changes what is actually being assessed, not just the flavor text. Covers designing for testability (polymorphism and interchangeable implementations so a test harness can substitute a fake), how exception-handling choices change what a test for that failure path looks like, choosing the right collection for test-result processing, methodology for testing concurrency correctness (writing a test that can actually reveal a race or a visibility bug, not just fixing one), serialization and diffing trade-offs for test fixtures and CI artifacts, detecting a hash/equality-contract violation through testing, memory-bounded generator and streaming I/O patterns for test data, and small test-engineering utilities (CI-config diffing, checksum-verified data sharding). Excludes writing or debugging the code inside a single automated test script (control flow, parameterization, translating a manual case, locator, wait, and retry mechanics), which belongs to test automation scripting. Excludes suite-wide or framework-wide structural and strategy decisions (Page Object Model, layering, tool or driver choice, CI wiring, governance, flaky-test-detection systems, scalable test infrastructure), which belongs to test automation framework architecture and design. Excludes classic array/string/graph technique problems with no real test-engineering judgment required, which belong to arrays, strings, and hashing. Also excludes generic OOP-principles surveys, generic concurrency-primitive implementation (singletons, thread pools, producer-consumer queues, lock-free structures) with no distinct testing angle, generic garbage-collection and memory-leak content, generic functional-programming surveys, and generic hash-function or hash-table design: each of these already has a dedicated, larger topic in the catalog (Object-Oriented Programming and Design, Concurrency Synchronization and Deadlock, Memory Management and Garbage Collection, Functional Programming, Hashing and Hash Tables), and a test-flavored costume on otherwise-identical content is not a reason to duplicate it here.
Compare when you would reach for an array/list, a set, a map/dictionary, a queue, and a priority queue. For each, give one concrete example drawn from processing test results or scheduling test execution, and explain the time-complexity and ordering trade-offs that matter most when this choice shows up in a coding round. Separately, compare how the same choice differs between Python and Java specifically when deduplicating test inputs or tracking already-seen records: complexity guarantees, ordering, and memory trade-offs for large test fixtures.
Sample Answer
Direct answer
Choose the collection by the operation you need to be fast, not by habit: an array/list for ordered, indexed access; a set for membership and uniqueness; a map/dictionary for key-based lookup; a queue for first-in-first-out processing order; a priority queue when the next item to process is determined by a score rather than arrival order.
Structured elaboration
| Structure | Best for | Example from test-result processing | Complexity |
|---|---|---|---|
| Array/List | Ordered iteration, indexed access | Raw test results in execution order | Append O(1) amortized, arbitrary index O(1), membership check O(n) |
| Set | Uniqueness, fast membership | Detecting duplicate failed-test IDs, or comparing planned vs executed test sets | Add/contains O(1) average |
| Map/Dictionary | Fast lookup by key | Test name -> latest status, or test name -> historical failure count | Get/put O(1) average |
| Queue | First-in-first-out processing | Feeding a worker pool tests in submission order | Enqueue/dequeue O(1) with a deque |
| Priority Queue | Process by score, not arrival order | Running historically flakier or slower tests first to surface failures sooner | Push/pop O(log n) |
The trade-off that actually shows up in a coding round is almost always membership/lookup speed versus ordering guarantees: a list gives you order for free but costs O(n) to check membership; a set or map gives O(1) membership/lookup but, in Java's HashSet/HashMap, gives up insertion order entirely (Python's set/dict also do not guarantee insertion order for set, though dict has guaranteed insertion order since Python 3.7, which is a genuine language difference worth naming explicitly if asked).
On the Python-versus-Java angle specifically, for deduplicating test inputs or tracking already-seen records: both languages reach for a hash-based set, but the practical difference is what you get "for free" alongside it. Python's dict preserves insertion order as a language guarantee, so a dict doubling as an ordered, deduplicated record of "first-seen order" needs no extra structure. Java's HashMap makes no such guarantee, so the equivalent pattern in Java is LinkedHashMap, which explicitly layers insertion-order tracking on top of hash-based lookup; reaching for plain HashMap when order matters is a common Java mistake.
Worked example
A compact illustration of the ordering difference:
# Python: dict preserves insertion order (language guarantee since 3.7)
seen = {}
for record in ["b", "a", "b", "c", "a"]:
seen.setdefault(record, True)
print("python dict insertion order preserved:", list(seen.keys()))
assert list(seen.keys()) == ["b", "a", "c"]
Output:
python dict insertion order preserved: ['b', 'a', 'c']
// Java: plain HashMap makes no ordering guarantee; LinkedHashMap does.
import java.util.LinkedHashMap;
import java.util.Map;
public class OrderDemo {
public static void main(String[] args) {
Map<String, Boolean> seen = new LinkedHashMap<>();
for (String record : new String[]{"b", "a", "b", "c", "a"}) {
seen.putIfAbsent(record, true);
}
System.out.println("java LinkedHashMap insertion order preserved: " + seen.keySet());
if (!seen.keySet().toString().equals("[b, a, c]")) {
throw new AssertionError("expected insertion order [b, a, c], got " + seen.keySet());
}
System.out.println("PASS");
}
}
Output (compiled and run with a real JDK):
java LinkedHashMap insertion order preserved: [b, a, c]
PASS
Trade-offs and pitfalls
- Reaching for a list and doing repeated
inchecks is the most common interview miss: it works on small inputs and silently becomes an O(n^2) bottleneck on large test suites, exactly the kind of bug that only shows up once a suite grows past a few hundred tests. - Priority queues are frequently the right answer to "how do you decide what to run first", but are frequently under-used in favor of manually sorting the whole list up front, which is more expensive when the run order needs to adapt as new information (a new failure) arrives mid-run.
- Do not reach for
HashMap/HashSetin Java when order matters, and do not assume Python'ssetpreserves order the waydictdoes; conflating the two is a common source of subtle, hard-to-reproduce ordering bugs in test-result aggregation code.
Explain the Java memory model's happens-before relationship and the role of volatile and synchronized, then contrast it with Python's GIL-based concurrency model. What are the practical implications for designing thread-safe code in each language, and specifically, how would you write a test that can actually reveal an ordering or visibility bug rather than passing by luck on a single run?
Sample Answer
Direct answer
The Java Memory Model's happens-before relationship defines the specific set of orderings the JVM guarantees between actions in different threads; without one of those orderings in place (via volatile, synchronized, or a small set of other constructs), the JVM is free to let one thread never observe another thread's write at all, not just observe it late. Python's GIL provides a much simpler, blunter guarantee: only one thread executes Python bytecode at a time, so most of the subtle reordering and caching effects the JMM has to name explicitly do not arise the same way in CPython, though the underlying data can still be corrupted by non-atomic multi-step operations, such as two threads racing on a plain, unlocked count += 1 (a read-modify-write race, which is a different failure than the visibility race discussed below).
Structured elaboration
volatile and synchronized both establish a happens-before edge, but for different purposes:
volatileguarantees that a write to the field is immediately visible to any thread that subsequently reads it, and prevents the compiler/JIT from reordering that specific field's reads and writes across the access. It does NOT make compound operations atomic (volatile int x; x++;is still a race), so it is the right tool for a simple flag or reference, not a counter.synchronizedadditionally provides mutual exclusion (only one thread executes the guarded block at a time) alongside the same happens-before guarantee, so it is the right tool when multiple related fields must be updated together consistently, not just made visible.
Without either, a JIT compiler (the Just-In-Time compiler, which translates JVM bytecode to optimized native machine code while the program runs) is permitted to cache a field's value in a CPU register or reorder instructions in ways that are entirely correct for a single thread but can leave a second thread spinning on a stale, cached value forever, because nothing ever tells that thread's CPU core to refresh its view of memory.
Designing a test that can reveal this, rather than merely explaining it: the standard technique is a spin-wait with a bounded iteration budget, not a fixed sleep. A writer thread sleeps briefly and then flips a field; a reader thread spins reading the field up to some large iteration cap and records whether it ever observed the flip within that budget. Run against the volatile field, the test should reliably observe the flip (that is the guarantee being asserted). Run against the plain field, the outcome is legitimately non-deterministic (a real visibility bug is not guaranteed to reproduce on every JVM, JIT, or run, much like the scheduling non-determinism seen when two Python threads race on an unlocked counter's count += 1, which can pass by luck on a single run) so a strong test asserts the volatile guarantee positively rather than trying to force a flaky failure out of the unsafe version on demand.
Worked example
Compiled and run with a real JDK:
public class VisibilityDemo {
static boolean plainFlag = false;
static volatile boolean volatileFlag = false;
static void writerPlain() {
try { Thread.sleep(20); } catch (InterruptedException ignored) {}
plainFlag = true;
}
static void writerVolatile() {
try { Thread.sleep(20); } catch (InterruptedException ignored) {}
volatileFlag = true;
}
static boolean readerSeesFlagFlip(boolean useVolatile, long spinBudget) throws InterruptedException {
Thread writer = new Thread(useVolatile ? VisibilityDemo::writerVolatile : VisibilityDemo::writerPlain);
writer.start();
long i = 0;
boolean seen = false;
while (i < spinBudget) {
boolean current = useVolatile ? volatileFlag : plainFlag;
if (current) { seen = true; break; }
i++;
}
writer.join();
return seen;
}
public static void main(String[] args) throws InterruptedException {
volatileFlag = false;
boolean volatileSeen = readerSeesFlagFlip(true, 200_000_000L);
System.out.println("volatile flag observed within spin budget: " + volatileSeen);
if (!volatileSeen) throw new AssertionError("expected the volatile write to be visible");
System.out.println("test_volatile_write_is_visible_to_reader: PASS");
plainFlag = false;
boolean plainSeen = readerSeesFlagFlip(false, 200_000_000L);
System.out.println("plain (non-volatile) flag observed within spin budget: " + plainSeen);
}
}
Output:
volatile flag observed within spin budget: true
test_volatile_write_is_visible_to_reader: PASS
plain (non-volatile) flag observed within spin budget: false
On this run, the plain, non-volatile field's write was in fact never observed by the spinning reader within a 200-million-iteration budget, which is a live demonstration of the exact visibility gap being discussed, though the test does not assert on this value either way (see trade-offs), since it is a real, timing-dependent phenomenon rather than a guaranteed one.
Trade-offs and pitfalls
- Do not write a test that asserts the plain field's visibility bug always reproduces. Doing so ties a test's pass/fail status to JIT and scheduler behavior outside your control; assert only the positive guarantee (
volatileis visible), and treat the negative case as informational at most. volatileis not a substitute forsynchronizedwhen more than one field must be updated consistently together, or when a compound read-modify-write (like an increment) is involved; conflating "visible" with "atomic" is one of the most common Java concurrency mistakes.- Python's GIL means this exact class of visibility bug does not arise the same way, which can lead engineers moving between the two languages to under-appreciate Java's memory model; the closest Python analog is the read-modify-write race (not a visibility race): two threads incrementing a plain, unlocked counter's
count += 1can silently lose updates for the same underlying reason (an unsynchronized multi-step operation), even though CPython's GIL prevents the register-caching/reordering visibility gap demonstrated above.
Given pseudocode for a data loader that uses a shared mutable buffer across multiple worker threads, identify the thread-safety and race-condition risks. Propose concrete fixes, then describe the unit and stress tests, in both Java and Python, that you would write to reliably detect the race condition and data corruption rather than relying on code review alone to catch it.
Sample Answer
Direct answer
A plain instance attribute updated with count += 1 from multiple threads is a read-modify-write operation, not an atomic one, so concurrent increments can silently lose updates. The fix is a lock around every read-modify-write access to the shared state, and the test that proves it must force the race window open on purpose rather than hope a natural stress test catches it.
Structured elaboration
A shared data loader with a mutable buffer commonly needs to do two things concurrently: append incoming items, and maintain some running aggregate (a count, a running total) alongside them. In CPython, list.append() is a single, GIL-atomic C-level call and is safe on its own. self.count += 1, however, compiles to three separate steps: load self.count, add 1, store the result back. The GIL is free to switch to another thread between any of those steps, so two threads can both load the same old value before either writes back, and one increment is lost.
The test-design challenge is that this exact race is scheduler-dependent: whether the interpreter happens to switch threads inside that narrow window depends on the OS scheduler and the GIL's switch interval, which is why a naive stress test (spin up many threads, do many increments, check the final count) can pass "by luck" even on genuinely racy code, especially on a machine that happens to let each thread run to completion before yielding. A reliable test for a race condition has to either (a) widen the race window deterministically to prove the mechanism exists, or (b) accept that a natural-contention stress test is a probabilistic detector at best and is not sufficient evidence on its own that the code is safe.
Worked example
Verified in three parts:
1. Deterministic proof the mechanism exists (widen the window on purpose):
import threading, time
class UnsafeCounterWithForcedWindow:
def __init__(self):
self.count = 0
def increment(self):
old = self.count # step 1: read
time.sleep(0.001) # force a context switch inside the window
self.count = old + 1 # step 2: write back the (possibly stale) value
def prove_the_mechanism_deterministically(n_threads=8):
counter = UnsafeCounterWithForcedWindow()
threads = [threading.Thread(target=counter.increment) for _ in range(n_threads)]
for t in threads: t.start()
for t in threads: t.join()
print(f"expected {n_threads}, got {counter.count} (lost {n_threads - counter.count})")
assert counter.count < n_threads
prove_the_mechanism_deterministically()
Output:
expected 8, got 1 (lost 7)
2. Honest disclosure: natural contention on this run did NOT reproduce the loss (the production-shaped code, no artificial delay, run 5 times at 16 threads x 20,000 increments each, even after lowering sys.setswitchinterval to force more frequent yielding):
import threading, sys
class UnsafeCounter:
def __init__(self):
self.count = 0
def increment(self, n):
for _ in range(n):
self.count += 1
def natural_contention_trial(n_threads=16, increments_per_thread=20000):
counter = UnsafeCounter()
threads = [threading.Thread(target=counter.increment, args=(increments_per_thread,)) for _ in range(n_threads)]
for t in threads: t.start()
for t in threads: t.join()
expected = n_threads * increments_per_thread
return counter.count, expected
sys.setswitchinterval(0.0001) # force more frequent yielding
for trial in range(1, 6):
count, expected = natural_contention_trial()
print(f"natural-contention trial {trial}: count={count} expected={expected} lost={expected - count}")
Output:
natural-contention trial 1: count=320000 expected=320000 lost=0
natural-contention trial 2: count=320000 expected=320000 lost=0
natural-contention trial 3: count=320000 expected=320000 lost=0
natural-contention trial 4: count=320000 expected=320000 lost=0
natural-contention trial 5: count=320000 expected=320000 lost=0
This is itself the point worth making in the room: a race-condition test that "usually passes" proves nothing about whether the underlying code is safe. Part 1's deterministic, forced-window proof is the reliable evidence; a passing natural-contention run is not.
3. The fix, stress-tested and verified to never lose an increment:
class SafeSharedBuffer:
def __init__(self):
self._lock = threading.Lock()
self.buffer = []
self.count = 0
def add(self, item):
with self._lock:
self.buffer.append(item)
self.count += 1
def stress_trial(n_threads=16, items_per_thread=20000):
buf = SafeSharedBuffer()
def worker(thread_id):
for i in range(items_per_thread):
buf.add((thread_id, i))
threads = [threading.Thread(target=worker, args=(t,)) for t in range(n_threads)]
for t in threads: t.start()
for t in threads: t.join()
expected = n_threads * items_per_thread
return buf.count, len(buf.buffer), expected
for trial in range(1, 4):
count, buffer_len, expected = stress_trial()
print(f"SafeSharedBuffer trial {trial}: count={count} buffer_len={buffer_len} expected={expected}")
Output across 3 trials of 16 threads x 20,000 increments each:
SafeSharedBuffer trial 1: count=320000 buffer_len=320000 expected=320000
SafeSharedBuffer trial 2: count=320000 buffer_len=320000 expected=320000
SafeSharedBuffer trial 3: count=320000 buffer_len=320000 expected=320000
Trade-offs and pitfalls
- A stress test that passes is not proof of thread safety. Part 2 above is the concrete demonstration: the exact same buggy code, under real concurrent load, did not lose a single increment on this run. Relying on that kind of test as your only safety net would ship a real bug with a green CI run.
- Widening the race window with a sleep, as in part 1, is a legitimate testing technique for PROVING a mechanism, but the widened code is not what ships. Do not confuse "I added a sleep to force the bug to show" with "the sleep is part of the fix"; the fix is the lock, verified separately in part 3.
- Locking too narrowly is a common half-fix: locking only around
self.count += 1but notself.buffer.append(item)looks safe (both individually become atomic) but can still let the two fall out of sync with each other if a reader observes them between the two separate lock acquisitions; locking both under the same critical section, as done here, avoids that entirely.
Implement a memory-efficient iterator, a Python generator or a Java Stream/Iterator, that parses and yields records from a very large newline-delimited JSON file. Include how you would unit test parsing errors on malformed records and how you would test backpressure or a slow consumer. Then discuss how the same streaming approach adapts to a typed, tabular format like CSV: reading a header row plus the first N rows as typed values, and testing malformed lines, missing files, and very large files without flaky filesystem behavior in CI.
Sample Answer
Direct answer
Parse the newline-delimited JSON file one line at a time with a generator, yielding a parsed record per non-empty line and surfacing malformed lines as either an immediate exception or an inline error object depending on the caller's needs; the same streaming shape adapts directly to a typed CSV table by reading the header once, then coercing and yielding rows incrementally.
Structured elaboration
Both formats share the same underlying shape: read one line/row at a time, never materialize the whole file, and decide what to do when a line does not parse.
For newline-delimited JSON, two error-handling modes matter for different callers: on_error="raise" stops the stream at the first malformed record (appropriate when any malformed record means the whole file is untrustworthy), while on_error="skip" yields a sentinel error object in place of a record and continues (appropriate when a consumer wants to process everything it can and separately collect diagnostics on what it could not). Testing backpressure or a slow consumer for either mode means confirming the generator does not read ahead of what the consumer actually asks for: since a Python generator only advances when next() is called, a slow consumer naturally throttles the producer for free, and the test below proves this directly by wrapping the line source in a counter and asserting the count of lines pulled matches the count of next() calls made, no more and no less.
For the CSV extension: read the header row once, then for each subsequent row, coerce each column's string value using a caller-supplied type map (falling back to str for unlisted columns), stop once n good rows have been collected, and separately collect rows whose column count does not match the header or whose coercion raises, so the caller sees both what parsed and exactly what did not, rather than the whole read aborting on the first bad row.
Worked example
Verified:
import csv, json
class MalformedRecordError(Exception):
def __init__(self, line_number, raw_line, cause):
super().__init__(f"line {line_number}: {cause}")
self.line_number = line_number
def stream_ndjson(path, on_error="raise"):
with open(path, "r", encoding="utf-8") as f:
for line_number, raw_line in enumerate(f, start=1):
stripped = raw_line.strip()
if not stripped:
continue
try:
yield json.loads(stripped)
except json.JSONDecodeError as e:
err = MalformedRecordError(line_number, raw_line, str(e))
if on_error == "raise":
raise err
yield err
# valid records stream correctly
import tempfile, os
def _write(content):
tf = tempfile.NamedTemporaryFile(delete=False, mode="w", encoding="utf-8")
tf.write(content); tf.close()
return tf.name
path = _write('{"a": 1}\n{"a": 2}\n\n{"a": 3}\n')
records = list(stream_ndjson(path))
assert records == [{"a": 1}, {"a": 2}, {"a": 3}]
os.unlink(path)
print("ndjson valid-records: PASS ->", records)
# skip mode continues past a malformed line, reporting it
path = _write('{"a": 1}\nNOT JSON\n{"a": 3}\n')
results = list(stream_ndjson(path, on_error="skip"))
assert results[0] == {"a": 1}
assert isinstance(results[1], MalformedRecordError) and results[1].line_number == 2
assert results[2] == {"a": 3}
os.unlink(path)
print("ndjson skip-mode: PASS, malformed line reported as line", 2, "stream continued")
# raise mode stops at the malformed line
path = _write('{"a": 1}\nNOT JSON\n{"a": 3}\n')
gen = stream_ndjson(path, on_error="raise")
assert next(gen) == {"a": 1}
try:
next(gen)
assert False
except MalformedRecordError as e:
print("ndjson raise-mode: PASS, raised at line", e.line_number)
os.unlink(path)
# backpressure / slow-consumer test: the generator must not read ahead of
# what the consumer has actually asked for
def gen_from_source(source):
for raw_line in source:
line = raw_line.strip()
if not line:
continue
yield json.loads(line)
def test_backpressure_no_read_ahead():
lines_read_count = {"n": 0}
class CountingLineSource:
"""Wraps an iterable of lines and counts how many the consumer has pulled."""
def __init__(self, lines):
self._lines = iter(lines)
def __iter__(self):
return self
def __next__(self):
line = next(self._lines)
lines_read_count["n"] += 1
return line
lines = ['{"a": 1}\n', '{"a": 2}\n', '{"a": 3}\n', '{"a": 4}\n']
source = CountingLineSource(lines)
gen = gen_from_source(source)
assert lines_read_count["n"] == 0, "nothing should be read before the consumer calls next()"
first = next(gen)
assert lines_read_count["n"] == 1, f"expected exactly 1 line read, got {lines_read_count['n']}"
assert first == {"a": 1}
second = next(gen)
assert lines_read_count["n"] == 2, f"expected exactly 2 lines read, got {lines_read_count['n']}"
assert second == {"a": 2}
print(f"backpressure test: PASS, reads={lines_read_count['n']} tracked 1:1 with next() calls, no read-ahead observed")
test_backpressure_no_read_ahead()
def stream_typed_csv(path, n, column_types):
typed_rows, errors = [], []
with open(path, "r", encoding="utf-8", newline="") as f:
reader = csv.reader(f)
header = next(reader, None)
if header is None:
return [], [], []
for row_number, row in enumerate(reader, start=2):
if len(typed_rows) >= n:
break
if len(row) != len(header):
errors.append((row_number, row, "column count mismatch"))
continue
try:
typed_rows.append({col: column_types.get(col, str)(val) for col, val in zip(header, row)})
except (ValueError, TypeError) as e:
errors.append((row_number, row, str(e)))
return header, typed_rows, errors
content = "name,age,score\nalice,30,9.5\nbob,not_a_number,8.0\ncarol,25,7.25\ndave,40,6.0\n"
path = _write(content)
header, rows, errors = stream_typed_csv(path, n=2, column_types={"age": int, "score": float})
assert header == ["name", "age", "score"]
assert rows == [{"name": "alice", "age": 30, "score": 9.5}, {"name": "carol", "age": 25, "score": 7.25}]
assert len(errors) == 1 and errors[0][0] == 3
os.unlink(path)
print("typed-csv: header=", header, "rows=", rows, "errors=", errors)
Output:
ndjson valid-records: PASS -> [{'a': 1}, {'a': 2}, {'a': 3}]
ndjson skip-mode: PASS, malformed line reported as line 2 stream continued
ndjson raise-mode: PASS, raised at line 2
backpressure test: PASS, reads=2 tracked 1:1 with next() calls, no read-ahead observed
typed-csv: header= ['name', 'age', 'score'] rows= [{'name': 'alice', 'age': 30, 'score': 9.5}, {'name': 'carol', 'age': 25, 'score': 7.25}] errors= [(3, ['bob', 'not_a_number', '8.0'], "invalid literal for int() with base 10: 'not_a_number'")]
Bob's row (line 3) is correctly excluded from rows and reported in errors with the exact coercion failure, while carol's row (which comes after the malformed one) still counts toward the requested n=2 good rows, confirming the malformed row does not consume the caller's row budget.
Trade-offs and pitfalls
- Skip-mode's sentinel-object approach mixes error objects into the same stream as data, which is simple to implement but requires every consumer to check
isinstance(item, MalformedRecordError); an alternative design (two separate generators, or a callback for errors) avoids that check at the cost of more complex control flow. - The CSV row-number bookkeeping starts at 2 because the header consumes row 1, a common off-by-one source; the test above explicitly checks the reported row number, not just that an error was recorded.
- Backpressure and slow-consumer behavior is a property of Python's generator protocol itself (nothing is read until
next()is called), so it is easy to assume it "just works," but it is worth an explicit test if the consumer additionally holds a lock or a bounded queue, since that is where the actual backpressure (not just laziness) could still be violated by a misbehaving consumer that reads ahead into a buffer.
Implement a class Sampler that wraps a fixed collection of items and provides a method sample(n, seed=None) returning n unique items without modifying the collection or any internal state. Given the same seed, two calls must return the same result. Explain why this kind of deterministic, seeded sampling matters for building reproducible test fixtures, and show a short unit test.
Sample Answer
Direct answer
Build Sampler around Python's random.Random(seed) (a private, independent random-number-generator instance), and never mutate the wrapped collection: copy it once at construction, and use random.sample, which itself returns a new list without touching its input.
Structured elaboration
Three requirements have to hold simultaneously, and each maps to a specific implementation choice:
- Determinism given the same seed. A module-level
random.random()call depends on global, mutable state that other code in the process can also perturb. Instantiating a freshrandom.Random(seed)inside the call isolates this sampler's randomness completely: two calls with the same seed see an identical generator state and produce identical output, regardless of what any other code in the process has done to the globalrandommodule. - No mutation of the wrapped collection. Store a private copy (
list(items)) at construction time so a caller mutating their own original list afterward cannot silently change what the sampler draws from.random.sampleitself also returns a new list and does not shuffle or consume its input in place. - No internal state changes across calls. Each
sample()call creates its own localrandom.Random(seed); nothing persists between calls, so callingsample()twice with the same arguments is safe and repeatable, and calling it withnlarger than the population is handled by clamping rather than raising.
Why this matters for test fixtures specifically: a flaky test that occasionally samples a different subset of fixture data produces failures that look unrelated to the actual change under test, and are expensive to triage because they cannot be reproduced on demand. A deterministic, seeded sampler turns "this test failed on CI but passes locally" into a reproducible, debuggable case.
Worked example
Verified:
import random
from typing import List, Optional, Sequence
class Sampler:
"""Wraps a fixed collection and returns deterministic, seeded samples
without ever mutating the wrapped collection or any internal state."""
def __init__(self, items: Sequence[str]) -> None:
self._items: List[str] = list(items) # private copy
def sample(self, n: int, seed: Optional[int] = None) -> List[str]:
if n < 0:
raise ValueError("n must be >= 0")
n = min(n, len(self._items))
rng = random.Random(seed)
return rng.sample(self._items, n)
def test_same_seed_same_result():
s = Sampler(["a", "b", "c", "d", "e"])
first = s.sample(3, seed=42)
second = s.sample(3, seed=42)
assert first == second
print("same seed ->", first)
def test_no_mutation_of_source():
original = ["a", "b", "c"]
s = Sampler(original)
s.sample(2, seed=7)
assert original == ["a", "b", "c"]
print("source unmutated:", original)
def test_n_larger_than_population():
s = Sampler(["x", "y"])
result = s.sample(10, seed=5)
assert sorted(result) == ["x", "y"]
print("n > population ->", result)
test_same_seed_same_result()
test_no_mutation_of_source()
test_n_larger_than_population()
Output:
same seed -> ['a', 'e', 'c']
source unmutated: ['a', 'b', 'c']
n > population -> ['y', 'x']
Trade-offs and pitfalls
random.Randomis not cryptographically strong. That is the correct trade-off here (speed and reproducibility matter far more than unpredictability for test fixtures), but do not reuse this pattern for anything security-sensitive.- Very large collections. Materializing a full copy at construction is fine for typical fixture sizes, but for a truly huge population, prefer reservoir sampling (single pass, O(1) extra memory) over copying the whole thing; that is a different, more complex algorithm with its own seeding subtleties.
- A common mistake is seeding the module-level
randommodule once at the top of a test session instead of per-call: that makes test order affect results, since every other call torandomin the process advances the same shared state. Isolating an independentrandom.Random(seed)per call, as done here, avoids that entirely.
Unlock Full Question Bank
Get access to all 15 Programming for Test Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.