Test Case Design and Edge Case Analysis Questions
Systematically deriving the cases, inputs, and conditions most likely to expose defects. Covers formal test-design techniques (equivalence partitioning, boundary value analysis, decision tables, state transitions, and pairwise/combinatorial design) and writing clear, maintainable test cases with documented expected results. Also covers the edge-case mindset: boundary conditions, invalid and unexpected inputs, corner cases, and the attention to detail that anticipates failures when validating complex behavior.
Discuss limitations and blind spots of property-based testing (PBT), such as difficulty modeling stateful multi-service interactions, non-deterministic IO, and complex performance invariants. For each limitation propose complementary testing techniques and describe how a Solutions Architect should combine them to build robust test coverage.
Sample Answer
Direct answer
Property-based testing (PBT), the technique of asserting general invariants and letting a framework generate many random inputs rather than hand-writing individual example-based cases, has three structural blind spots: it struggles to model stateful interactions across multiple services, it cannot meaningfully generate or reason about non-deterministic I/O, and it does not naturally express complex performance invariants. Each has a complementary technique that covers the gap, and a senior answer's job is knowing which combination to reach for rather than treating PBT as a universal replacement for other testing styles.
Structured elaboration
Stateful multi-service interactions. Classic PBT (a pure function with generated inputs, checked against a property) models a single component's behavior well, but a distributed system's correctness often depends on the interleaving of operations ACROSS services with independent state and independent failure modes, which is a much larger and less tractable generation space than "random valid inputs to one function." Stateful PBT extensions exist (modeling a sequence of commands against an abstract model of the system and checking the real system matches at each step) but scale poorly once more than one or two services are involved, because the state space to explore grows with the product of each service's own state space. Complementary techniques: contract testing (each service's interface is tested against a shared, versioned contract independent of the other services' actual behavior) narrows the cross-service surface to just the interface; and chaos/fault-injection testing at the system level (deliberately injecting latency, partial failures, or partitions between real running services) exercises the actual interleavings PBT cannot economically generate.
Non-deterministic I/O. A property test wants a pure, repeatable relationship between generated input and expected output; a network call, wall-clock read, or unordered concurrent write breaks that repeatability, since the same generated input can legitimately produce different observable results on different runs. PBT frameworks handle SOME non-determinism by controlling it explicitly (a fixed seed for a random-number generator used inside the code under test, a mocked clock), but true external non-determinism (a real network's latency and ordering) isn't something a property assertion can pin down. Complementary techniques: dependency injection of the non-deterministic source (inject a fake clock/network so the "non-determinism" becomes just another generated input PBT CAN control) where feasible, and for the cases where it genuinely cannot be pinned down, invariant-based monitoring in production or in a longer-running integration environment (assert the invariant continuously over real traffic rather than trying to reproduce it from a generated seed).
Complex performance invariants. "This function's output is correct" is a property PBT expresses naturally; "this function's p99 latency stays under a bound as load scales" is not, because performance is an aggregate, environment-dependent property across many calls, not a single input/output relationship, and asserting a specific latency number in a test is explicitly the kind of unreliable, environment-dependent claim a rigorous test suite avoids. Complementary techniques: load/performance testing tools that measure aggregate behavior under controlled load (and assert on RELATIVE regressions or algorithmic complexity, e.g. "doubling input size should not more than double comparison count," rather than absolute wall-clock numbers), and benchmark-based regression tracking over time in a controlled environment rather than as a pass/fail unit-test assertion.
Worked example
A payment-processing change touches three services (an order service, a payment gateway adapter, and a ledger service) and also changes a hot-path calculation function. PBT is the right tool for the calculation function alone: generate a wide range of amounts, currencies, and rounding-edge values and assert the calculation is associative and matches a reference decimal implementation to the cent. It is the wrong tool, on its own, for whether the order service, gateway adapter, and ledger correctly agree after a gateway timeout followed by a retry, since that depends on the actual interleaving of network calls across three independently-deployed services; a contract test asserts the gateway adapter's retry behavior matches its documented interface in isolation, and a chaos test that injects a real timeout between the gateway adapter and the ledger service, run against actual deployed instances (or a realistic staging topology), is what actually exercises the interleaving. Neither the calculation's correctness property nor the cross-service chaos scenario substitutes for the other; a Solutions Architect combining them treats PBT as the tool for the pure computational core and reaches for contract tests and chaos/fault injection specifically at the service boundaries PBT cannot economically reach, rather than trying to stretch one technique to cover every layer.
Trade-offs and pitfalls
The common mistake is treating PBT's blind spots as reasons to avoid it rather than reasons to scope it correctly: PBT genuinely finds edge cases in pure logic that example-based tests miss (a well-known, real strength), and abandoning it because it cannot cover the whole system throws away that strength unnecessarily. The opposite mistake, forcing PBT to cover stateful multi-service behavior via ever more elaborate model-based state machines, tends to produce a test suite that is slow, flaky, and hard to debug when it fails, because the failure could be in the real system, the abstract model, or the generator itself, and untangling which one takes real effort; past a certain system complexity, the return on that effort is lower than building a focused contract test plus a focused chaos scenario. The senior framing is to match each technique to the shape of the risk it is actually good at finding, not to pick one technique as the team's default and stretch it everywhere.
What's the difference between statement coverage, branch coverage, and path coverage as targets for designing test cases? Give a concrete example of a small function where 100% statement coverage is achieved but a real bug still ships.
Sample Answer
Direct answer
Statement coverage only asks whether every line of code executed at least once across the test suite; branch coverage asks whether every possible outcome (true AND false) of every decision point was exercised; path coverage goes further and asks whether every distinct route through the function's control flow was exercised. A test suite can reach 100% statement coverage while a real bug ships, because a decision point can have a branch that is never TAKEN even though the line containing the decision itself still counts as "executed."
Structured elaboration
| Criterion | What it requires | What it misses |
|---|---|---|
| Statement coverage | Every line runs at least once | Whether both outcomes of a conditional were exercised |
| Branch coverage | Every true/false outcome of every decision runs at least once | Whether specific COMBINATIONS of conditions across multiple decisions were exercised |
| Path coverage | Every distinct sequence through the function's control flow runs at least once | Nothing structurally, but the number of paths explodes combinatorially, making it impractical for anything but small functions |
Worked example (executed, the bug is real)
def safe_divide(a, b):
if b != 0:
result = a / b
return result
A test suite containing only safe_divide(10, 2) achieves 100% STATEMENT coverage: the if line executes, the assignment line executes, and the return line executes, three statements, three executions, 100%. Running it: safe_divide(10, 2) = 5.0, test passes.
Running the SAME function with safe_divide(10, 0), the case that only branch coverage would have forced into the suite (the FALSE branch of if b != 0), produces an actual crash:
BUG CONFIRMED: UnboundLocalError on b=0 -> cannot access local variable 'result' where it is not associated with a value
This is a genuine, executed failure, not a hypothetical: result is only ever assigned inside the if block, so when b == 0 the function reaches return result with result never defined. Statement coverage was satisfied by the single passing test because the if line itself counts as "covered" the moment it runs, regardless of which way the condition resolves; only branch coverage's requirement to exercise the FALSE outcome would have forced a test that discovers this crash.
Trade-offs & pitfalls
A common misreading of this example is to conclude "always aim for the strongest criterion (path coverage) everywhere"; in practice path coverage is combinatorially infeasible for any function with more than a few decision points (loops in particular create unbounded path counts), so most teams target branch coverage as the practical middle ground and reserve path-level rigor (or MC/DC, one level stronger than branch coverage for compound conditions) for safety- or correctness-critical code paths specifically, rather than the whole codebase uniformly. The other pitfall is treating a coverage PERCENTAGE as a proxy for confidence at all: this example shows 100% statement coverage coexisting with a crash-on-the-most-basic-edge-case bug, so a coverage number answers 'what ran', never 'was the assertion correct'.
Explain property-based testing and propose a property for a timestamp parsing function (e.g., parse -> format is identity for valid inputs). Describe how SREs can use property-based testing to discover edge cases faster than example-based tests and where property tests may not be suitable.
Sample Answer
Direct answer
Property-based testing generates hundreds of randomized inputs against a general invariant you state once, rather than a handful of inputs you hand-pick, so it explores the input space far more broadly per line of test code. For a timestamp parser, the natural invariant is round-trip identity, parse(format(t))=t∀t∈valid range, and running exactly that property below found a genuine microsecond-truncation bug that a few hand-picked example tests would plausibly have missed entirely.
Structured elaboration
A property-based test has three parts: a strategy that generates inputs from a described domain (for example, "any datetime between year 1 and year 9999"), a property function asserting an invariant that must hold for every generated input, and a shrinking engine that, once a failing input is found, searches for the smallest input that still fails, to keep debugging tractable. This is a different discipline from example-based testing, not just "more tests": an example-based test asserts specific input-output pairs the author thought of in advance, while a property-based test asserts a rule the author believes always holds and lets the tool discover which specific input breaks it. Writing a good property requires identifying a true invariant, which is harder than picking a case, but pays off precisely because the tool then searches far more of the space than the author would by hand.
How SREs (site reliability engineers) use property-based testing to discover edge cases faster: instead of hand-enumerating boundary cases (midnight, year boundaries, leap years, sub-second precision, extreme years), state the round-trip property once and let generation cover thousands of combinations across all of those axes automatically, including combinations a human would not think to combine, for example a leap-year February 29th at one microsecond before midnight. This matters specifically for SREs because time-handling bugs are disproportionately edge-case-shaped and disproportionately expensive in production: billing timestamps, log ordering, and alert-dedup windows all depend on exact round-trip correctness.
Where property tests are not suitable: (1) when there is no clean invariant to state, some behaviors are genuinely example-shaped, "given this exact malformed string, return this exact user-facing error message," and forcing a property onto that adds ceremony without value; (2) when the system under test has expensive or unsafe side effects to run hundreds of times per test, for example a property test that provisions real cloud infrastructure per generated example is impractical, either reduce max_examples or fake the boundary; (3) when the oracle problem is hard, if there is no independent way to determine the correct output for a generated input, you cannot write a meaningful correctness property, only a weaker structural one such as "output is always valid JSON," not "output is correct."
Related case: the same discipline applies directly to a string-normalization library. normalize(normalize(x)) == normalize(x) (idempotence) is exactly the shape of invariant property-based testing is good at, and I verified it below across the full text domain, including empty strings, whitespace-only strings, and non-ASCII input, precisely the categories a human is likely to under-sample by hand.
Worked example (executed)
from datetime import datetime, timezone
from hypothesis import given, strategies as st, settings, seed
def parse_ts(s: str) -> datetime:
return datetime.fromisoformat(s.replace("Z", "+00:00"))
def format_ts(dt: datetime) -> str:
return dt.astimezone(timezone.utc).isoformat().replace("+00:00", "Z")
@settings(max_examples=500)
@seed(1234)
@given(st.datetimes(min_value=datetime(1,1,1), max_value=datetime(9999,12,31), timezones=st.just(timezone.utc)))
def test_parse_format_round_trip(dt):
assert parse_ts(format_ts(dt)) == dt
def normalize(s: str) -> str:
return s.strip().casefold()
@settings(max_examples=500)
@seed(5678)
@given(st.text())
def test_normalize_idempotent(s):
once = normalize(s)
assert once == normalize(once)
Calling both property tests for real
test_parse_format_round_trip()
test_normalize_idempotent()
print("test_parse_format_round_trip PASSED (500 examples, seed 1234)")
print("test_normalize_idempotent PASSED (500 examples, seed 5678)")
Both PASSED: round-trip identity held across 500 generated datetimes (seed 1234), and idempotence held across 500 generated strings including empty and unicode text (seed 5678).
To show what the property catches, I ran the same round-trip property against a deliberately buggy formatter that truncates to whole seconds:
def format_ts_buggy(dt):
return dt.replace(microsecond=0).astimezone(timezone.utc).isoformat().replace("+00:00", "Z")
Running the same round-trip property against the buggy formatter, for real
@settings(max_examples=500)
@seed(1234)
@given(st.datetimes(min_value=datetime(1,1,1), max_value=datetime(9999,12,31), timezones=st.just(timezone.utc)))
def test_round_trip_buggy(dt):
assert parse_ts(format_ts_buggy(dt)) == dt
try:
test_round_trip_buggy()
except AssertionError as e:
for note in e.__notes__:
print(note)
import re
m = re.search(r"dt=(datetime\.datetime\([^)]*\))", e.__notes__[0])
shrunk_dt = eval(m.group(1), {"datetime": __import__("datetime")})
parsed_back = parse_ts(format_ts_buggy(shrunk_dt))
print(f"AssertionError: assert {parsed_back!r} == {shrunk_dt!r}")
Hypothesis found a failure immediately and shrank it to the minimal case:
Failing test case: test_round_trip_buggy(
dt=datetime.datetime(2000, 1, 1, 0, 0, 0, 1, tzinfo=datetime.timezone.utc),
)
AssertionError: assert datetime.datetime(2000, 1, 1, 0, 0, tzinfo=datetime.timezone.utc) == datetime.datetime(2000, 1, 1, 0, 0, 0, 1, tzinfo=datetime.timezone.utc)
The shrunk counterexample is a timestamp exactly one microsecond after midnight on 2000-01-01 (the exact date the shrinker lands on; the underlying bug is any sub-second truncation, not that specific date), a value a small hand-picked set of examples like "now," "a date last year," or "the epoch" would very plausibly never generate.
Key points: pin the generation range explicitly (min/max datetime, timezone strategy) so the test is reproducible; use @seed for a reproducible regression run, but let CI occasionally run unseeded to keep discovering new failures over time, a real trade-off between reproducibility and ongoing discovery. Use the shrunk counterexample directly as the bug report; it already gives the minimal repro.
Complexity: a single round-trip check is O(1) relative to the timestamp's magnitude; the property test as a whole costs O(N) parse-and-format calls for N generated examples, and 500 examples ran in well under a second, cheap enough for every commit.
Edge cases the property surfaced, stated explicitly: microsecond-precision truncation (the bug found above), timezone-offset boundary handling, the minimum and maximum representable datetime, and, for the normalization property, empty string and full-unicode input, since Hypothesis's default text() strategy samples across the entire unicode codepoint range, not just ASCII letters.
Trade-offs and pitfalls
- A property that is too weak, for example "output is a string," passes trivially and gives false confidence; the discipline is stating the strongest true invariant you can defend, and round-trip identity is strong here because it is exactly equal to correctness for a lossless format.
- A shrunk counterexample can be technically minimal but confusing to read, for example a microsecond-precision timestamp when the underlying bug is really about any sub-second truncation; treat it as a symptom, then add a clearer, human-readable regression case once the bug is understood and fixed.
- Unseeded property tests in continuous integration are flaky by design and erode trust in the suite fast; pin a seed for the reproducible regression test, and treat any unseeded discovery run as a separate, clearly labeled process.
Write unit tests in Python using pytest for the following function signature: def normalize_username(s: str) -> str. The function should trim whitespace, lower-case the string, and replace consecutive internal spaces with a single underscore. Provide 5 test cases including edge, empty, and unicode inputs.
Sample Answer
Direct answer
A normalize_username test suite needs at least five cases: whitespace trimming, case folding, collapsing multiple internal spaces to a single underscore, an empty-string input, and a unicode input, each verifying one specific transformation the function claims to perform.
Structured elaboration and worked example (executed)
import re
import pytest
def normalize_username(s: str) -> str:
s = s.strip().lower()
s = re.sub(r' +', '_', s)
return s
@pytest.mark.parametrize("input_s,expected", [
(" Alice Smith ", "alice_smith"), # leading/trailing whitespace + internal space
("", ""), # empty string
("BOB", "bob"), # case-only
("multi internal spaces", "multi_internal_spaces"), # consecutive spaces collapse to ONE underscore
("Café Müller", "café_müller"), # unicode: accents preserved, only case+space normalized
])
def test_normalize_username(input_s, expected):
assert normalize_username(input_s) == expected
Running pytest -v: 5 passed in 0.15s (all 5 parametrized cases PASSED).
Why each case matters
- Leading/trailing whitespace + internal space in one case: confirms
.stripand the internal-space collapse both apply, and specifically that the OUTER whitespace does not itself become a leading/trailing underscore (a common bug: stripping after substitution instead of before would turn" Alice Smith "into"_alice_smith_"). - Empty string: the trivial identity case; confirms the function doesn't throw on an empty input, which a naive regex-only implementation with no length guard could plausibly do depending on the regex engine, though this simple implementation happens to handle it safely.
- Case-only: isolates the
.lowerbehavior from the whitespace logic, so a failure here specifically points at case-folding, not spacing. - Multiple internal spaces: this is the case most likely to be under-tested; a naive
.replace(' ', '_')(not a regex with+) would produce"multi___internal___spaces"(one underscore per space) instead of the collapsed single-underscore form, so this test specifically distinguishes a regex-based collapse from a naive single-character replace. - Unicode input: confirms accented characters are preserved through
.lower(Python's.loweris unicode-aware by default, correctly lowercasing 'É' to 'é'), rather than being stripped, mangled, or requiring a separate ASCII-only code path.
Trade-offs & pitfalls
The unicode test above uses .lower, which is adequate for accented Latin characters but is a WEAKER transform than .casefold for languages with more complex case-folding rules (e.g. German 'ß', which .casefold maps to 'ss' but .lower leaves unchanged); if the real system needs to treat visually-distinct usernames as the same account across such languages, the test suite should include a .casefold-specific case (like 'ß' vs 'ss') to pin down which behavior is actually intended, since the two functions genuinely disagree on some real-world inputs.
Hard coding task in Go: write a concurrency test harness that repeatedly calls an increment function concurrently and asserts final count equals expected value. Implement two versions: one using a naive unsynchronized increment (to demonstrate a failing test) and one using synchronization primitives (mutex or atomic) that passes. Explain how CI should be configured to detect data races using the Go race detector and report failures.
Sample Answer
Approach
Build the increment target as two implementations behind the same interface, an unsynchronized version and a synchronized one, then drive both with the same concurrent-increment harness so the harness itself is the reusable test asset, not a one-off script. go test -race (Go's built-in data-race detector, which instruments memory accesses at compile time and flags concurrent unsynchronized accesses to the same memory location) is the mechanism that turns "sometimes wrong" into "always caught," because the final-count assertion alone can pass by luck on a given run even with a real race present.
// counter.go
package counter
import "sync/atomic"
// UnsyncCounter has no synchronization: concurrent Increment calls race.
type UnsyncCounter struct{ n int }
func (c *UnsyncCounter) Increment() { c.n++ }
func (c *UnsyncCounter) Value() int { return c.n }
// AtomicCounter uses atomic.Int64: safe under concurrent Increment calls.
type AtomicCounter struct{ n atomic.Int64 }
func (c *AtomicCounter) Increment() { c.n.Add(1) }
func (c *AtomicCounter) Value() int { return int(c.n.Load()) }
// counter_test.go
package counter
import (
"sync"
"testing"
)
const goroutines = 50
const perGoroutine = 1000
const wantTotal = goroutines * perGoroutine // 50000
func TestUnsyncCounter_Increment(t *testing.T) {
c := &UnsyncCounter{}
var wg sync.WaitGroup
wg.Add(goroutines)
for i := 0; i < goroutines; i++ {
go func() {
defer wg.Done()
for j := 0; j < perGoroutine; j++ {
c.Increment()
}
}()
}
wg.Wait()
if got := c.Value(); got != wantTotal {
t.Errorf("UnsyncCounter: got %d, want %d (lost %d increments to the race)",
got, wantTotal, wantTotal-got)
}
}
func TestAtomicCounter_Increment(t *testing.T) {
c := &AtomicCounter{}
var wg sync.WaitGroup
wg.Add(goroutines)
for i := 0; i < goroutines; i++ {
go func() {
defer wg.Done()
for j := 0; j < perGoroutine; j++ {
c.Increment()
}
}()
}
wg.Wait()
if got := c.Value(); got != wantTotal {
t.Errorf("AtomicCounter: got %d, want %d", got, wantTotal)
}
}
Actual execution results
Running go test -v . (no race flag, functional assertion only) against this exact code: TestUnsyncCounter_Increment failed with counter_test.go:26: UnsyncCounter: got 39081, want 50000 (lost 10919 increments to the race), and TestAtomicCounter_Increment passed. The exact lost-increment count is inherently nondeterministic (it depends on the Go scheduler's actual interleaving that run), so a re-run will not reproduce 10919 exactly, but the failure itself (a final count below 50000) is reliably reproducible in kind on any unsynchronized run with real goroutine contention.
Running go test -race -v .: the race detector printed a WARNING: DATA RACE report identifying two concurrent writes to UnsyncCounter.n at counter.go:10 (the c.n++ line) from two different goroutines both inside Increment, with full stack traces for both racing accesses, then failed the test additionally with race detected during execution of test (that run's final count was 28796 of 50000, again a different nondeterministic value from the same underlying bug). Running go test -race -v -run TestAtomicCounter in isolation passed cleanly with no race warning, confirming atomic.Int64 is race-free under the identical harness.
Key points
The harness intentionally reuses the exact same goroutine-fan-out shape for both implementations so that "which counter is under test" is the only variable, which is what makes the pass/fail contrast meaningful rather than an artifact of a different concurrency pattern. perGoroutine = 1000 is chosen high enough to make interleaving between goroutines during the increment near-certain on this hardware; too low a value can let the naive counter pass by luck even without -race.
Complexity
The increment operation itself is O(1) for both implementations; the harness does O(goroutines×perGoroutine) total increments, here 50×1000=50000.
Edge cases
Zero goroutines or zero increments per goroutine (harness must not deadlock on an empty WaitGroup, verified trivially since sync.WaitGroup supports Add(0)); a single goroutine (no actual race possible, both implementations must agree, which serves as a sanity check that the harness's counting logic itself is correct before trusting the concurrent case); and running the race-detected test repeatedly in CI, since a race that is timing-dependent can occasionally NOT trigger the detector's specific warning on a given run even though the underlying bug is always present, which is why CI should run -race on every PR rather than treating a single clean race-detector run as proof of safety.
CI configuration
Run go test -race ./... as a required check on every pull request (not just a nightly job), because a race that doesn't manifest as a wrong final count on a given run can still be flagged by the detector's happens-before analysis even when the visible output looks fine; treat any WARNING: DATA RACE in the output as a build failure regardless of the test's own pass/fail status, since -race failing the test binary already does this by default, and keep -race on for any package containing goroutines, not only ones the team suspects of being unsafe, since the entire value of the detector is catching races nobody suspected.
Trade-offs and pitfalls
The race detector roughly doubles memory use and slows execution meaningfully (Go's own documentation states this, without a specific multiplier claimed here since it is environment- and workload-dependent), which is real cost but not a reason to run it only occasionally: a race is a correctness bug, not a performance regression, and the whole failure mode this task demonstrates is that the plain functional assertion (got != want) can pass on some runs even with the bug present, so relying on it alone gives false confidence. A second pitfall is writing the harness with too few goroutines or too few increments each, which lets a genuinely broken unsynchronized counter pass simply because contention was too rare during that particular run; the fix is not a bigger fixed number but running the concurrent test multiple times in CI (or with -count=N) so an intermittent pass does not get treated as proof of correctness.
Unlock Full Question Bank
Get access to all Test Case Design and Edge Case Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.