Test Case Design and Edge Case Analysis Questions
Systematically deriving the cases, inputs, and conditions most likely to expose defects. Covers formal test-design techniques (equivalence partitioning, boundary value analysis, decision tables, state transitions, and pairwise/combinatorial design) and writing clear, maintainable test cases with documented expected results. Also covers the edge-case mindset: boundary conditions, invalid and unexpected inputs, corner cases, and the attention to detail that anticipates failures when validating complex behavior.
A distributed cache sometimes returns stale data after failover. Enumerate edge cases that can cause cache inconsistency (TTL, replication lag, race conditions, read-after-write, client-side caching, clock skew). For each, propose a concrete mitigation and discuss trade-offs.
Sample Answer
Direct answer
The failover scenario surfaces six distinct staleness mechanisms, and treating "the cache returned stale data" as one bug is the wrong turn: TTL lag, replication lag, races, read-after-write gaps, client-side caching, and clock skew each have a different root cause, so each needs its own mitigation with its own cost, not one generic fix.
Structured elaboration
| Edge case | Mechanism | Mitigation | Trade-off |
|---|---|---|---|
| TTL (time-to-live) expiry lag | An entry is served past its intended freshness window because passive expiry has not fired yet | Shorter TTL plus proactive background refresh, or versioned keys tied to a monotonic write counter instead of pure elapsed time | Shorter TTL raises backend load; a background refresh path adds its own complexity and can itself race with a write |
| Replication lag | Failover promotes a replica that has not caught up, and the cache still reflects the old primary's state | Track a replication-lag watermark and refuse to serve or cache reads from a replica beyond a lag threshold, or tag cached values with the source's log position so staleness is provable | Watermark checks add a round trip; an overly strict threshold increases read unavailability during ordinary lag |
| Race conditions | Concurrent cache writes or invalidations, an older value can overwrite a newer invalidation during a thundering-herd repopulation | Single-flight (lock-per-key) cache population, or compare-and-set the entry using a write version so an older write cannot clobber a newer one | Locking adds latency and, if only in-process, does nothing across cache nodes; a distributed lock adds an external dependency |
| Read-after-write | A client writes, then immediately reads and gets a stale cached copy because invalidation has not propagated yet | Read-your-writes via a session-scoped bypass to the source of truth for a bounded window after a write, or synchronous invalidate-before-ack | Session affinity adds infrastructure; synchronous invalidation slows every write and adds a new failure mode if the invalidation call itself fails |
| Client-side caching | Distributed clients (SDKs, browsers, HTTP caches) hold local copies the server-side invalidation never reaches | Short client TTLs with versioned ETag revalidation, or a push-based invalidation channel for high-value keys | Push invalidation only reaches connected clients and is more infrastructure; ETag revalidation still costs a round trip on every cache hit |
| Clock skew | TTL or version comparisons that rely on unsynchronized wall clocks expire early or late, or resolve write ordering incorrectly | Prefer a monotonic logical version/sequence number handed out by the source of truth for correctness-critical ordering; keep wall-clock TTL bounded by an NTP-tolerance window, not used for ordering | Logical versioning requires the source of truth to hand out monotonic versions, an added contract; an NTP-bounded window still leaves a gap where two close writes can resolve incorrectly |
The eventual-consistency framing maps directly onto two rows above rather than adding a new one: a "stale read" under eventual consistency is the read-after-write row generalized past a single writer (any reader, not just the writer itself, can observe the old value before propagation completes), and a "lost update" is the race-condition row, a concurrent write's effect gets silently overwritten. Naming this connection explicitly is the point: the same mitigations (session-scoped bypass, version-based compare-and-set) that fix read-after-write and races for a single failover also fix the general eventual-consistency versions of the same problems.
Worked example
A concrete failover timeline: at t=0ms a write commits to the primary in region 1. At t=50ms, a client in region 2 reads through a cache node whose backing replica is 120ms behind the primary (replication lag), well inside a 5000ms TTL, so passive TTL expiry has no chance to catch it either. A naive cache returns the stale pre-write value. The replication-lag mitigation above (tagging the cached value with the source replica's log position, and rejecting a read whose evidenced position is behind a required floor) is the row that specifically prevents this: at t=50ms the required floor (the primary's position at write time) is not yet met by the lagging replica, so the read is forced to the source of truth instead of the stale cache, exactly the case a TTL-only design cannot catch because 50ms is nowhere near a 5000ms TTL.
Trade-offs and pitfalls
- Every mitigation above trades latency, availability, or infrastructure complexity for a stronger consistency guarantee; there is no free fix, and a senior answer states which axis it is spending, not just which mitigation it picked.
- The common wrong turn is applying the strongest guarantee to every key uniformly. A financial-balance-style key needs the read-your-writes and version-based mitigations; a content-description key is often fine with plain TTL-based eventual consistency, and over-engineering the latter wastes latency budget for no real benefit.
- Failover itself amplifies the race-condition row specifically: a failover creates a burst of first-touch cache misses (a thundering herd) right when concurrent repopulation is most likely, so a test suite that only exercises races under steady-state traffic will miss the case that actually matters, testing must specifically include failover-timed concurrent access.
Given POST /orders accepting JSON {user_id, items: [{sku, qty}], total_amount}, enumerate edge cases that should map to specific HTTP responses (e.g., missing fields, negative qty, mismatched total, unauthorized user). For five distinct edge cases, state the expected HTTP status code and a one-line test assertion that verifies the behavior.
Sample Answer
Direct answer
Mapping edge cases to specific status codes forces a precision that a generic "assert it's not a success response" test case does not: a missing field, a semantically invalid value, and an authorization failure are all errors but deserve visibly different codes, and a test suite that only asserts "not 200" cannot catch a regression that silently swaps one error type for another, for example returning 400 where the API contract actually promises 422 for a specific case.
Structured elaboration
Convention used below, stated explicitly since status-code choice for validation errors is a design decision, not a single universal standard: 400 Bad Request for a structurally malformed or incomplete request; 401 Unauthorized for a missing or invalid credential (the caller has not proven who they are); 403 Forbidden for a caller who IS authenticated but is not permitted to perform this specific action (a distinction commonly gotten backwards in practice); 422 Unprocessable Entity for a request that is well-formed and type-valid but semantically wrong. Note that 422 is a widely-adopted REST convention rather than part of the original core HTTP status code set, so an API's actual choice between 400 and 422 for "the value is wrong" cases is itself a decision that needs to be made explicitly and applied consistently, not assumed.
Worked example: five distinct edge cases
| # | Edge case | Expected status | One-line test assertion | Rationale |
|---|---|---|---|---|
| 1 | items array is missing entirely from the payload | 400 | assert response.status_code == 400 | The request is structurally incomplete, not merely semantically wrong |
| 2 | An item's qty is negative | 422 | assert response.status_code == 422 | The payload is well-formed and type-valid; the value itself is semantically invalid |
| 3 | total_amount does not equal the server-computed sum of price * qty across items | 422 | assert response.status_code == 422 | A well-formed payload whose internal values are mutually inconsistent, a semantic defect rather than a missing/malformed field |
| 4 | Caller's authenticated identity does not match the user_id in the payload, and the caller holds no elevated role permitting that | 403 | assert response.status_code == 403 | The caller IS authenticated (a valid token was presented); they are simply not permitted to place this specific order, which is what distinguishes 403 from 401 |
| 5 | An item's sku (stock keeping unit, a product identifier) does not correspond to any known product | 422 | assert response.status_code == 422 | The request as a whole successfully addresses the order-creation resource; the unknown sku is a semantic defect within the payload, not a failure to locate the primary resource the URL points at, which is the conventional meaning of 404 |
Case 5 is genuinely a design decision, not the single correct answer: a team framing an unknown sku as "the referenced sub-resource was not found" could reasonably choose 404 instead, and that choice, once made, is exactly the kind of thing this table format exists to pin down and test explicitly rather than leave ambiguous per engineer.
Trade-offs & pitfalls
The 400-versus-422 boundary is a genuinely contested convention across the industry, so the real risk is not picking one in isolation, it is picking INCONSISTENTLY across similar endpoints within the same API, which breaks any client-side error-handling logic that dispatches behavior on the specific status code returned. A team that skips this design decision often lets it fall out ad hoc, one engineer per endpoint, producing an API where structurally identical validation failures return 400 on one endpoint and 422 on another; a test suite that asserts the SPECIFIC code, as the table above does, is exactly what catches that drift, while a test that only asserts status_code >= 400 would pass regardless. Confusing 401 and 403 is a second common and consequential mistake: returning 403 for "no credential was presented at all" (which should be 401) leaks information about resource existence to unauthenticated callers that 401 does not, while reserving 403 specifically for "we know who you are, and you may not do this" keeps that distinction meaningful for clients and for security review.
You are QA for a service that sorts large datasets and returns paged results (millions of records). What edge cases and test scenarios would you create to validate correctness and robustness: duplicate sort keys, stable vs unstable sorting, null/absent keys, comparator exceptions, inconsistent ordering across pages, serialization differences, and memory pressure? Define unit, integration, and end-to-end tests, and describe data generation approaches for large and pathological datasets.
Sample Answer
Direct answer
For a sort-and-page service over millions of records, the edge cases split into two families: correctness of the ordering itself (duplicate sort keys, stable versus unstable sorting, null/absent keys, comparator exceptions) and correctness of the PAGINATION on top of that ordering (inconsistent ordering across pages, serialization differences, memory pressure at scale). The single most dangerous edge case, and the one a code review will not catch by inspection, is heavy duplication in the sort key combined with offset-based paging: without an explicit tiebreaker, the exact same dataset can silently drop or duplicate rows across a page walk, even though each individual page looks correct in isolation.
Structured elaboration
Unit-level (comparator/ordering) tests
- Duplicate sort keys: many rows sharing the same key value. Stability matters here: a STABLE sort preserves the original relative order of equal-key rows across repeated runs on the same input; an UNSTABLE sort (a plain quicksort, or many distributed-engine parallel sorts) does not guarantee this, and the client-visible symptom is that two consecutive page requests for the same query can return duplicated-key rows in a DIFFERENT relative order, which breaks the natural expectation that "page 2 continues exactly where page 1 left off."
- Null/absent keys: decide and test an explicit placement rule (nulls first or nulls last), since most SQL and sort libraries default to a specific, sometimes surprising, placement that differs across databases and languages.
- Comparator exceptions: a custom comparator that assumes a specific type (e.g. always numeric) will raise on a heterogeneous or malformed value; test that the exception surfaces as a clear, typed validation error at ingestion time, not a raw crash deep inside the sort call during a customer-facing request.
Integration-level (pagination-consistency) tests
- Inconsistent ordering across pages: walk ALL pages of a query, concatenate every returned ID, and assert the concatenated set exactly equals the known full ID set for that dataset (no duplicates, no omissions), rather than eyeballing individual pages.
- Serialization differences: if the API paginates via an opaque cursor token, confirm the cursor's meaning is stable across a service redeploy (a cursor encoding a raw row offset breaks if the underlying storage is repartitioned; a cursor encoding a stable sort-key value plus a tiebreaker id survives it).
- Memory pressure: for millions of records, assert the service streams/paginates from the underlying store rather than materializing the full sorted result in memory before paging; a targeted test can seed a dataset sized to exceed a deliberately small memory limit in a test environment and confirm the service still completes without an out-of-memory failure, rather than asserting a specific memory number.
End-to-end tests
- A full client-simulated walk (as in the worked example below) over a realistic dataset shape, verifying the concatenated, deduplicated result matches an independent full-sort oracle.
Data generation for large and pathological datasets
- Generate at multiple scales (thousand, million-row smoke tests) with a FIXED seed for reproducibility.
- Deliberately skew the generator toward pathological shapes: a large fraction of rows sharing one sort-key value (the duplicate-key stress case), a long tail of unique keys, explicit NULL injection at a known rate, and adversarial comparator inputs (mixed types if the schema technically allows it).
Worked example (executed): duplicate-sort-key pagination bug, caught and fixed
import random
random.seed(42)
rows = [{"id": i, "score": i % 4} for i in range(500)] # 500 rows, only 4 distinct sort-key values
def unstable_source_page(all_rows, offset, limit, rng):
grouped = {}
for r in all_rows: grouped.setdefault(r["score"], []).append(r)
ordered = []
for k in sorted(grouped):
bucket = grouped[k][:]; rng.shuffle(bucket) # ties reordered on every call, simulating a real DB with no tiebreaker
ordered.extend(bucket)
return ordered[offset: offset + limit]
def stable_source_page(all_rows, offset, limit):
ordered = sorted(all_rows, key=lambda r: (r["score"], r["id"])) # explicit tiebreaker
return ordered[offset: offset + limit]
def walk_all(page_fn, total, limit):
seen = []
for offset in range(0, total, limit):
seen.extend(r["id"] for r in page_fn(offset, limit))
return seen
LIMIT = 37
rng = random.Random(7)
unstable_ids = walk_all(lambda o, l: unstable_source_page(rows, o, l, rng), 500, LIMIT)
stable_ids = walk_all(lambda o, l: stable_source_page(rows, o, l), 500, LIMIT)
expected = set(r["id"] for r in rows)
print(len(unstable_ids), len(set(unstable_ids)), len(unstable_ids) - len(set(unstable_ids)), len(expected - set(unstable_ids)))
print(len(stable_ids), len(set(stable_ids)), len(stable_ids) - len(set(stable_ids)), len(expected - set(stable_ids)))
Executed output: the unstable (no-tiebreaker) walk fetched 500 rows total but only 353 unique ids, 147 duplicated across page boundaries and 147 missing entirely from the walk, on the exact same 500-row dataset. The tiebreaker-stabilized walk fetched 500 rows, 500 unique, 0 duplicates, 0 missing. This is not a contrived pathological case, it is what happens on ANY dataset where the sort key alone does not uniquely order the rows and pagination crosses a tie boundary.
Trade-offs & pitfalls
The most common mistake is trusting that "the database sorts consistently" without an explicit tiebreaker column; many query engines make NO ordering guarantee for tied rows across separate query executions, especially under parallel execution plans, so relying on incidental stability is a latent bug that a small-scale manual test will not surface (the bug only shows up once ties outnumber a single page, which small test datasets rarely do). A second pitfall is testing pagination correctness only at the unit level (one page looks right) without the end-to-end concatenate-and-diff check across the FULL walk, which is the only way this class of bug is actually visible. Third, testing memory pressure by asserting a fixed byte count is fragile across environments; asserting that the operation SUCCEEDS under an artificially constrained limit, and that the service degrades to a documented, bounded failure mode (not an unbounded materialization) rather than crashing, is the more durable test design.
Discuss edge cases when using floating-point types for money in backend systems. Propose storage and computation strategies (integer cents, Decimal/BigDecimal), and list tests that verify correct rounding, accumulation across many transactions, and cross-service serialization/deserialization where languages differ (e.g., Python Decimal to Java BigDecimal).
Sample Answer
Direct answer
Storing money as a native floating-point type is unsafe because binary floating point cannot exactly represent most decimal fractions (the same root cause as the classic 0.1 + 0.2 != 0.3 trap), so backend systems should store and compute money either as integer minor units (cents) or as an arbitrary-precision decimal type (Python's Decimal, Java's BigDecimal), never as float/double.
Structured elaboration: the two safe strategies
| Strategy | How it works | Trade-off |
|---|---|---|
| Integer cents | Store $19.99 as the integer 1999; all arithmetic is integer arithmetic | Fast and exact, but every value needs an explicit, consistent 'divide by 100 to display' convention, and multi-currency systems need to track the minor-unit divisor per currency (not all currencies use 2 decimal places, e.g. JPY has 0) |
| Decimal/BigDecimal | Store an explicit base-10 decimal representation with defined precision and rounding rules | Exact for decimal arithmetic (no binary-fraction error), but slower than integer math and requires every language/service in the pipeline to use an equivalent decimal type consistently |
Worked example: tests to verify correct behavior
- Rounding test: compute a 7.25% tax on $19.99:
19.99 * 1.0725 = 21.439275exactly (verified:Decimal('19.99') * Decimal('1.0725')returnsDecimal('21.439275')), which rounds unambiguously to $21.44 under either round-half-up or round-half-to-even, since it is not a tie value. To specifically distinguish round-half-up from round-half-to-even (banker's rounding), the test needs a value that lands exactly on a half-cent tie, e.g. $21.425: round-half-up produces $21.43 (verified:Decimal('21.425').quantize(Decimal('0.01'), rounding=ROUND_HALF_UP) == Decimal('21.43')), while round-half-to-even produces $21.42, rounding to the nearest EVEN cent (verified:... rounding=ROUND_HALF_EVEN) == Decimal('21.42')). A rounding test suite needs both kinds of case: a routine value (to catch a broken rounding implementation generally) and an exact-tie value (to catch the wrong TIE-BREAKING rule specifically). - Accumulation test: sum 10,000 transactions of $0.01 using the system's chosen representation and assert the total is exactly $100.00. This is precisely the kind of test that would FAIL if float were used (accumulated rounding error across many additions), and pass reliably under integer-cents or Decimal.
- Cross-service serialization test: serialize a Decimal value from a Python service (e.g.
Decimal('19.99')) to JSON, deserialize it in a Java service into a BigDecimal, and assert the value survives the round trip EXACTLY. This matters because JSON has no native decimal type; a naive serializer that converts through a JSON number (which many JSON libraries parse as a double) can silently reintroduce floating-point error at the service boundary even if both services individually use a correct decimal type internally, so the test must exercise the actual serialization format used (e.g. a string-encoded decimal, not a bare JSON number) rather than trusting each service's internal type alone.
Trade-offs & pitfalls
A frequent partial-fix mistake is using Decimal/BigDecimal inside a service but still passing values through a numeric (not string) JSON field at the API boundary; this silently reintroduces the float-precision problem the moment ANY intermediate JSON parser in the pipeline (a proxy, a logging middleware, a client SDK) treats that field as a double, which is why the cross-service serialization test above is not optional if any part of the pipeline crosses a JSON boundary. A second pitfall is picking integer cents for a system that later needs to support currencies with different minor-unit conventions (JPY, or a hypothetical fractional-cent loyalty-points system) without having designed for a configurable minor-unit divisor from the start.
A DBA proposes adding a composite index to speed reads. Explain potential edge cases introduced by indexing (index bloat, stale statistics, different query plans, higher write latency, and the risk of index corruption). Design regression tests and benchmarks that ensure query correctness and no unacceptable performance regressions under high write load.
Sample Answer
Direct answer
A composite index changes five things at once: query plans (the optimizer may now pick a different, not always better, plan), write latency (every insert/update/delete must also maintain the index), storage (index bloat from dead tuples and page fragmentation), statistics freshness (a stale statistics snapshot can make the optimizer trust an index that no longer reflects the data's actual distribution), and, rarely, index integrity itself (corruption from a crash mid-write or a storage-layer bug). The test plan has to cover correctness (does the query still return the right rows) and performance (does write latency stay acceptable under load) as two SEPARATE concerns, because an index can pass one and silently fail the other.
Structured elaboration: each risk, why it happens, and how to design a test for it
| Risk | Mechanism | Regression test design |
|---|---|---|
| Different query plans | The optimizer now has a new access path (index scan/seek) alongside the old one (sequential/table scan) and chooses based on estimated selectivity | Capture the EXPLAIN plan for the target query BEFORE adding the index (baseline) and again AFTER, on a representative data snapshot; assert the AFTER plan uses the new index for the intended query shape, and separately assert a DIFFERENT, unrelated query's plan is UNCHANGED (a composite index can accidentally become "attractive" to the optimizer for a query it wasn't designed for, sometimes for the worse) |
| Higher write latency | Every INSERT/UPDATE/DELETE that touches an indexed column must also update the index's B-tree (or equivalent) structure, and updates to indexed columns can trigger index-page splits | A write-path test that inserts/updates a fixed, seeded batch of rows before and after the index exists, and asserts the operation still COMPLETES CORRECTLY (row counts, constraint checks) under a realistic concurrent write load; the actual latency number is environment-dependent and does not belong in the test as a fixed millisecond assertion, the test should assert correctness under load and flag a latency REGRESSION relative to a same-environment baseline run, not an absolute number |
| Index bloat | Frequent updates/deletes leave dead index entries that periodic vacuum/reindex operations are meant to reclaim; bloat accumulates between maintenance windows | A test that runs a representative churn workload (mixed insert/update/delete) against a copy of the schema, then checks the index's reported size/page count grows sub-linearly relative to live row count after a maintenance cycle runs, not that it stays flat forever without maintenance |
| Stale statistics | The query planner relies on a statistics snapshot (histogram of value distribution, cardinality estimates) that is only refreshed periodically or on an analyze/vacuum trigger; a bulk load or skewed insert pattern can invalidate it before the next refresh | A test that seeds a skewed distribution (e.g. 90% of rows sharing one value in an indexed column) WITHOUT triggering a manual statistics refresh, then confirms the query still returns correct results (correctness is unaffected by stale stats) even though the CHOSEN PLAN might degrade to a less optimal one; a separate test confirms an explicit statistics refresh restores the better plan |
| Index corruption | Rare, but real: a crash during an index page write, or a bug in a specific database version, can leave an index internally inconsistent, returning wrong or missing rows silently until a repair/reindex runs | An oracle-based correctness test: run the SAME query with the index enabled and with the index disabled (an optimizer hint, or a temporary index drop on a copy) and assert the RESULT SETS are identical; a corrupted index that still executes without erroring is exactly the failure mode this oracle catches, since a plan-only test (checking EXPLAIN output) would never notice |
Worked example: the correctness oracle in concrete steps
- Load a seeded dataset (fixed, deterministic, e.g. a specific 100k-row fixture with a known skew and a known set of edge values including NULLs in the indexed columns).
- Run the target query with the composite index present; capture the full result set as set A.
- Run the IDENTICAL query against the same data with the index dropped (or with the optimizer forced to ignore it), forcing a table scan; capture the result set as set B.
- Assert
A == Bas sets (order-independent unless the query has an explicitORDER BY, in which case assert exact sequence equality). - Repeat with NULL values specifically in the composite index's columns, since composite indexes often have surprising NULL-handling semantics (a row with NULL in a leading indexed column may or may not be included in an index range scan depending on the database and query shape), and NULL handling is exactly the kind of edge case a plan-diff alone would not surface.
Trade-offs & pitfalls
The most common mistake is validating ONLY that the new plan looks faster in EXPLAIN output and skipping the result-set equivalence check entirely; a plan that LOOKS more efficient can still be wrong if the index's column order interacts badly with the query's filter/sort order (a classic composite-index pitfall: an index on (status, created_at) does not efficiently serve a query that filters on created_at alone, and may not even get chosen, while a query that filters on status and sorts by created_at benefits enormously; testing only one query shape misses the other). A second pitfall is running the write-latency regression test against an idle database instead of a realistic CONCURRENT write load, since index maintenance cost under lock contention behaves very differently from a single isolated writer. A third pitfall, specific to the "benchmark" half of this question, is reporting an absolute wall-clock number ("writes now take 8ms") as the pass/fail bar; that number is only meaningful relative to a same-environment, same-load baseline captured in the same test run, an absolute millisecond target silently rots as hardware and load patterns change.
Unlock Full Question Bank
Get access to all 47 Test Case Design and Edge Case Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.