Pipeline Testing and Quality Gates Questions
Automated testing wired into the delivery pipeline: test orchestration and execution in CI, quality gates and gating criteria, test environments and test-data management for pipelines, and scaling test infrastructure. Covers deciding what must pass before a change advances and keeping pipeline test stages fast and reliable. Scoped to testing as a delivery-gating concern; test strategy and craft belong to Testing, Quality & Reliability.
What is selective (incremental) test execution, and why would a team adopt it? Describe two concrete heuristics for deciding which tests to run given a set of changed files, their expected accuracy, and the situations where you would need to fall back to running the broader suite.
Sample Answer
Direct answer
Selective (or incremental) test execution runs only the subset of your test suite that could plausibly be affected by a given code change, instead of the whole suite, so a small change gets fast feedback without giving up the safety net for changes it can't touch. Teams adopt it once the full suite is slow enough that running it on every commit is either too costly or too slow to be useful.
Structured elaboration
Two concrete heuristics, from simplest to more accurate:
- Path-based mapping: maintain a table mapping source file paths (or directories/modules) to the tests that cover them, then union the tests mapped from the files in the diff. Simple to build and reason about, but its accuracy depends entirely on how well the mapping is maintained; a change to a widely-shared utility file can map to "run everything," which is fine (safe) but limits how much time it actually saves.
- Tag-based selection: tests are tagged with the feature area or component they exercise, and a change is mapped to tags based on which components it touches (often via a simpler, coarser heuristic than full file-level mapping). Faster to compute but coarser-grained, so it tends to over-select (running more than strictly necessary) rather than under-select, which is the safer direction to err in.
Where these fail and need a fallback to the broader suite:
- A changed file has no entry in the mapping at all (new file, or the mapping fell out of date): the safe response is to widen to a broader default suite, not silently skip testing.
- A change to a shared dependency, build configuration, or something with wide, hard-to-enumerate blast radius, where any static mapping is likely to under-select.
- Generated code or files produced by a build step, where the "real" source of the change is upstream of what the mapping tracks.
Worked example
For a small repository, a workable version: at test-authoring time, each test file declares (via a comment or a lightweight registry) the source modules it exercises. A pre-submit script computes the changed files via a diff, looks each one up in the registry, and unions the matched test files; anything unmapped triggers a "run the full suite" fallback rather than being silently skipped, and the mapping is spot-checked periodically for drift.
Trade-offs & pitfalls
The main risk of selective execution is a stale or incomplete mapping silently under-selecting tests, which looks exactly like a passing pipeline until a real regression slips through; the mitigation is to always fail toward "run more, not less" when the mapping is uncertain, and to periodically validate the mapping against actual coverage data rather than trusting it was ever perfectly accurate.
Design a way to run many integration tests in parallel that each need their own isolated database state, for a suite that also depends on schema migrations. Compare provisioning an isolated database instance per run against isolated schemas/namespaces within a shared instance, and explain how you keep migration versions consistent and avoid expensive full-database snapshots while still guaranteeing isolation.
Sample Answer
Direct answer
For a large number of parallel integration tests that each need isolated database state, isolated schemas within a shared database instance are usually faster to provision and cheaper at scale than a full separate instance per test, provided migration versioning is kept consistent across schemas; a full instance per test buys stronger isolation at meaningfully higher provisioning cost and time, which matters when you're running 100 of them concurrently.
Structured elaboration
- Isolated schema per test (within a shared instance): fast to create (a schema-creation statement is much cheaper than spinning up a whole database process), and a single instance can host many schemas concurrently, which is why this scales better to high parallelism. The trade-off is weaker isolation between tests than a separate instance (they share the instance's resource limits, connection pool, and any instance-level configuration), and migration state has to be tracked per-schema to avoid drift.
- Isolated instance per test: stronger isolation (no shared resource contention at the instance level), but each instance is more expensive and slower to provision, which becomes the bottleneck at high concurrency (100+ simultaneous tests).
- Avoiding costly full snapshots: rather than restoring a full data snapshot into each schema/instance, apply the migration chain to a lightweight template (an empty or minimally-seeded schema) once, then clone that already-migrated template per test run, so migration cost is paid once rather than once per test.
- Guaranteeing reproducibility and isolation together: track the exact migration version each schema/instance was created against, and reject (or auto-remigrate) any that's behind the current version, so a stale schema template doesn't silently produce misleading test results.
Worked example
For 100 parallel PostgreSQL-schema-isolated tests: provision one shared Postgres instance, apply the full migration chain once to a template schema, then clone that template into 100 per-test schemas via a fast CREATE SCHEMA ... FROM TEMPLATE-style operation (or an equivalent schema-copy mechanism), each seeded with the same lightweight baseline data. Each test gets its own connection scoped to its schema, so concurrent tests don't see each other's data, while migration cost was paid exactly once instead of 100 times. This provisioning approach typically completes in well under a second per schema, versus tens of seconds for a full instance-per-test approach, at meaningfully lower disk usage.
Trade-offs & pitfalls
The main risk of schema isolation is under-estimating instance-level resource contention when many tests run concurrently against the same instance (connection pool exhaustion, lock contention on shared instance-level objects); the main risk of full instance-per-test is simply that it doesn't scale to 100 concurrent tests without a slow provisioning bottleneck. The migration-consistency risk applies to both: if the template isn't kept current with the latest migration, every clone silently tests against an outdated schema.
Given a list of tests with their historical average durations and an integer K, implement a deterministic algorithm that partitions the tests into K shards with balanced total runtime. Explain your approach and its time complexity, and describe what changes if some tests must always run together in the same shard (an affinity constraint) or if your duration estimates are noisy.
Sample Answer
Direct answer
Partitioning N tests into K balanced shards is a classic bin-packing problem; a greedy Longest-Processing-Time-first (LPT) heuristic (sort tests by duration descending, then repeatedly assign the next test to whichever shard currently has the smallest total) runs in O(n log n) and is provably within 4/3 of the optimal balance in the worst case, which is why it's the standard default rather than an exact (and much more expensive) optimal solution.
Structured elaboration
Why LPT specifically: sorting largest-first before greedily assigning matters because assigning small items first can leave large items to be dropped into already-unbalanced shards late; assigning large items first, when the most shard-balance-impacting decisions are made, then filling in with progressively smaller items, produces a materially better balance in practice and has the known 4/3-of-optimal worst-case guarantee.
Complexity: O(n log n) for the initial sort, plus O(n log K) for maintaining a min-heap of current shard totals during assignment (each of the n tests does one heap pop and one heap push), so O(n log n) overall since n log n dominates n log K for K much smaller than n.
Handling an affinity constraint (some tests must run together in the same shard): collapse each affinity group into one synthetic "unit" whose duration is the sum of its members, run ordinary LPT on the resulting mix of synthetic units and standalone tests, then expand each synthetic unit back into its member tests within whichever shard it landed in. This keeps the algorithm's shape identical while respecting the constraint.
Handling noisy or drifting duration estimates: since the algorithm only needs relative ordering and reasonably accurate magnitudes, small estimate noise doesn't break it, but persistent drift (a test that's grown much slower over time without the estimate being refreshed) degrades balance quality gradually; the practical mitigation is refreshing duration estimates from recent real run history periodically rather than relying on a stale one-time measurement.
Worked example
```python
import heapq
def shard_tests(tests, k):
shards = [[] for _ in range(k)]
heap = [(0.0, i) for i in range(k)]
heapq.heapify(heap)
for name, dur in sorted(tests, key=lambda t: t[1], reverse=True):
total, idx = heapq.heappop(heap)
shards[idx].append(name)
heapq.heappush(heap, (total + dur, idx))
return shards
```
Run against 8 tests (durations 3.2, 0.5, 2.1, 4.7, 1.0, 3.9, 0.8, 2.4) split into 3 shards, this produced totals of 6.2, 6.0, and 6.4 (a spread of only 0.4 against a mean of ~6.2), and adding an affinity constraint (two specific tests forced into the same shard, using the group-collapsing technique above) still produced every test assigned exactly once with no duplicates or losses, verified in a sandbox run.
Trade-offs & pitfalls
Greedy LPT is fast and good in practice but is not optimal; for a small number of tests with wildly varying durations, an exact solution (or a more expensive local-search refinement on top of the greedy result) can sometimes do meaningfully better, though for the shard counts and test counts typical in CI, the gap rarely justifies the extra complexity. The affinity-group technique is exact for hard "must be together" constraints, but doesn't help with softer preferences (tests that would merely benefit from co-location); those need a different, more heuristic approach.
Explain how consumer-driven contract testing works and how you would wire it into CI. Cover publishing a consumer's expected contract to a broker, running provider verification as part of the provider's own pipeline, and how verification results should gate a deployment or notify an impacted team.
Sample Answer
Direct answer
Consumer-driven contract testing lets a service's consumer define, via its own tests, the exact request/response shape it expects from a provider; that expectation becomes a shareable "contract" artifact, and the provider verifies its real implementation against every consumer's contract in its own pipeline, catching a breaking API change before deployment rather than after two real services are integrated together in a slower, more expensive environment.
Structured elaboration
The workflow, end to end:
- Consumer side: the consumer writes tests against a mock/stub of the provider, and those tests generate a contract file (a record of the exact requests it made and responses it expected) as a side effect of running.
- Publishing: that contract is published to a shared broker (a service that stores and versions contracts, tagged by consumer, provider, and version).
- Provider verification: the provider's own CI pipeline pulls every contract published against it and replays the recorded requests against its real implementation, checking that the real responses match what consumers expect.
- Gating on results: verification results determine whether the provider is safe to deploy a given change; a failing verification blocks the provider's deployment (or at minimum triggers a notification to the affected consumer team) rather than letting an incompatible change ship silently.
The key property that makes this valuable: neither side needs the other's real service running to test against, so both consumer and provider tests run fast, in their own pipelines, at unit-test-like speed, while still catching real cross-service incompatibility that isolated unit tests on either side would miss.
Worked example
A checkout service (consumer) expects an inventory service's (provider) /stock/{sku} endpoint to return a quantity field as an integer. The consumer's contract test records this expectation and publishes it to the broker. When the inventory team later changes that field to a string, their own CI pipeline pulls the checkout team's contract, replays the request against the real (changed) implementation, and fails the verification, catching the break in the inventory team's own pipeline before it ever reaches a shared environment where checkout would have broken.
Trade-offs & pitfalls
Contract testing verifies the shape and presence of what's actually being used, not full behavioral correctness beyond that; it won't catch every possible integration issue (timing, load behavior, business-logic bugs a matching shape can still hide). It complements, rather than replaces, some level of real integration or end-to-end testing for the flows where deeper behavioral correctness genuinely matters.
You run a nightly matrix of thousands of integration tests and must cut the cost by 50% while preserving at least 95% of the suite's historical bug-detection capability. Propose an optimization plan (prioritization, sampling, parallelization, caching, incremental runs), the metrics you would track, and how you would run the change as an experiment before fully committing to it.
Sample Answer
Direct answer
To cut a nightly test bill by 50% while preserving at least 95% of historical bug-detection capability, the right approach is measuring which specific tests have actually caught real regressions historically, prioritizing keeping those at full frequency, and reducing cost on the rest through a combination of sampling, smarter parallelization, and incremental (change-based) execution, validated as an experiment before fully committing rather than assumed to work.
Structured elaboration
Optimization plan:
- Measure current bug-detection contribution per test (or per test group): using historical data, identify which tests have actually caught real regressions versus which have never failed for a genuine reason; this is the foundation the rest of the plan is built on, since you can't safely cut cost without knowing what you'd be cutting.
- Prioritize based on that data: keep the highest-value tests (by historical detection contribution) running at full frequency; candidates for cost reduction are lower-value tests, especially ones that have never caught a genuine regression in the observed history.
- Parallelization and caching: ensure the suite is using its compute efficiently in the first place (proper sharding balance, dependency caching) before cutting scope, since inefficient parallelization can itself be a large, easily-recovered cost with zero coverage trade-off.
- Sampling for lower-value tests: rather than running every lower-value test every night, run a rotating sample (a different subset each night) so coverage is maintained over a window even if not every single night.
- Incremental/change-based execution for parts of the suite where it's safe: skip re-running tests unrelated to what's changed since the last run, where a reliable change-impact mapping exists.
- Metrics to track: total compute cost, and (critically) an ongoing measurement of actual bug-detection rate post-change, not just at the point of the initial decision, so a slow degradation in detection capability is caught rather than assumed away.
- Experimental rollout: run the reduced-cost configuration alongside the full nightly suite for an evaluation period (the same paired-comparison approach used for validating any test-suite change), confirming the 95% detection-rate target actually holds before fully retiring the more expensive configuration.
Worked example
Historical analysis of a year of nightly runs showed roughly 30% of the suite had never caught a genuine regression, while a specific 15% of tests accounted for the large majority of real catches. The team kept that high-value 15% running every night at full priority, applied rotating sampling to the historically-unproductive 30% (each running roughly once a week instead of nightly), and left the remaining 55% running nightly but with improved sharding balance that cut its wall-clock cost. Running both the old and new configurations in parallel for six weeks showed the new configuration caught 96% of the regressions the old one did, clearing the 95% bar, at roughly half the total compute cost.
Trade-offs & pitfalls
The critical discipline this plan depends on is measuring bug-detection contribution from real historical data rather than guessing which tests are "probably not that valuable"; cutting based on intuition alone risks silently removing exactly the test that would have caught the next real regression, which is precisely the failure mode a rigorous, measured, and experimentally-validated approach is designed to avoid.
Unlock Full Question Bank
Get access to all Pipeline Testing and Quality Gates interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.