Test Levels and the Test Pyramid Questions
How unit, integration, component, end-to-end and contract tests fit together and where each provides the most value. Covers the test pyramid and the competing shapes proposed against it (the testing trophy and the honeycomb), multi-layer test architecture, contract testing as the seam between services, choosing the right level to catch a given class of defect cheaply, and what to run per commit versus per release. Includes the cost and confidence trade-offs between fast low-level tests and slower, broader system tests. The scope is which level a test belongs at and why. Deciding how much to invest in testing and where to prioritize under time pressure is covered separately.
Many teams cite a '70/20/10' (unit/integration/end-to-end) or similar rule-of-thumb ratio for test distribution. Explain the rationale behind such ratios and the assumptions they make, and describe concrete scenarios where you would deviate from this guideline and why.
Sample Answer
A "70/20/10" (or similarly-shaped) ratio is a useful DEFAULT, not a target to hit for its own sake: it encodes the assumption that most of a typical system's defect risk lives in logic a unit test can isolate cheaply, a smaller amount lives in how components wire together, and only a small remainder needs the expensive proof that the whole assembled system works.
What the ratio assumes
- Most bugs are logic bugs, not integration bugs. If your system's core complexity is business logic (pricing rules, calculations, state machines), this holds well, and a heavy unit-test base pays off directly.
- Integration points are relatively few and stable. The ratio assumes there aren't so many service-to-service or component-to-component seams that integration testing alone would need to be a much larger share to give adequate confidence.
- The team can afford SOME slow, broad tests, but not many. The "10%" isn't zero: it assumes a small curated end-to-end layer is enough to catch whole-system wiring problems, which is only true if your riskiest journeys are few in number.
- Cost scales the way the model assumes. The whole justification for weighting the base so heavily rests on unit tests being drastically cheaper than integration and end-to-end tests; if that cost gap narrows (fast, hermetic integration tests via lightweight containers, for instance - hermetic meaning self-contained: no real network calls or shared external state, so the same test run always gets the same result), the "right" ratio shifts too.
When to deviate, and why
- A thin orchestration service whose logic is almost entirely "call service A, then call service B" has very little unit-testable logic of its own; here the risk genuinely concentrates at the integration boundary, so a heavier integration-test share (closer to something like 40/50/10) reflects reality better than forcing a 70% unit-test floor onto code that barely has any unit-testable branches.
- A frontend-heavy, interaction-driven product where most of the risk is "does clicking through this actually work for a user" may reasonably lean toward more integration-style component tests (a component test renders one UI component together with its real child components but fakes the network or backend, which is what separates it from a unit test, which isolates everything) - closer to the testing-trophy shape (an alternative to the pyramid that keeps a small unit-test base but makes these broader, more realistic tests the largest layer) - than a strict 70/20/10 pyramid, because the thing most likely to break is how components interact on screen, not isolated pure functions.
- A system with very few, very high-stakes end-to-end journeys (payment settlement, safety-critical workflows) may justify a larger-than-10% end-to-end share for those specific journeys, even while the rest of the system keeps the standard ratio, because the cost of an undetected wiring bug there is disproportionately high.
Trade-offs and pitfalls
The most common misuse of this heuristic is treating the numbers as a scorecard: chasing "70% unit tests" by writing large numbers of low-value unit tests for trivial getters, while under-investing in the harder work of a few well-chosen integration and end-to-end tests for the journeys that actually carry business risk. The ratio should be a lagging description of where your test investment naturally lands once you've tested the RIGHT things at each level, not a quota to satisfy directly.
Explain the differences between unit tests, integration tests, and end-to-end tests. For each level, give two concrete examples (functions, modules, services, or UI flows), state when it should run (on a pull request, at merge, or nightly) and its typical execution speed, and discuss the typical maintenance cost and failure modes. Conclude with the concrete trade-offs between speed, coverage, and flakiness for a web application.
Sample Answer
A unit test exercises a single function or class in complete isolation: every dependency is faked, stubbed, or simply absent, so the test runs in microseconds and its failure points at exactly one piece of logic. An integration test exercises how two or more real components work together, most often your code against a real (or near-real) database, queue, or external service, so it catches wiring and serialization bugs a unit test cannot see. An end-to-end test drives the system the way a real client would, through its actual entry point (an HTTP call, a UI click), with nothing faked, so it is the only level that proves the whole assembled system actually works.
What each level covers, with two concrete examples per level
| Level | Two example targets | Runs on | Speed | Maintenance cost | Failure mode it's good at catching |
|---|---|---|---|---|---|
| Unit | (1) a pure function, e.g. a discount calculator; (2) a class method with its collaborators mocked, e.g. an order-validation method tested with a fake repository | Every commit, on save | Microseconds to low milliseconds | Low, unless over-mocked | Wrong business logic, missed edge cases |
| Integration | (1) your code against a real database, e.g. does saving an order persist the right row; (2) your code against one real external service, e.g. a payment client against that gateway's sandbox | Pull request / merge | Tens of milliseconds to a few seconds | Medium: schema and API drift break these | Wiring bugs: wrong SQL, wrong serialization, a contract mismatch |
| End-to-end | (1) a full UI flow, e.g. add-to-cart through order confirmation in a real browser; (2) a full API flow, e.g. a real HTTP client driving create-then-fetch against the live server with nothing faked | Nightly or pre-release | Seconds to minutes | High: brittle to unrelated UI or infra changes | Environment and integration issues that only appear when everything runs together |
The QA-engineer angle on unit tests is collaborative, not just "who writes them": developers usually author the unit tests since they know the implementation, but QA should read them during review to spot missing edge cases the implementer didn't think of, and QA is often the one who notices a bug that unit tests theoretically should have caught but didn't (a coverage gap, not a process failure). "System testing" is a related but distinct idea: it validates the whole assembled system as one unit against requirements, similar in spirit to end-to-end testing, but typically owned by QA and run just before release with attention to environment parity and realistic test data, whereas end-to-end testing is often owned by whoever automates the user-facing flow and runs continuously.
The concrete tools differ by stack but the pattern holds everywhere: JUnit or pytest for the unit layer, pytest combined with testcontainers (spinning up a real, disposable database or service in a container) for the integration layer, and Selenium or Playwright for the end-to-end layer, with the same mock-vs-real-service decision applying at the integration boundary regardless of which tools you pick: mock a dependency when you're testing YOUR handling logic, use the real (or containerized) dependency when you're testing that the wiring itself is correct. For a payments microservice specifically, this maps onto where each level runs in the deployment pipeline: unit tests run locally on every save and in CI on every commit; integration tests run in CI against a containerized database and a sandboxed payment gateway; end-to-end tests run in a staging environment before a production deploy, and a small smoke subset may re-run immediately after reaching production to confirm the live deploy itself is healthy. This progression, more tests locally and in CI, fewer in staging, fewer still in production, is also how the pyramid should guide the ALLOCATION of engineering effort: invest the majority of new test-writing time at the level closest to the developer's own commit, not because the higher levels don't matter, but because that's where a fixed hour of effort buys the most coverage per dollar of CI time and developer attention.
Worked example: the same business rule, tested three ways
The clearest way to see the boundary is to test the identical rule at all three levels and watch what each level can and cannot catch.
# calculate_discount is pure logic: no I/O, so it belongs at the unit level.
def calculate_discount(price: float, tier: str) -> float:
if price < 0:
raise ValueError("price must be non-negative")
rate = {"standard": 0.0, "silver": 0.05, "gold": 0.15}.get(tier)
if rate is None:
raise ValueError(f"unknown tier: {tier}")
return round(price * (1 - rate), 2)
# UNIT TEST: no database, no network. Executed directly.
def test_calculate_discount_unit():
assert calculate_discount(100, "standard") == 100.0
assert calculate_discount(100, "silver") == 95.0
assert calculate_discount(100, "gold") == 85.0
Executed output: UNIT level: 4/4 assertions passed (pure function, no I/O, <1ms) (all four, including the negative-price ValueError case).
import sqlite3
class OrderRepository:
def __init__(self, conn):
self.conn = conn
self.conn.execute(
"CREATE TABLE IF NOT EXISTS orders (id INTEGER PRIMARY KEY, price REAL, tier TEXT, total REAL)"
)
def save(self, price, tier):
total = calculate_discount(price, tier)
cur = self.conn.execute(
"INSERT INTO orders (price, tier, total) VALUES (?, ?, ?)", (price, tier, total)
)
self.conn.commit()
return cur.lastrowid
# INTEGRATION TEST: a REAL SQLite database, catching serialization/wiring the unit test cannot see.
def test_order_repository_integration():
conn = sqlite3.connect(":memory:")
repo = OrderRepository(conn)
order_id = repo.save(200, "gold")
row = conn.execute("SELECT total FROM orders WHERE id = ?", (order_id,)).fetchone()
assert row[0] == 170.0
Executed output: INTEGRATION level: repository round-trip through real SQLite passed: {'id': 1, 'price': 200.0, 'tier': 'gold', 'total': 170.0}.
For the end-to-end level, the same repository was wired behind a real HTTP handler and hit with an actual socket-level POST followed by a GET, using Python's built-in http.server and urllib.request, no mocking anywhere in the path:
import json, threading, urllib.request, urllib.error
from http.server import BaseHTTPRequestHandler, HTTPServer
class Handler(BaseHTTPRequestHandler):
def log_message(self, format, *args):
pass
def do_POST(self):
length = int(self.headers.get("Content-Length", 0))
body = json.loads(self.rfile.read(length))
try:
order_id = repo.save(body["price"], body["tier"])
total = calculate_discount(body["price"], body["tier"])
except ValueError as e:
self.send_response(400)
self.end_headers()
self.wfile.write(json.dumps({"error": str(e)}).encode())
return
self.send_response(201)
self.end_headers()
self.wfile.write(json.dumps({"id": order_id, "total": total}).encode())
def do_GET(self):
order_id = int(self.path.rsplit("/", 1)[-1])
row = conn.execute("SELECT id, total FROM orders WHERE id = ?", (order_id,)).fetchone()
self.send_response(200)
self.end_headers()
self.wfile.write(json.dumps({"id": row[0], "total": row[1]}).encode())
# E2E TEST: real HTTP socket, real server thread, nothing faked.
def test_checkout_end_to_end():
server = HTTPServer(("127.0.0.1", 0), Handler)
port = server.server_address[1]
threading.Thread(target=server.serve_forever, daemon=True).start()
req = urllib.request.Request(
f"http://127.0.0.1:{port}/orders",
data=json.dumps({"price": 50, "tier": "silver"}).encode(),
headers={"Content-Type": "application/json"}, method="POST",
)
created = json.loads(urllib.request.urlopen(req).read())
fetched = json.loads(urllib.request.urlopen(f"http://127.0.0.1:{port}/orders/{created['id']}").read())
assert fetched == {"id": created["id"], "total": 47.5}
bad = urllib.request.Request(
f"http://127.0.0.1:{port}/orders",
data=json.dumps({"price": 50, "tier": "platinum"}).encode(),
headers={"Content-Type": "application/json"}, method="POST",
)
try:
urllib.request.urlopen(bad)
assert False, "expected HTTPError"
except urllib.error.HTTPError as e:
assert e.code == 400
server.shutdown()
Executed output: E2E level (valid order): POST+GET over real HTTP socket returned {'id': 1, 'total': 47.5} and E2E level (invalid tier): POST over real HTTP socket returned status 400. That is precisely the trade-off: the unit test told us the discount math is right in under a millisecond; the end-to-end test told us the whole pipe, JSON serialization, routing, and the network stack included, actually delivers that correct math to a real client, at the cost of running a live server and a real socket for the one test.
Trade-offs and pitfalls
The pyramid shape follows directly from this example: you want most of your assertions at the level that is cheapest to run and most precise about what broke, which is the unit level, and you want just enough integration and end-to-end coverage to prove the pieces are wired correctly, because that proof is expensive and comes with flakiness risk (a slow database, a stalled network call, a race in the test server) that a pure function can never have. A common pitfall is over-mocking at the unit level: if you replace so many collaborators that the "unit" test no longer exercises real logic, it stops earning its speed advantage and becomes a maintenance burden that breaks on every refactor without ever catching a real bug. The opposite pitfall is under-investing in unit tests and leaning on end-to-end tests to catch logic bugs, which works but means every logic bug takes minutes instead of milliseconds to surface, and a flaky end-to-end suite starts to be ignored by the team, which is worse than no suite at all.
Tell me about a time you discovered a production bug that was caused by missing or insufficient tests. Using the STAR structure, describe the situation, the task you were responsible for, the actions you took to fix the bug and improve the tests or process (including which test level was missing and why), and the measurable result afterward.
Sample Answer
A strong answer to this question names the specific test level that was missing, not just "we didn't have enough tests," because the level tells the interviewer exactly what you learned and whether your fix addressed the real gap.
How to structure the STAR response
Situation: describe the system and the context concisely, for example a checkout service where a promotional-discount feature shipped and, in production, allowed two discount codes to be combined when the business rule required only one to apply at a time.
Task: state your specific responsibility, for example being the engineer or QA owner responsible for the checkout service's test coverage and for triaging the incident once it was reported.
Action: this is the part worth being most concrete about. Two honest, common shapes:
- The bug existed at the UNIT level (the discount-combination rule itself was wrong) but no unit test covered that specific combination of inputs, only the single-discount case; the integration and end-to-end tests that existed happened to use test data that never exercised two codes together, so they passed without ever exercising the buggy path. The fix: add the missing unit test covering the combination case, verify it fails against the buggy code and passes against the fix, and then audit for other similarly-unexercised input combinations in the same rule.
- The bug existed at the INTEGRATION level (each discount's logic was individually correct, but the two didn't compose correctly once wired through the real order-total calculation, a case unit tests of each discount rule in isolation could not see). The fix: add an integration test that exercises the real combined path, not just each rule mocked in isolation.
Either shape is a legitimate, honest answer; the point is naming precisely which level was missing and why the existing suite's blind spot let the bug through, then closing that specific gap rather than adding coverage generically.
Result: describe the outcome concretely but honestly: for example, that specific discount-combination bug did not recur, and the broader audit of the same rule surfaced and closed a small number of similarly-unexercised input combinations before they caused an incident, giving you a real, verifiable before/after data point (the audit finding count) rather than an invented precision metric.
Trade-offs and pitfalls
The common way this answer goes wrong in an interview is staying vague ("we added more tests and it got better"), which gives the interviewer no way to judge your actual technical judgment; naming the specific test level, the specific gap in that level's coverage, and the specific fix is what turns a generic incident story into evidence of real testing judgment. A second pitfall is inventing a precise-sounding improvement metric you didn't actually measure; if you don't have a real number, describe the outcome qualitatively (no recurrence, caught in the audit before shipping) rather than fabricating one.
Discuss the trade-offs and return on investment between automating end-to-end UI tests versus API-level tests. Consider maintainability, flakiness, speed, coverage, and debugging ease, and how each should be used in release gating versus production monitoring. Conclude with a pragmatic hybrid approach and rules of thumb for deciding which user stories to automate at the UI level versus the API level.
Sample Answer
End-to-end UI tests and API-level tests both exercise real, wired-together behavior, but they trade realism against cost in opposite directions, and the right answer for most teams is not "pick one" but a deliberate hybrid.
The trade-offs, dimension by dimension
- Maintainability. UI tests are coupled to the rendered page: a renamed button, a reordered form field, or a redesigned layout can break many UI tests that were never testing that element, producing maintenance work disproportionate to any real regression. API tests are coupled only to the API contract, which typically changes far less often than the UI, so they need less ongoing repair.
- Flakiness. UI tests depend on rendering timing, animations, and browser quirks, all classic sources of non-determinism; API tests, hitting a server directly with no rendering step, are inherently more deterministic and far less prone to intermittent failure.
- Speed. API tests skip the browser entirely, so they typically run several times faster than the equivalent UI test, which matters directly for how large a suite you can afford to run per pull request.
- Coverage. A UI test is the only one of the two that can prove the interface itself renders correctly and reacts to real interaction; an API test proves the underlying business logic and data flow are correct but says nothing about whether a real user sees the right thing on screen.
- Debugging ease. When an API test fails, the failure usually points precisely at a request/response mismatch; when a UI test fails, you must first determine whether the underlying logic is wrong, the API contract changed, or the UI simply rendered slower than the test expected, which takes more investigation time per failure.
How each should be used in release gating versus monitoring
API tests are cheap and reliable enough to gate a release directly: a failing API test is a strong, low-noise signal that something is genuinely broken, so blocking on it costs little and catches real problems. UI tests, being slower and flakier, are better used more sparingly as a release gate (a small, curated smoke set for your most critical journeys) and more heavily as ongoing production monitoring (synthetic UI checks running continuously against production, alerting on failure) where an occasional false alarm is a minor cost rather than a blocked release.
A pragmatic hybrid approach and rules of thumb
Default new coverage to the API level; only add a UI test when the thing you're actually verifying is specifically about rendering or interaction (does a validation message appear where expected, does a button visibly disable during submission) rather than about the underlying business outcome (does the order total compute correctly), which an API test can verify just as well at a fraction of the cost. As a rule of thumb: if you could rewrite a proposed UI test to call the API directly and it would still prove the same business assertion, it belongs at the API level; if rewriting it that way would lose the actual thing being tested, it belongs at the UI level.
Trade-offs and pitfalls
The trap in a hybrid strategy is inconsistency: without an explicit rule like the one above, teams tend to add UI tests reflexively because they feel more "complete," slowly re-accumulating exactly the maintenance and flakiness burden the hybrid approach was meant to avoid. Make the API-first default explicit and require a stated reason (specifically about rendering or interaction) to add a UI test instead.
You are refactoring a legacy user interface. Compare unit, integration, and end-to-end tests for this specific situation: their relative maintenance cost, execution speed, flakiness risk, and the class of regression each is most likely to catch. Explain how you would prioritize which of these to write first during a large refactor, then propose a test pyramid for a typical medium-sized frontend application.
Sample Answer
Refactoring a legacy UI changes which test level is actually trustworthy: the code you're about to rewrite is exactly the code your existing tests were written against, so the SAME test that normally gives confidence can instead actively mislead you if it's coupled to implementation details rather than to observable behavior.
Comparing the three levels for a refactor specifically
- Unit tests: cheapest to run and fastest to give feedback, with the lowest flakiness risk of the three since they involve no real I/O, network, or rendering timing, but the ones most likely to be coupled to the OLD implementation's internal structure (specific component boundaries, internal state shape) rather than to genuinely observable behavior; a refactor that changes internal structure without changing behavior will break many of these even though nothing is actually wrong, producing false-negative noise exactly when you need signal most. That noise is a maintenance-cost problem, rewriting tests against the new structure, not a flakiness problem: a red unit test during a refactor is almost never a flake.
- Integration tests: a better middle ground for a refactor, since they typically assert on a slightly higher-level contract (does this component correctly call the API and update in response) that survives internal restructuring better than a narrow unit test does, while still being far cheaper and more precise than a full end-to-end test; flakiness risk sits between the other two levels, since touching one real dependency (a test server, a real local store) introduces some timing variance, but far less than a fully rendered UI does.
- End-to-end tests: the most trustworthy signal during a UI refactor specifically, because they assert on OBSERVABLE USER-FACING BEHAVIOR (does the page still do what a user needs it to do) rather than on any internal structure, so they are largely immune to being broken by the refactor itself; the cost is that they're slow, imprecise about WHERE a real regression is if one occurs, and the flakiest of the three, since real rendering, animation, and network timing introduce genuine non-determinism. During a refactor specifically that means a red end-to-end test needs a first pass to rule out an ordinary flake, by rerunning it, before you treat it as the trustworthy signal the rest of this comparison relies on it being.
Prioritizing which to write first during a large refactor
Write a small set of end-to-end tests FIRST, covering the critical user journeys the section being refactored supports, as a behavior-preserving safety net: these tests should pass before, during, and after the refactor, since they check outward behavior the refactor is not supposed to change. Only after that safety net exists should new unit and integration tests be introduced for the NEW internal structure, written against the refactored code's actual new boundaries, since writing them against the OLD structure just before deleting that structure would be wasted effort.
A test pyramid for a typical medium-sized frontend application
Once the refactor is complete and normal development resumes, return to a standard frontend shape: a large base of unit tests for pure logic and isolated component behavior, a solid middle layer of component-level integration tests (rendering a component with React Testing Library or similar, confirming it correctly calls its collaborators and updates state), and a small, curated top layer of full end-to-end tests for the handful of critical journeys, mirroring the same shape recommended for other frontend contexts, with the refactor-specific end-to-end safety net folded back into that small curated top layer rather than kept as a separate, larger set.
Trade-offs and pitfalls
The pitfall specific to a refactor is treating a failing unit test during the refactor as evidence something is broken, when it may simply be evidence the test was coupled to implementation details that were always going to change; before "fixing" such a test, first confirm via the end-to-end safety net whether user-facing behavior actually changed, and if it did not, the right fix is usually to rewrite the unit test against the new structure, not to change the refactored code to satisfy the old test.
Unlock Full Question Bank
Get access to all 20 Test Levels and the Test Pyramid interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.