Test Levels and the Test Pyramid Questions
How unit, integration, component, end-to-end and contract tests fit together and where each provides the most value. Covers the test pyramid and the competing shapes proposed against it (the testing trophy and the honeycomb), multi-layer test architecture, contract testing as the seam between services, choosing the right level to catch a given class of defect cheaply, and what to run per commit versus per release. Includes the cost and confidence trade-offs between fast low-level tests and slower, broader system tests. The scope is which level a test belongs at and why. Deciding how much to invest in testing and where to prioritize under time pressure is covered separately.
Explain the test pyramid and where UI and end-to-end tests fit within it. For a typical modern web application, explain how you would allocate automation effort across the unit, integration/service, and UI layers and justify your allocation. Describe the risks of over-emphasizing UI tests and how shifting testing earlier (shift-left) helps avoid that trap.
Sample Answer
UI and end-to-end tests sit at the TOP of the pyramid: they exercise the real, rendered interface and prove the whole assembled system delivers correct behavior to a real user, which is exactly why they should be the smallest, most carefully curated layer rather than the primary source of coverage.
Allocating automation effort across the layers
For a typical modern web application, allocate effort roughly as follows: the majority of effort goes into unit tests, since most individual bugs are logic bugs that a unit test catches fastest and most precisely; a meaningful but smaller share goes into integration/service tests, proving the pieces genuinely wire together (API to database, service to service); and only a small, deliberately curated slice goes into UI tests, reserved specifically for the handful of journeys where proving the RENDERED interface behaves correctly for a real user is the actual point (checkout, sign-up), not a general-purpose place to verify business logic that a lower level could check more cheaply.
The risk of over-emphasizing UI tests
A UI-heavy suite is slow (each test needs a real or simulated browser), flaky (rendering timing, animations, and selector brittleness introduce non-determinism no lower level has), and imprecise when it fails (a UI test failure could stem from broken logic, a changed API contract, or nothing more than an unrelated layout change, and diagnosing which requires real investigation time). At scale, this combination trains a team to distrust its own test suite: a slow, occasionally-flaky UI suite gets reflexively rerun on failure rather than investigated, which quietly defeats the entire purpose of having automated tests as a genuine safety net.
How shift-left helps avoid that trap
Shift-left means moving quality checks earlier in the development process, closer to where the code is written, rather than relying on late-stage UI or end-to-end tests to catch problems that a unit or integration test could have caught immediately and far more cheaply. Concretely, this means writing unit and integration tests alongside the code (or even before it) rather than backfilling UI tests after a feature ships, and using code review and static analysis to catch categories of bugs before any test even runs. The practical effect is that UI tests are left to do only the job unique to them, proving the rendered interface itself works, instead of being asked to carry logic-verification work that shifting earlier would have caught more cheaply and precisely.
Trade-offs and pitfalls
The trap teams fall into is treating UI test coverage as a proxy for overall confidence, since a UI test LOOKS the most reassuring (it exercises "the real thing"), which leads to writing UI tests for logic that a unit test would verify just as well. Push back on that instinct explicitly: ask, for any proposed UI test, whether the thing actually being verified requires the rendered interface at all, or whether it's really a logic assertion wearing a UI test's clothing.
Describe integration testing in depth: its purpose, and the common approaches to structuring it (big-bang, incremental, top-down, and bottom-up). Explain how you would decide whether to run integration tests against real third-party services, mocked responses, or recorded traffic, and the practical trade-offs of each choice.
Sample Answer
Integration testing exists to prove that two or more real components agree on how they interact, which unit tests, by testing each component alone, structurally cannot show.
Four common approaches to structuring it
- Big-bang: integrate and test all components together at once, only after every piece is individually complete. Simple to set up, but when it fails, it gives almost no information about WHICH interaction is broken, since everything is combined at the same time; best suited to small systems where "everything together" is a manageable scope.
- Incremental: integrate and test components a few at a time, growing the tested surface gradually. Failures are much easier to localize than big-bang, since you know which newly-added component caused a new failure, at the cost of more setup and more distinct test configurations to maintain.
- Top-down: start from the highest-level component (an API layer or orchestrator) and integrate downward, using stubs to stand in for lower components not yet integrated. Lets you validate the overall structure and control flow early, before every dependency is ready, at the cost of needing well-maintained stubs that can themselves drift from real behavior.
- Bottom-up: start from the lowest-level components (a data-access layer, a utility library) and integrate upward, using driver code to exercise components not yet wired to their real caller. Validates foundational pieces early and with high confidence, at the cost of not exercising the overall system structure until later in the process.
Deciding: real services, mocked responses, or recorded traffic
Use a REAL third-party service when the service is cheap or free to call, reliably available in a sandbox environment, and the specific behavior you need to verify (a genuine edge case in its real response) can't be faithfully reproduced any other way; the trade-off is speed, reliability, and cost, since your tests now depend on someone else's uptime and rate limits. Use MOCKED responses when you need fast, deterministic tests for your own code's handling logic (how do you react to a success, a specific error code, a timeout) and you're confident about the shape of the real service's responses; the trade-off is drift risk: the mock silently stops matching reality if the real service changes. Use RECORDED traffic (capturing real request/response pairs once, then replaying them) as a middle ground: it gives you realistic response bodies without a live network dependency on every test run, at the cost of the recordings themselves going stale if the real service changes and nobody re-records them.
Trade-offs and pitfalls
The most common mistake is picking one of these three uniformly for an entire integration suite rather than choosing per-test based on what that specific test needs to prove: a test verifying your error-handling logic rarely needs a real service call, while a test verifying your integration still matches the real service's current contract benefits from at least occasional real or recorded traffic, not a hand-maintained mock alone.
Explain the differences between smoke tests, regression tests, integration tests, system tests, and user-acceptance tests, and between functional and non-functional testing. For each, describe when it should be executed in a typical CI/CD pipeline and give one concrete example test appropriate for an e-commerce web application.
Sample Answer
These names describe two different axes, not one: smoke, regression, integration, system, and user-acceptance tests describe SCOPE and PURPOSE within a release process, while functional versus non-functional describes WHAT KIND of requirement is being verified. A single test can sit at one point on each axis at once (for example, a load test is a non-functional system test).
The five scope/purpose types
| Type | What it verifies | When it runs in CI/CD | Example for an e-commerce app |
|---|---|---|---|
| Smoke | The absolute basics work at all: the app starts, key pages load, nothing is catastrophically broken | Immediately after every deploy, before anything else runs | Confirm the homepage and checkout page both return HTTP 200 after a deploy |
| Regression | Previously-fixed bugs and previously-working behavior haven't broken again | On every pull request, or nightly for the full suite | Re-run the specific test that reproduces a past bug where applying two discount codes together double-discounted an order |
| Integration | Two or more real components agree on how they interact | Pull request / merge | Confirm the checkout API correctly writes a new order row to the real database |
| System | The whole assembled application behaves correctly as one unit against requirements | Pre-release, in a staging-like environment | Walk through browsing, adding to cart, and completing checkout as one continuous validation of the whole system, not just one flow |
| User-acceptance | The system satisfies what the business or the customer actually asked for | Just before release, often with a human sign-off | A product owner or customer confirms that the new "buy now, pay later" option behaves the way they specified in the requirements |
Regression testing's specific effect on release velocity
A solid regression suite is what lets a team ship frequently without re-manually-verifying everything that already worked: automated regression tests reliably catch a bug like the double-discount example above, where a change to one part of the pricing logic silently breaks a previously-correct interaction, the moment it's introduced, rather than after a customer reports it. What automated regression tests do NOT reliably catch is a bug that requires actual human judgment to notice, such as a new promotional banner rendering with confusing or misleading wording, which passes every automated check while still being wrong; that class of issue needs manual exploratory testing precisely because "correctness" here is a judgment call, not a fixed assertion.
Keeping a growing regression suite fast and reliable
As a regression suite grows, two problems compound: it gets slower, and it accumulates flaky tests (ones that fail intermittently for reasons unrelated to real regressions). Keep it fast by running only the subset of regression tests relevant to changed code on every PR, reserving the full suite for a nightly run. Keep it reliable by treating a flaky regression test as a bug in the test itself, not background noise to tolerate: track a rerun rate per test, and either fix or quarantine (temporarily exclude with an owner assigned to repair it) any test whose failures don't correlate with real code changes, since an ignored flaky test trains the team to distrust the whole suite.
Functional versus non-functional testing, as a separate axis
Functional testing asks "does the checkout flow correctly compute the total and complete the order," a direct check against a stated feature requirement. Non-functional testing asks a different kind of question entirely: for the same checkout flow, does it perform well under load (performance), does it protect payment data appropriately (security), is it usable by someone unfamiliar with the site (usability), and can someone using a screen reader complete a purchase (accessibility). These four non-functional concerns should be prioritized before release based on business risk, not treated as equally weighted: for a payment flow specifically, security and performance under peak load typically deserve the most pre-release attention, since a failure there has the most severe and hardest-to-reverse consequences, while usability and accessibility issues, though real and important, are more often caught and improved iteratively after release without the same acute risk.
In a typical CI/CD pipeline, functional tests run continuously as part of the regular suite on every commit or pull request, since a functional regression, like the checkout total being computed incorrectly, is valuable to catch immediately. Non-functional tests usually run on a slower, scheduled cadence: a load test simulating peak Black Friday traffic against the checkout API is a concrete non-functional example, and it typically runs nightly or pre-release rather than blocking every commit, since it needs a longer, resource-heavy run that would slow down PR feedback if it gated every merge.
Trade-offs and pitfalls
The most common confusion is treating "system test" and "end-to-end test" as interchangeable; they overlap heavily in practice but system testing traditionally emphasizes validating the WHOLE application against its requirements as one unit (often owned by QA, closer to release), while end-to-end testing more narrowly emphasizes a specific user JOURNEY through the real stack (often automated and run continuously). Naming this distinction explicitly, rather than treating the terms as synonyms, is itself a signal of depth in this space.
Explain the differences between unit tests, integration tests, and end-to-end tests. For each level, give two concrete examples (functions, modules, services, or UI flows), state when it should run (on a pull request, at merge, or nightly) and its typical execution speed, and discuss the typical maintenance cost and failure modes. Conclude with the concrete trade-offs between speed, coverage, and flakiness for a web application.
Sample Answer
A unit test exercises a single function or class in complete isolation: every dependency is faked, stubbed, or simply absent, so the test runs in microseconds and its failure points at exactly one piece of logic. An integration test exercises how two or more real components work together, most often your code against a real (or near-real) database, queue, or external service, so it catches wiring and serialization bugs a unit test cannot see. An end-to-end test drives the system the way a real client would, through its actual entry point (an HTTP call, a UI click), with nothing faked, so it is the only level that proves the whole assembled system actually works.
What each level covers, with two concrete examples per level
| Level | Two example targets | Runs on | Speed | Maintenance cost | Failure mode it's good at catching |
|---|---|---|---|---|---|
| Unit | (1) a pure function, e.g. a discount calculator; (2) a class method with its collaborators mocked, e.g. an order-validation method tested with a fake repository | Every commit, on save | Microseconds to low milliseconds | Low, unless over-mocked | Wrong business logic, missed edge cases |
| Integration | (1) your code against a real database, e.g. does saving an order persist the right row; (2) your code against one real external service, e.g. a payment client against that gateway's sandbox | Pull request / merge | Tens of milliseconds to a few seconds | Medium: schema and API drift break these | Wiring bugs: wrong SQL, wrong serialization, a contract mismatch |
| End-to-end | (1) a full UI flow, e.g. add-to-cart through order confirmation in a real browser; (2) a full API flow, e.g. a real HTTP client driving create-then-fetch against the live server with nothing faked | Nightly or pre-release | Seconds to minutes | High: brittle to unrelated UI or infra changes | Environment and integration issues that only appear when everything runs together |
The QA-engineer angle on unit tests is collaborative, not just "who writes them": developers usually author the unit tests since they know the implementation, but QA should read them during review to spot missing edge cases the implementer didn't think of, and QA is often the one who notices a bug that unit tests theoretically should have caught but didn't (a coverage gap, not a process failure). "System testing" is a related but distinct idea: it validates the whole assembled system as one unit against requirements, similar in spirit to end-to-end testing, but typically owned by QA and run just before release with attention to environment parity and realistic test data, whereas end-to-end testing is often owned by whoever automates the user-facing flow and runs continuously.
The concrete tools differ by stack but the pattern holds everywhere: JUnit or pytest for the unit layer, pytest combined with testcontainers (spinning up a real, disposable database or service in a container) for the integration layer, and Selenium or Playwright for the end-to-end layer, with the same mock-vs-real-service decision applying at the integration boundary regardless of which tools you pick: mock a dependency when you're testing YOUR handling logic, use the real (or containerized) dependency when you're testing that the wiring itself is correct. For a payments microservice specifically, this maps onto where each level runs in the deployment pipeline: unit tests run locally on every save and in CI on every commit; integration tests run in CI against a containerized database and a sandboxed payment gateway; end-to-end tests run in a staging environment before a production deploy, and a small smoke subset may re-run immediately after reaching production to confirm the live deploy itself is healthy. This progression, more tests locally and in CI, fewer in staging, fewer still in production, is also how the pyramid should guide the ALLOCATION of engineering effort: invest the majority of new test-writing time at the level closest to the developer's own commit, not because the higher levels don't matter, but because that's where a fixed hour of effort buys the most coverage per dollar of CI time and developer attention.
Worked example: the same business rule, tested three ways
The clearest way to see the boundary is to test the identical rule at all three levels and watch what each level can and cannot catch.
# calculate_discount is pure logic: no I/O, so it belongs at the unit level.
def calculate_discount(price: float, tier: str) -> float:
if price < 0:
raise ValueError("price must be non-negative")
rate = {"standard": 0.0, "silver": 0.05, "gold": 0.15}.get(tier)
if rate is None:
raise ValueError(f"unknown tier: {tier}")
return round(price * (1 - rate), 2)
# UNIT TEST: no database, no network. Executed directly.
def test_calculate_discount_unit():
assert calculate_discount(100, "standard") == 100.0
assert calculate_discount(100, "silver") == 95.0
assert calculate_discount(100, "gold") == 85.0
Executed output: UNIT level: 4/4 assertions passed (pure function, no I/O, <1ms) (all four, including the negative-price ValueError case).
import sqlite3
class OrderRepository:
def __init__(self, conn):
self.conn = conn
self.conn.execute(
"CREATE TABLE IF NOT EXISTS orders (id INTEGER PRIMARY KEY, price REAL, tier TEXT, total REAL)"
)
def save(self, price, tier):
total = calculate_discount(price, tier)
cur = self.conn.execute(
"INSERT INTO orders (price, tier, total) VALUES (?, ?, ?)", (price, tier, total)
)
self.conn.commit()
return cur.lastrowid
# INTEGRATION TEST: a REAL SQLite database, catching serialization/wiring the unit test cannot see.
def test_order_repository_integration():
conn = sqlite3.connect(":memory:")
repo = OrderRepository(conn)
order_id = repo.save(200, "gold")
row = conn.execute("SELECT total FROM orders WHERE id = ?", (order_id,)).fetchone()
assert row[0] == 170.0
Executed output: INTEGRATION level: repository round-trip through real SQLite passed: {'id': 1, 'price': 200.0, 'tier': 'gold', 'total': 170.0}.
For the end-to-end level, the same repository was wired behind a real HTTP handler and hit with an actual socket-level POST followed by a GET, using Python's built-in http.server and urllib.request, no mocking anywhere in the path:
import json, threading, urllib.request, urllib.error
from http.server import BaseHTTPRequestHandler, HTTPServer
class Handler(BaseHTTPRequestHandler):
def log_message(self, format, *args):
pass
def do_POST(self):
length = int(self.headers.get("Content-Length", 0))
body = json.loads(self.rfile.read(length))
try:
order_id = repo.save(body["price"], body["tier"])
total = calculate_discount(body["price"], body["tier"])
except ValueError as e:
self.send_response(400)
self.end_headers()
self.wfile.write(json.dumps({"error": str(e)}).encode())
return
self.send_response(201)
self.end_headers()
self.wfile.write(json.dumps({"id": order_id, "total": total}).encode())
def do_GET(self):
order_id = int(self.path.rsplit("/", 1)[-1])
row = conn.execute("SELECT id, total FROM orders WHERE id = ?", (order_id,)).fetchone()
self.send_response(200)
self.end_headers()
self.wfile.write(json.dumps({"id": row[0], "total": row[1]}).encode())
# E2E TEST: real HTTP socket, real server thread, nothing faked.
def test_checkout_end_to_end():
server = HTTPServer(("127.0.0.1", 0), Handler)
port = server.server_address[1]
threading.Thread(target=server.serve_forever, daemon=True).start()
req = urllib.request.Request(
f"http://127.0.0.1:{port}/orders",
data=json.dumps({"price": 50, "tier": "silver"}).encode(),
headers={"Content-Type": "application/json"}, method="POST",
)
created = json.loads(urllib.request.urlopen(req).read())
fetched = json.loads(urllib.request.urlopen(f"http://127.0.0.1:{port}/orders/{created['id']}").read())
assert fetched == {"id": created["id"], "total": 47.5}
bad = urllib.request.Request(
f"http://127.0.0.1:{port}/orders",
data=json.dumps({"price": 50, "tier": "platinum"}).encode(),
headers={"Content-Type": "application/json"}, method="POST",
)
try:
urllib.request.urlopen(bad)
assert False, "expected HTTPError"
except urllib.error.HTTPError as e:
assert e.code == 400
server.shutdown()
Executed output: E2E level (valid order): POST+GET over real HTTP socket returned {'id': 1, 'total': 47.5} and E2E level (invalid tier): POST over real HTTP socket returned status 400. That is precisely the trade-off: the unit test told us the discount math is right in under a millisecond; the end-to-end test told us the whole pipe, JSON serialization, routing, and the network stack included, actually delivers that correct math to a real client, at the cost of running a live server and a real socket for the one test.
Trade-offs and pitfalls
The pyramid shape follows directly from this example: you want most of your assertions at the level that is cheapest to run and most precise about what broke, which is the unit level, and you want just enough integration and end-to-end coverage to prove the pieces are wired correctly, because that proof is expensive and comes with flakiness risk (a slow database, a stalled network call, a race in the test server) that a pure function can never have. A common pitfall is over-mocking at the unit level: if you replace so many collaborators that the "unit" test no longer exercises real logic, it stops earning its speed advantage and becomes a maintenance burden that breaks on every refactor without ever catching a real bug. The opposite pitfall is under-investing in unit tests and leaning on end-to-end tests to catch logic bugs, which works but means every logic bug takes minutes instead of milliseconds to surface, and a flaky end-to-end suite starts to be ignored by the team, which is worse than no suite at all.
Describe the test pyramid and how you would apply it to a modern single-page-application stack (React frontend, Node API, PostgreSQL database). For each layer (unit, integration/component, and end-to-end), give concrete examples of what to test and recommended tooling, propose an approximate test-count ratio across the layers, and describe how you would validate and adjust that ratio over time as the product matures.
Sample Answer
For a React-frontend, Node-API, PostgreSQL-database SPA stack, the pyramid maps onto three layers whose boundary follows the technology seam as much as the logical one.
What to test at each layer, with tooling
- Unit: pure functions and isolated logic on both sides of the stack, for example a price-formatting helper or a validation function on the frontend, and a business-rule function on the Node API. Recommended tooling: Jest (or Vitest) for both the React frontend and the Node backend, since a single test runner across the stack keeps tooling simple.
- Integration/component: on the frontend, rendering a React component with React Testing Library and confirming it correctly calls a mocked API client and updates its own state and DOM in response, which proves the component's own logic and rendering without needing the real backend running; on the backend, hitting the real Node API with Supertest against a real (test) PostgreSQL database, proving the route, the query, and the schema all agree, which no frontend-only or backend-only unit test can show.
- End-to-end: driving the real React app in a real browser against the real API and database (or a close staging equivalent) using Playwright or Cypress, proving the whole assembled stack delivers a correct user-facing outcome, such as a full checkout flow from click to confirmation.
Guidance on test-count ratio
A reasonable starting ratio for this stack is roughly 65-70% unit tests (split across frontend logic and backend logic), 20-25% integration/component tests (split between frontend component tests and backend API-to-database tests), and 5-10% end-to-end tests covering only the handful of journeys where the whole assembled stack matters most (checkout, authentication). The SPA's heavy client-side interaction pushes the integration/component share slightly higher than a pure backend service would need, since a meaningful share of this stack's real risk lives in how React components manage state and respond to user interaction, which a backend-only pyramid wouldn't need to account for.
Validating and adjusting the ratio over time
Track, per release, which layer actually caught each regression found either in code review, staging, or production, and compare that distribution to your current test-count ratio: if end-to-end tests are catching bugs that a component test could have caught more cheaply, that's a signal to push more coverage down a layer; if production bugs keep slipping through despite full coverage lower in the pyramid, that's a signal the end-to-end layer, not the lower layers, needs to grow for that specific journey. Revisit the ratio on a fixed cadence (quarterly is common) rather than continuously, since a ratio that reacts to every single incident tends to overfit to the most recent bug rather than reflecting the system's actual steady-state risk.
Trade-offs and pitfalls
The most common mistake on this specific stack is testing React component behavior primarily through end-to-end browser tests, because it's the most "realistic," when a React Testing Library component test at the integration/component layer can prove the same interaction logic in a small fraction of the time and with far less flakiness. Reserve full end-to-end coverage for the journeys where the point genuinely is proving the whole stack, frontend, API, and database together, works correctly.
Unlock Full Question Bank
Get access to all 22 Test Levels and the Test Pyramid interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.