Test Levels and the Test Pyramid Questions
How unit, integration, component, end-to-end and contract tests fit together and where each provides the most value. Covers the test pyramid and the competing shapes proposed against it (the testing trophy and the honeycomb), multi-layer test architecture, contract testing as the seam between services, choosing the right level to catch a given class of defect cheaply, and what to run per commit versus per release. Includes the cost and confidence trade-offs between fast low-level tests and slower, broader system tests. The scope is which level a test belongs at and why. Deciding how much to invest in testing and where to prioritize under time pressure is covered separately.
Explain the differences between smoke tests, regression tests, integration tests, system tests, and user-acceptance tests, and between functional and non-functional testing. For each, describe when it should be executed in a typical CI/CD pipeline and give one concrete example test appropriate for an e-commerce web application.
Sample Answer
These names describe two different axes, not one: smoke, regression, integration, system, and user-acceptance tests describe SCOPE and PURPOSE within a release process, while functional versus non-functional describes WHAT KIND of requirement is being verified. A single test can sit at one point on each axis at once (for example, a load test is a non-functional system test).
The five scope/purpose types
| Type | What it verifies | When it runs in CI/CD | Example for an e-commerce app |
|---|---|---|---|
| Smoke | The absolute basics work at all: the app starts, key pages load, nothing is catastrophically broken | Immediately after every deploy, before anything else runs | Confirm the homepage and checkout page both return HTTP 200 after a deploy |
| Regression | Previously-fixed bugs and previously-working behavior haven't broken again | On every pull request, or nightly for the full suite | Re-run the specific test that reproduces a past bug where applying two discount codes together double-discounted an order |
| Integration | Two or more real components agree on how they interact | Pull request / merge | Confirm the checkout API correctly writes a new order row to the real database |
| System | The whole assembled application behaves correctly as one unit against requirements | Pre-release, in a staging-like environment | Walk through browsing, adding to cart, and completing checkout as one continuous validation of the whole system, not just one flow |
| User-acceptance | The system satisfies what the business or the customer actually asked for | Just before release, often with a human sign-off | A product owner or customer confirms that the new "buy now, pay later" option behaves the way they specified in the requirements |
Regression testing's specific effect on release velocity
A solid regression suite is what lets a team ship frequently without re-manually-verifying everything that already worked: automated regression tests reliably catch a bug like the double-discount example above, where a change to one part of the pricing logic silently breaks a previously-correct interaction, the moment it's introduced, rather than after a customer reports it. What automated regression tests do NOT reliably catch is a bug that requires actual human judgment to notice, such as a new promotional banner rendering with confusing or misleading wording, which passes every automated check while still being wrong; that class of issue needs manual exploratory testing precisely because "correctness" here is a judgment call, not a fixed assertion.
Keeping a growing regression suite fast and reliable
As a regression suite grows, two problems compound: it gets slower, and it accumulates flaky tests (ones that fail intermittently for reasons unrelated to real regressions). Keep it fast by running only the subset of regression tests relevant to changed code on every PR, reserving the full suite for a nightly run. Keep it reliable by treating a flaky regression test as a bug in the test itself, not background noise to tolerate: track a rerun rate per test, and either fix or quarantine (temporarily exclude with an owner assigned to repair it) any test whose failures don't correlate with real code changes, since an ignored flaky test trains the team to distrust the whole suite.
Functional versus non-functional testing, as a separate axis
Functional testing asks "does the checkout flow correctly compute the total and complete the order," a direct check against a stated feature requirement. Non-functional testing asks a different kind of question entirely: for the same checkout flow, does it perform well under load (performance), does it protect payment data appropriately (security), is it usable by someone unfamiliar with the site (usability), and can someone using a screen reader complete a purchase (accessibility). These four non-functional concerns should be prioritized before release based on business risk, not treated as equally weighted: for a payment flow specifically, security and performance under peak load typically deserve the most pre-release attention, since a failure there has the most severe and hardest-to-reverse consequences, while usability and accessibility issues, though real and important, are more often caught and improved iteratively after release without the same acute risk.
In a typical CI/CD pipeline, functional tests run continuously as part of the regular suite on every commit or pull request, since a functional regression, like the checkout total being computed incorrectly, is valuable to catch immediately. Non-functional tests usually run on a slower, scheduled cadence: a load test simulating peak Black Friday traffic against the checkout API is a concrete non-functional example, and it typically runs nightly or pre-release rather than blocking every commit, since it needs a longer, resource-heavy run that would slow down PR feedback if it gated every merge.
Trade-offs and pitfalls
The most common confusion is treating "system test" and "end-to-end test" as interchangeable; they overlap heavily in practice but system testing traditionally emphasizes validating the WHOLE application against its requirements as one unit (often owned by QA, closer to release), while end-to-end testing more narrowly emphasizes a specific user JOURNEY through the real stack (often automated and run continuously). Naming this distinction explicitly, rather than treating the terms as synonyms, is itself a signal of depth in this space.
You are adding a small new feature to an existing web app: a per-user settings toggle that changes UI behavior. Which tests would you write first, and why? Specify the concrete unit, integration, and end-to-end tests you would create, which tests could reasonably be postponed or executed manually instead, and how you would weigh risk, regression likelihood, and return on investment in that decision.
Sample Answer
For a small, contained change like a per-user settings toggle, the order you write tests in should mirror how cheaply each level can rule out a class of bug, not an arbitrary checklist.
Which tests to write first, and why
- Unit test first: does the function that decides the toggle's effective value (given the stored setting, any default, and any override) return the right value for on, off, and unset states. This is the cheapest possible check on the actual logic and should exist before anything else, since if the logic itself is wrong, no amount of higher-level testing will reliably catch every case.
- Integration test second: does saving the toggle's new value actually persist it correctly (a real database write and read-back), and does the API endpoint that exposes the setting return the correct value after a save. This proves the logic from step 1 is correctly wired to real storage, which the unit test cannot show.
- One end-to-end test: a single test that toggles the setting through the real UI and confirms the resulting UI behavior actually changes, proving the whole path (UI action, API call, storage, and the resulting render) is wired together correctly for a real user. One is enough here: additional UI-level variations (different starting states, different pages) are better covered at the unit or integration level instead.
What to postpone or handle manually
Visual polish around the toggle (exact spacing, animation) is better handled by manual or exploratory testing rather than an automated test, since automating a purely visual detail for a low-risk, easily-reverted feature rarely pays back its maintenance cost. Similarly, exhaustive combinations of the toggle with every other unrelated setting are not worth automating up front for a small, isolated feature; a manual spot-check covers that risk more cheaply until real usage shows a combination actually matters.
Weighing risk, regression likelihood, and ROI
The toggle's blast radius (does it interact with billing, security, or just cosmetic UI behavior) should set how much automated coverage it earns: a purely cosmetic toggle justifies exactly the three tests above and nothing more, while a toggle that gates paid functionality would justify more integration-level coverage of its interaction with the billing system specifically. Regression likelihood also matters: an isolated, rarely-touched feature is unlikely to be broken by unrelated future changes, so its test investment can stay minimal; a setting that many other features read is worth more integration coverage precisely because future unrelated changes are more likely to break it.
A harder version of the same judgment: a runtime, per-tenant feature flag
The same reasoning scales up for a runtime feature flag that can be toggled per tenant and rolled out to a percentage of users, but the stakes and the required levels both grow. Here you need coverage at four levels rather than three: unit tests for the flag-evaluation logic itself (given a tenant and a rollout percentage, does it correctly decide on or off); integration tests confirming flag state changes are correctly read and cached; a canary-level check confirming a partial rollout percentage is honored in practice (roughly the right proportion of requests see the new behavior, no more); and end-to-end tests confirming no leakage occurs between tenants (a flag enabled for tenant A never leaks its effect to tenant B) and that a rollback of the flag takes effect promptly. The extra canary level and the explicit tenant-isolation and rollback checks exist because the blast radius of getting a runtime, partial rollout wrong (accidentally exposing new behavior to the wrong users, or being unable to roll it back quickly) is categorically larger than a simple settings toggle's.
Trade-offs and pitfalls
The pitfall for the small-toggle case is over-testing a low-risk feature out of habit, writing UI-level tests for every state combination when a single end-to-end test plus solid unit coverage would give equivalent confidence at a fraction of the cost. The pitfall for the feature-flag case is the opposite: under-testing because "it's just a flag," when in practice a flag with per-tenant, partial-rollout semantics is closer in risk to a small distributed system than to a simple toggle, and deserves the fuller four-level treatment above.
Explain the test pyramid concept. Describe its tiers (unit, integration, and end-to-end), the primary goal of each tier, and why the pyramid recommends many more low-level tests than high-level tests. For a typical web application, give concrete examples of test types and common tools at each tier (for instance, unit tests for helpers, integration tests for API-to-database interactions, end-to-end tests for a checkout flow), and briefly mention limitations or scenarios where the pyramid shape may not apply.
Sample Answer
The test pyramid is a shape you aim for when deciding how many tests to write at each level: many fast, narrow unit tests at the base, a smaller number of integration tests in the middle, and very few, broad end-to-end tests at the top. The core claim is not "unit tests are better," it is that the ratio should be inverted from what a naive test-writer defaults to: most bugs are logic bugs that a unit test finds cheaply, so you want the bulk of your assertions living where they are cheap to write, fast to run, and precise about what broke, and you reserve the slow, broad, more failure-prone end-to-end tests for the small number of things only they can prove: that the assembled system, wired together for real, actually works.
The tiers
- Unit: a function or class tested alone, dependencies faked. Goal: prove the logic is correct, in isolation, in milliseconds.
- Integration: your code against one real neighbor (a database, a queue, one real service). Goal: prove the wiring and serialization between two real things is correct.
- End-to-end: the system driven through its real entry point, nothing faked. Goal: prove the whole thing actually delivers the right behavior to a real caller.
Why more low-level tests than high-level
Three forces push the shape into a pyramid rather than a rectangle or its inverse:
- Cost. An end-to-end test typically needs a running environment, real data, and real network calls; a unit test needs none of that. If a unit test costs 1 unit of setup and run time, an integration test might cost 10-50x that, and an end-to-end test 100-1000x that, so a rectangle-shaped suite (equal counts at every level) would make your CI/CD pipeline unusably slow and destroy the fast feedback a pipeline exists to provide.
- Feedback precision. When a unit test fails, you already know which function is wrong. When an end-to-end test fails, you know the system as a whole is broken but not where, and diagnosing that costs real engineering time.
- Flakiness. The more real infrastructure a test touches (network, clock, shared state), the more opportunities it has to fail for reasons unrelated to the code under test. A large end-to-end suite tends to accumulate intermittent failures that erode trust in the whole pipeline.
Worked example (web application)
For a typical web application: unit tests for pure helper functions (a discount calculator, a date formatter), commonly written with a plain test runner like pytest or Jest; integration tests for the API-to-database path (does saving an order actually persist the right row?), commonly using an HTTP-assertion library such as Supertest against a real test database; end-to-end tests for a checkout flow driven through the real UI or a real HTTP client, confirming a user can go from "add to cart" to "order confirmed," commonly using a browser-automation tool such as Playwright or Cypress. A healthy team might run thousands of unit tests in under a minute, a few hundred integration tests in several minutes, and a few dozen end-to-end tests in tens of minutes, matching the pyramid's shape to the cost curve above.
Trade-offs and limitations of the model
The pyramid assumes most defect risk lives in logic that a unit test can isolate. That assumption weakens for systems whose main risk is integration itself, such as a thin orchestration layer that mostly calls other services and has little logic of its own: here, integration and contract tests carry more of the confidence burden, and a strict pyramid ratio would under-test the actual risk. This is the same observation that motivates alternative shapes like the testing trophy (an alternative shape that keeps a small unit-test base but makes integration tests the largest layer, on the idea that tests resembling real usage give more confidence), which is worth naming as a caveat even in a definitional answer: the pyramid is a strong default, not a law. The common pitfall in applying it is treating the shape as a hard quota (chasing a specific unit-test count) rather than as a description of where investment should land once you've correctly identified where a given system's real risk lives.
Describe the test pyramid and how you would apply it to a modern single-page-application stack (React frontend, Node API, PostgreSQL database). For each layer (unit, integration/component, and end-to-end), give concrete examples of what to test and recommended tooling, propose an approximate test-count ratio across the layers, and describe how you would validate and adjust that ratio over time as the product matures.
Sample Answer
For a React-frontend, Node-API, PostgreSQL-database SPA stack, the pyramid maps onto three layers whose boundary follows the technology seam as much as the logical one.
What to test at each layer, with tooling
- Unit: pure functions and isolated logic on both sides of the stack, for example a price-formatting helper or a validation function on the frontend, and a business-rule function on the Node API. Recommended tooling: Jest (or Vitest) for both the React frontend and the Node backend, since a single test runner across the stack keeps tooling simple.
- Integration/component: on the frontend, rendering a React component with React Testing Library and confirming it correctly calls a mocked API client and updates its own state and DOM in response, which proves the component's own logic and rendering without needing the real backend running; on the backend, hitting the real Node API with Supertest against a real (test) PostgreSQL database, proving the route, the query, and the schema all agree, which no frontend-only or backend-only unit test can show.
- End-to-end: driving the real React app in a real browser against the real API and database (or a close staging equivalent) using Playwright or Cypress, proving the whole assembled stack delivers a correct user-facing outcome, such as a full checkout flow from click to confirmation.
Guidance on test-count ratio
A reasonable starting ratio for this stack is roughly 65-70% unit tests (split across frontend logic and backend logic), 20-25% integration/component tests (split between frontend component tests and backend API-to-database tests), and 5-10% end-to-end tests covering only the handful of journeys where the whole assembled stack matters most (checkout, authentication). The SPA's heavy client-side interaction pushes the integration/component share slightly higher than a pure backend service would need, since a meaningful share of this stack's real risk lives in how React components manage state and respond to user interaction, which a backend-only pyramid wouldn't need to account for.
Validating and adjusting the ratio over time
Track, per release, which layer actually caught each regression found either in code review, staging, or production, and compare that distribution to your current test-count ratio: if end-to-end tests are catching bugs that a component test could have caught more cheaply, that's a signal to push more coverage down a layer; if production bugs keep slipping through despite full coverage lower in the pyramid, that's a signal the end-to-end layer, not the lower layers, needs to grow for that specific journey. Revisit the ratio on a fixed cadence (quarterly is common) rather than continuously, since a ratio that reacts to every single incident tends to overfit to the most recent bug rather than reflecting the system's actual steady-state risk.
Trade-offs and pitfalls
The most common mistake on this specific stack is testing React component behavior primarily through end-to-end browser tests, because it's the most "realistic," when a React Testing Library component test at the integration/component layer can prove the same interaction logic in a small fraction of the time and with far less flakiness. Reserve full end-to-end coverage for the journeys where the point genuinely is proving the whole stack, frontend, API, and database together, works correctly.
You are testing a RESTful web application built from a React single-page application, a Node.js REST API, a PostgreSQL database, and an external payment gateway. For each test-pyramid tier (unit, integration, end-to-end), list four concrete example tests you would create, name a common tool or library for each example, and justify why each test belongs at that tier: what it verifies, and what it depends on.
Sample Answer
For a RESTful web application (React SPA, Node.js REST API, PostgreSQL, external payment gateway), each pyramid tier should own a different, non-overlapping slice of confidence, and the tests below make that concrete.
Unit tier (four examples)
- Discount/price calculator: a pure function
calculateDiscount(price, tier); verifies core business math, e.g. Jest for the frontend or a plain test runner on the backend. - React component render logic: does the checkout form component render a validation error when the card field is empty; React Testing Library.
- Request validator: does the order-creation handler reject a negative price before touching the database; a plain Jest unit test with mocked input, no HTTP or DB involved.
- Payment-gateway response parser: given a sample JSON response from the gateway, does your parser extract the correct transaction ID and status; a pure-function Jest unit test with a hard-coded fixture, no real network call.
Each of these verifies one piece of logic in isolation and depends on nothing external, which is why they can run in milliseconds.
Integration tier (four examples)
- API-to-database write path: POST an order to the real Node API running against a real (test) PostgreSQL instance, then query the database directly to confirm the row and its computed total are correct; Supertest plus a real Postgres test container.
- Repository layer against Postgres: an ORM query (e.g., a Prisma or TypeORM query) that joins orders and customers, run against a seeded test database via Testcontainers, to catch a wrong join or a migration mismatch a mocked-DB unit test would miss.
- Payment-gateway client against the gateway's sandbox: using an HTTP client library such as Axios (or the gateway's official Node.js SDK if one is provided), call the gateway's real sandbox endpoint (not your parser in isolation) to confirm your client sends a well-formed request and correctly handles the sandbox's real success and decline responses.
- React SPA against a mocked API layer: render the checkout page and confirm it calls the real API client code (not the component logic alone) and correctly updates state on a real HTTP response, using a tool like MSW (Mock Service Worker) to intercept only the network boundary, not the application code.
Each of these proves two real components agree on a contract (route shape, SQL schema, gateway request format) that a unit test, by construction, cannot check because it never invokes the second component for real.
End-to-end tier (four examples)
- Full checkout journey: drive the real React SPA in a real browser through add-to-cart, checkout form, and payment, against the real (or sandboxed) full stack, using Playwright or Cypress, to prove the whole assembled system delivers a working checkout.
- Payment failure path end-to-end: using Playwright or Cypress, submit a card the gateway's sandbox is configured to decline, and confirm the SPA shows the correct user-facing error, proving the failure path is wired correctly all the way through, not just handled by the parser in isolation.
- Session and auth flow: using Playwright or Cypress, log in, add an item, refresh the page, and confirm the cart persists, exercising the real session/cookie mechanism no lower-level test touches.
- Order confirmation and receipt: using Playwright or Cypress to complete the purchase, combined with a test email-capture tool such as Mailhog or Mailtrap, confirm a confirmation email or receipt page reflects the correct final total, proving the pricing logic, the database write, and the presentation layer all agree once wired together for real.
Why each test belongs where it does
The dividing line is what would have to be REAL for the test to fail the way it's meant to: the unit tests fail only if the pure logic is wrong; the integration tests fail only if two real components disagree, even when each one's internal logic is correct in isolation; the end-to-end tests fail only if something in the full assembly, including things no lower test can see (routing, session state, real gateway behavior), is broken.
Trade-offs and pitfalls
The most common mistake with this stack specifically is testing the payment-gateway integration primarily at the end-to-end level because "it's the riskiest part": that inflates the slowest, flakiest tier with coverage that a much cheaper integration test against the gateway's sandbox could provide almost as well. Reserve end-to-end for the few journeys where the VALUE is specifically in proving the pieces are wired together, and push everything else down a tier.
Unlock Full Question Bank
Get access to all 19 Test Levels and the Test Pyramid interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.