Mocking, Stubbing, and Test Isolation Questions
Isolating the unit under test from its dependencies. Covers mocks, stubs, fakes, and spies, when to use test doubles versus real dependencies, and controlling external services and time. Includes designing for isolation so tests are fast, deterministic, and focused.
As a technical lead, craft a stakeholder-facing proposal to reduce brittle, over-specified mocks in the test suite for a mission-critical service. Include a cost/benefit case, a phased technical plan, a rollback plan if the change causes problems, and the metrics you would use to measure success over the next quarter.
Sample Answer
Direct answer
A production bug escaped because an over-specified mock hid the real integration behavior; the right response is to name that concrete failure as the trigger, then propose a phased, measured plan (add real integration and contract tests, tighten what mocks are allowed to assert, and track defect-escape metrics) with an explicit rollback path if the change itself causes disruption.
Structured elaboration
- Lead with the trigger, not the abstraction: a stakeholder-facing proposal lands better grounded in the specific incident (what broke, why the existing mock-heavy tests didn't catch it) than in a general claim that "our tests are too mock-heavy."
- Cost/benefit case: quantify what the incident cost (customer impact, engineering time to diagnose and fix) against the ongoing cost of the proposed change (slower test suite, engineering time to build real/contract tests), so the trade-off is concrete rather than asserted.
- Phased technical plan over roughly a quarter: start by adding a small number of real integration tests or consumer-driven contract tests specifically covering the class of interaction that broke; in parallel, audit existing mocks for over-specification (asserting on internal call details rather than observable behavior) and simplify the worst offenders; avoid attempting to rewrite the entire suite's mocking strategy at once.
- Rollback plan: if the added integration/contract tests turn out to be too slow or flaky to run on every commit, have a fallback (running them on a slower cadence, or gating only specific high-risk services) rather than reverting the whole initiative; state this explicitly so the proposal doesn't read as an irreversible, high-risk bet.
- Metrics for success: track the defect-escape rate (bugs that reach production versus caught pre-merge) and integration-test coverage growth over the quarter, reporting back against these numbers rather than declaring success by feeling.
- Migration steps that acknowledge different audiences: individual contributors need concrete guidance on which mocks to simplify and how to write a new contract test; engineering leadership needs the cost/benefit framing and the metrics; both need to see this connected back to the specific incident that motivated it.
Worked example
After a mock hid a change in a downstream service's error-response shape (the mock always returned the old shape, so the calling code's new error-handling branch was never actually exercised, and shipped broken), the proposal names that incident explicitly, then commits to: adding 3 consumer-driven contract tests for that specific downstream service in month one, auditing and simplifying the 12 most over-specified mocks in month two, and reporting defect-escape rate and contract-test coverage at the end of the quarter. If the contract tests prove too slow for the standard per-commit pipeline, the rollback path moves them to a nightly gate rather than abandoning them.
Trade-offs and pitfalls
A proposal that argues from principle ("mocking is a code smell") without anchoring in the specific incident and its concrete cost is a harder sell to stakeholders who weigh engineering time against other priorities; leading with the real trigger event, the real cost, and a bounded, reversible phased plan is what actually gets this kind of initiative funded and, more importantly, followed through on rather than abandoned halfway.
Your CI pipeline is unstable because several tests depend on flaky or slow third-party services. Propose a hybrid strategy that decides, dependency by dependency, whether to mock it, virtualize it, or keep a small number of real integration tests. Give criteria for which tests should stay live, and explain how you would know if that mix is drifting the wrong way over time.
Sample Answer
Direct answer
Stabilize a flaky CI pipeline by deciding, dependency by dependency, whether to mock it, virtualize it, or keep a small number of real calls, rather than applying one blanket policy; the criteria are how flaky the dependency has actually been, how safety-critical it is, and how expensive it is to call for real.
Structured elaboration
A practical migration plan:
- Classify each external dependency by observed failure rate over the last few weeks of CI runs and by how central its real behavior is to what you're trying to validate.
- Mock-first for chronically flaky, low-risk dependencies: anything with a high enough failure rate that it's already causing "just re-run it" behavior, and where the business logic under test doesn't depend on the dependency's exact real-world quirks.
- Keep or add a small number of real integration tests for high-risk dependencies: run them less frequently (nightly, or gated on a separate slower pipeline) rather than on every commit, so their occasional real flakiness doesn't block every PR.
- Migrate incrementally, dependency by dependency, verifying after each migration that the pipeline's overall flakiness rate actually drops and that no regression slipped through because a mock is now hiding real behavior.
- Watch for the mix drifting the wrong way: track what fraction of tests exercise a real dependency over time; if it trends toward zero, the suite is losing its ability to catch real integration bugs, and if a chronically-mocked dependency's real API changes, nothing will tell you until it breaks in production.
Worked example
A pipeline has three external dependencies: a currency-conversion API (occasional slowness, low business risk, used in many tests), a fraud-check API (occasional slowness, high business risk), and an email-sending service (occasional slowness, low risk, used in far fewer tests). The right migration: mock currency conversion everywhere except in one or two integration tests that revalidate the response shape weekly; keep the fraud-check API real in a small, nightly-run suite specifically because its exact decision logic is what the business needs the tests to protect; fully mock the email service, since almost no test needs to verify that email actually sends, only that the code attempted to send it.
Trade-offs and pitfalls
The plan fails if it's applied as an all-or-nothing switch: flipping every test to mocks in one pass removes the flakiness immediately but can silently remove real-integration coverage for months before anyone notices a drift. Track the migration with a simple metric, like "count of tests exercising each real dependency," reviewed periodically, so the team can see the mix drifting and course-correct before it becomes invisible.
Explain the roles of unit, integration, and end-to-end tests in a delivery pipeline. For each layer, describe what should typically be mocked versus run against a real dependency, and walk through a concrete example for a payment flow showing where mocks belong and why.
Sample Answer
Direct answer
As a rough default: unit tests mock everything outside the function or class under test; integration tests run against real (or lightly virtualized) adjacent components but still mock the furthest-out third parties; end-to-end tests run against real dependencies wherever it's safe and affordable to do so.
Structured elaboration
- Unit tests: the fastest, most numerous layer. Everything the unit under test calls, database, network, other services, gets mocked, so the test isolates and exercises only the logic in that one unit.
- Integration tests: verify that two or more of YOUR OWN components wire together correctly (your service and your database, your service and an internal message queue). Real (or a very high-fidelity fake of) your own infrastructure is used here, but external third parties are usually still mocked or virtualized, since the point is validating your own integration code, not the third party's uptime.
- End-to-end tests: exercise a full user-facing flow across real systems. Real external dependencies are used where safe (sandboxed accounts, staging environments); anything unsafe, costly, or nondeterministic to call for real (an actual bank transfer, an actual SMS bill) still gets stubbed even at this layer.
Worked example
For a payment flow: at the unit level, mock the PaymentGateway interface entirely and test that the order service calls charge() with the right amount and handles a thrown exception by not saving the order. At the integration level, run the order service against a real local database to confirm the SQL actually persists correctly, while still mocking the payment gateway (a third party, out of scope for this layer). At the end-to-end level, run the full checkout flow against the payment gateway's real sandbox environment, so the test confirms the actual network contract and response shapes match what your code expects, without charging a real card.
Trade-offs and pitfalls
The three-layer split breaks down if a team pushes everything into unit tests with mocks and skips the integration layer entirely: unit tests can all pass while the wiring between your own components is broken, because no test ever exercised your actual database or actual message-queue configuration. Conversely, pushing too much into end-to-end tests makes the suite slow and flaky without necessarily catching bugs any earlier or more precisely than a well-placed integration test would.
Explain approaches to intercept and modify network requests and responses during a Selenium test in order to simulate backend conditions. Compare using an HTTP proxy, browser-level network interception, and a service-virtualization tool, and provide a short example showing how you would stub a single JSON API response.
Sample Answer
Direct answer
Three approaches intercept network traffic during a Selenium test at different layers: an HTTP proxy sits between the browser and the network and can rewrite any request/response; browser-level network interception (via the Chrome DevTools Protocol) hooks directly into the browser's own network stack; and a service-virtualization tool replaces the backend entirely with a configurable fake server the browser talks to normally.
Structured elaboration
| Approach | How it works | Pros | Cons |
|---|---|---|---|
| HTTP proxy (e.g. a local proxy the browser is configured to route through) | Sits outside the browser, intercepts and can rewrite any request/response crossing it | Works across any browser or client, language-agnostic | Extra process to run and configure, and HTTPS interception needs a trusted certificate installed in the browser |
| Browser-level interception via the Chrome DevTools Protocol | The test driver registers request/response handlers directly with the browser's own devtools connection | No separate process, no certificate trust issues, very precise control per-request | Chromium-specific (or needs an equivalent for other engines), and ties the test to the devtools API surface |
| Service virtualization | The backend the app calls is a real, separately configurable server | Exercises the app's real network code path against a realistic backend, reusable across UI and API tests | Slower to set up, requires the app to be pointed at the virtual server's URL instead of production |
A short example stubbing a single JSON API response (shown here with an HTTP-level interception library so it is genuinely runnable without a browser; the identical idea applies whether the interception happens via a proxy or the Chrome DevTools Protocol):
import responses
import requests
def fetch_profile(user_id):
resp = requests.get(f"https://app.example.com/api/profile/{user_id}")
resp.raise_for_status()
return resp.json()
@responses.activate
def test_stub_a_single_json_api_response():
responses.add(
responses.GET,
"https://app.example.com/api/profile/77",
json={"id": 77, "plan": "pro"},
status=200,
)
result = fetch_profile(77)
assert result == {"id": 77, "plan": "pro"}
assert len(responses.calls) == 1
Executed with pytest:
test_s13_network_interception.py::test_stub_a_single_json_api_response PASSED
1 passed in 0.03s
Worked example
Testing a profile page that should show a "Pro" badge only for pro-tier users: stub the GET /api/profile/77 call to return {"plan": "pro"} and assert the badge renders; stub the same endpoint to return {"plan": "free"} in a second test and assert the badge does not render. Neither test depends on a real backend being up, and both are deterministic regardless of what the real profile service currently returns for user 77.
Trade-offs and pitfalls
An HTTP proxy adds a genuinely separate moving part (the proxy process itself, HTTPS certificate trust) to the test environment; browser-level interception avoids that but is coupled to whichever browser engine's devtools protocol you're using; service virtualization is the heaviest but also the most representative of real end-to-end behavior. Whichever layer is chosen, the stubbed response shape needs to be kept realistic, if the real API adds a required field the stub never includes, the UI test can pass while the real integration is actually broken.
Explain the differences between mocks, stubs, fakes, spies, and dummies. For each kind of test double, give a concrete example from a typical web application (an HTTP API call, a database, a message queue, or a cache) and a short guideline for when you would prefer that double over the others.
Sample Answer
Direct answer
A test double is a stand-in for a real dependency in a test. The five common kinds differ in what they do when called and what the test asserts about them: a dummy is passed around but never actually used, a stub returns canned data, a fake is a working but simplified implementation, a mock records calls so the test can verify they happened correctly, and a spy wraps a real object while also recording calls to it.
Structured elaboration
| Double | Behavior when called | What the test checks | Typical web-app example |
|---|---|---|---|
| Dummy | Does nothing meaningful; only fills a required parameter slot | Nothing about it directly | Passing a placeholder database-connection object into a constructor that requires one but is never queried on the code path under test |
| Stub | Returns a pre-programmed value | The return value the code under test produces | Making a "get exchange rate" HTTP client return a fixed 1.08 so a pricing calculation is deterministic |
| Fake | Runs real logic, just a lighter-weight implementation | The behavior/output, same as against the real thing | An in-memory key-value store used instead of a real Redis instance |
| Mock | Returns canned data AND records how it was called | That specific calls happened, with specific arguments, in the right order | Verifying that PaymentGateway.charge() was called exactly once with the correct amount |
| Spy | Wraps a real object, forwards calls through, and records them | Both real behavior and the interaction | Wrapping a real email-sending service so you can assert "send was called" while letting a test double catch the actual network call underneath |
The line between stub and mock is really about what the test asserts on: a stub only shapes the INPUT to the code under test (state verification), while a mock is used to verify OUTPUT in terms of interactions (behavior verification). The same test-double library (Mockito, unittest.mock, Sinon) is usually used to build all five; the taxonomy names roles, not library features.
Worked example
For a database example: a stub database client always returns a fixed list of three users for find_active_users(), regardless of what was inserted, so a report-generation test has predictable input. A mock database client would instead be used to verify that save(user) was called exactly once with a User object whose status field is "active", when testing the code that is supposed to activate a user. A fake database would be a real, lightweight in-memory dictionary-backed store that actually persists and retrieves records within the test, useful when the test needs realistic query behavior (like "insert then find") that a stub's fixed answer can't provide.
Trade-offs and pitfalls
Reach for a dummy when a dependency is only required to satisfy a constructor or function signature and is never actually exercised on the path under test; building anything more elaborate (a stub, a fake) for it would be wasted setup effort. Reach for a stub when you need to control an input; reach for a mock when the ACT of calling the dependency (with the right arguments, in the right order) is itself part of the behavior being tested, such as making sure a payment is charged exactly once. Reach for a spy when you want the dependency's real behavior to genuinely happen (unlike a mock, which fully replaces it) while still asserting that a specific interaction occurred, such as confirming a real cache was actually written to while also checking the write call's arguments. Overusing mocks where a stub would do makes tests brittle: every internal refactor that doesn't change observable behavior can still break a mock-heavy test, because the test is coupled to how the code calls its dependency rather than what the code produces.
Unlock Full Question Bank
Get access to all 16 Mocking, Stubbing, and Test Isolation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.