Microservices Architecture and Service Decomposition Questions
Decomposing a system into services along bounded contexts, defining service boundaries and ownership, and managing the tradeoffs against a monolith. Covers cohesion, coupling, data ownership per service, the distributed-monolith anti-pattern, and when a modular monolith beats microservices. Emphasizes decomposition reasoning rather than any single framework.
Explain the strangler fig pattern for migrating a monolith to microservices: the role of a routing/intercept layer, the role of anti-corruption layers when the new service must still talk to the old monolith's data, and how you would manage shared-database access during the transition period. Describe how you'd prioritize which components to extract first.
Sample Answer
Direct answer
The strangler fig pattern migrates a monolith to microservices incrementally: you place a routing layer in front of the monolith, redirect one slice of functionality at a time to a newly-built service behind that layer, and let the monolith's role shrink slice by slice until, eventually, it can be retired, rather than attempting a single big-bang rewrite.
Structured elaboration
The routing (or intercept) layer is the mechanism that makes this incremental: it sits in front of the monolith and inspects each incoming request, forwarding requests for the not-yet-migrated functionality to the monolith as before, and requests for the newly-extracted slice to the new service instead. Because the routing decision happens per-request, the cutover for any one slice can be gradual (a percentage of traffic, or a specific subset of users) and reversible (route back to the monolith if the new service misbehaves), rather than an all-at-once switch.
An anti-corruption layer handles the case where the new service still needs to talk to the old monolith, either because the monolith still owns some data the new service needs, or because other parts of the monolith still call into the functionality that's now been extracted. Rather than letting the new service's clean domain model get contaminated by the monolith's legacy data shapes and conventions, the anti-corruption layer translates between the two, so the new service's internal model stays coherent even while it's still dependent on the old system underneath.
Managing shared database access during migration is the trickiest part: in the transition window, both the monolith and the new service may need to read or write data that hasn't fully moved yet. A common approach is dual-write (the monolith continues writing to its tables while also publishing changes the new service consumes) with reconciliation to catch drift, or a change-data-capture stream from the monolith's database that the new service consumes without writing to the monolith's tables directly, avoiding a genuine two-way dependency on the same schema.
Worked example
Prioritizing which components to extract first: start with a slice that's relatively self-contained (few dependencies on the rest of the monolith), has a clear owning team ready to take it on, and delivers a visible win (either operational, like removing a frequent source of incidents, or business, like unblocking a team that's currently release-blocked by the monolith's shared deploy train). Extracting the riskiest, most deeply-entangled part of the monolith first, even if it's the part causing the most pain, tends to produce the highest-risk first migration when the team has the least experience running this pattern; a smaller early win builds the operational muscle (routing, dual-write, monitoring the new service) that the harder extractions will need later.
Trade-offs and pitfalls
The most common failure is leaving the routing layer and the dual-write/reconciliation logic in place indefinitely instead of treating them as temporary migration scaffolding; if a slice never gets fully cut over (the monolith keeps a code path alive "just in case"), the system ends up permanently carrying the complexity of both the old and new implementations, which is worse than either the original monolith or a clean microservice on its own.
How would you quantify and present the technical risk and business cost of having many microservices with overlapping responsibilities, versus consolidating some of them into fewer services? Describe the metrics you would gather (deployment coordination overhead, on-call load, infra cost per service, cross-service change frequency), any lightweight experiments you might run, and how you would present the trade-off to executives who are not engineers.
Sample Answer
Direct answer
To quantify the cost of too many overlapping-responsibility microservices, measure deployment coordination overhead (how often a single logical change requires touching multiple services in lockstep), on-call and incident load per service, infrastructure cost per service (even a nearly-idle service carries a fixed baseline cost), and cross-service change frequency (how often a change to one service requires a corresponding change to another); present the trend in those numbers to executives rather than an architectural opinion about service count.
Structured elaboration
Deployment coordination overhead is measurable directly: track how many recent releases required coordinating changes across two or more services that, if consolidated, would have been a single deploy, and how much calendar time that coordination added compared to a single-service change. On-call load is measurable from incident data: total pages per service per month, and specifically how many of those incidents were caused by an inter-service contract mismatch (a caller and callee disagreeing about a field's meaning or an API version) rather than a genuine bug in one service's own logic, since contract mismatches are a direct symptom of over-decomposition. Infrastructure cost is measurable from the cloud bill: baseline compute, monitoring, and logging cost per service, multiplied by the number of services that could plausibly be consolidated without losing a real scaling or ownership benefit. Cross-service change frequency is measurable from version-control history: how often a pull request in one service's repo is immediately followed by a corresponding pull request in another service's repo within a short window, a proxy for services that are more tightly coupled in practice than their separate deployability suggests.
Worked example
A lightweight experiment to gather this evidence without a large upfront investment: instrument the deploy pipeline to tag any release that required a coordinated multi-service change, and run that for a month before making a consolidation recommendation, rather than relying on anecdotes about "it feels like everything requires touching three services." Present the finding to executives in business terms: "N% of releases in the last quarter required coordinating three or more services, adding an average of X days to the release; consolidating these two specific services would remove that coordination cost for roughly Y% of those releases," rather than a purely technical argument about service-count aesthetics.
Trade-offs and pitfalls
The risk of presenting this case poorly is framing it as "we have too many microservices" in the abstract, which invites a debate about architectural philosophy instead of a decision grounded in measured cost; naming the SPECIFIC services with the worst coordination and incident numbers, and proposing a targeted consolidation of just those, is both more persuasive and less risky than a broad "let's reduce our service count" initiative. The countervailing risk is consolidating services that look similar on paper but actually have a real, measured difference in scaling or team ownership; the same data-gathering discipline that justifies a consolidation should also be used to rule one out when the signals don't actually support it.
When would you choose synchronous request/response calls between services versus asynchronous messaging? For each choice, discuss the impact on end-to-end latency, coupling between services, error handling and retry behavior, and the operational implications for on-call and SLOs.
Sample Answer
Direct answer
Choose synchronous calls when the caller genuinely needs the result before it can proceed and can tolerate the callee's latency and availability becoming part of its own; choose asynchronous messaging when the caller can proceed without waiting for the result, or when decoupling the caller's availability from the callee's is more important than getting an immediate answer.
Structured elaboration
Latency: synchronous calls put the callee's latency directly on the critical path of the caller's response time, and a chain of several synchronous calls compounds that (each hop adds its own latency, and the caller waits for the slowest one). Asynchronous messaging removes the callee's latency from the caller's response time entirely, since the caller doesn't wait for the message to be processed. Coupling: synchronous calls create a direct availability dependency (if the callee is down, the caller's request fails or blocks); asynchronous messaging decouples availability, since a message can sit in a queue until the consumer is back up, at the cost of the consumer's effect on the world happening later, not immediately. Error handling: a synchronous call gives the caller an immediate, explicit success-or-failure signal it can act on right away (retry, show an error, fall back); an asynchronous message's failure needs a different mechanism entirely (a dead-letter queue, a retry policy on the consumer side, and some way for the ORIGINAL caller to eventually learn the outcome if it needs to, since it already moved on). Operational implications: synchronous chains make on-call debugging comparatively straightforward (a single request trace shows the whole call chain and where it failed) but make service-level objectives (SLOs, the reliability/latency targets a service commits to) harder to hit as the chain gets longer, since the end-to-end latency and availability are the product of every hop's; asynchronous flows make individual components easier to keep within their own SLOs independently, but debugging "why didn't this eventually happen" requires tracing through queues and consumers rather than a single linear request.
Worked example
A checkout flow illustrates both: charging a customer's card needs a synchronous call to the payment processor, because the checkout page genuinely can't tell the customer "success" until the charge is confirmed, and the caller needs an explicit success-or-failure signal to act on immediately. Sending the order-confirmation email, by contrast, is a good fit for asynchronous messaging: the checkout flow doesn't need to wait for the email to send before showing the customer a success page, and decoupling it means an email-service outage doesn't block checkout at all, only delays the email itself.
Trade-offs and pitfalls
The most common mistake is defaulting to synchronous calls for everything because it's simpler to reason about in the moment, which quietly makes every downstream service's availability and latency a dependency of the caller's SLO, even for work that didn't need an immediate answer. The opposite mistake is making something asynchronous that the caller actually needed an immediate answer for (like the payment charge above), which either forces an awkward polling loop on the caller's side or produces a confusing user experience where the system says "success" before it actually knows whether the operation succeeded.
Discuss the pros and cons of a monorepo versus multiple repositories when organizing the code for many independently-owned services at a large company. Focus on the impact on CI build times, release cadence, cross-service refactors, and code-ownership boundaries.
Sample Answer
Direct answer
A monorepo (all services' code in one repository) makes cross-service refactors and dependency management easier because everything is visible and changeable atomically in one place, at the cost of CI build times and tooling that need to scale with the size of the whole repository, not just one service; multiple repositories give each service team full independence over its own build and release cadence, at the cost of cross-service changes requiring coordinated, separately-versioned pull requests across several repos.
Structured elaboration
CI impact: a monorepo's CI system needs to be smart enough to build and test only what actually changed (otherwise every commit anywhere in the repo triggers a full-fleet rebuild, which doesn't scale past a modest number of services); this typically requires investment in build-graph-aware tooling that many organizations don't have off the shelf. A polyrepo's CI is naturally scoped to one service per repo, so each build stays fast and simple by default, without needing that investment. Release cadence: a monorepo makes it straightforward to enforce that everyone is always building against the latest version of a shared library (there's only one version in the repo at any time), which prevents version-skew problems but also means a breaking change to a shared library needs to either fix every consumer in the same commit or be rolled out very carefully; a polyrepo lets each service pin its own version of a shared dependency and upgrade on its own schedule, which is more flexible per-team but means version skew (different services depending on different, sometimes old, versions of the same library) becomes a real and recurring problem to track. Cross-service refactors: a monorepo makes a refactor that touches five services a single atomic commit, visible and reviewable as one change; a polyrepo makes the same refactor five separate pull requests across five repos, each independently reviewed and merged, which is harder to keep consistent and harder to roll back as a single unit if something goes wrong.
Worked example
A large company with genuinely independent teams and a strong platform investment in build tooling (build-graph-aware CI, code-ownership enforcement at the directory level) can make a monorepo work well at significant scale, gaining the refactor and dependency-consistency benefits without paying an unbounded CI cost. A company without that tooling investment, or with teams that genuinely want full autonomy over their own release process (including the freedom to lag behind on a shared dependency's latest version while they finish their own migration), is often better served by a polyrepo, accepting the coordination cost of cross-repo changes as the price of that independence.
Trade-offs and pitfalls
The common mistake with a monorepo is adopting it without the CI tooling investment to make it scale, producing painfully slow builds that engineers start working around (skipping tests locally, batching unrelated changes to amortize a slow CI run), which erodes the discipline the monorepo was meant to encourage. The common mistake with a polyrepo is underestimating how much coordination overhead accumulates from routine, small cross-service changes (a shared library security patch, for example, that now needs a separate PR, review, and release in every consuming repo), which can end up costing more engineering time in aggregate than the monorepo's CI investment would have.
A company is moving from roughly 20 to 200 services. Explain Conway's Law's practical impact on the resulting architecture and reliability, and propose an organizational structure and set of team boundaries (platform/infra teams, service-owning teams, shared libraries) that improves ownership clarity and reduces cross-team coupling at that scale.
Sample Answer
Direct answer
Going from 20 to 200 services multiplies the coordination surface roughly with the number of services, not linearly with headcount, so the organizational structure that worked at 20 services (loose conventions, informal coordination) breaks down well before 200; the fix is introducing explicit platform and infrastructure teams that own shared concerns, clear per-service ownership with no orphaned services, and enough standardization that a new service doesn't require reinventing deployment, observability, and on-call practices from scratch.
Structured elaboration
Conway's Law's practical impact at this scale: at 20 services, informal cross-team communication (a Slack message, a quick sync) is usually enough to coordinate a shared concern; at 200, the number of possible pairwise team interactions grows far faster than the team count itself, and informal coordination stops scaling, showing up as duplicated effort (multiple teams independently solving the same infrastructure problem), inconsistent practices (some services have solid observability, others none), and slower cross-cutting changes (a security fix that needs to land in every service takes far longer to propagate without a shared mechanism). The organizational fix mirrors the technical one: introduce dedicated platform/infrastructure teams whose job is providing the shared capabilities every service team would otherwise reimplement (a deployment pipeline template, a standard observability stack, a shared authentication library), reducing the coordination surface from "every team talks to every other team" to "every team talks to the platform team."
Worked example
A concrete structure: product-facing teams each own a small, clear set of services end to end (their own on-call, their own release cadence), a platform team owns the shared deployment pipeline, service templates, and core infrastructure every other team builds on, and a smaller number of specialist teams (security, data platform) own concerns that genuinely need central expertise and shouldn't be duplicated 200 times. Team boundaries at this scale should be reviewed periodically (not fixed forever at whatever they were when the org had 20 services), since a boundary that worked well at 20 services can become a bottleneck at 200 if, for example, one team ends up owning far more services than it can operate well.
Trade-offs and pitfalls
The most common failure at this scale is under-investing in the platform team's capacity relative to how many product teams depend on it, turning the platform team itself into the new coordination bottleneck; the platform team's own roadmap needs to be resourced and prioritized as seriously as any product team's, since if it can't keep up with demand, product teams start working around it with one-off solutions, which recreates the inconsistency the platform team existed to prevent. Reliability at scale also depends on this structure: a shared, well-maintained deployment and observability platform, rather than 200 independently-invented ones, is what makes it possible to have a consistent incident-response process across the whole fleet.
Unlock Full Question Bank
Get access to all 34 Microservices Architecture and Service Decomposition interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.