Legacy Modernization and Architecture Evolution Questions
Evolving an existing system rather than designing greenfield. Covers the modernization patterns (strangler fig, anti-corruption layers, facades and protocol adapters), choosing between rehosting, replatforming, incremental refactoring and a full rewrite, data migration and coexistence (dual-running, change-data-capture versus bulk cutover, reconciliation and drift), cutover readiness and decommissioning, recovering undocumented behavior from legacy code and stored procedures, instrumenting a migration so you can tell in real time whether it is working, and the organizational and risk management of long migrations across many teams. The scope is the migration itself: not quantifying or prioritizing technical debt, not code-level refactoring craft, not how to decompose a system into microservices, and not cloud migration or deployment and rollback mechanics as topics in their own right.
How does Conway's Law show up when you're migrating a system and deciding how to carve it up? When does aligning the new architecture to existing team boundaries help, and when does it lock in a bad design?
Sample Answer
Direct answer
Conway's Law says a system's architecture ends up mirroring the communication structure of the organization that builds it, and during a migration that cuts both ways: aligning new service boundaries to existing team boundaries accelerates the migration because ownership and decision-making stay clear, but it locks in whatever boundaries the org happens to have today, which may not be the technically sound ones if the org structure itself grew for historical or political reasons rather than around the actual domain.
Structured elaboration
When alignment helps: if a team already owns a coherent piece of business capability and communicates efficiently internally, extracting a service along that exact boundary means the people best positioned to understand its behavior are the ones building and owning it, decisions can be made quickly without cross-team negotiation, and the resulting service boundary has a natural, sustainable owner from day one. This is usually the fastest path during an active migration, because you're not fighting the org's existing communication patterns.
When it creates suboptimal boundaries: if the org structure predates the domain understanding you have now (teams split by historical accident, by who was hired first, or by a previous reorg that no longer reflects how the business actually works), forcing the new architecture to match that structure bakes a bad boundary into the system for the sake of migration speed. A classic version: a team owns "everything related to orders" because that's how the org chart happened to be drawn years ago, even though order creation and order fulfillment are genuinely separate domains with different scaling and consistency needs, and splitting them would produce a better architecture at the cost of also having to split (or renegotiate) team ownership.
Governance and organizational techniques to balance this:
- Decide the target service boundaries from the domain first, independent of current team structure, then explicitly evaluate the gap between that target and the org's current shape, rather than letting the current org shape silently define the target.
- Where the gap is small, align and move fast. Where it's large, make a deliberate call: either accept a suboptimal near-term boundary to keep the migration moving and plan to fix it later (a real cost, since re-splitting a service after the fact is its own migration), or invest in the reorg alongside the technical migration, accepting it will be slower.
- Use a cross-functional architecture review (not owned by any single team) to catch the case where a proposed service boundary looks convenient for the current org chart but doesn't hold up as a coherent domain boundary, before it's built and becomes expensive to undo.
Worked example
A team migrating an e-commerce monolith discovers that "the orders team" currently owns order creation, payment processing, and fulfillment tracking, three domains that grew together historically because one early engineer built all of it. Aligning the new service boundaries to this existing team would mean one large "orders service" doing all three, fast to build because ownership is unambiguous, but it recreates a version of the same coupling problem the migration was meant to solve, since payment processing has very different compliance and scaling needs than fulfillment tracking.
The team instead decides the domain-correct boundary is three separate services, and explicitly plans the organizational change alongside the technical one: the orders team splits into two teams (payments, and order-and-fulfillment) over the following two quarters, with the split sequenced to match which service gets extracted first, so a team never owns a service that doesn't exist yet, and the migration and the reorg reinforce each other rather than one blocking the other.
Trade-offs and pitfalls
The trade-off is migration speed against long-term architectural soundness: aligning to the current org is almost always faster in the near term, and for many boundaries that's a perfectly reasonable trade to make deliberately. The pitfall is making that trade unconsciously, letting the org chart silently define the architecture without ever asking whether it's actually the right domain boundary, which is exactly how Conway's Law produces an architecture nobody consciously designed.
After twelve months and real engineering investment, you conclude a major platform rewrite has to be abandoned. Walk through how you'd handle the decision itself: what you tell engineering, customers, and executives, and what happens to the system the rewrite was supposed to replace.
Sample Answer
Direct answer
The real question when abandoning a major rewrite after significant investment isn't just how to communicate it, it's how to convert a sunk-cost decision into a credible plan that stabilizes what exists, protects the team's credibility, and gets real value out of whatever was learned and built, rather than the abandonment reading as pure loss.
Structured elaboration
- Be honest and specific about why, without relitigating the original decision endlessly. Stakeholders (engineering, customers, executives) need to understand what changed or what was learned that makes continuing not worth it, stated plainly, not buried in caveats that make it sound like the team is unsure whether this was the right call. The depth and framing differ by audience: engineering gets the full technical root cause in detail, since they need it to avoid repeating the same failure mode and to trust the assessment is honest rather than face-saving; executives get the same root cause summarized against business impact and the go-forward ask; customers get only what's changing for them and by when, without the internal post-mortem detail that isn't theirs to carry.
- Stabilize the existing system first, as an explicit, separate work stream. The system the rewrite was meant to replace still needs to keep running, and often has been getting less investment while the rewrite was underway; a stabilization plan for the current system needs its own resourcing and timeline, not an assumption that it'll be fine because it's been fine so far.
- Re-allocate resources deliberately, not just disperse the team. Decide explicitly what the freed-up engineering capacity goes toward: incremental improvement of the existing system (the strangler-fig alternative, keep the old system running and move one capability at a time behind a routing layer instead of a single big-bang cutover, that might have been the better initial call), a different priority entirely, or a narrower, better-scoped follow-up effort that learned from what went wrong the first time.
- Salvage what's genuinely reusable from the abandoned work. A rewrite that's cancelled rarely produced zero value: some components, some newly-gained team expertise, or some architectural insight is usually worth carrying forward, and naming that explicitly helps the narrative be "we're redirecting effort" rather than "we wasted a year."
- Define KPIs and timelines that support the decision and minimize disruption, so stakeholders see a credible go-forward plan (what stabilizes by when, what capacity is available for what) rather than open-ended uncertainty following an admission of failure.
Worked example
A team recommending abandoning a 12-month internal platform rewrite:
- Communication: leadership is told directly that the rewrite's scope grew to underestimate the complexity of a specific legacy subsystem that turned out to encode far more undocumented business logic than the original estimate accounted for, the same class of risk that makes rewrites harder to estimate than incremental work, and that continuing under the current approach would extend the timeline past what the business can absorb. Customers relying on features the rewrite was meant to eventually deliver are told what's changing about the roadmap, in plain terms, with a revised realistic timeline for anything they were expecting. The engineering team that built the rewrite gets that same root cause in full technical depth, including exactly which assumptions about the legacy subsystem turned out wrong, is told explicitly this is being treated as a scoping and estimation failure rather than a reflection on individual or team performance, and is given clarity up front on where they're being reassigned so the news isn't followed by weeks of uncertainty about what happens to them.
- Stabilization: the existing platform, which had been getting minimal maintenance investment during the rewrite effort, gets a dedicated two-engineer stabilization track for the next quarter, addressing the highest-risk known issues that accumulated during the period of reduced investment.
- Re-allocation: rather than a full second rewrite attempt, the team pivots to an incremental, strangler-fig-style approach targeting the same underlying problems the rewrite was meant to solve, using the domain understanding gained during the rewrite attempt (which is real and valuable, even though the code itself is being set aside) to inform which capability to extract first.
- KPIs and timeline: the revised plan commits to a specific, smaller first milestone (one capability incrementally modernized) within two months, giving stakeholders a concrete, near-term proof point that the new approach is actually delivering, rather than another open-ended multi-quarter commitment right after the previous one failed to land.
Trade-offs and pitfalls
The trade-off is the short-term cost to team and organizational credibility of admitting a major effort didn't work out against the much larger cost of continuing to sink resources into an approach that's already shown it isn't working; the earlier this decision is made once the evidence is clear, the smaller that cost is. The pitfall is under-investing in the stabilization work stream because attention naturally gravitates to "what's next" rather than "what got neglected while we were focused on the now-cancelled effort," leaving the existing system in worse shape than before the rewrite was even attempted.
When is a full rewrite of a legacy system actually the right call instead of an incremental refactor, and what has to go right for it to work? Walk through the risks a rewrite introduces that an incremental approach avoids, and vice versa.
Sample Answer
Direct answer
A full rewrite wins when the legacy system's structure fights every incremental change so hard that the accumulated cost of working around it, in time, in bugs, in the parts of the team's attention it consumes, exceeds a realistic (padded) estimate for building it again from what you now know. What has to go right: the team genuinely understands the system's current behavior well enough to not silently drop requirements, the business can tolerate a real delivery gap, and there's a credible plan for the parts of the estimate that are hardest to predict, data migration and the long tail of edge cases nobody remembers exist until they break in production.
Structured elaboration
The risks a rewrite introduces that incremental work avoids:
- The estimate is a guess dressed as a plan. You cannot fully know what a legacy system does until you've read every path through it, and if you could do that cheaply, you probably wouldn't need a rewrite. Rewrites reliably underestimate the "long tail," the 20% of behavior that's undocumented, weird, and only matters for edge cases, because that's exactly the part that's hardest to discover in advance.
- All-or-nothing delivery. An incremental effort can stop, ship partial value, and reassess. A rewrite typically can't ship real value until most of it is done, which means the business is exposed to the full cost of the effort before seeing any of the benefit, and a rewrite that stalls at 80% delivers zero value for the investment made.
- Data migration risk concentrates at the end. Incremental approaches migrate data piece by piece, catching problems early on a small blast radius. A rewrite often defers data migration to a single cutover at the very end, which is exactly when the team has the least remaining slack to absorb a surprise.
The risks incremental work introduces that a rewrite avoids:
- Running two systems for longer than planned, with the ongoing cost and the risk of the migration simply never finishing.
- The seam itself becoming a source of bugs, translation errors at the boundary between old and new that a from-scratch rewrite wouldn't have to deal with.
- Slower overall delivery of the target end state, since incremental work is deliberately paced to be safe rather than fast.
Concrete mitigations for the rewrite risks: build characterization tests against the legacy system's actual behavior before writing a line of the replacement, so the estimate is grounded in observed behavior rather than assumed behavior; plan the data migration and cutover as a first-class, staged piece of the project rather than a final step; and set an internal checkpoint (not a public commitment) partway through where the team honestly reassesses whether the estimate is holding, with permission to fall back to a hybrid approach if it isn't.
Worked example
A retrospective example: a team decided a legacy inventory system's coupling to a proprietary rules engine made incremental extraction impractical, so they chose a rewrite. What made it work: they spent the first month purely on characterization testing against the legacy system, capturing behavior for every product category and edge case they could enumerate, before writing any new code. That testing surfaced a rounding rule for one product category that looked like a bug but turned out to be a deliberate accommodation for a regulatory requirement in one region, exactly the kind of hidden logic a rewrite risks silently dropping. Because they'd captured it in a test before starting, the new system preserved it correctly, and the team could point to a passing test suite, not just confidence, when they cut over.
An architectural decision that limited scale, brought up in the same retrospective: the original system had used a single shared table for all regions' inventory data, which made the regulatory rounding rule and a dozen similar region-specific rules invisible in the schema and discoverable only by reading application code path by path. The rewrite's replacement schema made region-specific rules an explicit, first-class concept, directly because the team had been burned by how hard the old shape made those rules to find.
Trade-offs and pitfalls
The single biggest risk-reducer for a rewrite is treating "we understand the current behavior" as a deliverable to prove, via characterization tests against real behavior, rather than an assumption to proceed on. Teams that skip this step and start writing new code based on what they believe the system does, rather than what it's actually observed to do, are the ones whose rewrites silently change behavior and discover it in production months later.
You inherit a system with almost no documentation, dependencies nobody wrote down, and the one engineer who understood it just left. What do you actually do in the first weeks, and how do you avoid either freezing all feature work or making the risk worse?
Sample Answer
Direct answer
The first weeks are about reducing risk cheaply, not about rewriting anything: map what the system actually does and who depends on it, put a safety net under it (monitoring and tests) before touching the code, and only then start making small, reversible changes. The goal by day 90 is not "the system is modernized," it is "the system is no longer a black box, and the org has evidence about what's safe to touch," while feature work continues in parallel rather than freezing entirely.
Structured elaboration
A workable structure:
- Days 1 to 30, discovery. Build a dependency map using a combination of static analysis (what does the code call), dynamic tracing and log correlation (what actually happens in production, which is often different from what the code suggests), and conversations with anyone who has touched the system, since undocumented tribal knowledge is itself a source you have to capture before it walks out the door. This is also when you look for lightweight, low-production-impact ways to instrument the system if it has no observability at all.
- Days 30 to 60, safety net. Add monitoring and alerting so you would actually notice if something broke, and add characterization tests, tests that pin down the system's current behavior (correct or not) so a later change that alters that behavior gets caught immediately, rather than tests that assert what the behavior should be, which requires understanding the system better than you do yet.
- Days 60 to 90, first small changes. Make a handful of low-risk, reversible improvements, fixing the most painful operational issue, extracting the most clearly separable piece, and use these as a way to validate that your dependency map and safety net actually work, before committing to anything bigger.
Throughout, the discovery approach itself matters: static analysis alone misses runtime-only dependencies (a job triggered by a cron entry nobody documented, a hidden call made only under a rare condition), so combining it with dynamic tracing, network traffic capture, and log correlation catches what static analysis alone would miss, at low production impact since these are observational techniques, not changes to the system itself.
To avoid freezing feature work: communicate explicitly that discovery and safety-net work is happening in parallel with, not instead of, feature delivery, and pick the first few changes specifically because they are small enough not to threaten the delivery timeline while still proving the approach works.
Worked example
An engineer inherits a legacy payment-reconciliation service: no documentation, three flaky integration tests, and the one person who understood it left six months ago.
- Week 1 to 2: they run static analysis to find every internal call path, and separately turn on request logging (a low-impact, purely observational change) to see what actually gets called in production. The two do not fully agree: static analysis misses a nightly batch job triggered by an external cron system nobody had documented, which the logs reveal because it shows up as unexplained traffic at 2am.
- Week 3 to 6: they build monitoring on the reconciliation service's key outputs (does the daily reconciliation total match expectations) so a regression would actually be visible, and write characterization tests around the three most business-critical code paths, capturing current behavior rather than guessing at intended behavior.
- Week 7 to 12: with the safety net in place, they fix the most painful operational issue (a memory leak that forces a weekly manual restart) as the first real change, verify the characterization tests and monitoring both catch the change as expected (a sanity check that the safety net itself works), and report back to stakeholders with an actual map of the system's dependencies and risk areas, rather than a vague "it's better now."
Trade-offs and pitfalls
The tempting shortcut is to skip straight to fixing the most obviously bad code, but without a dependency map and a safety net first, you cannot tell whether a "fix" broke something you didn't know depended on the old behavior, which is exactly how well-intentioned early changes to an undocumented system make things worse rather than better. The other trap is treating discovery as a one-time exercise rather than an ongoing habit; legacy systems that have been running for years often have dependencies that only surface under conditions (end of quarter, a specific customer's data shape) you will not see in the first 90 days no matter how thorough you are.
What is the anti-corruption layer pattern, and what job is it actually doing when you put one between a legacy system and a new one? Walk through a concrete example of translating legacy data into a new service's model.
Sample Answer
Direct answer
An anti-corruption layer (ACL) is a translation boundary you deliberately put between a legacy system and a new one so that the new system's domain model never has to bend to accommodate the legacy system's quirks. It is not just an adapter that converts data formats; its job is to protect the new model's integrity by absorbing all the legacy system's inconsistencies, missing fields, and outdated assumptions on the legacy side of the boundary, so nothing about the old system's design leaks into the new one.
Structured elaboration
Concretely, an ACL is responsible for:
- Translation: converting the legacy system's data shapes, field names, and units into the new system's domain model, not the other way around.
- Mapping semantic gaps: the legacy system might represent a concept the new system does not have an exact equivalent for (a status enum with legacy-only values, a field that means two different things depending on another field). The ACL is where you decide how those map, once, in one place, instead of every consumer inventing its own interpretation.
- Isolation: consumers on the new side never call the legacy system directly or see its raw shapes. If the legacy system changes (or if you eventually replace it), only the ACL has to change.
The benefit is that your new services get to have a clean domain model that reflects how the business actually works today, not how a fifteen-year-old system happened to represent it. The cost is that the ACL itself becomes a piece of infrastructure someone has to own, test, and keep in sync as both sides evolve, and a badly maintained ACL can become exactly the kind of tangled legacy code it was meant to prevent.
Worked example
Say a legacy order system represents order status as an integer code (0, 1, 2, 9) where 9 means "cancelled" but also gets reused for "refunded" depending on an unrelated flag elsewhere in the record, a real and common kind of legacy inconsistency. A new order service wants a clean OrderStatus enum: PENDING, CONFIRMED, SHIPPED, CANCELLED, REFUNDED.
The ACL sits at the boundary and does the translation:
def translate_legacy_status(legacy_code: int, refund_flag: bool) -> str:
if legacy_code == 9 and refund_flag:
return "REFUNDED"
mapping = {0: "PENDING", 1: "CONFIRMED", 2: "SHIPPED", 9: "CANCELLED"}
return mapping[legacy_code]
Every consumer on the new side calls the ACL and gets back a clean OrderStatus, never the raw integer code or the refund flag. If the legacy system later adds a sixth status code, only this one function needs to change.
To validate this kind of adapter, the concrete test strategy is a contract test against a fixture of known legacy inputs paired with their expected new-model outputs, covering every legacy value (including the ambiguous ones like the reused code 9) and every combination the ACL has to disambiguate, not just the happy path. That test suite is what tells you the ACL is behaving correctly before any real traffic depends on it, and it is what catches the translation silently breaking if someone touches the mapping later.
Trade-offs and pitfalls
The main pitfall is letting the ACL grow into a second copy of legacy logic instead of a thin translation boundary: if it starts encoding business rules of its own rather than just mapping shapes, you have created a new piece of legacy code, not protected against the old one. The other common mistake is skipping the ACL for "just this one caller" because it seems faster, which reliably ends with the legacy system's quirks leaking into the new domain model through that one exception, and everyone downstream having to account for it.
Unlock Full Question Bank
Get access to all 16 Legacy Modernization and Architecture Evolution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.