Legacy Modernization and Architecture Evolution Questions
Evolving an existing system rather than designing greenfield. Covers the modernization patterns (strangler fig, anti-corruption layers, facades and protocol adapters), choosing between rehosting, replatforming, incremental refactoring and a full rewrite, data migration and coexistence (dual-running, change-data-capture versus bulk cutover, reconciliation and drift), cutover readiness and decommissioning, recovering undocumented behavior from legacy code and stored procedures, instrumenting a migration so you can tell in real time whether it is working, and the organizational and risk management of long migrations across many teams. The scope is the migration itself: not quantifying or prioritizing technical debt, not code-level refactoring craft, not how to decompose a system into microservices, and not cloud migration or deployment and rollback mechanics as topics in their own right.
You inherit a system with almost no documentation, dependencies nobody wrote down, and the one engineer who understood it just left. What do you actually do in the first weeks, and how do you avoid either freezing all feature work or making the risk worse?
Sample Answer
Direct answer
The first weeks are about reducing risk cheaply, not about rewriting anything: map what the system actually does and who depends on it, put a safety net under it (monitoring and tests) before touching the code, and only then start making small, reversible changes. The goal by day 90 is not "the system is modernized," it is "the system is no longer a black box, and the org has evidence about what's safe to touch," while feature work continues in parallel rather than freezing entirely.
Structured elaboration
A workable structure:
- Days 1 to 30, discovery. Build a dependency map using a combination of static analysis (what does the code call), dynamic tracing and log correlation (what actually happens in production, which is often different from what the code suggests), and conversations with anyone who has touched the system, since undocumented tribal knowledge is itself a source you have to capture before it walks out the door. This is also when you look for lightweight, low-production-impact ways to instrument the system if it has no observability at all.
- Days 30 to 60, safety net. Add monitoring and alerting so you would actually notice if something broke, and add characterization tests, tests that pin down the system's current behavior (correct or not) so a later change that alters that behavior gets caught immediately, rather than tests that assert what the behavior should be, which requires understanding the system better than you do yet.
- Days 60 to 90, first small changes. Make a handful of low-risk, reversible improvements, fixing the most painful operational issue, extracting the most clearly separable piece, and use these as a way to validate that your dependency map and safety net actually work, before committing to anything bigger.
Throughout, the discovery approach itself matters: static analysis alone misses runtime-only dependencies (a job triggered by a cron entry nobody documented, a hidden call made only under a rare condition), so combining it with dynamic tracing, network traffic capture, and log correlation catches what static analysis alone would miss, at low production impact since these are observational techniques, not changes to the system itself.
To avoid freezing feature work: communicate explicitly that discovery and safety-net work is happening in parallel with, not instead of, feature delivery, and pick the first few changes specifically because they are small enough not to threaten the delivery timeline while still proving the approach works.
Worked example
An engineer inherits a legacy payment-reconciliation service: no documentation, three flaky integration tests, and the one person who understood it left six months ago.
- Week 1 to 2: they run static analysis to find every internal call path, and separately turn on request logging (a low-impact, purely observational change) to see what actually gets called in production. The two do not fully agree: static analysis misses a nightly batch job triggered by an external cron system nobody had documented, which the logs reveal because it shows up as unexplained traffic at 2am.
- Week 3 to 6: they build monitoring on the reconciliation service's key outputs (does the daily reconciliation total match expectations) so a regression would actually be visible, and write characterization tests around the three most business-critical code paths, capturing current behavior rather than guessing at intended behavior.
- Week 7 to 12: with the safety net in place, they fix the most painful operational issue (a memory leak that forces a weekly manual restart) as the first real change, verify the characterization tests and monitoring both catch the change as expected (a sanity check that the safety net itself works), and report back to stakeholders with an actual map of the system's dependencies and risk areas, rather than a vague "it's better now."
Trade-offs and pitfalls
The tempting shortcut is to skip straight to fixing the most obviously bad code, but without a dependency map and a safety net first, you cannot tell whether a "fix" broke something you didn't know depended on the old behavior, which is exactly how well-intentioned early changes to an undocumented system make things worse rather than better. The other trap is treating discovery as a one-time exercise rather than an ongoing habit; legacy systems that have been running for years often have dependencies that only surface under conditions (end of quarter, a specific customer's data shape) you will not see in the first 90 days no matter how thorough you are.
You inherit a legacy system running on infrastructure (a mainframe, an unsupported platform, or similar) that the business cannot tolerate downtime or data loss from, with real compliance exposure if you get the migration wrong. How do you approach modernizing it?
Sample Answer
Direct answer
When the legacy system runs somewhere modern tooling barely reaches (a mainframe, an unsupported OS, proprietary hardware) and the business genuinely cannot absorb downtime or data loss, the approach is staged and conservative by design: rehost or wrap first to buy breathing room without touching the risky internals, then incrementally refactor or replace behind that wrapper, with retirement criteria and compliance sign-off treated as first-class deliverables, not an afterthought once the technical work is done.
Structured elaboration
A workable staged plan looks like:
- Rehost or API-wrap before you refactor. Get the system onto supportable infrastructure, or put a modern API in front of it, without changing its internal logic. This buys time and reduces operational risk (older hardware failing, staff who understand the system leaving) without touching the highest-risk code.
- Cost and risk analysis before committing to a target architecture. For a system with real compliance exposure, the honest options are usually rehost, wrap-and-strangle (wrap the legacy system behind a new API now, then incrementally re-implement and cut pieces over behind that wrapper later), or replace with a vendor product, and the choice should be driven by where the audit and revenue risk actually sits, not by which option is technically most interesting.
- Refactor incrementally behind the wrapper, treating each extracted piece with the same rigor as the strangler-fig capability extraction described above (the same wrap-and-strangle approach, now executed piece by piece): parallel-run it against the legacy path, verify identical output, then cut over.
- Define retirement criteria up front. For a compliance-sensitive system this usually means: a defined retention period for historical data in a queryable archive, a compliance sign-off checklist, and proof that every downstream consumer has migrated off the legacy interface, not just that the new system works.
- Build governance into the plan, not around it: who approves each cutover stage, what the audit trail requirements are for the migration itself (not just the resulting system), and how revenue-affecting changes get an extra review gate.
If the system is used globally, add explicit handling for data sovereignty (which regions' data can legally move where), multi-region active-active requirements if the system serves users with low-latency expectations across regions, and a per-region cutover sequence with its own compliance validation, since a single global cutover date is rarely realistic once regional regulatory approval is in the critical path.
On rehost-first ("forklift") specifically: virtualizing or rehosting the legacy application before refactoring reduces operational risk quickly, but it does not reduce the technical or compliance risk buried in the application logic itself, so it should be paired with parallel testing and a prioritized backlog of what still needs refactoring, not treated as the finish line.
Worked example
A team modernizing a mainframe billing system doing monthly batch settlements, with zero tolerance for revenue leakage:
- Phase 1 (rehost): they move the mainframe workload to an emulated environment on modern infrastructure, changing nothing about the COBOL logic, primarily to stop depending on aging hardware nobody can source parts for anymore.
- Phase 2 (API wrap): they build a service that exposes the settlement process through a modern API, so new reporting and reconciliation tools can integrate without touching the COBOL directly.
- Phase 3 (incremental refactor): they identify the least revenue-sensitive settlement rule (a fee calculation used by a small customer segment) and reimplement just that rule in a new service, running it in parallel against the legacy output for two full billing cycles before trusting it, because "the numbers matched in staging" is not sufficient evidence for a system with zero tolerance for revenue leakage.
- Retirement criteria: the legacy COBOL code for a given rule is only removed once its replacement has matched legacy output for a full compliance-relevant reporting period, and the audit function has signed off that the new system's controls satisfy the same requirements the old one did.
Trade-offs and pitfalls
The trade-off is speed versus provable correctness: the staged, parallel-run-everything approach is slow, but for a system where a mistake means real revenue leakage or a compliance failure, that slowness is the entire point, not an inefficiency to optimize away. The most damaging pitfall is treating the rehost step as the modernization itself and declaring victory once the hardware problem is solved, leaving the actual compliance and technical debt in the COBOL logic completely untouched.
Walk through how you'd run a staged, dual-run migration where the old and new systems operate side by side for a while. What has to be true before you cut over the last piece of traffic?
Sample Answer
Direct answer
A staged, dual-run migration keeps the old and new systems live at the same time, moving traffic (or workloads) over gradually while continuously checking that the new side agrees with the old one, so you never bet everything on a single cutover moment. What has to be true before the last piece of traffic moves: the new system has matched the old one's behavior consistently, for long enough, across enough real traffic, that the remaining risk is genuinely small, not just "it's worked so far."
Structured elaboration
A concrete staged approach:
- Stand up the new system alongside the old one, with both able to receive traffic, but only the old one authoritative at the start.
- Shadow or dual-run a slice of traffic: send a copy of real requests to the new system without letting its response count, and compare its output against the old system's real response. This validates behavior with zero user-facing risk, since nothing depends on the new system's answer yet.
- Move real traffic incrementally, starting with a small, low-risk percentage or cohort, watching error rates, latency, and any business-correctness metrics you can measure automatically, before increasing.
- Reconciliation checks throughout, not just at each stage boundary: continuously compare outputs (or data state) between the two systems so drift is caught early rather than accumulating silently.
- Cutover gating: define explicit criteria for moving to the next stage, error rate under a threshold, reconciliation clean for a minimum window, no open incidents tied to the new system, rather than a subjective "seems fine" call.
For workloads with hard data-consistency requirements (a payments ledger, an inventory system, anything a customer-facing application depends on), the reconciliation step needs to specifically check data consistency, not just functional correctness, since the two systems could return the same answer to a query while their underlying data has quietly diverged.
A permanent variant of this pattern shows up whenever a piece of the old system is deliberately kept running for good, rather than the coexistence period being purely transitional. That's a different commitment than a staged cutover, and it needs a different kind of discipline: a stable, explicit integration boundary between the old and new pieces (an anti-corruption layer, not an ad hoc set of point-to-point calls that grows over time), clear, permanent ownership assigned to whoever maintains the piece that stays, so it doesn't quietly rot once attention moves to the new system, and a periodic, scheduled re-evaluation of whether "stays forever" is still the right call, since the reasons a capability was kept back (cost, risk, a hard dependency) can change well after the original migration project has wrapped up and everyone has moved on.
Worked example
Migrating a customer-facing web application from a legacy monolith to a new service-based platform in stages, preserving data consistency:
- Stage 0: the new version is deployed and can serve traffic, but all real traffic still goes to the old system. A shadow pipeline sends a copy of production requests to the new version and logs differences without affecting real users.
- Stage 1: shadow comparison runs clean for two weeks (response equivalence above 99.9%, with every discrepancy investigated and either fixed or explained). The team starts routing 5% of real traffic to the new version, chosen as a specific low-risk customer segment, with an instant rollback flag.
- Stage 2 through 4: traffic ramps 5% to 25% to 100% over several weeks, gated at each stage by error rate staying flat and a data-consistency check (comparing a sample of database state between the old and new systems) staying clean.
- Cutover complete: once at 100% for a defined period with no reconciliation issues, the old system is marked as the fallback rather than primary, and eventually decommissioned following the same discipline as any other legacy retirement.
Trade-offs and pitfalls
The main trade-off is time against certainty: a staged rollout with real gating criteria takes meaningfully longer than a single cutover, but it converts "we hope this works" into "we have evidence this works," which is exactly the trade a business accepting real downtime risk is making when it agrees to this approach. The pitfall to watch for is gating criteria that are too loose to catch a real problem (an error-rate threshold set so high that a meaningful regression still passes) or reconciliation checks that only validate the easy cases, functional correctness is necessary but not sufficient if the underlying data can still silently diverge.
Explain the strangler fig pattern for retiring a legacy system incrementally. What are the most common ways teams get it wrong in practice, and what signals tell you strangling is the right call versus a full rewrite?
Sample Answer
Direct answer
Strangler fig retires a legacy system by growing a new one around it: you put a routing layer (a proxy, gateway, or facade) in front of the legacy system, move one capability at a time behind that layer to a new implementation, and let the old and new code paths coexist until every capability has moved and the legacy system can be turned off. The name comes from the strangler fig vine, which grows around a host tree until the host is no longer needed. It is popular precisely because it avoids the two failure modes of a big-bang rewrite: shipping nothing for a year, and cutting over everything at once with no way back.
Structured elaboration
The mechanics, in order:
- Put a seam in front of the legacy system. Usually an API gateway, reverse proxy, or a facade inside the monolith itself, so callers do not know or care whether a request lands on old or new code.
- Pick the first capability to extract, usually the one that is both low-risk and high-pain (a module that changes often but is not the most business-critical, so mistakes are cheap to learn from).
- Build the new implementation and route a slice of traffic to it, verifying its output matches the old path before trusting it fully.
- Repeat, capability by capability, until nothing is left behind the seam pointing at the legacy system.
- Decommission the legacy code only once nothing routes to it and you have confirmed there are no hidden callers.
The most common ways teams get this wrong, in rough order of how often they show up:
- Never actually finishing. The easy 80% of capabilities get strangled in the first few months, and the hard 20% (the ones with the messiest coupling to legacy state) get deferred indefinitely. Two years later the org is paying to run and secure both systems forever, which is worse than either a rewrite or leaving the legacy system alone. This is the single most common failure and the reason a strangler effort needs a decommission target with a rough date attached, not just a start.
- Building a permanent adapter instead of a temporary one. The seam is meant to shrink as capabilities move; if it keeps growing new special cases instead, it has quietly become a second system to maintain, not a migration path.
- Not enforcing a hard boundary between the shared state. If the new service and the legacy system both write to the same tables without a clear ownership rule, you get silent data corruption long before anyone notices a functional bug.
- Treating the seam as free. Every hop through a translation layer costs latency and adds a new thing that can fail; teams who never measure this get surprised when the "temporary" adapter becomes the slowest part of the system.
Strangling is usually the right call when the system has to keep serving traffic throughout the change (most production systems), when you can identify genuinely separable capabilities, and when the org can tolerate running two systems for a while. A full rewrite becomes more attractive when the legacy system's capabilities are too entangled to peel apart one at a time, when the business can tolerate a real code freeze, or when the legacy code is so far from correct that incrementally wrapping it just preserves its bugs behind a nicer API.
Worked example
Say a monolith handles catalog, cart, checkout, and recommendations for an e-commerce site. A team decides to strangle it:
- They put an API gateway in front of all four capabilities.
- They pick recommendations first: it changes often, has no write path into the order/payment data, and a bug there degrades the experience rather than losing money.
- They build a new recommendations service, route 5% of traffic to it behind a flag, compare its output to the legacy path for a few weeks, then ramp to 100% and delete the legacy recommendations code.
- They repeat for catalog, then cart, and leave checkout, the most state-heavy and highest-risk capability, for last, once the team has practiced the pattern three times on lower-stakes capabilities.
- Eighteen months in, nothing routes to the legacy monolith and it is decommissioned.
The order matters: doing checkout first, before the team has proven the pattern on anything, is exactly the kind of decision that produces the abandoned-halfway failure mode above.
Trade-offs and pitfalls
Strangling trades speed for safety: you ship value continuously and can stop or reverse at almost any point, but you pay for it in the ongoing cost of running two systems and maintaining the seam between them, and in the discipline required to actually finish rather than stall. A full rewrite is the opposite bet: faster in principle if nothing goes wrong, but an all-or-nothing wager on a fixed-price estimate for a system whose exact behavior nobody has fully mapped, which is exactly the situation legacy modernization starts from. The senior mistake to watch for is choosing strangling for the safety story and then never applying the same rigor to actually retiring the legacy code, which converts a migration strategy into permanent architectural debt.
Design a decommissioning plan for shutting down a legacy system after its replacement has taken over. What has to be true before you actually delete anything?
Sample Answer
Direct answer
Before you delete anything, you need proof the migration actually succeeded (not just that the new system is live), a defined retention and archival plan for whatever legal or audit obligations outlive the system itself, a rollback path in case something surfaces after decommission that the pre-cutover testing missed, and a communication plan that reaches every stakeholder who might still depend on the old system, including the ones you do not already know about.
Structured elaboration
A decommissioning plan has four parts, and skipping any of them is how "the migration is done" turns into "we deleted data we needed":
- Verification that the migration is actually complete. Not "the new system works," but "nothing depends on the old one anymore." This means auditing traffic and access logs on the legacy system for weeks after the cutover, not just at the moment of cutover, because low-frequency dependencies (a monthly batch job, a quarterly report) will not show up in a one-week traffic sample.
- Legal and audit retention requirements. Many systems have data that has to remain queryable for years after the system itself is gone, for regulatory or contractual reasons. That means an archival strategy decided before decommission, not scrambled together after someone asks for five-year-old records the week after you deleted the database.
- A rollback plan for the decommission itself, distinct from the rollback plan for the original migration. If something surfaces after you have shut the legacy system down (a caller nobody knew about, a data discrepancy only visible under a rare condition), you need a defined path back, even if that path is "restore from the last verified backup and re-enable the legacy code path for a bounded window," not "we have no idea, we deleted it."
- Third-party integrations and stakeholder communication. External partners often integrate with systems in ways your internal traffic logs cannot see (a partner polling an API you exposed to them specifically). The communication plan needs to reach them with enough lead time to migrate on their side, not just notify internal teams.
Worked example
A team decommissioning a legacy order-management system after a successful migration:
- They keep the legacy system read-only (not deleted) for 90 days post-cutover, monitoring access logs the whole time. In week six, they find a quarterly compliance report job still reading directly from the legacy database, which nobody had flagged as a dependency because it only runs four times a year.
- They export the full historical dataset to a queryable archive with a retention period matching the company's seven-year audit requirement, and verify a sample of archived records against the live system before the live system goes away, since an archive nobody has tested is not actually a safety net.
- They notify the three external partners who integrate with the legacy system's API directly, giving them 60 days' notice and a migration guide, rather than assuming internal migration alone covers everyone with a dependency.
- Only after all of this, and after a final confirmed zero-traffic week on the legacy system, do they actually shut it down, with the archived data and a documented restore procedure kept in case something surfaces later.
Trade-offs and pitfalls
The tempting shortcut is to declare victory the moment the new system handles 100% of live traffic and decommission immediately, but "no traffic this week" is not the same as "no dependencies," and the cost of being wrong (deleted data you needed, a partner integration silently broken) is far higher than the cost of a monitored grace period before deletion. The other common mistake is treating archival as a technical afterthought rather than a compliance requirement with its own sign-off, which is how companies end up unable to produce records a regulator or auditor asks for.
Unlock Full Question Bank
Get access to all 31 Legacy Modernization and Architecture Evolution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.