Cloud Migration Strategy and Execution Questions
Planning and executing a move to the cloud: the migration strategies (rehost, replatform, refactor, repurchase, retire, retain), legacy assessment, dependency mapping, cutover planning, and rollback. Covers phased migration roadmaps, workload modernization, risk management during cutover, and validating success post-migration. The end-to-end migration lifecycle, not steady-state operations.
Compare the common cutover strategies used in migrations (big-bang cutover, phased migration, blue-green deployment, canary rollout, parallel-run). For each strategy describe typical use cases, advantages, risks, and rollback considerations. Provide guidance on choosing a cutover approach for an enterprise-facing customer service application with moderate traffic and strict SLA requirements.
Sample Answer
Direct answer: Big-bang, phased, blue-green, canary, and parallel-run cutover strategies trade off risk exposure against complexity and duration; for an enterprise customer-facing application with moderate traffic and strict SLA requirements, blue-green or a phased approach usually wins over a big-bang cutover, since the SLA requirement makes rollback speed a priority the simpler big-bang approach doesn't provide as cleanly.
Structured elaboration. Big-bang cutover: switch everything at once at a scheduled moment; simplest to plan and execute, but highest risk (if something's wrong, the FULL user base is affected immediately) and rollback means reversing the whole cutover, not a subset. Best for low-traffic, low-SLA-risk applications where the simplicity is worth the concentrated risk. Phased migration: move distinct components or user segments over time (e.g., migrate by geography or by feature area); reduces blast radius per phase, but requires the old and new systems to coexist and interoperate correctly for the duration, adding architectural complexity. Rollback means reverting one phase's component or segment at a time, back to the old system, which is only safe if the not-yet-migrated phases can still interoperate correctly with whichever phase just got reverted. Blue-green deployment: stand up the full new environment (green) alongside the still-live old one (blue), validate it, then switch traffic over via a routing change (DNS or load balancer); rollback is simply switching the router back, making it one of the fastest-to-rollback options, at the cost of running two full environments simultaneously (cost overhead) for the transition period. Canary rollout: shift a small percentage of traffic to the new environment first, watching metrics closely, before progressively increasing; lowest risk exposure per increment (only a small user slice is affected by any given step), but takes the longest to reach full cutover and requires infrastructure capable of fine-grained traffic splitting. Rollback is simply reducing the traffic percentage back toward zero, which is why canary is often held up as the lowest-regret option: the blast radius of a bad step is capped by construction, and reversing it is a routing change, not a data operation. Parallel-run: run both systems processing the same input simultaneously WITHOUT cutting user traffic over, comparing outputs to validate correctness before any real cutover happens; adds the most validation confidence but the most complexity (need a way to compare outputs meaningfully) and doesn't reduce cutover risk itself, just the UNCERTAINTY about whether the new system is correct before that risk is taken. Rollback in the traditional sense doesn't apply here: since real user traffic was never cut over to the new system in the first place, there is nothing user-facing to revert, only a decision to keep the new system in shadow mode longer or promote it once its outputs are trusted.
Guidance for the stated scenario. For an enterprise customer-facing app with moderate traffic and strict SLA: blue-green is a strong default (full environment validated in advance, fast rollback via a routing switch, moderate complexity), and can be COMBINED with a brief canary step (shift a small percentage through the new environment first, even within a blue-green setup) to catch any issue before the full switch, giving the fast-rollback property of blue-green with some of canary's exposure-limiting benefit.
Worked example. Stand up the new (green) environment fully, run a parallel-run validation period comparing outputs against production traffic without affecting users, then execute the cutover as a canary-within-blue-green: shift 5% of traffic to green, hold and monitor against the SLA's specific metrics for a defined period, then 25%, then 100%, with an automated or fast-manual rollback (routing back to blue) available at every step.
Trade-offs & pitfalls. Choosing big-bang purely because it's operationally simpler to plan, without weighing that simplicity against a strict-SLA application's actual risk tolerance, is a common mismatch; the right cutover strategy should be selected against the application's stated risk tolerance, not the team's preference for a simpler runbook.
Describe a practical, repeatable framework you would use to prioritize workloads for migration across an enterprise portfolio. Include technical and business criteria, how you'd score or weight them, how to include dependency and risk, and an example of how that prioritization changes across pilot, early waves, and later waves.
Sample Answer
Direct answer: A repeatable prioritization framework scores each workload on two independent axes, business value and migration complexity/risk, then sequences waves to front-load low-complexity/high-learning applications before tackling high-value, high-risk ones once the process is proven.
Structured elaboration. Technical criteria: dependency count and coupling tightness (more dependencies = harder to isolate for a wave), data volume and sensitivity, current infrastructure age/support risk (an unsupported OS is a forcing function), and technical migration effort (does an obvious managed-service or lift-and-shift path exist, or does it need bespoke work). Business criteria: revenue/criticality impact if something goes wrong, compliance exposure, and stakeholder appetite for near-term change (some business units are more risk-tolerant than others). Scoring: a simple 1-5 scale per criterion, weighted (business criticality often weighted higher than raw technical effort, since the cost of a botched migration on a critical app dwarfs the engineering time saved on an easy one), summed into two composite scores: business value and migration risk. Dependency and risk inclusion: any application with high fan-in/fan-out dependency count (fan-in = how many other systems depend on it; fan-out = how many other systems it depends on) gets an automatic risk bump regardless of its own technical simplicity, since its blast radius on failure is larger than its own footprint suggests.
Worked example. Plot every application on a 2x2 of business value vs. risk. Pilot wave: low-risk, low-to-medium value apps (proves the migration process cheaply). Early waves: low-risk, high-value apps (captures value fast while the process is still being refined) plus a few medium-risk apps to stress-test the process before the hardest work. Later waves: high-risk applications regardless of value, since by then the team has the most operational experience, and high-risk/low-value apps are re-examined for Retire instead of migrating them at all.
Trade-offs & pitfalls. The common mistake is sequencing purely by technical ease, which defers the highest-business-value applications indefinitely and never proves the process against real complexity until the deadline is close; the framework above deliberately forces some medium-risk exposure into the early waves specifically to avoid that trap.
Explain how DNS cutover works during migration. Describe the role of TTL, phased cutover strategies (blue/green, canary), and one method to minimize client-side caching issues during DNS-based migration.
Sample Answer
Direct answer: DNS cutover during migration works by changing which IP address(es) a domain resolves to, but because DNS resolutions are cached (by resolvers, browsers, and ISPs) according to the record's TTL, the cutover isn't instantaneous: clients using a cached (stale) resolution keep hitting the old system until their cache expires, which is the central operational fact every DNS-based cutover has to plan around.
Structured elaboration. Role of TTL: TTL (time-to-live) controls how long a DNS resolver caches a record before re-querying; a HIGH TTL means fewer DNS lookups (good for normal operation, lower latency/load) but a SLOW cutover (clients keep using the old answer for up to the TTL duration after the record changes); a LOW TTL means faster cutover propagation but more frequent DNS queries. The standard practice is to LOWER the TTL well before the planned cutover (e.g., reduce it from a normal 24 hours down to 60 seconds, days ahead of the actual cutover), wait for the old high-TTL value to fully expire from caches, THEN perform the cutover, so the actual traffic-shift propagates quickly once it happens. Phased cutover strategies (blue/green, canary): DNS itself can support a phased shift via weighted routing (some DNS providers support returning different answers to different resolvers with configurable weights, effectively canarying at the DNS layer) rather than an all-or-nothing record change; alternatively, a load balancer or traffic-management layer BEHIND a single DNS name can do the fine-grained blue/green or canary shifting, with DNS itself only pointing at that layer and never needing to change again for that specific cutover. Minimizing client-side caching issues: beyond server-side TTL, some clients (older browsers, some corporate proxies, some OS-level DNS caches) may not perfectly respect a lowered TTL and can cache longer than instructed; the practical mitigation is the same TTL-lowering discipline PLUS keeping the old environment available and functioning for a bake period well beyond the nominal TTL, so any slow-to-update client isn't hitting a dead endpoint even if it's slower than expected to pick up the new DNS answer.
Worked example. A cutover planned for a specific date: 1 week prior, lower the domain's TTL from 24 hours to 60 seconds; wait at least 24 hours (the OLD TTL's duration) to ensure any cached copies of the high-TTL record have expired everywhere; on cutover day, change the DNS record to point at the new environment; within roughly 60-120 seconds (the new low TTL, plus some propagation slack), the large majority of traffic is hitting the new environment, though a small tail of non-compliant caches may take longer, which is why the old environment stays live and functional for a bake period afterward rather than being decommissioned the moment the DNS record changes.
Trade-offs & pitfalls. Changing the DNS record WITHOUT first lowering the TTL well in advance is the classic mistake: if the record was at a 24-hour TTL when changed, a meaningful fraction of traffic keeps hitting the OLD environment for up to 24 hours after the team believes the cutover is complete, which is confusing at best and can mean the old environment needs to keep functioning (and being monitored) far longer than planned.
Explain hybrid/coexistence patterns used during migration: database replication/CDC for data sync, dual-write, strangler pattern for gradual refactor, API gateways/proxies for routing between on-prem and cloud components, and active-active vs active-passive modes. For each pattern describe when it is appropriate and the main operational considerations.
Sample Answer
Direct answer: Hybrid/coexistence patterns during migration exist to let old and new systems interoperate correctly while a migration is in progress; the main ones are database replication / change-data-capture (CDC) for data sync, dual-write, an API gateway/proxy for routing between on-prem and cloud components, and active-active vs active-passive as an operating mode for the coexistence period itself.
Structured elaboration. Database replication/CDC for data sync: keeps a target database current with an ongoing source of truth during migration, appropriate when the two systems don't both need to accept writes at the same time (one is clearly the source of truth while the other is a synchronized read target, until cutover flips which one is authoritative). Dual-write: the application writes to both old and new systems as part of normal operation during the coexistence period; useful when BOTH systems genuinely need to be current and usable (e.g., some traffic is already being served by the new system while some is still on the old one), at the cost of coordination risk (if one write succeeds and the other fails, the systems diverge) that needs an explicit reconciliation process to catch. Strangler pattern for gradual refactor: routes an increasing share of functionality to the new system over time while the old system still handles what hasn't yet been migrated, typically via a routing/proxy layer in front of both; most relevant when the migration is functionally incremental (moving one capability at a time) rather than a wholesale data-store swap. API gateways/proxies for routing between on-prem and cloud components: the mechanical layer that makes several of the above patterns possible, directing a given request to whichever system (old or new) currently owns that functionality or that user/tenant, and providing a single point to adjust routing as migration progresses. Active-active vs active-passive modes: active-active means both old and new systems are live and serving real traffic simultaneously (highest operational complexity, but enables gradual, low-risk traffic shifting); active-passive means one system is fully authoritative while the other is a synchronized standby not yet serving traffic (simpler to reason about, but the cutover moment is more of a discrete event rather than a gradual shift).
When each is appropriate and main operational considerations. Replication/CDC: appropriate for a straightforward data-store migration with a clear before/after cutover moment; operational consideration is monitoring replication lag and validating parity continuously. Dual-write: appropriate when gradual, traffic-based cutover (not a single data-store swap) is needed; operational consideration is building and monitoring reconciliation to catch write-coordination failures. Strangler pattern: appropriate for a functionally-incremental migration; operational consideration is maintaining the routing layer's correctness as the split between old/new functionality shifts over time. API gateways/proxies: appropriate whenever any of the other patterns need a single, well-defined place to route requests between old and new (it's the mechanical enabler behind strangler-pattern routing and behind gradual traffic shifting generally, not a competing alternative to them); the main operational consideration is keeping the routing rules themselves correct and low-latency as the old/new split shifts over time, since a bug in the gateway's routing logic can silently misroute traffic in either direction. Active-active: appropriate when a gradual traffic-percentage cutover is the goal; operational consideration is the doubled infrastructure cost and the complexity of keeping both systems genuinely consistent while both are live. Active-passive: appropriate when a discrete cutover event is acceptable; operational consideration is that a longer time to build confidence before cutover is needed, since there's no gradual traffic-based signal along the way.
Worked example. A migration moving traffic gradually to a new backend: an API gateway routes requests based on a rollout percentage (active-active), the new backend's database stays synchronized via CDC from the still-authoritative old database (until a later point where dual-write or a full cutover flips authority), giving a combination of patterns rather than a single one in isolation, which is typical of a real migration.
Trade-offs & pitfalls. Reaching for dual-write by default because it "keeps both systems current" without first asking whether a simpler replication/CDC pattern (single source of truth, one-directional sync) would satisfy the actual requirement is a common overcomplication; dual-write's coordination risk should be accepted only when the migration genuinely needs both systems live and authoritative simultaneously, not as a default choice.
Design a test strategy for validating a migrated application. Cover unit/integration/acceptance tests, performance and load testing, security scans, and user-acceptance testing. Explain the sequencing and pass/fail criteria for each test type.
Sample Answer
Direct answer: A migrated application's test strategy needs layered coverage in a specific sequence: unit and integration tests first (confirm the migrated components work correctly in isolation and together), then automated acceptance tests (confirm end-to-end business scenarios work correctly against the migrated system, before involving performance or human validation), then performance/load testing (confirm the new environment meets the same or better performance characteristics), then security scans (confirm the new environment's configuration doesn't introduce vulnerabilities), and user-acceptance testing last (confirm real users/business stakeholders agree the migrated system meets their needs), with clear pass/fail criteria defined for each layer before testing starts.
Structured elaboration. Unit/integration tests: re-run the application's existing automated test suite against the migrated environment (if it doesn't pass at this level, nothing downstream matters); pass/fail criterion is straightforward: the existing suite passes at the same rate as it did pre-migration, with any NEW failures investigated as migration-introduced regressions. Acceptance tests (automated, distinct from both the unit/integration layer below it and the human-driven UAT above): run the existing suite of automated end-to-end/business-scenario tests (e.g., Gherkin/BDD-style scenarios or API-level contract tests exercising full user flows) against the migrated environment; pass/fail criterion is that every previously-passing acceptance scenario still passes unchanged, since this layer exists specifically to catch integration-level regressions across service boundaries that unit/integration tests, being narrower in scope, can miss. Performance and load testing: benchmark the migrated system against the SAME load profile and SLA targets as the pre-migration baseline (not a fresh, arbitrary target), since "is it fast enough" only means something relative to what it needs to replace; pass/fail criterion is defined against specific latency/throughput targets tied to the actual SLA, not a vague "seems fine" assessment. Security scans: run the org's standard security scanning (dependency vulnerabilities, configuration scanning for the new cloud environment specifically, since cloud misconfigurations are a distinct risk category from application-level vulnerabilities) against the migrated environment; pass/fail criterion is typically zero new HIGH/CRITICAL findings introduced by the migration itself (pre-existing findings in the application code aren't newly introduced by the migration and can be tracked separately). User-acceptance testing: business stakeholders/end users validate the migrated system against real workflows, ideally using the SAME acceptance scenarios used when the system was originally built/accepted, if those exist; pass/fail criterion is explicit sign-off from the designated business owner, not just "no complaints so far." Sequencing and pass/fail criteria: each layer gates the next (don't run expensive load tests against a build that's still failing unit tests; don't ask business stakeholders to validate a system that hasn't yet passed security scanning), and every layer's criteria should be defined and agreed BEFORE testing starts, not retrofitted to whatever results come back.
Worked example. For a migrated e-commerce checkout flow: unit/integration suite must pass at 100% parity with pre-migration results; the automated acceptance suite (e.g., "place an order with a saved card," "apply a discount code at checkout") must pass every previously-passing scenario against the migrated environment before load testing begins; load testing must sustain the peak historical Black-Friday-level traffic profile within the existing p99 latency SLA; security scanning must show zero new critical findings in the cloud environment's configuration (IAM policies, network exposure); UAT requires explicit sign-off from the product owner walking through the actual checkout flow, including edge cases like a failed payment and a cart abandonment, not just the happy path.
Trade-offs & pitfalls. Skipping or compressing the performance/load-testing layer under time pressure, reasoning that "if unit tests pass, it should be fine," is a common and risky shortcut: functional correctness and performance characteristics are genuinely independent, and a migrated system can pass every unit test while performing meaningfully worse under real production load due to the new environment's different resource characteristics or network topology.
Unlock Full Question Bank
Get access to all 34 Cloud Migration Strategy and Execution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.