Infrastructure Strategy and Technology Selection Questions
Setting technical direction for infrastructure and deciding what to build on. Covers infrastructure vision and long-term roadmap, modernization and technical-debt strategy, platform decisions, and organizational and governance considerations, alongside the decision framework for adopting or retiring technology: build-versus-buy-versus-cloud-versus-on-premises trade-offs, vendor and platform evaluation, technology-portfolio rationalization, requirements-driven service selection, and weighing total cost of ownership and risk. It also covers managing an existing vendor relationship: quantifying and mitigating lock-in, negotiating contract terms, and planning an exit from a service the organization depends on. The leadership-altitude discipline of directing an infrastructure estate over time, building consensus among stakeholders with conflicting priorities, and persuading leadership or a team to accept a major technology change.
You're planning a multi-quarter rollout to replace a legacy, widely-used internal tool with a modern replacement (for example replacing static legacy reporting with a self-service platform). Provide a sequencing plan across quarters, migration criteria for cutting over existing users, training and enablement activities, and risk mitigation.
Sample Answer
A multi-quarter replacement of a widely used legacy tool succeeds or fails on whether users can opt in gradually with an easy way back, so the plan sequences by user cohort risk, gates each cutover on measured adoption and parity rather than a calendar date, and treats training as a first-quarter activity, not a launch-week afterthought.
Sequencing plan across quarters
Quarter 1: parallel availability and a pilot cohort. Stand up the replacement alongside the legacy tool with no forced migration. Recruit a pilot cohort deliberately, choosing a mix of a few power users (who will find gaps fast and are motivated to help fix them) and a few average users (whose experience predicts the general rollout better than power users' does). Build the training material and a feedback channel in this quarter, since a self-service platform replacing static legacy reporting typically fails on discoverability and confidence, not raw feature parity.
Quarter 2: expand to early-adopter teams and close parity gaps. Widen access to teams that opted in after seeing the pilot cohort's results, and use their usage data (not just their stated satisfaction) to find the specific reports or workflows still missing from the replacement. Close the highest-usage gaps first, ranked by how many legacy-tool sessions touch that specific report or feature, not by how interesting the feature is to build.
Quarter 3: default-on for new usage, migration push for remaining teams. Make the replacement the default for anyone starting fresh, and begin actively migrating remaining teams on a schedule they help set, with a named migration criterion per team (below) that must be met before their legacy access is scheduled for removal.
Quarter 4: sunset legacy access team by team, not all at once. Turn off legacy access for each team individually once their migration criteria are met, keep the legacy system in a read-only archival state for a defined retention period afterward, and do not decommission it entirely until every team has confirmed cutover and the archival window has passed.
Migration criteria for cutting over an existing user
A team is ready to cut over from legacy access when all of the following hold: their top N most-used legacy reports or workflows (by their own usage logs) exist and produce matching results in the replacement, at least one full business cycle (commonly a month-end or quarter-end cycle, whichever the team's real usage pattern is) has been run successfully on the replacement without a fallback to legacy, and a named team contact has completed the training track and confirmed readiness in writing. These criteria are deliberately usage-based and team-specific rather than a single global date, because "everyone must be ready by June 1" is exactly the plan that forces a rushed cutover on the team with the most complex legacy usage.
Training and enablement activities
Build role-based training rather than one generic walkthrough: a report-consumer track (how to find and read what you used to get from the legacy tool) and a report-builder track (how to construct what the legacy tool's admins used to build for you), since a self-service replacement shifts work from a central team to end users, and that shift is the actual behavior change being asked of people, not just a new user interface. Run live office hours weekly during quarters 1 and 2 when confusion is highest, tapering to on-demand documentation and an asynchronous help channel by quarter 3 once the early-adopter cohort's questions have mostly been answered and recorded.
Risk mitigation
Keep a fast, well-known fallback path to the legacy tool available through quarter 3 for any team that hits a genuine blocker, since forcing a team off legacy before their blocker is fixed just pushes the failure into a production incident instead of a tracked migration risk. Track a leading indicator, not just a lagging one: weekly active usage of the replacement per team, so a team that is technically "cut over" but has quietly reverted to spreadsheets or the legacy tool through some side channel shows up in the data before their formal migration criteria are checked. Keep an explicit rollback plan per team (if the replacement breaks after their cutover, how fast can they get legacy access restored) since removing that safety net too early is what turns a manageable rollout into a trust-destroying incident for the team it happens to.
Trade-offs and pitfalls
The main trade-off is speed versus adoption durability: a faster, calendar-driven cutover looks better on a project timeline but produces teams that comply on paper while quietly working around the new tool, which resurfaces as a second, harder migration later. The most common pitfall is treating training as a one-time launch event rather than a sustained enablement track through the quarters where real usage volume is ramping, and a second is picking a pilot cohort of only power users, whose tolerance for rough edges does not predict how the average user will react once the rollout widens.
A company has accumulated a dozen overlapping point solutions (CRM, billing, analytics, support) across several years. Leadership wants a recommendation: integrate the existing landscape, replace it with a single suite, or build custom middleware. Present a structured analysis and recommendation.
Sample Answer
Direct answer
Diagnose the actual pain before picking a path; leadership's framing of "too many tools" tends to jump straight to "replace everything with one suite," when the real driver is very often a single broken data flow between two of the twelve tools, which a targeted integration fixes at a fraction of the cost and risk of a full suite replacement. Default to integrating over an existing landscape unless a specific, measurable trigger argues otherwise.
Structured elaboration
Diagnose first
Determine whether the actual complaint is data inconsistency across tools (a customer record disagreeing between systems), duplicate manual work (someone re-keying data between systems by hand), pure cost redundancy (paying for genuinely overlapping features), or a real capability gap. Each of these has a different correct answer, and none of them is answered by counting the number of tools in use.
Score the three options
Score integrate, replace with a single suite, and build custom middleware against: switching cost (data migration, retraining, and contract-termination cost for whatever gets replaced), the risk of disrupting revenue-generating workflows during the transition, time-to-value, and ongoing total cost of ownership.
Recommendation logic
- If the pain is data inconsistency or duplicate manual work, and the underlying tools are otherwise adequate, integrate: adopt or build a middleware or integration layer (an integration-platform-as-a-service, or an event-bus pattern that publishes changes from one system for others to consume) as the lowest-risk, fastest-time-to-value fix.
- If the pain is a genuine capability gap across most of the landscape, and a mature single suite covers the actual required capability set at acceptable cost, replace with a single suite, but budget for a multi-quarter phased cutover, never a big-bang replacement of a dozen tools at once.
- Build custom middleware beyond a thin integration layer is rarely the right default; recommend it only when the integration need is genuinely bespoke to the business's own data model and no viable off-the-shelf integration option covers it, because a custom middleware layer simply becomes a thirteenth system someone has to maintain indefinitely.
Worked example
Investigation traces the large majority of the reported pain, in this scenario, to one specific problem: the customer relationship management (CRM) system and the billing system disagree on customer status, because there is no synchronization between them, which causes revenue leakage from continuing to bill customers who have already churned according to the CRM. That single finding argues strongly for integrate: a lightweight event-bus or integration-platform connector synchronizing customer status from the CRM to billing, rather than a full twelve-tool suite replacement, which would cost far more, carry far more transition risk, and take far longer to deliver than fixing the one broken data flow that is actually driving the complaint.
Trade-offs and pitfalls
The most common overreach is letting leadership's "too many tools" framing drive straight to "buy one suite to replace them all" without first diagnosing what is actually broken. A close second is underestimating a suite replacement's true switching cost across a dozen already-integrated systems, since migration risk compounds with each additional connected system rather than adding up linearly. And custom middleware quietly built because it "feels like just glue code" has a well-documented failure mode: a few years later it is an unmaintained thirteenth system that nobody wants to own, which is worth naming explicitly as the reason it is not the default recommendation.
Design a proof-of-concept to compare a managed platform service against a self-managed equivalent before committing to one. Specify the experiments and metrics you'd collect (throughput, latency, operational overhead, mean time to recover), the failure-injection scenarios you'd run, and how you'd reach a go/no-go decision.
Sample Answer
A proof of concept (PoC) that actually settles a managed-versus-self-managed decision has its success and failure thresholds written down before any test runs, deliberately tests failure modes and not just peak throughput, and ends in a documented go/no-go against those pre-committed numbers rather than a subjective impression of "it felt fine."
Designing the PoC
- Metrics to collect: sustained throughput at a defined load, tail latency (p95/p99, the 95th and 99th percentile response time, which shows what your slowest real users experience, not just the average), operational overhead (hours of hands-on toil per week to run and tune it during the pilot), and mean time to recovery (MTTR) from an injected failure.
- Failure-injection scenarios, not just load tests: kill a node or process mid-traffic, saturate the CPU or network on a dependency, simulate an upstream dependency outage, and simulate a bad deploy that has to be rolled back. A platform that only gets load-tested on the happy path will pass a PoC and then fail in its first real incident, because the PoC never asked the question that actually matters operationally.
- Latency measured from the vantage point that matters: measure from the client's perspective (the caller actually waiting on the response), not just server-side processing time, since network hops and queuing between the two can dominate the difference between "the backend is fast" and "the user experienced it as fast." A benchmark that only reports server-side numbers is measuring the wrong thing for a latency-sensitive decision.
- Pre-committed go/no-go thresholds: set numeric thresholds for each metric before the test begins, plus at least one "kill criterion," a failure so serious it overrides every other metric (for example, any data loss during failover is an automatic no-go regardless of how good throughput looked).
- A lighter-weight variant for lower-stakes decisions: when the decision's blast radius doesn't justify a multi-week experiment-and-metrics exercise, a lightweight PoC template, objectives, scope, timeline, and deliverables, with a shorter and less rigorous set of go/no-go criteria, is a legitimate and honest alternative. The rigor of the PoC should match the cost of getting the decision wrong, not be applied uniformly to every choice.
Worked example (illustrative)
Consider comparing a managed message queue against a self-hosted equivalent for an order-processing pipeline that needs to sustain 2,000 messages per second at peak with a 99.9% availability target. That target implies roughly 8.8 hours of allowed downtime per year (0.001 x 365.25 x 24 ≈ 8.8), which is a useful number to have in hand before setting the PoC's failure-recovery threshold.
A three-week PoC plan: week one runs a sustained load test at 2,000 messages/second for at least an hour, recording p99 latency measured client-side; week two runs failure injection (kill a broker node, saturate disk I/O) and measures message loss and recovery time; week three tracks operational load, how many alerts fired and how many required a human to act. A team running this exercise might set thresholds along these lines before starting: p99 latency under 250ms at the target load, recovery from a node failure in under 2 minutes with zero message loss, and less than 2 hours per week of hands-on operational toil during the pilot. Those specific numbers are illustrative of how a team sets its bar, not a report of an actual measured run; the discipline that matters is committing to numbers like these in writing before the test, not the particular values chosen.
Trade-offs and pitfalls
- Testing only the happy path is the single most common way a PoC gives false confidence; the PoC's job is specifically to find the failure modes production would otherwise find for you, at a worse time.
- Moving the goalposts after seeing results, quietly relaxing a threshold because the preferred option narrowly missed it, defeats the entire purpose of pre-committing them. If a threshold needs to change, that change should be justified and documented separately from the result it happens to rescue.
- A PoC environment that doesn't match production's real data shape (message size distribution, traffic burstiness, request skew) can pass cleanly and still mislead, because throughput and failure behavior both depend heavily on that shape, not just on the average load.
Design a process for evaluating and approving new cloud services or technologies before they're adopted anywhere in the organization (for example a new managed storage product). Detail the steps, who needs to be involved, the risk checks, and the approval gates.
Sample Answer
An approval process for adopting new cloud services needs to be fast enough that teams use it instead of routing around it, and rigorous enough that a genuinely risky choice (a new database with no support for your compliance regime, for example) gets caught before it is embedded in production. The design that achieves both is a tiered gate: low-risk, well-understood categories move through a lightweight self-service check, and only higher-risk categories escalate to a full review board.
Steps, participants, and gates
Step 1: classify the request by risk tier before anything else. A short intake form captures what the service is, what data it would touch, and which of a small set of risk categories apply (handles regulated data, is a new infrastructure-as-a-service/platform-as-a-service/software-as-a-service (IaaS/PaaS/SaaS) category the org has no existing vendor relationship in, requires a new identity and access management (IAM) integration, or is a low-risk addition within an already-approved vendor's product line). This classification, not the requester's own risk assessment, determines which gate applies next.
Step 2, low-risk tier: self-service approval against a published checklist. For a request that stays within an already-approved provider and touches no regulated data (for example a new managed queue from a cloud provider you already use broadly), the requesting team self-certifies against a short published checklist (cost estimate attached, no new data classification, no new external network exposure) and proceeds without a review meeting, logging the decision for later audit.
Step 3, higher-risk tier: a review board with named, specific participants. For anything touching regulated data, introducing a genuinely new vendor relationship, or requiring new IAM integration, convene a small standing review board rather than an ad hoc group assembled per request: a security representative (checks data handling and access model), a cost or finance representative (checks the pricing model and long-term cost exposure, especially usage-based pricing with no cap), an architecture representative (checks fit against existing standards and lock-in exposure), and the requesting team's technical lead. Keep this board small and standing, meeting on a fixed cadence (for example biweekly) with a defined turnaround service-level objective (SLO), so it functions as a fast gate rather than a queue that discourages teams from asking.
Step 4: an explicit decision-criteria matrix for the specific case of choosing among IaaS, PaaS, and SaaS. Because this is one of the most common request types, give it a named sub-process rather than a generic risk conversation: does the team need control over the underlying runtime or operating system (favors IaaS), does the team want to own only application code and let the provider manage the runtime (favors PaaS), or does an off-the-shelf product already do the job with no custom logic needed (favors SaaS)? Score each option against control needed, time to value, and total cost of ownership, using the org's standard vendor scorecard so this decision is not re-litigated from scratch for every request.
Step 5: a documented, revisitable decision, not a one-time yes. Every approval, at either tier, is logged with the reasoning and a review date (commonly annual, or triggered by a material pricing or feature change), so approval is a point-in-time judgment the org can revisit, not a permanent endorsement.
Responsibilities matrix
| Role | Low-risk tier | Higher-risk tier |
|---|---|---|
| Requesting team | Self-certify against checklist | Present business case and technical fit to the board |
| Security | Publishes the checklist criteria; audits samples after the fact | Reviews data handling and access model directly, every request |
| Finance | Publishes cost-checklist thresholds | Reviews pricing model and cost exposure directly |
| Architecture | Not directly involved | Reviews fit against standards and lock-in exposure |
| Review board (standing) | Not convened | Convenes on fixed cadence, gives a yes/no/conditional decision within the SLO |
Worked example
A team requests a new managed vector database from a provider the org has never used, to store data that includes customer personal information. Intake classifies this as higher-risk on two counts: new vendor relationship and regulated data. It goes to the standing board, which requests a data processing addendum from the vendor (security), checks the pricing model for a usage-based cost that could scale unpredictably with data volume (finance), and confirms whether an existing approved vendor's product could serve the same need before approving a net-new relationship (architecture). The board approves conditionally: proceed, with a mandatory annual cost and security review, rather than a blanket unconditional yes.
Trade-offs and pitfalls
The tiered design trades some rigor on the low-risk path for speed, on the bet that most requests are genuinely low-risk and forcing every request through a full board either creates a bottleneck that teams learn to route around (using a service without asking) or trains the board to rubber-stamp requests out of review fatigue. The main pitfall is letting classification itself become political, with teams downplaying risk to reach the faster self-service tier; mitigate this with periodic security audits of a sample of self-certified requests, not just trust. A second pitfall is treating an approval as permanent, which is why the review-date field above exists: a vendor that was the right choice at approval time can become the wrong one after a pricing change or a compliance-scope change, and a process with no revisit trigger will not catch that until an incident forces the conversation.
You inherit an infrastructure modernization program that is over budget and behind schedule. Walk through your first 30 days of triage, a 90-day stabilization plan, and a 12-24 month recovery and re-baselining roadmap, including how you'd communicate the reset to stakeholders.
Sample Answer
Inheriting a modernization program that's over budget and behind schedule calls for stopping new commitments before doing anything else, rebuilding the plan from the program's actual observed velocity and spend rather than its original estimates, and being explicit with stakeholders about what changed, why, and what specifically prevents the same failure from repeating, since a second broken promise is far more costly to the relationship than the first one was.
The first 30 days: triage
Freeze new scope immediately. Build an honest inventory of actual state: what's genuinely shipped, what's reported as "90% done" but actually stuck, and what's been spent against what was budgeted. Talk directly to the people doing the work to find the real blockers, a skills gap, unclear requirements, a dependency on another team, scope creep, rather than trusting the reported status alone. Identify anything actively making the situation worse (a half-finished migration now running two systems in parallel, doubling cost) and stop that bleeding first. Do not commit to a new date yet; committing before the inventory is done just produces a second wrong estimate.
The 90-day stabilization plan
Pick two or three concrete near-term wins that reduce the biggest unknowns, and re-baseline the budget and schedule bottom-up from the program's actual observed burn rate and delivery velocity, not the original plan's assumptions. Stand up a lightweight, visible governance cadence (a biweekly checkpoint with a real burndown chart) so that the next status report is trustworthy by construction, not by reassurance.
12-24 months: recovery and re-baseline roadmap
Sequence the remaining scope using the same risk-adjusted value framework used for any competing set of investments, and be willing to explicitly de-scope or phase out anything that doesn't clear the bar under the new, more honest estimates. A re-baseline without contingency built in tends to blow through the new baseline for the same reason it blew through the old one; build margin in deliberately this time.
Communicating the reset
Separate three things cleanly: what happened (own it without relitigating blame in front of stakeholders), what's different now (the specific mechanism, like the new checkpoint cadence, that prevents a repeat, not just a promise to try harder), and what's being asked for now (ideally a smaller, more specific ask than before, backed by a track record of two or three already-delivered milestones rather than asking for trust up front again).
Worked example
Say the program was budgeted at $2.4M over 18 months, and at month 14 has spent $2.1M (87.5% of budget) while delivering roughly 40% of scope by an honest function-point count (a standard technique for sizing software work by counting its inputs, outputs, and data interactions, rather than trusting self-reported percent-complete). Bottom-up re-baselining from that observed rate, rather than the original schedule, projects the remaining 60% of scope taking roughly 14 x (60/40) ≈ 21 more months at the same delivery rate, a 35-month total program against an 18-month original plan. The cost side follows the same logic: if $2.1M bought 40% of scope, the full scope at that rate implies a total cost near $2.1M / 0.4 ≈ $5.25M, more than double the original $2.4M budget. Presenting that honest number alongside a de-scoped alternative gives leadership a real choice instead of a single, much larger bill with no alternative attached: for example, if the lowest-value half of that remaining 60% (30 percentage points of the total scope) turns out to be disproportionately expensive too (about 39% of the projected $5.25M, or roughly $2.05M, for only those 30 points of value, because it is also the riskiest and most effort-heavy part of the backlog), cutting it leaves a reduced-scope path that delivers the higher-value 70% of scope, the 40% already delivered plus the remaining 30 points, for roughly $5.25M minus $2.05M, or about $3.2M.
Trade-offs and pitfalls
- Re-baselining once and then blowing through the new baseline too is the most common failure, usually because the same optimistic estimation habit that produced the first bad plan produced the second one. Requiring a banked track record (two or three real, delivered milestones) before asking for the big number again is a structural defense against repeating it.
- A 30-day triage that turns into a victory-lap status presentation instead of an honest gap analysis wastes the one period where stakeholders are primed to hear hard truths.
- De-scoping decisions made unilaterally by the recovering team, without the business at the table, risk cutting something more valuable than what was kept; the trade-off should be made jointly, with the honest numbers in front of everyone.
Unlock Full Question Bank
Get access to all Infrastructure Strategy and Technology Selection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.