Infrastructure Strategy and Technology Selection Questions
Setting technical direction for infrastructure and deciding what to build on. Covers infrastructure vision and long-term roadmap, modernization and technical-debt strategy, platform decisions, and organizational and governance considerations, alongside the decision framework for adopting or retiring technology: build-versus-buy-versus-cloud-versus-on-premises trade-offs, vendor and platform evaluation, technology-portfolio rationalization, requirements-driven service selection, and weighing total cost of ownership and risk. It also covers managing an existing vendor relationship: quantifying and mitigating lock-in, negotiating contract terms, and planning an exit from a service the organization depends on. The leadership-altitude discipline of directing an infrastructure estate over time, building consensus among stakeholders with conflicting priorities, and persuading leadership or a team to accept a major technology change.
A non-functional requirement states an API must meet a strict latency target (for example 50ms p95) for users spread across the globe. Describe how you'd map that requirement to compute choice, network design, caching, and regional deployment strategy, and how you'd validate the choice actually meets the target.
Sample Answer
A strict p95 (95th percentile) latency target for a globally distributed user base is really a latency budget that has to be allocated across every hop a request takes, and the design work is deciding, hop by hop, which portion of that budget goes to network distance, which goes to compute, and which is eliminated entirely by caching, then validating the whole chain against real client locations rather than trusting the design on paper.
Mapping the requirement to concrete choices
Network design: put compute close to the user before optimizing anything else. For a global 50 millisecond p95 target, the single largest lever is geography: a round trip between two points on opposite sides of the globe over the public internet routinely costs 150 to 250 milliseconds on its own, which already blows the entire budget before any application logic runs. This makes multi-region deployment a requirement, not an optimization: deploy the service into regions close to where users actually are (informed by real traffic geography, not assumption), and use each cloud provider's global load-balancing or anycast routing to send each user's request to their nearest healthy region automatically.
Regional deployment strategy: choose regions by user density, and pair each with at least one nearby fallback availability zone (AZ, an isolated failure domain within a region) for resilience without adding cross-region latency. A common pattern for three major user population centers (for example North America, Europe, and Asia-Pacific) is one primary region per population center, each with two or more AZs for intra-region failover, so a single AZ failure does not force traffic across an ocean and blow the latency budget along with the availability target.
Compute choice: pick a runtime whose worst-case latency, not just its average, fits the remaining budget. Once network distance is minimized by regional placement, the remaining budget (perhaps 30 to 40 milliseconds after accounting for realistic in-region network and load-balancer overhead) has to cover application processing. A compute platform with unpredictable cold-start latency, such as a serverless function that scales to zero, can spend that entire remaining budget on a single cold start; for a strict p95 target, a warm, always-on compute tier (or a serverless configuration with provisioned concurrency to avoid cold starts) is usually the safer choice, even at higher baseline cost.
Caching: remove work from the request path entirely wherever the data allows it. Any data that is read far more often than it changes (for example a product catalog entry or a user's configuration settings) should be served from a content delivery network (CDN) or regional cache rather than computed fresh per request, since the fastest way to meet a strict latency target is to avoid needing compute or a database round trip at all for the majority of requests. Cache invalidation strategy (short time-to-live for frequently changing data, explicit invalidation for rarely changing data) determines how much of the traffic this actually removes from the latency-critical path.
Validating the choice actually meets the target
Design intent is not evidence. Run synthetic latency probes from real geographic vantage points, ideally using a third-party monitoring service with probe locations in each target region, not just from your own cloud provider's network, since testing from inside the same provider's backbone can understate real end-user latency. Measure p95 and p99, not average, since a strict target stated as p95 is explicitly about the tail, and instrument the measurement to break down where time is spent (network, load balancer, application, database) so a target miss can be diagnosed rather than just observed.
Worked example
A service with a global 50 millisecond p95 target deploys to three regions (US-East, EU-West, AP-Southeast), each with two AZs. A user in Frankfurt routes to EU-West via anycast, adding roughly 5 to 15 milliseconds of network overhead to the nearest AZ pair depending on the user's own last-mile connection, leaving roughly 35 to 45 milliseconds for everything else. The catalog lookup that used to hit the database on every request is moved behind a CDN cache with a 5-minute time-to-live, cutting the typical request's server-side work from a 25 millisecond database query plus a 10 millisecond application layer down to a 2 millisecond cache read for the roughly 90% of requests that hit the cache; the remaining 10% (cache misses) still take the full database path and are measured separately as a distinct p95 within the cache-miss population, confirmed to also fit inside the remaining budget on its own.
Trade-offs and pitfalls
Multi-region deployment and always-on compute both cost more than a single-region, scale-to-zero design, so the trade-off is explicit: this target cannot be met cheaply, and presenting a cheaper design as meeting a strict global p95 without disclosing that trade-off is the most common way this kind of requirement gets silently violated in production. The most common pitfall in validation is measuring only from the cloud provider's own internal network or from a single test location near one of the deployed regions, which can look like it meets the target while users on the far side of the global footprint experience something very different.
Design a proof-of-concept to compare a managed platform service against a self-managed equivalent before committing to one. Specify the experiments and metrics you'd collect (throughput, latency, operational overhead, mean time to recover), the failure-injection scenarios you'd run, and how you'd reach a go/no-go decision.
Sample Answer
A proof of concept (PoC) that actually settles a managed-versus-self-managed decision has its success and failure thresholds written down before any test runs, deliberately tests failure modes and not just peak throughput, and ends in a documented go/no-go against those pre-committed numbers rather than a subjective impression of "it felt fine."
Designing the PoC
- Metrics to collect: sustained throughput at a defined load, tail latency (p95/p99, the 95th and 99th percentile response time, which shows what your slowest real users experience, not just the average), operational overhead (hours of hands-on toil per week to run and tune it during the pilot), and mean time to recovery (MTTR) from an injected failure.
- Failure-injection scenarios, not just load tests: kill a node or process mid-traffic, saturate the CPU or network on a dependency, simulate an upstream dependency outage, and simulate a bad deploy that has to be rolled back. A platform that only gets load-tested on the happy path will pass a PoC and then fail in its first real incident, because the PoC never asked the question that actually matters operationally.
- Latency measured from the vantage point that matters: measure from the client's perspective (the caller actually waiting on the response), not just server-side processing time, since network hops and queuing between the two can dominate the difference between "the backend is fast" and "the user experienced it as fast." A benchmark that only reports server-side numbers is measuring the wrong thing for a latency-sensitive decision.
- Pre-committed go/no-go thresholds: set numeric thresholds for each metric before the test begins, plus at least one "kill criterion," a failure so serious it overrides every other metric (for example, any data loss during failover is an automatic no-go regardless of how good throughput looked).
- A lighter-weight variant for lower-stakes decisions: when the decision's blast radius doesn't justify a multi-week experiment-and-metrics exercise, a lightweight PoC template, objectives, scope, timeline, and deliverables, with a shorter and less rigorous set of go/no-go criteria, is a legitimate and honest alternative. The rigor of the PoC should match the cost of getting the decision wrong, not be applied uniformly to every choice.
Worked example (illustrative)
Consider comparing a managed message queue against a self-hosted equivalent for an order-processing pipeline that needs to sustain 2,000 messages per second at peak with a 99.9% availability target. That target implies roughly 8.8 hours of allowed downtime per year (0.001 x 365.25 x 24 ≈ 8.8), which is a useful number to have in hand before setting the PoC's failure-recovery threshold.
A three-week PoC plan: week one runs a sustained load test at 2,000 messages/second for at least an hour, recording p99 latency measured client-side; week two runs failure injection (kill a broker node, saturate disk I/O) and measures message loss and recovery time; week three tracks operational load, how many alerts fired and how many required a human to act. A team running this exercise might set thresholds along these lines before starting: p99 latency under 250ms at the target load, recovery from a node failure in under 2 minutes with zero message loss, and less than 2 hours per week of hands-on operational toil during the pilot. Those specific numbers are illustrative of how a team sets its bar, not a report of an actual measured run; the discipline that matters is committing to numbers like these in writing before the test, not the particular values chosen.
Trade-offs and pitfalls
- Testing only the happy path is the single most common way a PoC gives false confidence; the PoC's job is specifically to find the failure modes production would otherwise find for you, at a worse time.
- Moving the goalposts after seeing results, quietly relaxing a threshold because the preferred option narrowly missed it, defeats the entire purpose of pre-committing them. If a threshold needs to change, that change should be justified and documented separately from the result it happens to rescue.
- A PoC environment that doesn't match production's real data shape (message size distribution, traffic burstiness, request skew) can pass cleanly and still mislead, because throughput and failure behavior both depend heavily on that shape, not just on the average load.
What are the common forms of vendor lock-in when adopting managed cloud services or SaaS (for example proprietary data formats, APIs, egress costs)? As the technical decision-maker, how would you detect that risk early, and what mitigations would you put in place before committing?
Sample Answer
Direct answer
Lock-in shows up as proprietary data formats, proprietary APIs, and asymmetric egress pricing, but also less obviously as specialized team skills and multi-year contractual commitments. The right response is not to avoid managed services (that trades real velocity for a hypothetical future migration), it is to detect the degree of lock-in during evaluation, before signature, using a small set of concrete signals, and to build in specific mitigations proportional to what you find.
Structured elaboration
Common forms
- Data format lock-in: your data is stored in a format only that vendor's tooling can read efficiently (a proprietary binary format, a non-standard query dialect).
- API lock-in: your application code calls a vendor-specific SDK with no equivalent interface elsewhere.
- Egress cost lock-in: getting data out is priced high enough to functionally deter leaving, even though getting it in was free.
- Skills/operational lock-in: your team is trained and tooled only around this vendor's specific console and workflows.
- Contractual lock-in: multi-year commitments or steep early-termination penalties.
Six early-detection indicators
- No standard export format is offered, or the export exists but drops fidelity (loses relationships, metadata, or history).
- There is no documented, general-purpose API; the only interface is a proprietary SDK.
- The pricing model concentrates cost in data transfer or egress rather than compute or storage.
- The vendor's own roadmap is steering you toward more of its proprietary managed features over time (a "gravity well" pattern).
- The contract requires a multi-year term or a steep early-termination penalty.
- No credible second vendor or open-source alternative exists for the same capability today.
Four mitigations
- Abstraction layer: put an internal interface or adapter between your application code and the vendor SDK, so a future swap touches one layer, not every call site.
- Data-format standards: store the canonical copy of your data in an open, portable format (for example Parquet for analytical data, standard SQL for relational data) even if the vendor also keeps its own optimized copy.
- Contract terms: negotiate explicit data-export rights, a reasonable termination clause, and, where possible, a cap on egress pricing before signing, not after a dispute starts.
- Exit-readiness drills: periodically test that you actually could export and reload your data elsewhere; an untested export right is not a real mitigation.
Worked example
Evaluating a managed analytics platform: during the trial, you export 10 gigabytes of your own event data and time how long it takes and whether the schema survives intact. You then check the vendor's published egress pricing: at $0.09 per gigabyte, moving a hypothetical 10 terabytes (10,000 gigabytes) out later would cost 10,000 times $0.09, or $900, which by itself is a mild deterrent, not a wall. Combined with indicator 4 (their roadmap page shows three new proprietary connectors launched in the last two quarters, each deepening the coupling) and indicator 2 (the only way to query historical data is their SQL dialect with several non-standard extensions), the composite picture is moderate-to-high lock-in, which argues for mitigation 2 (also land a Parquet copy of raw events in your own object storage on ingest) even if you proceed with the vendor.
Trade-offs and pitfalls
There is no zero-lock-in option, only priced lock-in; the real decision is how much you are willing to accept for a given amount of velocity or capability, not whether to accept any. The opposite failure is also real: teams that build an elaborate abstraction layer to avoid a small, cheap-to-reverse decision end up carrying the ongoing cost and complexity of that layer indefinitely, which can exceed the lock-in risk it was meant to prevent. Size the mitigation to the actual indicator score, not to a general anxiety about vendors.
A company needs a specific business-critical capability (for example billing and invoicing, payment processing, or a logging/analytics platform) and is weighing a vendor/SaaS product against building it in-house. Walk through how you'd run that build-vs-buy evaluation end to end: criteria, a proof-of-concept, and how you'd present the recommendation.
Sample Answer
Run a build-versus-buy evaluation for a specific business capability as a five-part process: define the requirements before looking at any vendor, score the finalists on a shared set of criteria, validate the top choices with a real proof of concept against your own workload, present a memo that names a recommendation, and define up front how you'll check after adoption that the decision was actually right. The same process applies whether the capability in question is billing and invoicing, payment processing, a CRM (customer relationship management system), an analytics dashboard, or an observability platform; only the specific requirements list changes.
The process
- Define requirements first. Write down the capability's actual requirements (for a logging and observability platform: log volume, retention period, query latency, alerting integration, any compliance retention rule) before looking at a single vendor, so requirements aren't quietly reverse-engineered from whatever a vendor's feature list happens to include.
- Score against shared criteria. Total cost of ownership (TCO) at your real or projected volume, feature velocity (how much faster the option gets you the capability versus building it yourself), lock-in and exit cost, staffing and operational burden to run the option, and whether the capability is genuinely "core" (your differentiator) or "context" (necessary but not differentiating), a logging platform is almost always context, which pushes hard toward buying or adopting unless there's a specific reason otherwise.
- Validate with a real proof of concept. Run the actual workload's shape, log volume and query patterns, or transaction volume for a billing system, against the top one or two shortlisted options for a defined period, using the same operational-overhead and reliability metrics a platform-level PoC would use. A practical checklist for this step covers: functional coverage against the written requirements, a real data-migration dry run, an integration test against your actual authentication system, performance under real (not demo) load, a documented rollback plan, and the total cost at your real volume, not the vendor's list price.
- Present a named recommendation. A short memo with the scored comparison, the PoC evidence, and TCO at one-year and three-year horizons that commits to a specific choice, not a survey of options with no call made.
- Define the post-adoption check. Before adoption, name the specific signals that will confirm the decision was right, reduced time spent maintaining the old tool, faster incident detection, less on-call toil, and revisit them at a fixed checkpoint (six months is typical) instead of assuming the decision was correct just because it shipped.
Worked example
A company running a homegrown logging and observability stack (self-hosted log aggregation with a hand-built alerting layer) evaluates replacing it with a managed observability platform. Requirements: 500 GB/day of log volume, 30-day retention, sub-5-second query latency for on-call use, and integration with the existing paging tool. Scoring shows the managed option winning heavily on staffing, the in-house stack currently consumes an estimated 0.4 full-time-equivalent (FTE) per quarter in patching and capacity management, and on feature velocity, built-in anomaly detection the in-house tool lacks. On raw cost, though, the managed option is actually more expensive: at this ingest volume it runs roughly $18,000/month, or $216,000/year, versus the in-house stack's roughly $9,000/month in infrastructure cost. Adding the 0.4 FTE at a fully loaded $160,000/year engineer cost ($64,000/year) brings the true in-house total to about 9,000 x 12 + 64,000 = $172,000/year. The managed option costs about $44,000/year more in raw dollars ($216,000 versus $172,000), and the recommendation memo says so explicitly rather than hiding it, arguing instead that redeploying the freed 0.4 FTE to differentiating work and reducing on-call toil is worth that premium. The post-adoption check, six months later, verifies whether observability-related pages actually dropped and whether that freed engineer time was genuinely redeployed, rather than quietly absorbed back into maintaining the new tool's edge cases.
Trade-offs and pitfalls
- A proof of concept that tests a vendor's demo data instead of your own workload's real cardinality, volume, and query patterns is exactly what makes observability and logging evaluations go wrong; insist on your own data.
- A recommendation memo that lists pros and cons without committing to a choice is a survey, not a recommendation, and is the dominant weak answer on this kind of question.
- Skipping the post-adoption check means a wrong call made two years ago is still quietly costing the business, and nobody ever re-examines it because nothing was defined up front to trigger that look.
What does a multi-year technical vision and infrastructure roadmap actually mean in practice, and how does it differ from a shorter execution plan? Describe the components you would expect to see (architecture direction, sequencing, governance, cost, risk) and the audiences each is written for.
Sample Answer
A technical vision is a directional statement over a multi-year horizon (typically 3-5 years): where the platform needs to end up and why, in enough detail to make trade-offs but not enough to lock in implementation. A roadmap sequences that vision into named phases with rough timing. An execution plan is the short-horizon (quarter or sprint) commitment of specific work items to specific owners. The three exist because they answer different questions for different audiences, and collapsing them into one document is the most common failure mode in practice.
The five components a real roadmap includes
- Architecture direction (the pillars): the small number of durable architectural bets the organization is making (for example "consolidate onto one data platform," "move from a monolith to domain-owned services"). This is the part that changes slowest and should survive most quarterly re-plans.
- Sequencing (milestones): the order phases happen in and why that order was chosen, usually because one phase is a technical or organizational prerequisite for the next. Milestones are the checkpoints a person outside the work can use to tell if it is on track.
- Governance: who approves scope changes, what triggers a re-plan, and what checkpoint cadence keeps the roadmap honest instead of becoming a document nobody revisits.
- Cost: a multi-year budget envelope (order of magnitude, not a line-item quote) tied to each phase, so a funding conversation can happen before the phase starts rather than mid-execution.
- Risk: the handful of things that could derail the plan (a key vendor's viability, a skills gap, a dependency on another team's roadmap) named explicitly, with an early-warning signal for each, not buried in prose.
A roadmap without an explicit KPI (key performance indicator) attached to each milestone is not falsifiable: nobody can tell six months in whether the phase actually succeeded or just shipped.
Audiences
- The vision is written for the people who fund and sponsor the direction: the CTO's peers, the board, other VPs whose roadmaps need to line up with it. They need the "why" and the trade-offs made, not the sequencing detail.
- The roadmap is written for engineering leadership and an architecture review body: people who need to plan their own team's work against the milestones and dependencies, and who will push back on sequencing if it conflicts with something they already know.
- The execution plan is written for the team doing the work this sprint: specific tickets, specific owners, specific acceptance criteria. Nobody outside the team needs this level of detail, and putting it in the roadmap document buries the signal a leadership audience actually needs.
Worked example
Suppose the vision is: "consolidate three regionally siloed customer data stores into one global platform within three years, to cut cross-region query latency and stop paying for duplicated storage and duplicated on-call coverage." The roadmap breaks that into Year 1 (build the unified platform's core and migrate the smallest region as a pilot, a $400K-order-of-magnitude phase), Year 2 (migrate the two remaining regions, contingent on the pilot's migration tooling proving reliable), and Year 3 (decommission the legacy stores and retire the redundant on-call rotation). Governance names the architecture review board as the approver of any scope change and sets a quarterly checkpoint against the milestone dates. Risk names "the pilot region's migration tooling doesn't generalize to the larger regions' data volume" as the single biggest threat, with a leading indicator (migration throughput measured during the pilot) to catch it early.
Now a new data-residency requirement lands mid-Year-1. It does not change the vision (the destination is still one global platform); it changes the roadmap's sequencing (the pilot region may need to move earlier or later depending on where the requirement bites hardest) and it is absorbed entirely inside the execution plan for the current sprint boundary, without anyone needing to touch the vision document at all. That containment, a change rippling only as far up as it has to, is the actual test of whether these three artifacts were built as separate things in the first place.
Trade-offs and pitfalls
- The most common failure is writing a "vision" that is really a roadmap with more adjectives: it names activities instead of making an actual directional bet with an explicit trade-off (what the organization is choosing not to do, and why).
- A vision that is never revisited ages badly. Review it on a fixed cadence (annually is typical), because the business strategy it's supposed to serve will have moved even if the technology hasn't.
- A roadmap with no stated non-goals turns into a wish list that absorbs every stakeholder's pet request, because nothing in the document says what's explicitly out of scope.
- Skipping the execution-plan layer and asking engineers to work directly off roadmap-level milestones leaves ambiguity about who owns what this sprint, which shows up later as missed dependencies nobody flagged in time.
Unlock Full Question Bank
Get access to all 19 Infrastructure Strategy and Technology Selection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.