Infrastructure Strategy and Technology Selection Questions
Setting technical direction for infrastructure and deciding what to build on. Covers infrastructure vision and long-term roadmap, modernization and technical-debt strategy, platform decisions, and organizational and governance considerations, alongside the decision framework for adopting or retiring technology: build-versus-buy-versus-cloud-versus-on-premises trade-offs, vendor and platform evaluation, technology-portfolio rationalization, requirements-driven service selection, and weighing total cost of ownership and risk. It also covers managing an existing vendor relationship: quantifying and mitigating lock-in, negotiating contract terms, and planning an exit from a service the organization depends on. The leadership-altitude discipline of directing an infrastructure estate over time, building consensus among stakeholders with conflicting priorities, and persuading leadership or a team to accept a major technology change.
Construct a 5-year TCO model for running an enterprise analytics platform in the cloud. Walk through the variables you'd model (on-demand vs. reserved/committed pricing, sustained-use discounts, storage tiers and lifecycle, network egress growth, managed-service premiums, staffing and tooling costs, and the depreciation of migration effort), and explain how you'd run sensitivity analysis on your key assumptions.
Sample Answer
A 5-year total cost of ownership (TCO) model for an enterprise analytics platform has to go well beyond compute list price, and the model is only as useful as its sensitivity analysis, since the single biggest source of TCO surprise in analytics workloads is a variable nobody stress-tested, most often data growth driving both storage and egress far past the year-one assumption.
Variables to model
Compute pricing structure. Model on-demand pricing as the baseline, then layer in reserved or committed-use discounts (typically 30 to 50% off on-demand for a 1-to-3-year commitment) for the portion of workload that is steady-state rather than bursty, and sustained-use discounts (automatic discounts some providers apply once a resource runs continuously past a threshold within a billing period) for anything that runs near-continuously without an explicit reservation.
Storage tiers and lifecycle. Analytics platforms accumulate data far faster than they delete it. Model hot storage (for actively queried recent data) separately from a lifecycle policy that automatically moves older data to cooler, cheaper tiers (typically 3 to 5x cheaper per gigabyte) after a defined age threshold, and model the accumulating volume over the full 5 years, not a static year-one snapshot, since storage cost in an analytics platform is a growth curve, not a flat number.
Network egress growth. Model the cost of data leaving the platform, for example results delivered to downstream business intelligence tools or replicated to a disaster-recovery region, and grow this cost proportionally with the same data-growth assumption used for storage, since egress and storage typically grow together.
Managed-service premiums. A managed data-warehouse or analytics service typically costs more per unit of compute than raw infrastructure, but recover this premium in the model against the staffing cost it displaces (below), rather than treating it as pure additional cost with no offset.
Staffing and tooling costs. Include the fully loaded cost of the engineers who operate the platform, plus any third-party tooling (monitoring, data governance, cataloging) layered on top, since for a 5-year enterprise platform, staffing is frequently a larger share of total cost than the cloud bill itself, and omitting it from a TCO model is one of the most common ways a cloud migration's business case ends up wrong.
Depreciation of migration effort. Amortize the one-time cost of the initial migration (engineering time, parallel-running old and new systems during cutover, data validation) across the 5-year horizon rather than booking it entirely in year one, since a 5-year TCO comparison against an alternative should compare like-for-like ongoing cost, with the migration cost spread the way a capital expenditure would be.
Worked example: baseline model and sensitivity analysis
Baseline assumptions, year 1: compute (with a 3-year committed-use discount on the steady-state 70% of workload) $480,000 per year; storage, starting at 100 terabytes hot plus 200 terabytes in a cooler tier, $180,000 per year; egress $60,000 per year; managed-service premium over raw infrastructure, $150,000 per year; staffing, 2 full-time engineers fully loaded at $180,000 each, $360,000 per year; migration effort, a one-time $500,000 amortized over 5 years, or $100,000 per year. Year-1 total: $480,000 + $180,000 + $60,000 + $150,000 + $360,000 + $100,000 = $1,330,000.
Growth assumption: data volume, and therefore storage and egress cost, grows 25% per year in this example, a planning assumption in the range commonly cited for an actively used analytics platform rather than a fixed law, which is exactly why this is the assumption to stress-test in the sensitivity pass below and to validate against the organization's own trailing data-growth trend before treating it as authoritative; compute grows more slowly at 10% per year as query volume rises with the business, and staffing and the amortized migration cost stay flat. Year 5 storage: $180,000 times 1.25 to the 4th power, or roughly $439,000. Year 5 egress: $60,000 times 1.25 to the 4th power, or roughly $146,000. Year 5 compute: $480,000 times 1.10 to the 4th power, or roughly $703,000. Year 5 total: $703,000 + $439,000 + $146,000 + $150,000 + $360,000 + $100,000 = $1,898,000. Summing years 1 through 5 with each year's figures compounding at the same growth rates (compute at 1, 1.10, 1.21, 1.331, 1.4641 times its year-1 rate; storage and egress at 1, 1.25, 1.5625, 1.953, 2.441 times theirs) gives a 5-year total of approximately $7.95 million, about 20% above a naive year-one-times-five estimate of $6.65 million, precisely because storage and egress compound rather than staying flat.
Sensitivity analysis on the key assumption. Data growth is the assumption to stress-test, since it drives two of the largest compounding line items. Running the same model at 15% annual growth instead of 25% drops the year-5 storage and egress figures by roughly 28% relative to the base case, moving the 5-year total down to approximately $7.6 million; running it at 35% annual growth raises the 5-year total to approximately $8.4 million. Presenting this as a range (a low, base, and high case tied explicitly to the data-growth assumption) rather than a single number is what lets a stakeholder see how much of the total cost's uncertainty rides on one variable they can actually influence, for example by tightening the data-retention lifecycle policy.
Trade-offs and pitfalls
The most common modeling mistake is running the whole 5 years at year-one unit costs, which understates the true total by ignoring exactly the compounding growth that analytics workloads reliably exhibit; always project the growth-sensitive line items forward rather than multiplying a static year. A second pitfall is omitting staffing and treating the model as a pure cloud-bill exercise, which produces a number that looks favorable for a fully managed option but hides where the real cost differential against a cheaper, less-managed alternative actually lives. Sensitivity analysis is not optional polish here: presenting a single point estimate for a 5-year projection implies a precision the underlying growth assumption cannot support, and a stakeholder who is only shown the base case has no way to judge how much risk the number is carrying.
You're evaluating vendor platforms for a technology roadmap decision. Describe a scoring rubric covering technical fit, operational fit (SLAs, monitoring, upgrade cadence), security/compliance, cost model and vendor viability, and how you'd run the process with a cross-functional group of stakeholders.
Sample Answer
Build one shared, weighted scoring rubric across five or six dimensions, run it with the actual cross-functional stakeholders who will live with the choice, and normalize each scorer's numbers before combining them, so the process is debated on criteria and evidence rather than on whoever argues loudest or scores most senior in the room.
The dimensions
- Technical fit: does the platform do the actual job against your real requirements, not the vendor's demo scenario. For different platform types the concrete criteria shift, embedding support and single sign-on (SSO) integration for an embeddable analytics tool, chart-rendering performance and dashboard-sharing for a visualization tool, but the underlying question is the same.
- Operational fit: SLAs (service-level agreements, the vendor's contractual uptime and response-time commitments), monitoring hooks into your existing tooling, and upgrade cadence. Support-model readiness belongs here specifically, and it should be demanded as evidence during the PoC (proof of concept), not taken as a sales claim: a named, dedicated technical account manager, a real 24x7 on-call escalation path, documented escalation SLAs, and, for the most critical systems, a track record or contractual commitment to a joint incident war room for severity-1 issues.
- Security and compliance: relevant certifications (SOC 2, an independent audit of a vendor's security controls; ISO 27001, an international information-security management standard; PCI-DSS attestation, certification against the payment-card industry's data-security standard; HIPAA business associate agreement availability, a contract required under US health-privacy law when a vendor handles protected health information) and data-residency options for your regulatory context.
- Cost model: not the list price. Model the true multi-year cost at your projected usage, including likely support-tier upsells and typical renewal increases.
- Vendor viability: financial stability and roadmap credibility, plus ecosystem and partner-network maturity as its own named sub-criterion, marketplace offerings, certified partners, ISV (independent software vendor) certifications, and the maturity of the consulting ecosystem around the product. A vendor with a thin partner ecosystem leaves you doing more integration work yourself and signals a smaller, riskier install base than the sales pitch suggests.
Running the process
Score independently first, each function scores before seeing anyone else's numbers, to avoid anchoring on the loudest voice in the room, then normalize (rescale each scorer's average to a common midpoint, or convert to a common scale) before averaging across scorers, since raw averaging lets one systematically harsh or lenient scorer distort the group result. Someone needs to own convening and driving the process to an actual decision, not just collecting opinions and hoping consensus appears on its own.
Worked example
The same five-dimension rubric produces different recommendations depending on organizational context, which is exactly what a well-built rubric should do. A regulated bank's weight profile might be: technical fit 0.15, operational fit 0.15, security/compliance 0.35, cost 0.15, vendor viability 0.20. A high-growth consumer startup's weight profile might instead be: technical fit 0.30, operational fit 0.15, security/compliance 0.10, cost 0.20, vendor viability 0.25. Score two vendors identically on a 1-5 scale: Vendor A scores 4, 4, 5, 3, 3 across the five dimensions; Vendor B scores 5, 3, 3, 4, 4.
Bank profile: A = (4x0.15)+(4x0.15)+(5x0.35)+(3x0.15)+(3x0.20) = 0.60+0.60+1.75+0.45+0.60 = 4.00; B = (5x0.15)+(3x0.15)+(3x0.35)+(4x0.15)+(4x0.20) = 0.75+0.45+1.05+0.60+0.80 = 3.65. The bank profile favors Vendor A, driven by A's security/compliance strength being weighted heavily.
Startup profile: A = (4x0.30)+(4x0.15)+(5x0.10)+(3x0.20)+(3x0.25) = 1.20+0.60+0.50+0.60+0.75 = 3.65; B = (5x0.30)+(3x0.15)+(3x0.10)+(4x0.20)+(4x0.25) = 1.50+0.45+0.30+0.80+1.00 = 4.05. The startup profile flips the recommendation to Vendor B, driven mainly by B's edge on technical fit (5 vs. 4, contributing 0.30 of the swing at the startup's heaviest weight) and on vendor viability (4 vs. 3, contributing 0.25), with B's smaller edge on cost (0.20) adding a third push in the same direction; B's weaker scores on operational fit and security/compliance are not enough to offset those three. Same two vendors, same raw scores, opposite recommendation, purely from the weight profile matching organizational context, which is the whole point of building the rubric with the stakeholders rather than importing a generic one.
Trade-offs and pitfalls
- A rubric with weights chosen after seeing which vendor everyone already favors is reverse-engineered to justify a foregone conclusion; lock the weights before scoring, the same discipline that keeps a proof-of-concept's thresholds honest.
- Letting the most senior person's raw score dominate without normalization quietly turns a group process back into one person's opinion with extra steps.
- Treating vendor viability and ecosystem maturity as a soft, easily ignored line item is a mistake; it's often the best available predictor of long-term support quality, which only becomes visible after the contract is signed and the honeymoon period ends.
Tell me about a time you persuaded stakeholders to change a tool, vendor, or infrastructure component they were attached to. Walk through the situation, how you built consensus, the objections you encountered, and the measurable outcome.
Sample Answer
The strongest version of this story names the specific stakeholder's underlying interest (not just their stated position), describes how a low-risk proof was used to make the case with evidence instead of argument, and states a measurable outcome the stakeholder themselves would recognize as the win. The same underlying competency shows up in several shapes: recommending a managed service over building something in-house, leading a platform evaluation and adoption effort, or persuading a team that was skeptical of a specific new tool. Pick whichever one actually happened and tell it with that specificity.
Building the story
- Situation: name the incumbent tool, vendor, or component, and be honest about why it was sticky, usually a mix of sunk cost (years of runbooks and tribal knowledge built around it), legitimate risk aversion (the current thing, however dated, is known to work), and sometimes a personal stake (the person being asked to change owns that system's reputation).
- Task: state the concrete reason a change was worth pursuing (a cost, reliability, or velocity problem the incumbent was causing).
- Action: the persuasion mechanism matters more than the argument. A credible version usually includes: scoping a low-risk pilot on something that wasn't business-critical, defining success criteria with the skeptical stakeholder before running it (so the result can't be disputed as biased afterward), and bringing that stakeholder in as a co-owner of the transition rather than someone being overruled.
- Result: state the outcome in terms the objecting stakeholder's own scorecard would recognize, and be ready to say what would have happened if the pilot had failed, since a strong answer names the walk-away criteria, not just the win.
A related but distinct version of this story is not about overcoming resistance at all: it is a neutral, end-to-end evaluation process, defining selection criteria up front, running a proof of concept or benchmark against real workload data, quantifying the relevant non-functional requirements (NFRs: qualities like latency, availability, or security posture, as opposed to a specific feature), and presenting a recommendation to stakeholders who did not have a strong prior attachment either way. Both are legitimate answers to this question; the difference is whether the hard part of the story was building evidence or building consensus against attachment, and a candidate should tell whichever one is actually true rather than manufacturing conflict that wasn't there.
Worked example (skeleton)
A team owned an aging configuration-management tool that had years of custom runbooks built around it, and the person who had built most of those runbooks was, understandably, the most skeptical voice in the room about replacing it. Rather than opening with a comparison deck, the candidate proposed running the new tool on a single non-critical internal service for one release cycle, with the skeptical engineer helping define what "working" meant before the pilot started: deployment lead time, number of manual interventions needed, and rollback count, measured the same way on both old and new tooling over a comparable number of deployments. When the pilot's numbers held up, the skeptical engineer, now a co-author of the migration runbook rather than someone being overruled, became one of the people advocating for the wider rollout. The measurable outcome cited in an interview should be the metric that was agreed on up front (fewer manual interventions per deployment, say), not a number invented after the fact to sound impressive.
Trade-offs and pitfalls
- Presenting the story as "I convinced them with a better argument" and skipping the relationship work (co-ownership, agreeing on success criteria in advance) makes the story sound like it worked through logic alone, which interviewers are right to be skeptical of.
- Quoting an improvement number without describing how it was measured invites the obvious follow-up: measured against what baseline, over what sample size. Have that answer ready.
- Not naming what would have counted as failure suggests the "evaluation" was really a foregone conclusion looking for validation, which undercuts the credibility of the whole story.
When you review a vendor's Service Level Agreement, what specific clauses and metrics do you pay closest attention to? Walk through a prioritized checklist covering availability targets, measurement windows, credit calculations, scheduled-maintenance exclusions, and support response times, and explain why each one matters.
Sample Answer
Reviewing a vendor's service-level agreement (SLA) is about finding the gap between what the headline availability number implies and what the actual clauses commit the vendor to, since the number on the marketing page and the number that survives the fine print are frequently different, and the difference only shows up if you read the measurement, exclusion, and remedy clauses closely.
Prioritized checklist
1. Availability target, and what it is actually measured against. A headline like "99.95% uptime" is meaningless without knowing what counts as "up": is it measured at the load balancer, per individual service component, or only at a data-center level that could mask a specific feature being down for hours while the overall platform is technically "available." Push for the SLA to define availability at the granularity that matters to your actual usage, not the vendor's most favorable aggregate.
2. Measurement window. Check whether availability is measured monthly, quarterly, or annually, since a longer measurement window lets a vendor absorb a serious multi-hour outage without breaching the SLA, as long as the rest of the window was clean; a monthly window is meaningfully more protective for you than an annual one, because a single bad incident cannot be diluted across as much good uptime.
3. Credit calculations. Read exactly how a breach translates into a remedy: what percentage service credit applies at each tier of missed availability, whether that credit is capped, and critically, whether the credit is your sole and exclusive remedy (meaning you contractually give up the right to claim actual damages beyond the credit). A small service credit as your only recourse for an outage that caused real business damage is a common, easy-to-miss asymmetry.
4. Scheduled-maintenance exclusions. Check how much scheduled maintenance time the vendor can take before it counts against the availability number, and how much advance notice you get. A generous, loosely bounded maintenance exclusion (for example unlimited maintenance windows with only 24 hours notice) can let a vendor take meaningful planned downtime that never touches the SLA calculation at all.
5. Support response times. Distinguish between response time (how fast the vendor acknowledges your ticket) and resolution time (how fast the actual problem is fixed); many SLAs commit only to the former, which can technically be met by an automated acknowledgment while the real issue sits unresolved for hours. Check whether response and resolution commitments vary by severity tier, and whether your account's support tier actually includes the fastest tier at all.
Why each one matters
Availability granularity and measurement window together determine whether the SLA would actually catch the kind of outage your business cares about; a coarse measurement can technically comply with the SLA while your users experienced a real, damaging outage. Credit calculations determine your actual financial recourse, which is frequently far smaller than the business impact of the outage it is meant to compensate. Scheduled-maintenance exclusions determine how much planned downtime is invisible to the SLA entirely, which matters for any workload with limited tolerance for planned interruption. Support response versus resolution time determines whether "we responded" during an active incident actually means anything for how fast the problem gets fixed.
Worked example
A vendor's SLA states 99.9% monthly availability with a 10% service credit for the following month if breached, capped at 30% of monthly fees, credits stated as sole remedy; scheduled maintenance excluded from the calculation with 48 hours notice, up to 8 hours per month; and a 1-hour response time for severity-1 tickets, with no stated resolution-time commitment. A careful reviewer notes: 99.9% monthly still permits about 43 minutes of unplanned downtime per month before breach, which may or may not be acceptable depending on the workload; the sole-remedy language means a costly outage yields, at most, a 30%-of-fees credit regardless of actual business impact, worth negotiating up or striking if this workload is business-critical; 8 hours of monthly maintenance with only 48 hours notice is a meaningful, SLA-invisible planned-downtime allowance worth negotiating down or requiring more advance notice for; and the missing resolution-time commitment for severity-1 issues means a 1-hour acknowledgment could be followed by an open-ended repair time with no contractual pressure on the vendor to move faster.
Trade-offs and pitfalls
Negotiating every one of these clauses tighter will cost negotiating leverage or price, so prioritize based on the workload's actual risk profile: a business-critical, revenue-generating system justifies pushing hard on the sole-remedy and resolution-time gaps, while a lower-stakes internal tool may reasonably accept the vendor's standard terms. The most common pitfall is anchoring entirely on the headline availability percentage and never reading past it to the measurement window, exclusions, and remedy clauses, which is precisely where the real difference between a protective SLA and a decorative one lives.
List and justify the top 8 evaluation criteria you would use when selecting a cloud platform for a large enterprise. For each one, define it precisely, explain why it matters at enterprise scale, and propose one or two measurable metrics or tests you'd use to evaluate vendors against it.
Sample Answer
Selecting a cloud platform for a large enterprise comes down to eight criteria that each protect against a distinct failure mode at scale, not a longer wish list of nice-to-haves; below is each one defined precisely, why it matters specifically at enterprise scale, and one or two measurable ways to test a vendor against it.
The eight criteria
1. Regulatory and compliance coverage. Defined as the set of certifications, attestations, and contractual data-handling terms (for example data residency guarantees, signed data processing agreements) the provider holds for every jurisdiction and industry regulation the enterprise operates under. At enterprise scale, a single unaddressed jurisdiction can block an entire market expansion. Measure it by requesting the provider's current certification list and cross-checking it directly against your specific regulatory obligations, not against a general marketing claim of "enterprise-grade compliance."
2. Total cost of ownership (TCO) at your actual usage pattern. Defined as the full 3-to-5-year cost including compute, storage, egress, support tier, and the staffing needed to operate the platform, not the list price of any single service. At enterprise scale, usage-based pricing components (egress especially) can dwarf the headline compute price. Measure it by building a cost model from your own projected usage and running it against the provider's actual published rate card, not their sales estimate.
3. Managed-service maturity for your specific workload categories. Defined as how far along the reliability, feature-completeness, and operational-simplicity curve the provider's managed offering is for the categories you actually need (databases, container orchestration, machine learning infrastructure), not their catalog's total breadth. Enterprise workloads cannot absorb the operational risk of an immature managed service. Measure it by checking the service's general availability date and searching its own status-history page for the frequency and duration of past incidents.
4. Identity and access management (IAM) depth and enterprise directory integration. Defined as how well the platform's access-control model integrates with your existing enterprise identity provider and supports fine-grained, auditable permissions at the scale of thousands of employees and hundreds of teams. At enterprise scale, a weak IAM model becomes either a security gap or an operational bottleneck as the number of people needing access grows. Measure it by running a proof of concept (a small-scale trial to validate a specific capability) that provisions and audits access for a representative slice of your actual org structure, not a demo account with five users.
5. Global network footprint and inter-region performance. Defined as the provider's number and geographic distribution of regions and availability zones, and measured latency between the regions your business actually operates across. This matters because an enterprise with a genuinely global footprint needs the provider's backbone to match its user geography, not just its headquarters location. Measure it with synthetic latency tests between your specific target regions, run over a representative time window rather than a single snapshot.
6. Exit and migration cost. Defined as the technical and financial cost of moving a representative workload off the platform, including egress fees and how proprietary the managed services you would depend on are. Enterprise contracts run for years, and an option that looks best today can become the expensive one to leave once you are deep into it. Measure it by pricing a hypothetical export of your largest dataset at the provider's published egress rate, and by counting how many of the services in your intended architecture have no equivalent open standard.
7. Support tier and incident response commitment. Defined as the contractual response-time commitments for your support tier, and how those commitments are enforced (service credits, escalation paths). At enterprise scale, a production incident's cost accrues by the minute, and a support tier with a slow contractual response time is a real business risk, not a convenience feature. Measure it by reading the actual SLA (service-level agreement) document's response-time table for your specific proposed tier, not the marketing page's summary.
8. Ecosystem and partner support in your operating markets. Defined as the availability of certified system integrators, local account management, and third-party tooling compatibility in the specific regions and industries you operate in. This matters at enterprise scale because a platform with a thin local ecosystem leaves you dependent entirely on the vendor's own support for every implementation challenge. Measure it by asking the vendor for named reference customers in your industry and region, and independently contacting at least one.
Trade-offs and pitfalls
No provider will lead on all eight criteria simultaneously, so the enterprise's own priority order among these (which is likely to be regulatory coverage and total cost of ownership for a heavily regulated industry, or managed-service maturity and global footprint for a fast-scaling consumer product) should be decided before scoring vendors, not derived afterward to justify a preferred choice. The most common pitfall is trusting a vendor's own marketing claim as the measurement for a criterion instead of running the independent test described for each one above; a certification list, a support SLA, and a synthetic latency benchmark are each verifiable in a way that a sales conversation is not.
Unlock Full Question Bank
Get access to all 19 Infrastructure Strategy and Technology Selection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.