Infrastructure Strategy and Technology Selection Questions
Setting technical direction for infrastructure and deciding what to build on. Covers infrastructure vision and long-term roadmap, modernization and technical-debt strategy, platform decisions, and organizational and governance considerations, alongside the decision framework for adopting or retiring technology: build-versus-buy-versus-cloud-versus-on-premises trade-offs, vendor and platform evaluation, technology-portfolio rationalization, requirements-driven service selection, and weighing total cost of ownership and risk. It also covers managing an existing vendor relationship: quantifying and mitigating lock-in, negotiating contract terms, and planning an exit from a service the organization depends on. The leadership-altitude discipline of directing an infrastructure estate over time, building consensus among stakeholders with conflicting priorities, and persuading leadership or a team to accept a major technology change.
Engineering leadership worries about vendor lock-in as the organization standardizes on one cloud provider. Propose a framework to quantify lock-in risk in business-relevant terms (not just technical ones) - covering egress fees, proprietary managed-service dependence, and estimated cost/effort to migrate later - and the architectural mitigations (abstraction layers, data-format standards, periodic migration drills) you'd recommend.
Sample Answer
Direct answer
Convert lock-in from a vague worry into a weighted numeric score built from a handful of business-relevant dimensions, then convert the score into an actual migration-cost-and-effort estimate in dollars and months. A number a finance stakeholder can compare against the value of staying is far more useful than a qualitative "high" or "low" risk label, and the scoring dimensions and mitigation catalog both need to flex by technology category, because a serverless function and an analytics platform fail differently.
Structured elaboration
Weighted scoring dimensions
Score each candidate service 1 (low risk) to 5 (high risk) on each dimension, multiply by its weight, and sum:
LockInScore=20×∑i=16wi⋅si
| Dimension | Weight | What it captures |
|---|---|---|
| Data egress cost exposure | 0.25 | Egress is the data that leaves the vendor's network; this dimension is the cost to move your data volume out at current pricing |
| Proprietary managed-service dependence | 0.25 | How much business logic lives inside vendor-specific managed features, not portable code |
| API/interface uniqueness | 0.15 | Whether an open standard or common interface exists as an alternative |
| Contractual constraints | 0.15 | Commitment length, termination penalties, minimum spend |
| Skills/tooling specialization | 0.10 | How much team expertise is tied to vendor-specific tooling |
| Ecosystem interdependency | 0.10 | How many other services from the same vendor this one is wired into |
The weights sum to 1, so the weighted average itself never leaves the 1-5 scoring range; the 20x multiplier rescales that 1-5 average into a 20-100 business-facing score, 20 if every dimension scores the safest possible 1, 100 if every dimension scores the riskiest possible 5, not a 0-100 score, since a weighted average of numbers that are never below 1 can never reach 0.
Technology-specific indicator sets
- Serverless/FaaS (function-as-a-service, a model where you deploy individual functions rather than servers): a proprietary event-trigger schema, an orchestration layer (a step-function-style workflow engine) that embeds business logic in vendor-specific configuration, and a per-invocation billing model that is hard to replicate elsewhere for cost comparison.
- Cloud analytics platform: a query dialect with non-standard SQL extensions, a storage format tied to the vendor's own compute engine, and data gravity (the sheer volume and the time and egress cost required to move it) as its own lock-in force independent of the API surface.
Cost and effort to migrate
MigrationCost=M⋅Ceng+E+D
where M is estimated engineer-months to rebuild the integration, Ceng is fully loaded monthly cost per engineer, E is one-time data egress cost, and D is the cost of running both systems in parallel during a dual-run validation window.
PaaS-specific scoring rubric
For a platform-as-a-service (PaaS) dependency such as a proprietary workflow-orchestration platform, score these four dimensions the same 1-5 way: API compatibility (can an equivalent platform run the same interface), data exportability (can workflow definitions and history be exported intact), custom runtime dependence (does your code rely on a runtime only this platform provides), and vendor-specific tooling dependence (deployment, monitoring, or debugging tools with no equivalent elsewhere).
Mitigations as a 24-month roadmap
| Phase | Action | Checkpoint |
|---|---|---|
| Months 0-6 | Introduce an abstraction layer around the highest-risk dependency; land canonical data in an open format | Abstraction layer covers the top-scoring service |
| Months 6-12 | Stand up a multi-cloud orchestration layer (for example, infrastructure-as-code and a container runtime that both major clouds support) for new workloads | New workloads default to the portable pattern |
| Months 12-24 | Run a live migration drill on the highest-risk existing dependency; refactor it if the drill reveals gaps | Drill completes within the estimated M months from the cost model |
Worked example
Scoring a proprietary FaaS platform: egress exposure 4, managed-service dependence 5, API uniqueness 4, contractual constraints 2, skills specialization 3, ecosystem interdependency 4. Weighted sum: 0.25(4)+0.25(5)+0.15(4)+0.15(2)+0.10(3)+0.10(4)=1.0+1.25+0.6+0.3+0.3+0.4=3.85, times 20 gives a LockInScore of 77 out of 100, solidly high risk. Migration estimate: rebuilding the function layer on a portable container runtime is estimated at 5 engineer-months at a fully loaded $18,000 per engineer-month ($90,000), plus $12,000 in one-time data egress, plus a 6-week dual-run window costing roughly $15,000 in duplicated infrastructure, for a total estimated migration cost of $117,000. That number, not the abstract "77," is what goes in front of a budget owner.
Trade-offs and pitfalls
The most common failure is scoring the dimension that is easiest to measure (egress cost, which is a line item) while under-scoring the one that matters more (managed-service dependence, which requires actually reading how deeply the code is coupled). A second failure is treating the score as permanent: usage and dollar exposure grow with adoption, so a service scored low at pilot scale should be re-scored as it becomes load-bearing. Finally, do not spend more on the scoring exercise than the position justifies; a low-weighted, low-volume service does not need the full rubric.
Your company must satisfy data-residency and encryption requirements that differ across the countries you operate in. As the technical decision-maker, how do you factor that multi-jurisdiction compliance requirement into your infrastructure choices (regions, key management, backup location) without fragmenting your architecture unnecessarily?
Sample Answer
Direct answer
Keep the application architecture shared wherever the law actually allows it, and localize only the specific data stores and encryption keys that a jurisdiction genuinely constrains; standing up an entirely separate regional stack per country by default is the expensive reflex most teams reach for first, and it is usually broader than what the law actually requires.
Structured elaboration
Classify by legal constraint, not by convenience
For each data category (personally identifiable information, payment data, health data, and so on), determine per country whether it faces a true residency requirement (data must physically stay in that region), a processing or transfer-adequacy constraint (data can leave, but access and processing must meet certain controls), or no special constraint at all. These are legally distinct, and conflating them is the most common cause of unnecessary fragmentation.
Architectural response by layer
- Regions: deploy a dedicated regional presence only where a genuine residency requirement exists; everything else stays on the existing shared regions.
- Key management: use a key management system supporting regional key isolation for the residency-constrained data specifically, so its encryption keys never leave that jurisdiction, while non-constrained data continues on the existing shared key hierarchy; this is usually far cheaper than replicating an entire database.
- Backup location: backups of residency-constrained data must also stay in-region; a commonly missed gap is solving primary-region residency and then routing backups to a shared, out-of-region disaster-recovery vault, which is often still a violation of the same requirement.
- Application tier: keep a shared control plane (the layer that manages configuration, deployment, and orchestration, not the customer data itself) and application logic where the law allows, and isolate only the data plane (where the actual customer data is stored and processed) where required, using a routing layer that directs a given tenant's data-plane calls to the correct jurisdiction's data store, so most of the system, application code, the deployment pipeline, and observability, stays unfragmented.
Worked example
A company serves customers in the European Union, which falls under the General Data Protection Regulation (GDPR); GDPR itself does not strictly mandate in-EU storage, it mandates adequate protection for any transfer outside the EU, an important distinction from a hard residency law. The same company also serves a customer in a country with an explicit data-localization law requiring certain data categories to stay in-country, a stricter model than GDPR's transfer-adequacy approach.
| Jurisdiction | Actual legal requirement | Architecture response |
|---|---|---|
| European Union (GDPR) | Transfer-adequacy, not strict residency | Keep data in the nearest compliant region and rely on an adequacy mechanism (such as standard contractual clauses) rather than forcing a dedicated EU-only deployment |
| Country with strict data-localization law | True in-country residency for the constrained data categories | Stand up a dedicated regional data store, regional key management, and a regional backup vault only for that jurisdiction's constrained categories; the application tier remains shared |
Trade-offs and pitfalls
Over-fragmenting by default, treating every country with any privacy law as if it required a full dedicated regional stack, multiplies the total cost of ownership and the ongoing operational burden for no actual legal benefit. Getting the legal distinction wrong yourself, rather than engaging counsel per jurisdiction, is a serious risk here specifically because whether a law requires strict residency versus transfer-adequacy is a legal determination, not an engineering judgment call; treat your own read of the requirement as a working hypothesis for legal review, never as the final answer. And the backups-in-the-wrong-region gap described above is, in practice, the most common real finding when this area gets audited.
Design a process for evaluating and approving new cloud services or technologies before they're adopted anywhere in the organization (for example a new managed storage product). Detail the steps, who needs to be involved, the risk checks, and the approval gates.
Sample Answer
An approval process for adopting new cloud services needs to be fast enough that teams use it instead of routing around it, and rigorous enough that a genuinely risky choice (a new database with no support for your compliance regime, for example) gets caught before it is embedded in production. The design that achieves both is a tiered gate: low-risk, well-understood categories move through a lightweight self-service check, and only higher-risk categories escalate to a full review board.
Steps, participants, and gates
Step 1: classify the request by risk tier before anything else. A short intake form captures what the service is, what data it would touch, and which of a small set of risk categories apply (handles regulated data, is a new infrastructure-as-a-service/platform-as-a-service/software-as-a-service (IaaS/PaaS/SaaS) category the org has no existing vendor relationship in, requires a new identity and access management (IAM) integration, or is a low-risk addition within an already-approved vendor's product line). This classification, not the requester's own risk assessment, determines which gate applies next.
Step 2, low-risk tier: self-service approval against a published checklist. For a request that stays within an already-approved provider and touches no regulated data (for example a new managed queue from a cloud provider you already use broadly), the requesting team self-certifies against a short published checklist (cost estimate attached, no new data classification, no new external network exposure) and proceeds without a review meeting, logging the decision for later audit.
Step 3, higher-risk tier: a review board with named, specific participants. For anything touching regulated data, introducing a genuinely new vendor relationship, or requiring new IAM integration, convene a small standing review board rather than an ad hoc group assembled per request: a security representative (checks data handling and access model), a cost or finance representative (checks the pricing model and long-term cost exposure, especially usage-based pricing with no cap), an architecture representative (checks fit against existing standards and lock-in exposure), and the requesting team's technical lead. Keep this board small and standing, meeting on a fixed cadence (for example biweekly) with a defined turnaround service-level objective (SLO), so it functions as a fast gate rather than a queue that discourages teams from asking.
Step 4: an explicit decision-criteria matrix for the specific case of choosing among IaaS, PaaS, and SaaS. Because this is one of the most common request types, give it a named sub-process rather than a generic risk conversation: does the team need control over the underlying runtime or operating system (favors IaaS), does the team want to own only application code and let the provider manage the runtime (favors PaaS), or does an off-the-shelf product already do the job with no custom logic needed (favors SaaS)? Score each option against control needed, time to value, and total cost of ownership, using the org's standard vendor scorecard so this decision is not re-litigated from scratch for every request.
Step 5: a documented, revisitable decision, not a one-time yes. Every approval, at either tier, is logged with the reasoning and a review date (commonly annual, or triggered by a material pricing or feature change), so approval is a point-in-time judgment the org can revisit, not a permanent endorsement.
Responsibilities matrix
| Role | Low-risk tier | Higher-risk tier |
|---|---|---|
| Requesting team | Self-certify against checklist | Present business case and technical fit to the board |
| Security | Publishes the checklist criteria; audits samples after the fact | Reviews data handling and access model directly, every request |
| Finance | Publishes cost-checklist thresholds | Reviews pricing model and cost exposure directly |
| Architecture | Not directly involved | Reviews fit against standards and lock-in exposure |
| Review board (standing) | Not convened | Convenes on fixed cadence, gives a yes/no/conditional decision within the SLO |
Worked example
A team requests a new managed vector database from a provider the org has never used, to store data that includes customer personal information. Intake classifies this as higher-risk on two counts: new vendor relationship and regulated data. It goes to the standing board, which requests a data processing addendum from the vendor (security), checks the pricing model for a usage-based cost that could scale unpredictably with data volume (finance), and confirms whether an existing approved vendor's product could serve the same need before approving a net-new relationship (architecture). The board approves conditionally: proceed, with a mandatory annual cost and security review, rather than a blanket unconditional yes.
Trade-offs and pitfalls
The tiered design trades some rigor on the low-risk path for speed, on the bet that most requests are genuinely low-risk and forcing every request through a full board either creates a bottleneck that teams learn to route around (using a service without asking) or trains the board to rubber-stamp requests out of review fatigue. The main pitfall is letting classification itself become political, with teams downplaying risk to reach the faster self-service tier; mitigate this with periodic security audits of a sample of self-certified requests, not just trust. A second pitfall is treating an approval as permanent, which is why the review-date field above exists: a vendor that was the right choice at approval time can become the wrong one after a pricing change or a compliance-scope change, and a process with no revisit trigger will not catch that until an incident forces the conversation.
When evaluating a managed platform vendor, what security and compliance questions should you make sure to ask? Cover encryption and key management, identity and access controls, audit logging, data residency, breach notification, and relevant certifications.
Sample Answer
Direct answer
Ask the vendor to show evidence, not describe intentions, across six areas: encryption and key management, identity and access controls, audit logging, data residency, breach notification, and certifications. A vendor that answers fluently but cannot point to a document, a report, or a contract clause for each one has not actually answered.
Structured elaboration
| Area | What to ask | Why it matters |
|---|---|---|
| Encryption and key management | Is data encrypted at rest and in transit by default? Can we bring or manage our own keys through a key management system (a service that generates, stores, and rotates encryption keys), and can we revoke a key to make our data unreadable to the vendor? | Default-on encryption is table stakes; customer-managed keys are the difference between "the vendor promises not to look" and "the vendor is technically unable to." |
| Identity and access controls | Do you support single sign-on and short-lived federated credentials (OIDC, OpenID Connect, a standard for proving identity across systems without sharing a password) rather than only long-lived API keys? Is access role-based and can we scope a role to least privilege? | Long-lived shared credentials are the most common root cause of a third-party breach; federation lets you revoke access centrally the moment an employee leaves. |
| Audit logging | Do you log administrative and data-access actions, how long are logs retained, and can we export them to our own SIEM (security information and event management system)? | If the vendor is breached, your own retained logs are often the only way to determine what of yours was touched, independent of what the vendor tells you. |
| Data residency | Which physical region(s) does our data (including backups) live in, and can we pin it? | A verbal "we're global" answer often means backups silently replicate to a region your regulator does not permit. |
| Breach notification | What is the contractual notification window, and does it match or beat the legal requirement you are subject to (for example, the General Data Protection Regulation's 72-hour notification requirement to the relevant authority)? | A vendor's marketing page and its actual signed contract can promise different numbers; only the contract is enforceable. |
| Certifications | Which certifications (SOC 2 Type II, an independent auditor's report on a vendor's security controls; ISO/IEC 27001; and for payment data PCI DSS) do you hold, and does the audit's stated scope actually cover the specific product or service we are buying? | Vendors with a large portfolio sometimes hold a certification for one product line and imply it covers all of them. |
Worked example
Evaluating a payments-adjacent vendor for a fintech company serving European Union customers: you ask for the SOC 2 Type II report and find its scope explicitly excludes the newer API product you actually intend to use, which is the real finding, not "do they have SOC 2." You ask about breach notification and the contract says 30 days; your own regulatory obligation under GDPR (the EU's General Data Protection Regulation) is to notify the supervisory authority within 72 hours of becoming aware, so a 30-day vendor clause would make you non-compliant before you even know there was an incident, and that gap has to be renegotiated or the vendor rejected. You ask about key management and learn keys are vendor-managed only, meaning a subpoena or an internal vendor incident gives them technical ability to read your data regardless of their policy promises.
Trade-offs and pitfalls
Treating a certification badge as the finding instead of reading the report's actual scope is the most common mistake; a badge tells you an audit happened, not that the product you're buying was inside it. Pre-sales teams will often answer these questions optimistically without engineering sign-off, so ask for the security or compliance team directly rather than accepting an account manager's paraphrase. Finally, resist writing a purely defensive checklist with no way to fail a vendor: decide in advance which answers are disqualifying (no customer-managed keys for regulated data, no contractual notification window at or better than your legal obligation) so the exercise produces a decision, not just a document.
A major cloud provider region suffers a prolonged outage. How would prior platform choices (managed vs. self-managed, single-region vs. multi-region) have affected recovery options and customer impact? What concentration-risk mitigations would you have put in place beforehand, knowing this could happen?
Sample Answer
Direct answer
Prior technology-selection choices set your recovery ceiling long before any outage happens; you cannot architect a fast recovery during the incident if the underlying choice (a single region, a fully managed service with no verified cross-region path) never allowed for one. The recovery you get during a real regional outage is a direct, mechanical consequence of decisions made months or years earlier, not something improvised on the day.
Structured elaboration
| Prior choice | Consequence during a regional outage |
|---|---|
| Single region, fully managed | Total dependency on that vendor's recovery of that one region; your recovery time objective (RTO) equals the vendor's, and your recovery point objective (RPO, how much data you can afford to lose) is whatever the vendor's own cross-region durability guarantee happens to be, which is often none for a regional-only service. Customer impact is a full outage for the incident's entire duration. |
| Single region, self-managed on the same cloud | Same regional blast radius as above; you control your own failover mechanics once the region's underlying capacity returns, which does not help at all during the outage itself. |
| Multi-region, fully managed | Recovery is bounded by how the managed service's cross-region replication and failover actually work, which needs to have been verified beforehand, not assumed. Customer impact shrinks to a failover window, potentially minutes, if that failover was tested in advance. |
| Multi-region, self-managed or multi-cloud | Highest resilience, and also the highest ongoing cost and complexity, since someone has to build and operate the data-consistency and failover engineering; this is the same cost-versus-risk trade-off that applies to any resilience investment. |
Concentration-risk mitigations to have in place beforehand
Identify true single points of failure at the region level specifically, not just the service level; a setup spread across multiple availability zones (physically separate data centers within the same region, each with independent power and networking) within one region is still single-region and offers no protection against a regional outage, a very common false-confidence gap. Define and actually test cross-region failover before you need it; an untested runbook fails exactly when it matters. Understand the gap between the vendor's own service-level agreement credits and your customer-facing commitment, since they are rarely the same number, and that gap is your uncovered risk. Finally, decide explicitly, in dollars, how much permanent multi-region cost you are willing to carry against the quantified probability and impact of a regional outage, treating it as an insurance premium you are choosing to pay rather than a vague fear you are reacting to.
Worked example
A company ran a fully managed database service in a single region. During a roughly 6-hour regional outage, its recovery point objective turned out in practice to be whatever the vendor's last cross-region backup snapshot happened to be, which, since it had never been verified, turned out to be 18 hours stale. Recovery therefore meant a full outage for the 6-hour incident, plus 18 hours of transactions that had to be manually reconciled afterward. The fix implemented afterward was a tested multi-region read replica with an automated failover runbook drilled quarterly, cutting recovery time to under 15 minutes and recovery point to under 1 minute of replication lag going forward.
Trade-offs and pitfalls
Mistaking a multi-availability-zone setup for regional resilience is the single most common false-confidence gap in this area, and it is worth naming explicitly rather than assuming. Never testing the failover path is the second most common failure, echoing the same pattern as an untested vendor-exit runbook. And treating multi-region as free insurance ignores its real, ongoing dollar and operational cost, which needs to be sized against the actual probability and business impact of the specific outage scenario, not against a general anxiety about outages.
Unlock Full Question Bank
Get access to all 8 Infrastructure Strategy and Technology Selection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.