Microsoft Staff DevOps Engineer Interview Preparation Guide
Microsoft's Staff-level DevOps Engineer interview process typically consists of a recruiter screening phase, followed by two technical phone screens, and concludes with 6 onsite interview rounds. The process evaluates expertise in large-scale infrastructure design, cloud platform mastery (with emphasis on Azure), CI/CD architecture, system reliability, and the ability to influence and mentor across teams. Staff-level candidates are expected to demonstrate strategic thinking, ownership of complex systems, and cross-functional leadership.
Interview Rounds
Recruiter Screening
What to Expect
Initial phone conversation with a Microsoft recruiter to assess background, career motivations, compensation expectations, and general fit. This round combines the initial screen and recruiter follow-up into a single interaction. The recruiter will verify your experience matches Staff-level expectations (12+ years, senior project ownership), understand your interest in DevOps at Microsoft, and confirm availability for the interview process.
Tips & Advice
Be clear and concise about your career trajectory and why you're targeting a Staff role at Microsoft. Prepare a 2-minute elevator pitch highlighting your largest infrastructure projects, impact (metrics are critical—cost savings, uptime improvements, deployment frequency gains), and what excites you about DevOps at scale. Ask about the team structure, current challenges they're facing, and the role's scope—this shows genuine interest. Be honest about compensation expectations upfront. Emphasize your experience with cloud platforms (Azure preferred for Microsoft), modern DevOps practices, and team leadership. Show enthusiasm for continuous learning and Microsoft's mission.
Focus Topics
Role Expectations & Team Dynamics
Asking informed questions about team structure, current infrastructure challenges, and how the Staff role influences technical direction.
Practice Interview
Study Questions
Microsoft Culture & Growth Mindset Alignment
Understanding and articulating alignment with Microsoft's growth mindset, learning culture, and collaborative values.
Practice Interview
Study Questions
Career Narrative & Impact Quantification
Crafting a compelling story of your 12+ year DevOps journey with specific, measurable outcomes from major projects.
Practice Interview
Study Questions
Technical Phone Screen 1: Infrastructure Architecture & Cloud Platform Strategy
What to Expect
A 45-60 minute technical phone screen focused on your ability to design and justify infrastructure decisions at enterprise scale. You'll be asked to discuss a complex infrastructure challenge you've solved, make trade-off decisions between cloud services, and explain how you'd architect systems for reliability, cost, and scalability. This round assesses your strategic thinking and depth of cloud platform knowledge (AWS, Azure, GCP).
Tips & Advice
Approach this as a conversational architecture discussion, not a script. Have 2-3 detailed infrastructure projects ready, especially any multi-cloud or large-scale projects. When discussing architecture, explicitly mention trade-offs: 'We chose managed Kubernetes (EKS/AKS/GKE) over self-managed because it reduced operational toil, despite higher unit costs.' Use architecture diagrams verbally—'Let me walk you through the data flow: requests hit the load balancer, route to the Kubernetes cluster in multiple AZs, data persists in a managed database with read replicas for resilience.' For Microsoft/Azure roles, emphasize Azure-specific knowledge: AKS, Azure DevOps, Azure Cost Management. Discuss how you've optimized costs in cloud—tagging strategy, reserved instances, autoscaling, non-prod environment shutdowns. When asked about challenges, be honest and show learning: 'We initially over-engineered with too many services; after incident analysis, we consolidated to reduce complexity.' Practice explaining complex topics simply—recruiters may ask follow-up questions to verify understanding.
Focus Topics
Azure Platform Mastery (for Microsoft context)
Deep knowledge of Azure services (AKS, Azure DevOps, ARM templates, Azure Cost Management, Azure Security Center) and how Azure differs from AWS/GCP.
Practice Interview
Study Questions
Incident Analysis & Infrastructure Improvement
Discussing specific production incidents, root cause analysis, systemic improvements made, and how you prevented recurrence.
Practice Interview
Study Questions
Cost Optimization Strategies
Implementing tagging strategies, reserved instance planning, autoscaling tuning, egress cost minimization, and right-sizing approaches across cloud platforms.
Practice Interview
Study Questions
Large-Scale System Reliability & SRE Principles
Applying SLO/SLI/error budget concepts, designing for high availability across regions, and implementing chaos engineering practices.
Practice Interview
Study Questions
Multi-Cloud Architecture Design & Trade-Off Analysis
Designing resilient infrastructure across AWS, Azure, and GCP with explicit justification of cost vs. managed service trade-offs and vendor lock-in considerations.
Practice Interview
Study Questions
Technical Phone Screen 2: CI/CD Automation & Deployment Pipelines
What to Expect
A 45-60 minute technical phone screen focused on your expertise in designing and implementing enterprise CI/CD pipelines. You'll discuss pipeline architecture, deployment strategies (blue-green, canary, rolling), rollback mechanisms, security scanning, and how you've automated complex deployment workflows. This round evaluates your ability to design pipelines that balance speed, safety, and reliability.
Tips & Advice
Have a detailed CI/CD project ready—ideally one that deployed to multiple environments (dev, staging, prod) with automated testing, security scanning, and rollback capabilities. Walk through the pipeline stages: code commit triggers → automated tests → security/quality gates → approval → infrastructure provisioning → application deployment → monitoring/alerting. Discuss tool choices with justification: 'We migrated from Jenkins to GitLab CI because GitLab provided better Kubernetes integration and reduced operational overhead.' When discussing deployment strategies, be specific: 'For critical services, we use canary deployments with automatic rollback triggered by error rate > 1% or latency p99 > 500ms.' Address these pain points: how do you handle blue-green deployments with stateful services? How do you manage database migrations safely? How do you prevent failed deployments from cascading? For Staff-level, discuss how you've influenced CI/CD culture across teams—standardization, best practices, training. If asked about tools like Jenkins, GitHub Actions, GitLab CI, or Azure Pipelines, show depth: configuration as code, state management, secrets handling, conditional deployments.
Focus Topics
GitOps & Declarative Deployment Patterns
Using ArgoCD, Flux, or similar tools to manage Kubernetes applications declaratively; managing configuration drift; synchronization and promotion workflows.
Practice Interview
Study Questions
Infrastructure as Code (Terraform, ARM Templates, CloudFormation)
Provisioning infrastructure through CI/CD pipelines using IaC; managing Terraform/ARM state; testing infrastructure code; drift detection and remediation.
Practice Interview
Study Questions
CI/CD Security & Compliance Automation
Integrating security scanning (SAST, DAST, container scanning), compliance checks, secrets management, and audit trails into pipelines.
Practice Interview
Study Questions
Deployment Strategies & Rollback Mechanisms
Implementing blue-green, canary, and rolling deployments; designing automated rollback triggered by metrics and alerts; handling stateful services.
Practice Interview
Study Questions
Enterprise CI/CD Pipeline Architecture
Designing multi-stage pipelines with automated testing, security scanning, approval gates, and deployment orchestration across environments.
Practice Interview
Study Questions
Onsite Round 1: System Design – Distributed Infrastructure at Enterprise Scale
What to Expect
A 60-minute hands-on system design round where you'll be asked to architect a complex infrastructure system (e.g., 'Design a global SaaS platform serving millions of requests daily,' 'Migrate a monolithic on-premises application to cloud with zero downtime,' 'Build a disaster recovery strategy for a multi-region deployment'). You'll draw architecture diagrams, discuss component selection, explain trade-offs, estimate costs, and defend your decisions against follow-up challenges. This round evaluates your ability to think strategically about infrastructure at scale.
Tips & Advice
Start by clarifying requirements: scale (requests/second, data volume, users), latency expectations, availability targets, budget, and compliance constraints. Structure your answer: compute (Kubernetes, Lambda, VMs), networking (load balancing, CDN, DNS), storage (databases, caches, object storage), messaging (event streaming), monitoring, and disaster recovery. Draw your architecture step-by-step on the whiteboard—talk through it as you draw. For each component, justify the choice: 'We chose Kubernetes on EKS because we need flexibility to run heterogeneous workloads, but AKS would be equivalent on Azure.' Discuss trade-offs explicitly and quantify them: 'Managed databases cost 3x more than self-managed but reduce operational toil from 2 FTE to 0.2 FTE.' Address failure scenarios: 'If the primary region fails, traffic automatically routes to the secondary region; we maintain warm standby to minimize RTO.' Be prepared for aggressive follow-ups: 'What if your database can't handle the write throughput?' (sharding, write replicas, event sourcing). Stay calm, think aloud, and iterate. For Staff-level, interviewers expect you to consider organizational factors too: 'This design requires 3 teams to collaborate on deployment; we'd implement standardized IaC and CI/CD to enable independent releases.' Use metrics: discuss how you'd measure success (latency percentiles, error rates, cost per request).
Focus Topics
Operational Complexity & Team Enablement
Designing for observability, operability, and team scalability; considering documentation, runbooks, automation, and training requirements.
Practice Interview
Study Questions
Cost Estimation & Resource Optimization
Estimating infrastructure costs for proposed designs, identifying cost optimization opportunities, and balancing cost against reliability and performance.
Practice Interview
Study Questions
Disaster Recovery & Business Continuity
Designing RTO/RPO targets, failover automation, backup strategies, and recovery procedures; testing and validating DR plans.
Practice Interview
Study Questions
Network Architecture & Load Balancing
Designing network topology with load balancers, CDNs, DNS strategies, network segmentation, and latency optimization.
Practice Interview
Study Questions
Database Architecture & Data Storage Strategy
Selecting appropriate databases (relational, NoSQL, time-series), designing for scale (sharding, replication), ensuring consistency, and optimizing query patterns.
Practice Interview
Study Questions
Scalable System Architecture Design (Multi-Region, Multi-AZ)
Architecting systems for millions of concurrent users with high availability, disaster recovery, and low-latency access across geographic regions.
Practice Interview
Study Questions
Onsite Round 2: Kubernetes & Container Orchestration at Scale
What to Expect
A 60-minute deep-dive into Kubernetes expertise. You'll discuss how you'd architect Kubernetes clusters for enterprise use, handle operational challenges (node failures, pod evictions, resource contention), design multi-cluster strategies, implement networking and security policies, and troubleshoot real issues under pressure. Expect both conceptual questions ('How would you design a Kubernetes platform for internal teams?') and hands-on scenarios ('A pod is crash-looping; walk me through debugging'). This round evaluates your operational depth with Kubernetes and ability to support teams building on top of it.
Tips & Advice
Demonstrate expertise across cluster design, application deployment, and troubleshooting. For design questions: discuss cluster sizing (number of nodes, resource allocation), multi-AZ/multi-region strategies, ingress architecture, and admission controllers. Explain differences between managed services (EKS, AKS, GKE) and self-managed: 'AKS manages the control plane, so we focus on node management and network policies. GKE adds Workload Identity for RBAC, which is more secure than IAM role assumption in AKS.' For application deployment, show mastery of Helm, Kustomize, and GitOps: 'We standardize on Helm charts for releases and ArgoCD for GitOps-driven promotion across environments.' For troubleshooting scenarios, use a systematic approach: 'First, I'd check pod status (kubectl get pods -o wide, kubectl describe pod), then look at logs (kubectl logs), then resource usage (kubectl top). If the pod is crash-looping, I'd investigate application logs for startup errors. If it's pending, I'd check node resources and scheduling constraints.' Show understanding of Kubernetes internals: resource requests/limits, QoS classes, eviction policies, and how these interact under resource pressure. For Staff-level, discuss how you've designed Kubernetes platforms for multiple teams, standardized deployment practices, and scaled cluster management. Address multi-cluster strategies: 'We run workloads across regions for DR; cross-cluster communication uses service mesh (Istio) to enable resilience.' Be ready for edge cases: 'Pods are pending despite available nodes—why?' (node affinity, pod disruption budgets, resource fragmentation).
Focus Topics
Multi-Cluster & High-Availability Kubernetes
Designing multi-cluster deployments for disaster recovery, managing cluster upgrades with zero downtime, and implementing cross-cluster networking.
Practice Interview
Study Questions
Kubernetes Networking & Service Mesh
Designing ingress architecture, implementing network policies for segmentation, using service meshes (Istio, Linkerd) for resilience, and managing multi-cluster communication.
Practice Interview
Study Questions
Kubernetes Security & RBAC
Implementing role-based access control, pod security policies, network segmentation, workload identity management, and auditing practices.
Practice Interview
Study Questions
Kubernetes Operational Troubleshooting
Debugging pod startup failures, investigating resource contention, resolving scheduling issues, interpreting Kubernetes events and logs, and diagnosing networking problems.
Practice Interview
Study Questions
Kubernetes Cluster Architecture & Platform Design
Designing clusters for scale, multi-AZ/multi-region deployments, managed vs. self-managed trade-offs, node autoscaling, and resource management.
Practice Interview
Study Questions
Application Deployment & Release Management in Kubernetes
Deploying applications using Helm, Kustomize, GitOps (ArgoCD, Flux), managing ConfigMaps/Secrets, and implementing progressive delivery patterns.
Practice Interview
Study Questions
Onsite Round 3: CI/CD Pipeline Design & Deployment Automation
What to Expect
A 60-minute technical interview focused on your ability to design and implement end-to-end CI/CD pipelines that enable organizational velocity while maintaining safety and reliability. You'll be presented with a scenario (e.g., 'A company wants to increase deployment frequency from weekly to daily; design the pipeline changes needed') and asked to discuss pipeline architecture, tool selection, security scanning integration, deployment strategies, rollback mechanisms, and metrics for success. This round evaluates your understanding of how to balance developer experience, automation, safety, and organizational constraints.
Tips & Advice
Start by clarifying the scenario: current state (tool, frequency, success rate), desired state (frequency, SLOs), constraints (compliance, team skills, budget). Then design the pipeline logically: trigger conditions → source control integration → build and test → security scanning → approval gates → staging deployment → production deployment → monitoring/rollback. For tool selection, show reasoning: 'We chose GitHub Actions because it's natively integrated with GitHub and provides good Kubernetes integration, vs. Jenkins which requires more operational overhead.' Discuss test strategy: 'We run unit tests on every commit, integration tests on pull requests, and end-to-end tests in staging before production.' Address security: 'We scan container images for vulnerabilities with Trivy, run SAST with SonarQube, and require manual approval for production deployments.' For deployment strategy, match the risk profile: 'For critical services, we use canary deployments (5% → 25% → 100%) with automatic rollback if error rate exceeds 1%. For lower-risk services, we use rolling deployments for speed.' Discuss metrics: 'We track deployment frequency, lead time, change failure rate, and MTTR. We target 10 deployments/day, <1 day lead time, <5% failure rate, and <1 hour MTTR.' For Staff-level, discuss team enablement: 'We provide self-service deployment dashboards so developers don't need to understand the full pipeline.' Address scaling challenges: 'As we grew from 10 to 50 microservices, we standardized pipeline templates to avoid duplication and enable faster onboarding.' Show understanding of cost implications: 'Moving CI to serverless reduced costs by 40% and improved build times.'
Focus Topics
Metrics, Observability & Deployment Feedback
Defining deployment success metrics (lead time, deployment frequency, change failure rate, MTTR), implementing automated rollback triggers, and providing visibility into pipeline performance.
Practice Interview
Study Questions
Testing Strategy & Test Automation
Designing test layers (unit, integration, end-to-end, contract), managing test data, optimizing test execution time, and balancing coverage with speed.
Practice Interview
Study Questions
Infrastructure as Code Deployment
Provisioning infrastructure through CI/CD using Terraform/ARM/CloudFormation, managing state safely, testing infrastructure code, and rolling back infrastructure changes.
Practice Interview
Study Questions
Security Scanning & Compliance in CI/CD
Integrating SAST, DAST, container scanning, secrets detection, and compliance checks into pipelines without creating bottlenecks.
Practice Interview
Study Questions
Deployment Automation & Progressive Delivery
Implementing deployment strategies (blue-green, canary, rolling), automating rollback based on metrics, and managing complex deployments with dependencies.
Practice Interview
Study Questions
End-to-End CI/CD Pipeline Architecture
Designing comprehensive pipelines with source control integration, automated testing, security scanning, approval gates, deployment stages, and observability.
Practice Interview
Study Questions
Onsite Round 4: Infrastructure Troubleshooting & Incident Response
What to Expect
A 60-minute practical round where you'll be given real or simulated infrastructure issues and asked to troubleshoot them under time pressure. Scenarios include: 'A Kubernetes service is unreachable,' 'Pod deployment is stuck pending,' 'Database performance has degraded,' 'API latency has spiked,' or 'A deployment failed unexpectedly.' You may be given terminal access to a live environment or asked to walk through your debugging process verbally. This round evaluates your systematic troubleshooting methodology, knowledge of diagnostic tools, communication during investigation, and ability to identify root causes rather than treating symptoms.
Tips & Advice
Approach troubleshooting systematically—never jump to solutions without understanding the problem. (1) Understand the symptom: 'When does the issue occur? Is it consistent or intermittent? What changed recently?' (2) Check the obvious: 'kubectl get pods, kubectl get services, check DNS resolution.' (3) Narrow the scope: 'Is it a networking issue, compute issue, or application issue?' Use structured logging: 'I'd check application logs first for errors, then infrastructure metrics (CPU, memory, disk I/O) to identify resource constraints.' For Kubernetes issues: master kubectl commands (get, describe, logs, port-forward, exec, debug), understand the control plane (scheduler, API server, kubelet), and know how to read events. For networking: understand DNS resolution, service discovery, network policies, and load balancer logs. For database issues: identify slow queries, connection pool exhaustion, and lock contention. For API issues: check latency distribution (p50, p95, p99), error rates, and traffic patterns. Show communication: 'Based on the symptoms, my hypothesis is X. I'd test this by checking Y. If confirmed, the root cause is Z.' Be honest about unknowns: 'I haven't seen this before, but here's my investigative approach.' For Staff-level, discuss how you'd prevent the issue: 'This failure could be caught by better alerting on [metric]. We should implement a health check that detects this within 5 minutes.' Discuss post-incident actions: 'We'd write a blameless postmortem, identify systemic improvements, and track follow-ups.' Practice with real tools if possible (minikube, Docker, Linux utilities); don't just talk about them.
Focus Topics
Incident Response & Post-Incident Learning
Managing incidents under pressure, communicating with stakeholders, driving blameless postmortems, identifying systemic improvements, and tracking action items.
Practice Interview
Study Questions
Database Troubleshooting & Performance Diagnosis
Identifying slow queries, connection pool exhaustion, lock contention, replication lag, and resource constraints; reading database metrics and logs.
Practice Interview
Study Questions
Application Performance & Latency Analysis
Identifying latency sources (application, network, I/O), interpreting latency distributions (percentiles), and using distributed tracing to correlate issues across services.
Practice Interview
Study Questions
Network Troubleshooting & Connectivity Issues
Diagnosing DNS resolution failures, network policy issues, load balancer problems, routing issues, and latency sources using tools like dig, ping, traceroute, tcpdump.
Practice Interview
Study Questions
Linux & Container Troubleshooting
Using Linux utilities (ps, netstat, strace, lsof) to investigate system issues, understanding process behavior, diagnosing network and I/O problems.
Practice Interview
Study Questions
Kubernetes Debugging & Root Cause Analysis
Systematically debugging pod crashes, pending pods, networking issues, and resource constraints using kubectl, logs, and metrics.
Practice Interview
Study Questions
Onsite Round 5: Infrastructure as Code & Cloud Architecture Deep Dive
What to Expect
A 60-minute technical interview focused on your expertise in infrastructure as code (Terraform, CloudFormation, ARM templates) and cloud platform architecture. You'll be asked to design infrastructure using IaC for a given scenario, discuss module design and code organization, address state management challenges, explain testing strategies for infrastructure code, and discuss how you scale IaC across teams. This round evaluates your ability to treat infrastructure with the same engineering rigor as application code, including versioning, testing, review, and safe deployment.
Tips & Advice
Demonstrate that you treat infrastructure code with the same rigor as application code. Start by discussing your IaC tool choice: 'We chose Terraform because it's cloud-agnostic and has a mature module ecosystem, vs. CloudFormation which is AWS-specific or ARM templates for Azure-specific work.' For module design, show architectural thinking: 'We organize modules by abstraction level—low-level (VPC, security group) and high-level (application environment with all dependencies). Each module has clear inputs/outputs and semantic versioning.' Address the state management challenge: 'We use remote state stored in S3 with versioning and locking to prevent concurrent modifications. State is sensitive, so it's encrypted and access is restricted via IAM.' Discuss testing infrastructure code: 'We use Terratest for integration testing (provision → verify → destroy) and Checkov for static analysis (scanning for misconfigurations). Drift detection runs daily to identify infrastructure changes not in code.' For scaling across teams: 'We provide curated modules that enforce organizational standards—security groups with required tags, RDS instances with automated backups, etc. Teams consume modules rather than writing raw Terraform, reducing errors.' Show understanding of cloud differences: 'In AWS, we manage networking with Terraform. In Azure, ARM templates have limitations; we often use Terraform or a higher-level tool. In GCP, we use Google Cloud Deployment Manager for some resources.' For multi-cloud scenarios: 'We avoid cloud-specific features; we use Terraform instead of CloudFormation to avoid lock-in. For Azure, we use Terraform rather than ARM templates for the same reason.' Address common pitfalls: 'We avoid storing secrets in state; we use AWS Secrets Manager or Azure Key Vault. We don't hardcode region or environment; we use variables or workspaces.' Discuss cost implications: 'IaC enables self-service infrastructure, but uncontrolled provisioning can spiral costs. We implement cost policies in Terraform (instance type restrictions, tagging requirements) and conduct regular cost audits.'
Focus Topics
Multi-Cloud & Cloud-Agnostic Infrastructure
Designing infrastructure that's portable across clouds, managing cloud-specific variations, and selecting tools and patterns that minimize lock-in.
Practice Interview
Study Questions
Infrastructure Code Testing & Validation
Testing infrastructure code using tools like Terratest, Checkov, or tflint; implementing policy-as-code; detecting drift and implementing remediation.
Practice Interview
Study Questions
Cloud Platform Architecture (AWS, Azure, GCP)
Deep knowledge of cloud services (compute, networking, storage, managed services) and their trade-offs; designing for multi-cloud portability.
Practice Interview
Study Questions
Infrastructure Change Management & Safe Deployments
Implementing plan/apply workflows in CI/CD, requiring approvals for infrastructure changes, managing dependencies, and safely rolling back infrastructure changes.
Practice Interview
Study Questions
Terraform State Management & Remote State
Managing Terraform state remotely, implementing locking to prevent concurrent modifications, handling secrets in state, and recovering from state corruption.
Practice Interview
Study Questions
Terraform Module Design & Code Organization
Designing reusable, well-documented Terraform modules with clear inputs/outputs; organizing modules hierarchically; managing module versions and dependencies.
Practice Interview
Study Questions
Onsite Round 6: Behavioral Interview & Technical Leadership
What to Expect
A 45-60 minute behavioral and culture-fit round focused on your leadership approach, how you influence across teams, your approach to learning and failure, and alignment with Microsoft's values. You'll be asked about significant infrastructure projects you've led, how you've mentored or influenced colleagues, how you've handled disagreements with stakeholders, and how you approach solving ambiguous problems. This round evaluates whether you're a good cultural fit for Microsoft's collaborative environment and whether you demonstrate the growth mindset and learning orientation Microsoft values.
Tips & Advice
Use the STAR method (Situation, Task, Action, Result) to structure answers, but focus on quantified outcomes that demonstrate impact. Prepare stories about: (1) Leading a significant infrastructure project end-to-end; (2) Mentoring or influencing engineers on your team or across teams; (3) A disagreement with stakeholders (product, finance, security) and how you resolved it; (4) A major failure and what you learned; (5) How you stay current with technology; (6) A time you advocated for a unpopular decision and why it was right. For each story, quantify impact: 'I led the migration to containerized infrastructure, which reduced deployment time from 45 minutes to 8 minutes (82% improvement) and enabled 3x more daily deployments, directly supporting the company's product velocity.' When discussing leadership: 'I mentored 3 junior engineers on infrastructure practices; two of them led infrastructure projects independently within 6 months.' Show learning orientation: 'The migration had unexpected challenges with database replication lag; we investigated root causes, implemented caching, and documented lessons to prevent recurrence.' Address Microsoft-specific values: discuss growth mindset—'I approach problems as learning opportunities and encourage my team to experiment and learn from failures.' Discuss collaboration—'I work closely with application teams to understand their needs; sometimes infrastructure and product goals conflict, and I facilitate discussions to find optimal solutions.' Address any potential concerns about your background directly. Keep answers concise (2-3 minutes per story); let the interviewer drive deeper with follow-ups. Show genuine curiosity about Microsoft's challenges: 'I'm excited about Microsoft's AI initiatives and how they're applying ML to platform engineering. Tell me about how your team is addressing [specific challenge].'
Focus Topics
Learning from Failures & Incident Management Mindset
Discussing significant failures you've experienced, how you analyzed them, what you learned, and how you applied those lessons to improve systems and processes.
Practice Interview
Study Questions
Navigating Ambiguity & Advocating for Technical Decisions
Describing times you've made decisions with incomplete information, advocated for positions that weren't initially popular, and influenced stakeholders to support infrastructure investments.
Practice Interview
Study Questions
Microsoft Culture Alignment: Growth Mindset & Collaboration
Expressing understanding of Microsoft's values (growth mindset, collaboration, innovation, respect for diversity) and providing examples of how you embody these values.
Practice Interview
Study Questions
Growth Mindset & Continuous Learning
Showing how you stay current with rapidly evolving DevOps practices, pursue learning opportunities, and adapt to new technologies and challenges.
Practice Interview
Study Questions
Infrastructure Leadership & Cross-Team Influence
Demonstrating how you've led infrastructure initiatives that spanned teams, influenced organizational practices, and built consensus across stakeholders with different incentives.
Practice Interview
Study Questions
Mentorship & Technical Development of Team Members
Describing how you've mentored junior engineers, helped them grow into independent contributors, and created learning opportunities within your team.
Practice Interview
Study Questions
Frequently Asked DevOps Engineer Interview Questions
You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?
Sample Answer
Whether to fully automate, keep human-in-the-loop, or leave manual comes down to four questions applied to each specific action: how often does it happen, how bad is it if it goes wrong, can it be safely retried, and can success be verified automatically. High frequency, low blast radius, idempotent, and observable pushes toward full automation; anything destructive or hard to verify stays manual or gated behind a human, no matter how routine it feels.
The criteria
| Criterion | Favors automation | Favors manual / human-in-the-loop |
|---|---|---|
| Frequency | Happens often enough that manual toil adds up | Rare enough that automation investment doesn't pay back |
| Blast radius | Failure is contained (one instance, easily reverted) | Failure can be irreversible or affect data integrity broadly |
| Idempotency | Running it twice is harmless | Running it twice causes a different, possibly worse outcome |
| Verifiability | Success can be checked automatically (health check, row count) | Success requires human judgment to confirm |
Applying it to the three actions
Restarting a cached worker instance: high frequency, low blast radius (stateless, replaceable), fully idempotent, and easily verified with a health check. This is a strong automate candidate: drain connections, spin up a replacement, run a health check, cut traffic over, roll back automatically if the health check fails.
Reattaching a detached volume: lower frequency, meaningfully higher blast radius (attaching to the wrong instance or double-attaching can corrupt data), and only moderately idempotent, reattaching twice isn't necessarily safe. This sits in the middle: automate the pre-checks and the mechanical steps (verify volume ID, verify target instance, snapshot before attaching), but require a human to confirm before the final attach executes.
Running a database schema migration: low frequency, high blast radius (can be destructive and hard to reverse), low idempotency for anything involving DDL (data definition language: schema-altering SQL statements like ALTER TABLE), and success often isn't verifiable by a simple automated check, it needs someone to look at whether the data actually came out right. This stays manual, or more precisely, human-gated: automation handles the mechanical parts (schema diff, pre-migration validation, backup, staged rollout to a canary), but a person approves the production apply.
The underlying argument for phasing automation in gradually
Automating a step doesn't just remove toil, it also removes the moment a human would have caught something unusual about this particular instance of the problem. That's fine for the worker-restart case, where "unusual" mostly doesn't exist, but risky for the migration case, where every migration is a little different. The practical path is phasing: run a new automation in shadow mode first (it proposes the action but a human executes), then human-in-the-loop (it executes after one-click approval), and only promote to full automation once it has a track record across enough real incidents that its false-positive and false-negative rate are actually known, not assumed.
Guarding automated actions with least privilege
Whatever is automated should run with only the permissions that specific action needs, a worker-restart automation shouldn't hold credentials that could also run a schema migration, and every automated action should be logged with who (or what) triggered it and why. For the human-in-the-loop tier, the approval step itself should require a specific person's action (not a shared bot token anyone can trigger), so there's a real approval trail, not a rubber stamp.
Trade-offs and pitfalls
The common wrong turn is automating based on how annoying a task feels rather than how safe it is, restarting workers manually is annoying but safe to automate; migrations are also annoying, but the annoyance is not the variable that should decide it. The other pitfall is leaving a human-in-the-loop step gated behind an approval that nobody actually reads before clicking, if the approval doesn't include enough context (what will run, what's the blast radius, what's the rollback) to make a real judgment, it's automation with an extra click, not a genuine safety gate.
Compare reserved instances, savings plans, and committed-use discounts across the major cloud providers. What is the mechanical difference between them in commitment scope, term, and flexibility across instance types, and how would you decide what percentage of a steady-state workload's capacity to commit?
Sample Answer
Direct answer
All three mechanisms trade a usage commitment for a lower price, but they commit to different things: AWS Reserved Instances (RIs) commit to a specific instance configuration, AWS Savings Plans commit to a dollar-per-hour spend level that flexes across instance types, Google Cloud committed-use discounts (CUDs) commit to either a resource quantity or a dollar-per-hour spend depending on which CUD type you buy, and Azure Reservations commit to a specific VM configuration similar to AWS RIs. The general pattern across every provider is the same trade-off: the more precisely you commit to a specific instance shape, the bigger the discount; the more flexibility you keep, the smaller the discount but the lower your risk if the workload changes shape.
Structured elaboration
Mechanism comparison
| Mechanism | Provider | Commits to | Term | Flexibility |
|---|---|---|---|---|
| Standard Reserved Instance | AWS | Specific instance family, size, region | 1 or 3 yr | Least flexible: can change availability zone and, within limits, instance size in the same family, but not family or OS |
| Convertible Reserved Instance | AWS | Instance family (exchangeable) | 1 or 3 yr | Can exchange for a different family, size, or OS during the term, at a lower discount than Standard |
| Compute Savings Plan | AWS | Dollar-per-hour compute spend | 1 or 3 yr | Most flexible: applies across instance family, size, OS, tenancy, and region, and across EC2, Fargate, and Lambda |
| EC2 Instance Savings Plan | AWS | Dollar-per-hour spend, locked to one instance family and region | 1 or 3 yr | Flexible on size and OS within that family and region only; typically a larger discount than Compute Savings Plans for the same term because it's narrower |
| Resource-based CUD | Google Cloud | A quantity of vCPUs, memory, GPUs, or similar, on Compute Engine | Typically 1 or 3 yr | Locked to the committed resource type and quantity; scope can be a single project or shared across a billing account |
| Flexible (spend-based) CUD | Google Cloud | Dollar-per-hour spend | 1 or 3 yr | Pools eligible spend across only three services, Compute Engine, Google Kubernetes Engine (GKE), and Cloud Run, similar in spirit to an AWS Compute Savings Plan in that the discount follows a dollar-per-hour spend level rather than a specific SKU. BigQuery and Cloud SQL are NOT part of this pool: each has its own separate, service-specific spend-based commitment, purchased and applied independently |
| Reserved VM Instance | Azure | Specific VM series, size, and region | 1 or 3 yr | Instance-size flexibility within the same VM size-flexibility group; can be rescoped after purchase to a subscription, resource group, shared billing scope, or management group without a new commercial transaction |
All four providers offer some form of upfront, partial-upfront, or no-upfront (pay monthly) payment on these commitments at the same total cost, so the payment option is a cash-flow decision, not a discount-size decision on most of these products.
Why the scope difference matters in practice
A resource-level commitment (Standard RI, resource-based CUD, Azure Reservation) only pays off if the workload keeps needing that exact shape for the whole term; if the team migrates to a different instance family six months in, the commitment sits partially wasted (though AWS and Azure both allow some exchange or resale mechanisms to recover part of that). A spend-based commitment (Compute Savings Plan, flexible CUD) survives an instance-family change automatically, because the discount is applied to dollars spent on eligible usage, not to a specific SKU, at the cost of a somewhat smaller discount than the narrowest resource-level option.
Deciding what percentage of steady-state capacity to commit
Start from the floor, not the average: pull 3 to 6 months of utilization history for the workload, and find the usage level that held true on the worst week, not the typical week. That floor, not the mean, is the safe commitment baseline, because a commitment above the actual steady floor pays for idle capacity on every low-usage day. From there, the commitment size is a risk trade-off, not a fixed rule: a stable, mature workload with a long recent history of holding above that floor supports committing close to the full floor, while a workload still changing shape (recent re-architecture, aggressive growth, planned migration) justifies leaving more of the floor on-demand or covering it with a flexible, spend-based commitment instead of a rigid resource-level one, specifically because the risk being managed is "commitment outlives the workload's actual shape," not "commitment size in the abstract."
Worked example
A team's steady-state EC2 fleet held at a minimum of 40 instances of a given family over the last 4 months, with normal weekday peaks around 55 and occasional bursts to 70. The 40-instance floor is the commitment candidate, not the 55-instance average and not the 70-instance peak: committing at 55 would mean paying the commitment rate for capacity that isn't reliably used on quieter days, and any spike above 40 (up to and including the 70-instance bursts) is served by on-demand or spot capacity regardless of the commitment size. If this workload is expected to stay on the same instance family for the full term, an EC2 Instance Savings Plan or Standard RI sized to 40 instances captures the largest discount available on that stable floor; if a re-platforming project is likely to change instance family within the year, a Compute Savings Plan sized to the equivalent dollar-per-hour spend protects the same floor's discount while surviving the family change.
Trade-offs and pitfalls
The most common mistake is committing to the peak or the average instead of the floor, which either overpays for capacity that isn't reliably used or, worse, sizes a "safe" commitment so conservatively it captures almost none of the available discount. The second is choosing the narrowest, highest-discount resource-level commitment for a workload that's still changing shape, and then discovering the commitment doesn't match the new instance family, wasting real money for the rest of the term. The third, specific to the flexible/spend-based products, is assuming "flexible" means "no attention needed": a spend-based commitment still needs the underlying usage to stay above the committed dollar level, or the unused portion is still paid for and simply not applied to any usage.
You are defining SLOs for an HTTP JSON API used by a billing product. Describe how you would pick SLO targets and error budgets, which latency and availability metrics to use, and how to translate business impact (e.g., lost revenue, customer churn) into SLO thresholds. Explain how error budgets should influence release cadence and incident response playbooks.
Sample Answer
Picking SLO targets
Start from what actually affects the business and users, not an arbitrary round number: for a billing API, a slow or failed checkout call directly risks lost revenue and support load, so the SLO (Service Level Objective: the internal target you set for a metric like latency or availability) should track what customers actually feel. Anchor the target against your own baseline distribution first (don't set a target far tighter than what the system currently, realistically achieves without major investment), and set it off percentiles, not the average, since the average can look healthy while a real slice of checkout attempts are painfully slow.
Metrics to use
- Latency: p95 and p99 response time for the billing endpoints, being explicit about whether failed or timed-out requests are included in the latency population or tracked separately (excluding them silently is a common way an SLI (Service Level Indicator: the actual measured metric, e.g. p99 latency, that gets compared against the SLO) quietly stops meaning what people think it means).
- Availability: the percentage of requests that complete successfully (not 5xx, not timed out), often tracked as its own SLO alongside latency rather than folded into it.
Translating business impact into thresholds
Look at your own funnel or conversion data for where checkout abandonment starts climbing as latency increases, and set the latency ceiling with real headroom below that point, not right at the edge of where users start leaving. The same logic applies to availability: estimate what a percentage point of failed billing requests costs in lost transactions and support tickets, and use that to justify how tight (and how expensive) the target should be, since tighter targets generally cost more engineering effort and infrastructure to sustain.
Error budgets
The error budget is simply what's left over: 1 minus the SLO. A 99.9% availability SLO over a 30-day month leaves a 0.1% error budget:
30 days = 43,200 minutes
0.1% of 43,200 minutes = 43.2 minutes
That's roughly 43 minutes of full-equivalent downtime allowed per month, whether spent as one incident or many smaller partial-degradation periods that add up.
How the error budget should drive release cadence and incident response
- Release cadence: while budget remains healthy, ship at normal (or even accelerated) velocity. As the budget gets consumed, slow down: require canary rollouts (releasing a change to a small slice of traffic or users first, before rolling out to everyone), extra review, or freeze risky changes entirely until the budget recovers. This turns "reliability" from a vague value into a concrete, negotiable resource that product and engineering both watch.
- Incident response: alert on burn RATE (how fast the budget is being consumed) rather than only on absolute SLO breach, so a fast, severe burn pages immediately while a slow, minor burn becomes a lower-urgency ticket. Postmortems and remediation priority should also weigh directly against how much budget an incident consumed, giving the team a shared, quantitative way to decide whether to prioritize a reliability fix over the next feature.
Write a Terraform module that provisions a VPC with a configurable list of availability zones and subnets, using for_each rather than count. It should accept the VPC CIDR, the AZ list, and a flag for whether to create NAT gateways, and expose the resulting subnet and route table IDs. Walk through your variable and output design.
Sample Answer
Direct answer
Build one VPC, then iterate AZs with for_each (not count) so each subnet is keyed by AZ name instead of a list index. That matters because with count, removing an AZ from the middle of the list shifts every subsequent index and Terraform proposes destroying and recreating subnets that didn't actually change. Keying the resource by AZ name only fixes half of that: the CIDR block assigned to each subnet also has to be tied to the AZ itself, not to that AZ's position in a list, so each AZ gets a stable, explicitly-assigned subnet index (a small az => netnum map) instead of index(var.azs, each.value). NAT gateway creation is gated behind a boolean input and created per AZ only when enabled. Outputs expose subnet IDs and route table IDs as maps keyed by AZ, so a caller can look up exactly the subnet it needs by name instead of guessing a list position.
Approach
- One
aws_vpc, sized byvar.vpc_cidr. - One
aws_subnetper AZ viafor_each = var.az_netnums(a map of AZ name to a stable subnet index), with each subnet's CIDR carved out viacidrsubnet(var.vpc_cidr, var.subnet_newbits, each.value), whereeach.valueis that AZ's fixed index, not its position in a list. - One public route table (shared, routes to the internet gateway) and one private route table per AZ (routes to that AZ's NAT gateway only if NAT is enabled; otherwise no default route at all).
- NAT gateways and their EIPs are
for_each-keyed the same way as the subnets, and only exist whenvar.enable_nat_gateway = true.
The code below is illustrative HCL meant to be read, not executed output from a real terraform apply (no CLI available here to run it against real state).
variables.tf
variable "vpc_cidr" {
type = string
description = "CIDR block for the VPC, e.g. 10.0.0.0/16."
}
variable "az_netnums" {
type = map(number)
description = "Map of AZ name to a stable subnet index used by cidrsubnet, e.g. { \"us-east-1a\" = 0, \"us-east-1b\" = 1 }. Assign each AZ its number once, in the order it was first added, and treat this map as append-only: never renumber or reuse an existing AZ's index, even after that AZ is removed."
}
variable "subnet_newbits" {
type = number
description = "Additional bits for cidrsubnet when carving subnets out of vpc_cidr (e.g. 8 turns a /16 into /24s)."
}
variable "enable_nat_gateway" {
type = bool
default = false
description = "Whether to create one NAT gateway per AZ for private-subnet egress."
}
main.tf
resource "aws_vpc" "this" {
cidr_block = var.vpc_cidr
enable_dns_support = true
enable_dns_hostnames = true
tags = { Name = "vpc-module" }
}
resource "aws_subnet" "this" {
for_each = var.az_netnums
vpc_id = aws_vpc.this.id
availability_zone = each.key
cidr_block = cidrsubnet(var.vpc_cidr, var.subnet_newbits, each.value)
tags = { Name = "subnet-${each.key}" }
}
resource "aws_internet_gateway" "this" {
vpc_id = aws_vpc.this.id
}
resource "aws_eip" "nat" {
for_each = var.enable_nat_gateway ? var.az_netnums : {}
domain = "vpc"
}
resource "aws_nat_gateway" "this" {
for_each = var.enable_nat_gateway ? var.az_netnums : {}
allocation_id = aws_eip.nat[each.key].id
subnet_id = aws_subnet.this[each.key].id
tags = { Name = "nat-${each.key}" }
}
resource "aws_route_table" "public" {
vpc_id = aws_vpc.this.id
}
resource "aws_route" "public_internet" {
route_table_id = aws_route_table.public.id
destination_cidr_block = "0.0.0.0/0"
gateway_id = aws_internet_gateway.this.id
}
resource "aws_route_table_association" "public" {
for_each = aws_subnet.this
subnet_id = each.value.id
route_table_id = aws_route_table.public.id
}
resource "aws_route_table" "private" {
for_each = var.az_netnums
vpc_id = aws_vpc.this.id
tags = { Name = "rt-private-${each.key}" }
}
resource "aws_route" "private_egress" {
for_each = var.enable_nat_gateway ? aws_route_table.private : {}
route_table_id = each.value.id
destination_cidr_block = "0.0.0.0/0"
nat_gateway_id = aws_nat_gateway.this[each.key].id
}
resource "aws_route_table_association" "private" {
for_each = aws_subnet.this
subnet_id = each.value.id
route_table_id = aws_route_table.private[each.key].id
}
outputs.tf
output "subnet_ids" {
value = { for k, s in aws_subnet.this : k => s.id }
}
output "route_table_ids" {
value = {
public = aws_route_table.public.id
private = { for k, rt in aws_route_table.private : k => rt.id }
}
}
Key design decisions
for_eachkeyed by AZ name, notcount, and the CIDR index keyed the same way. The AZ string is a stable, meaningfulfor_eachkey, but that alone isn't enough:cidrsubnet's index argument also has to come from a fixedaz => netnummap (var.az_netnums) instead ofindex(var.azs, each.value), or removing/reordering an AZ silently recomputes every other AZ'sindex()result and reshuffles their CIDR blocks, forcing a destroy/recreate on subnets that didn't actually change (cidr_blockis aForceNewattribute onaws_subnet). With both the resource key and the CIDR index pinned to the AZ itself, adding or removing an AZ only touches the resources tied to that specific key.az_netnumsis append-only: never reuse or renumber an existing AZ's index, even after that AZ is removed.- A "private" subnet with NAT disabled gets no default route at all, rather than falling back to routing straight to the internet gateway. Routing a "private" route table to the IGW when NAT is off would silently make those subnets not actually private, which is a common mistake worth naming explicitly if you see it in someone else's module.
- Outputs are maps keyed by AZ, not lists. A consuming module or root config wants
module.vpc.subnet_ids["us-east-1a"], notmodule.vpc.subnet_ids[0]and a separate assumption about ordering.
Complexity
Resource count scales linearly with the number of AZs: for n AZs you get n subnets, n private route tables, and (if NAT is enabled) n EIPs and n NAT gateways, so the plan/apply graph is O(n) in both size and, roughly, cost. cidrsubnet with subnet_newbits = 8 splits a /16 into /24s, giving up to 28=256 possible subnet blocks, of which this module only allocates n of them (one per AZ), leaving room to grow the AZ list later without re-carving existing subnets, since each AZ's block is computed from its own fixed netnum in az_netnums, not from a list position.
Edge cases
- Fewer than 2 AZs: the module still works, but a single-AZ deployment has no real availability benefit; worth flagging back to whoever is calling it.
- The
az_netnumsmap is reordered in the config file, or an AZ is removed: since both thefor_eachkey and thecidrsubnetindex come fromaz_netnumsrather than a list position, every remaining AZ keeps its assigned netnum and therefore its exactcidr_block; HCL map literals have no meaningful order in the first place, so reordering entries in the file is purely cosmetic. Removing an AZ destroys only that AZ's own subnet, route table, and (if enabled) NAT gateway and EIP; its netnum must not be reassigned to a different AZ afterward, or that AZ's CIDR (and, sincecidr_blockisForceNew, its subnet) would change too. enable_nat_gatewayflips from true to false on an existing deployment: this destroys every NAT gateway and EIP and removes the private route tables' default routes, cutting off private-subnet egress. That is a real production-impacting change, not just a cosmetic one, and belongs behind the same "review the plan's destroy count" discipline as any other production change.subnet_newbitstoo small for the number of AZs:cidrsubneterrors out at plan time if you ask for more subnet bits than the parent CIDR has room for, so this fails fast rather than silently.
Trade-offs and pitfalls
Making this production-ready and reusable pulls in a few things beyond the base module:
- VPC Flow Logs: add an
aws_flow_logresource pointed at CloudWatch Logs or an S3 bucket, gated behind its own optional variable, so network visibility is available to callers who want it without forcing it (and its cost) on every caller. - An S3 gateway endpoint: add an
aws_vpc_endpointof typeGatewayfor S3, associated with the private route tables. This lets private-subnet workloads reach S3 without traversing the NAT gateway, which avoids NAT data-processing charges and keeps that traffic off the public path entirely. - Versioning, testing, and documentation for reuse: tag the module with semantic versions in git (
v1.2.0) and have consumers pin to a version (source = "git::...?ref=v1.2.0"or a registry version constraint) rather than tracking a branch. Write a README documenting every input and output with an example call. Test withterraform test(native since Terraform 1.6) or Terratest against a real or LocalStack-backed provider, asserting that togglingenable_nat_gatewayactually changes the private route tables' routes and that the exposed output maps contain exactly the AZ keys passed in. - One-shared-NAT-gateway alternative: creating one NAT gateway per AZ (as above) maximizes availability but multiplies NAT cost by the AZ count; a single shared NAT gateway in one AZ is cheaper but makes every other AZ's egress depend on that one AZ staying healthy, worth surfacing as an explicit cost/resilience trade-off rather than a silent default.
Design a secure and scalable CI/CD pipeline for deploying to production Kubernetes clusters that enforces image signing and verification, vulnerability scanning, and policy-as-code gates. Recommend tools (for example: Tekton or ArgoCD, cosign/notation for signing, Trivy for scanning, OPA/Gatekeeper for policies), explain how signing keys and secrets are managed, and describe automated rollback and audit trails.
Sample Answer
For a Kubernetes production deployment specifically, the pipeline's job is to make sure nothing gets scheduled onto the cluster unless it's been scanned, signed, and can be verified as such at the moment Kubernetes actually tries to run it, using named, concrete tooling rather than abstract components.
Recommended tools and the flow
flowchart LR
Src[Source] --> Tekton[Tekton pipeline]
Tekton --> Trivy[Trivy scan]
Trivy --> Cosign[cosign/notation sign]
Cosign --> Registry[Container registry]
Registry --> ArgoCD[ArgoCD sync]
ArgoCD --> Gatekeeper[OPA/Gatekeeper admission]
Gatekeeper -->|signature + scan verified| Cluster[K8s cluster]
Tekton (or an equivalent Kubernetes-native pipeline engine) builds and orchestrates the pipeline itself, running natively alongside the cluster it deploys to. Trivy scans the built image for known vulnerabilities before it's allowed to be signed; a CRITICAL finding blocks the pipeline before signing ever happens, since there's no point cryptographically attesting to an image you already know is unsafe to ship. cosign (or notation, Microsoft's equivalent) signs the scanned image, again preferably keylessly via OIDC (OpenID Connect) rather than a long-lived key the pipeline has to protect. ArgoCD reconciles the desired state from Git to the cluster, and OPA/Gatekeeper enforces, at admission time, that any image being scheduled carries a valid signature from the expected signer and comes from an approved registry; an unsigned or unverifiable image is rejected at the Kubernetes API layer itself, not merely flagged.
Key and secret management
Signing keys (or, for keyless signing, the OIDC trust configuration binding which CI identity is allowed to produce a valid signature) are the highest-value asset in this whole chain, since anything with signing authority can make an attacker-controlled image look legitimate; keyless signing via Sigstore's Fulcio removes the long-lived-key management burden entirely, tying signing authority to short-lived, OIDC-verified certificates instead. Application secrets needed at runtime are handled separately from signing, typically via the Secrets Store CSI driver or a Vault Agent injecting them at pod startup, scoped per namespace.
Automated rollback and audit trails
ArgoCD's own sync history gives a built-in rollback mechanism (reverting to a prior, known-good commit's manifests triggers a resync to that state); pairing this with the admission controller's own decision log (every allow and deny decision, with the reason) gives a complete audit trail of both what was deployed and why any given deployment was allowed or blocked.
Trade-offs
Enforcing signature verification at the admission-controller layer, rather than earlier in the pipeline alone, is what actually prevents an unsigned or tampered image from running even if it somehow bypassed the CI pipeline entirely (a compromised registry push, for instance); the cost is that a misconfigured admission policy can block legitimate deploys just as effectively as it blocks malicious ones, so the policy itself needs its own careful testing and staged rollout before being made mandatory cluster-wide.
Explain what a Platform as a Service (PaaS) offering is and how it differs from IaaS and serverless. Give concrete examples of PaaS products suited to web applications and to container-based workloads, then walk through the operational trade-offs a team takes on by choosing PaaS instead of one of the other two models.
Sample Answer
Direct answer
PaaS sits between IaaS and serverless on a single spectrum of "how much of the running instance do I still manage," rather than off to the side as an unrelated third category. IaaS gives you a virtual machine and you build everything above it. PaaS gives you a managed runtime and deployment target: you push code and the platform runs it continuously. Serverless goes one step further than PaaS by removing the concept of a continuously running instance altogether, scaling to zero and billing per invocation instead.
How PaaS differs from each neighbor
Versus IaaS. On IaaS you decide the OS, patch it, install a runtime, and configure a process manager, and an instance typically stays running whether or not it's serving traffic. On PaaS you supply application code or a container image and a deployment manifest; the platform handles OS patching, runtime installation, health checks and restarts, and autoscaling policy configuration replaces manual instance management.
Versus serverless. PaaS still asks you to think in terms of how many instances and what size, even if it autoscales them for you, and instances generally stay warm between requests. Serverless removes the instance concept entirely, scales per invocation including to zero, and bills per invocation and duration rather than per provisioned hour.
Examples for the two workload shapes the question asks about
Web applications: AWS Elastic Beanstalk, Azure App Service, and Google App Engine let you push source code and get a managed, autoscaled web service with a built-in load balancer.
Container-based workloads: managed container-runtime platforms such as Google Cloud Run or AWS App Runner take a container image, giving you the middleware and dependency control of a container, while the platform still handles the OS, orchestration, and autoscaling, including scaling to zero when idle. This is why these products sit right at the boundary between PaaS and serverless.
Operational trade-offs of choosing PaaS
Instead of IaaS: you give up fine-grained OS and kernel control and the ability to install arbitrary system-level software, and you gain freedom from patch management and instance-level monitoring and alerting.
Instead of serverless: you give up the cost benefit of scaling fully to zero for spiky or idle workloads and take on slightly more capacity-planning thinking, how many instances, what size, but you gain a simpler mental model for stateful or long-running processes and typically fewer cold-start surprises (the extra delay a request hits when it wakes an idle instance or function that has to initialize before it can respond).
Worked example
A small team is choosing between the three for a customer-facing web app with steady, moderate daily traffic that also needs a custom authentication middleware wrapping every request. Raw IaaS would mean owning OS patches for no real benefit, since nothing about the workload needs kernel-level control. Pure serverless functions would technically work but would mean re-architecting the middleware into a function-per-endpoint shape, adding real engineering cost for a workload that isn't particularly spiky. A container-based PaaS product lets the team package the existing middleware into the container image unchanged, deploy it once, and let the platform handle scaling and OS management, which is the best fit given steady traffic and existing middleware investment.
Trade-offs and pitfalls
The main pitfall is picking PaaS reflexively as "the safe middle choice" without checking whether the workload is actually spiky, where serverless's per-invocation billing would save real money, or genuinely needs OS-level control, where PaaS's restrictions become a real blocker, for example a workload needing raw socket access a managed platform won't expose. A second pitfall is assuming all PaaS products are equally restrictive: a container-based PaaS product gives you a much larger slice of control over the runtime and middleware than a source-code-only PaaS product, so the choice between PaaS flavors matters almost as much as the IaaS-versus-PaaS-versus-serverless choice itself.
Explain the differences between Azure SQL Database (single database / elastic pools), Azure SQL Managed Instance, and running SQL Server on an Azure VM. For each option discuss operational responsibilities, high availability patterns (failover groups, active geo-replication, Always On), maintenance control, compatibility, and scenarios where each is preferable (SaaS, lift-and-shift, legacy apps).
Sample Answer
Direct answer
Azure SQL Database (single database or elastic pool) is the right default for a new application that doesn't need instance-level SQL Server features: Microsoft owns patching, backups, and most high availability (HA) entirely. Azure SQL Managed Instance (MI) is the migration-friendly middle ground: near-complete SQL Server surface-area compatibility (SQL Server Agent, Common Language Runtime/CLR, cross-database queries and transactions, linked servers, Service Broker) inside your own Virtual Network (VNet), still as a managed service. SQL Server on an Azure virtual machine (VM) is the last resort for workloads that need something even MI doesn't support, or need an exact SQL Server version, edition, or operating-system-level control pinned, at the cost of owning patching and HA yourself.
Comparison
| Option | Operational responsibility | High availability pattern | Maintenance control | Best for |
|---|---|---|---|---|
| SQL Database, single database or elastic pool | Microsoft: patching, backups, most HA | Zone-redundant configuration (where available) or geo-replication/auto-failover groups for cross-region | None; Microsoft schedules it | New cloud-native software-as-a-service (SaaS) applications; elastic pools specifically for many databases with unpredictable, non-overlapping usage sharing one resource pool |
| SQL Managed Instance | Microsoft: patching, backups, most HA | General Purpose: remote-storage-based redundancy; Business Critical: a 4-node Always On availability group under the hood | None; Microsoft schedules it | Lift-and-shift of an app depending on SQL Agent, cross-database transactions, linked servers, or Service Broker, without rewriting it |
| SQL Server on an Azure VM | You: patching, backups, HA architecture | You build it yourself (Always On availability groups, failover cluster instances) | Full; you choose the patch level and timing | Legacy features even MI doesn't support, or a pinned SQL Server version/edition an application depends on |
Evaluation checklist
When choosing among the three for a given workload, walk through, in order: (1) which SQL Server compatibility features the application actually calls, since that alone can rule out SQL Database; (2) the team's appetite for owning patching and HA operations; (3) the target Recovery Time Objective (RTO, how long the application can stay down before it must be back up) and Recovery Point Objective (RPO, how much recent data you can afford to lose); (4) the budget model, Database Transaction Unit (DTU) or virtual core (vCore) consumption pricing versus a VM plus a SQL Server license.
Worked example
A legacy on-premises application uses SQL Server Agent jobs, several cross-database transactions, and one linked server to a partner system. Azure SQL Database supports none of that by default (cross-database access there is limited to a workaround called elastic query, not true cross-database transactions), so a rewrite would be required to land it there. SQL Managed Instance supports SQL Agent, cross-database transactions, and linked servers directly, letting the migration proceed largely unchanged; the real migration effort shifts to the VNet integration itself, MI requires a dedicated subnet, correctly sized and with no other resource types placed in it, plus Domain Name System (DNS) and network security group planning before cutover. If the application also depended on something MI still doesn't support (an unusual Component Object Model/COM object hosted in-process, for example), SQL Server on a VM would be the fallback, accepting full patching and HA ownership in exchange for full compatibility.
Trade-offs and pitfalls
- Picking SQL Database first because it's the simplest to provision, without checking for SQL Agent or linked-server dependencies, is a common way to discover a rewrite is needed mid-migration.
- MI's VNet requirement is not optional and has real subnet-sizing and DNS-planning work attached; underestimating that work is the most common MI migration surprise.
- Running SQL Server on a VM to dodge MI's constraints reintroduces the exact patching and HA burden the other two options were built to remove; only choose it when a specific, named feature gap forces it.
Write a concise Go CLI program that accepts three inputs: (1) a JSON array of build inputs (file paths + SHA256), (2) a JSON array of outputs (file paths + SHA256), and (3) a PEM-format private key file path. The program should produce a JSON provenance attestation containing inputs, outputs, timestamp, builder ID (from BUILDER_ID env var), and a base64 signature field signing the attestation. Use only Go standard library packages. Include comments to explain deterministic JSON serialization choices.
Sample Answer
A provenance attestation is fundamentally a signed statement, so the implementation has two halves that both have to be exactly right: building a deterministic payload to sign, and signing it in a way a verifier can independently reproduce and check.
package main
import (
"crypto"
"crypto/rand"
"crypto/rsa"
"crypto/sha256"
"crypto/x509"
"encoding/base64"
"encoding/json"
"encoding/pem"
"fmt"
"os"
"time"
)
type FileDigest struct {
Path string `json:"path"`
Sha256 string `json:"sha256"`
}
// Field order is fixed by the struct definition, which is what makes
// json.Marshal deterministic for this payload: the same logical content
// always produces the same bytes to sign and to re-verify against.
type Attestation struct {
Inputs []FileDigest `json:"inputs"`
Outputs []FileDigest `json:"outputs"`
Timestamp string `json:"timestamp"`
BuilderID string `json:"builderId"`
Signature string `json:"signature"`
}
func signPayload(payload []byte, priv *rsa.PrivateKey) (string, error) {
digest := sha256.Sum256(payload)
sig, err := rsa.SignPKCS1v15(rand.Reader, priv, crypto.SHA256, digest[:])
if err != nil {
return "", fmt.Errorf("signing attestation payload: %w", err)
}
return base64.StdEncoding.EncodeToString(sig), nil
}
func loadPrivateKey(pemPath string) (*rsa.PrivateKey, error) {
data, err := os.ReadFile(pemPath)
if err != nil {
return nil, fmt.Errorf("reading key file: %w", err)
}
block, _ := pem.Decode(data)
if block == nil {
return nil, fmt.Errorf("no PEM block found in %s", pemPath)
}
key, err := x509.ParsePKCS1PrivateKey(block.Bytes)
if err != nil {
return nil, fmt.Errorf("parsing private key: %w", err)
}
return key, nil
}
func buildAttestation(inputsPath, outputsPath, keyPath string) (*Attestation, error) {
inputsRaw, err := os.ReadFile(inputsPath)
if err != nil {
return nil, fmt.Errorf("reading inputs file: %w", err)
}
var inputs []FileDigest
if err := json.Unmarshal(inputsRaw, &inputs); err != nil {
return nil, fmt.Errorf("parsing inputs JSON: %w", err)
}
outputsRaw, err := os.ReadFile(outputsPath)
if err != nil {
return nil, fmt.Errorf("reading outputs file: %w", err)
}
var outputs []FileDigest
if err := json.Unmarshal(outputsRaw, &outputs); err != nil {
return nil, fmt.Errorf("parsing outputs JSON: %w", err)
}
builderID := os.Getenv("BUILDER_ID")
if builderID == "" {
return nil, fmt.Errorf("BUILDER_ID environment variable is not set")
}
priv, err := loadPrivateKey(keyPath)
if err != nil {
return nil, err
}
unsigned := Attestation{
Inputs: inputs,
Outputs: outputs,
Timestamp: time.Now().UTC().Format(time.RFC3339),
BuilderID: builderID,
}
payload, err := json.Marshal(unsigned)
if err != nil {
return nil, fmt.Errorf("canonicalizing payload: %w", err)
}
sig, err := signPayload(payload, priv)
if err != nil {
return nil, err
}
unsigned.Signature = sig
return &unsigned, nil
}
func main() {
if len(os.Args) != 4 {
fmt.Fprintln(os.Stderr, "usage: provenance <inputs.json> <outputs.json> <private_key.pem>")
os.Exit(2)
}
att, err := buildAttestation(os.Args[1], os.Args[2], os.Args[3])
if err != nil {
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
out, _ := json.MarshalIndent(att, "", " ")
fmt.Println(string(out))
}
Deterministic JSON serialization
The struct's field order is fixed at compile time by its definition, so json.Marshal always emits fields in the same order for the same logical content; this is why the payload is built as a typed Attestation struct rather than a map[string]interface{}, since Go's JSON encoding of a map sorts keys alphabetically by default but a struct preserves declaration order, either of which is deterministic on its own, but mixing the two within one payload risks a verifier reconstructing a different byte layout than the signer used. The signature field itself is excluded from the signed payload (set to empty string, or omitted, at signing time) since the signature obviously can't be part of what it's signing over.
Verified
Compiled with go build. Generated a real RSA keypair with OpenSSL, ran the program to produce a signed attestation, then wrote an independent Go verifier that re-reads the attestation, strips the signature field, re-marshals the remaining struct exactly as the signer did, and calls rsa.VerifyPKCS1v15 against the public key: verification succeeded. A tamper test, flipping one character in an input's SHA-256 digest and re-running verification, correctly failed with a signature-mismatch error, confirming the scheme actually detects tampering rather than passing regardless of content.
Trade-offs
Using PKCS1v15 padding and a single RSA key here is simple and fully supported by the Go standard library alone (no external dependency), matching the question's constraint; a production system would more likely use Sigstore's keyless signing (short-lived, OIDC (OpenID Connect)-backed certificates) to avoid the operational burden of protecting a long-lived private key file: that file has to be generated once and then protected for its entire lifetime, encrypted at rest, access-restricted, rotated on a schedule, and revoked immediately if it is ever exposed, and if it does leak, whoever holds it can forge valid attestations indefinitely, or at least until the leak is discovered and the key revoked, with no built-in record of who actually produced a given signature. Keyless signing removes that persisted secret entirely: at sign time the CI runner exchanges a short-lived OIDC token (proof that this exact signing request came from this exact workflow run) for a certificate from a public certificate authority (Sigstore's Fulcio), signs with an ephemeral key that is discarded the moment signing completes, and records the signature alongside the certificate and a transparency-log entry (Rekor) that anyone can later audit. There is no long-lived key file to steal, back up, or rotate, at the cost of depending on the OIDC identity provider and Sigstore's infrastructure being reachable at sign time. Container image signing, discussed separately in this topic, makes the identical trade-off: a persisted key you must guard forever versus an ephemeral, OIDC-backed identity with nothing to steal.
List and explain five key metrics you would monitor to assess the health and performance of a production relational database. For each metric, describe one alert condition you might set and why.
Sample Answer
Direct answer
Five metrics cover the core questions a responder needs answered first: is the database slow (a high-percentile query or write latency), is it actually serving load correctly (throughput and error rate), is it running out of room (connection and CPU or memory saturation), is its working set still fitting in memory (buffer or cache hit rate), and, if replicas exist, is it falling behind (replication lag). Pick alert thresholds relative to the workload's own historical baseline and its SLO (service-level objective, an internal target for how the system should perform), not a generic industry number, and expect an OLTP workload (online transaction processing, many small, fast operations) and an OLAP workload (online analytical processing, fewer, larger, longer-running queries) to need different baselines for the same metric.
Structured elaboration
| Metric | What it tells you | Example alert condition | Why that condition |
|---|---|---|---|
| p99 query or write latency | The tail experience, the slowest 1% of operations, catches problems an average hides | p99 latency more than 3x its trailing 7-day baseline for 5 consecutive minutes | Relative-to-baseline avoids one fixed number being wrong for every workload; 5 minutes filters transient blips from real degradation |
| Throughput and error rate (queries or transactions per second, and error or rollback rate) | Whether the database is serving load, and serving it correctly | Error rate above 1% of transactions over a rolling 5-minute window | A latency-only view can miss a database that's fast because it's failing fast, for example rejecting connections outright |
| Connections in use (percent of max) and CPU or memory saturation | Whether the instance is running out of room to accept more work | Connections in use above 80% of the configured maximum for 5 minutes, or CPU above 85% sustained for 10 minutes | Headroom-based thresholds catch the problem before the hard ceiling (connection refusals, out-of-memory) rather than after |
| Buffer or cache hit rate (percent of reads served from memory rather than disk) | Whether the working set still fits in memory; a falling hit rate predicts rising I/O-driven latency before it fully arrives | Hit rate drops more than 10 percentage points below its 30-day baseline | A sudden drop usually means a growing working set or a large one-off query evicting the cache, both worth knowing early |
| Replication lag (if replicas exist) | How far behind a replica is, and how much recent data a failover would risk losing | Lag exceeds the defined RPO (recovery point objective, the maximum acceptable amount of recently written data you could afford to lose), or a fixed threshold like 30 seconds absent a stated RPO | Lag is meaningless without a target; the condition should be "will this violate our stated data-loss tolerance," not an arbitrary number |
On a managed database
A managed relational database service typically removes OS shell access, but the same five categories are still visible, just organized through the provider's own console or API instead of OS commands: host-level metrics (CPU, memory, disk, network, exposed as instance-level monitoring), database-level metrics (connections, buffer hit ratio, via the engine's own views proxied through the console), query-level metrics (top SQL by load, via a performance-insight feature layered on top rather than raw direct access to the engine's own statistics views), and replication metrics (lag, as a first-class console metric since the provider manages the replica topology). Losing shell access doesn't mean losing observability; it means the same five things live in a different namespace. Know where each one lives before an incident, not during one.
OLTP versus OLAP baselines
The same five metrics need structurally different thresholds depending on workload shape. An OLTP system, many small, fast, latency-sensitive operations, should alert tight on p99 latency, since milliseconds matter, and treat connection saturation as urgent, since a starved OLTP system fails user-facing requests immediately; its throughput baseline is usually fairly steady and predictable. An OLAP system, fewer, larger, longer-running analytical queries, should alert loose on p99 latency, since a single query legitimately taking minutes is normal, but tight on sustained resource saturation, CPU, memory, or I/O held near capacity for an extended period is the actual OLAP health signal, since a single slow query is expected but the whole system staying pinned means queuing is building, and on queue depth (how many queries are waiting to run) rather than on individual query latency. Applying an OLTP-shaped p99-latency alert to a warehouse would page constantly on completely normal behavior, and applying an OLAP-shaped hourly-average view to an OLTP system would miss a real user-facing latency spike entirely.
Worked example
Turning the p99-latency alert condition into concrete numbers, showing why a baseline-relative threshold rather than a fixed one is the right shape:
p99_baseline_ms = 42 # trailing 7-day baseline
multiple = 3
threshold_ms = p99_baseline_ms * multiple
print(threshold_ms)
126
At a 42ms trailing baseline, the alert fires above 126ms. If legitimate growth naturally moves the baseline to, say, 55ms next quarter, the threshold recalculates to 165ms automatically, whereas a fixed ">100ms" rule set today would already be uncomfortably close to normal, healthy behavior by then and would need someone to remember to revisit it.
Trade-offs and pitfalls
- A fixed threshold copied from a blog post or another team's dashboard, rather than derived from this workload's own baseline, is the most common reason alerts are either constantly noisy (too tight for this workload's normal variance) or dangerously quiet (too loose, so real incidents don't page). Baseline-relative thresholds avoid this, but need enough historical data to compute a meaningful baseline; don't set them from one week of a system still finding its steady state right after launch.
- Alerting only on averages instead of a high percentile hides exactly the tail-latency problems users actually feel; a p99 spike can coexist with a perfectly healthy-looking average if it affects a small fraction of requests, which is often the fraction most likely to matter most.
- Five metrics is a floor, not a ceiling. This is the minimum viable health picture, not a complete observability strategy; deeper investigation, specific slow-query identification, lock waits, log growth, still needs its own dedicated tooling once one of these five flags a problem.
After a release with repeated friction between design and engineering, how would you run the retrospective, and what would you want to come out of it that actually changes how the two teams work together going forward?
Sample Answer
Direct answer
A retro after a release with repeated design-engineering friction should produce two things: an honest, specific account of where the handoff actually broke down, not a vague 'communication issues,' and a small number of concrete process changes, each with an owner and a way to tell in a quarter whether it worked. Running it well means separating fact-finding from diagnosis, and diagnosis from blame.
Structured elaboration
Design principles for the session
- Facts before diagnosis: start from a timeline of what actually happened (spec dates, handoff dates, bug counts, points where implementation and design diverged), not from opinions about who was at fault.
- Root cause, not the nearest symptom: 'engineering didn't follow the spec' is a symptom; the root cause might be that the spec didn't capture edge-case states, or that both sides were working from different versions of a shared design system mid-migration.
- Few, high-leverage commitments: two or three process changes people will actually do beat ten action items that quietly get dropped.
- Everyone leaves with the same understanding of what changed, not just what went wrong.
A workable structure
One illustrative shape, adaptable to a team's own rhythm:
| Segment | Goal |
|---|---|
| Shared timeline | Ground the room in what happened, not opinions |
| Perspective mapping | Small mixed groups surface where the handoff broke, from each side's view |
| Root-cause discussion | Push past the first symptom to the structural cause |
| Prioritize and commit | Pick a small number of changes, each with an owner and a way to check later whether it worked |
What 'actually changes how the two teams work' looks like
The output isn't a list of intentions, it's a specific artifact or habit that exists after the meeting and didn't before: a shared checklist embedded in the handoff process, an automated check that catches a class of mismatch before it ships, or a standing short sync during implementation windows. Whatever it is, it needs a way to tell if it worked, not just that it happened.
Worked example
One team's root cause turned out to be that design tokens (colors, spacing values) were maintained in the design tool but hand-copied into code, so drift was inevitable and nobody could tell which side was 'correct' when they disagreed. The concrete fix was an automated export from the design tool into the codebase, checked by both a design reviewer and a frontend reviewer before merge, plus a short recurring sync during active implementation. A quarter later, the team had a real signal that it worked: noticeably fewer visual-mismatch comments on pull requests and less late-stage rework than the release that triggered the retro. The same root-cause pattern shows up in other domains as a hand-copied data contract or config value instead of a design token, so the same fix shape (automate the handoff, add a lightweight check, add a short sync during the risky window) generalizes well beyond design and engineering specifically.
Trade-offs and pitfalls
- A retro that produces ten action items usually produces zero completed ones; prioritizing ruthlessly matters more than being thorough.
- If the room jumps straight to solutions or blame instead of facts first, the real root cause, often structural or tooling-related rather than a person's failure, never surfaces.
- A retro that isn't revisited becomes theater. Put the check-in on the calendar before the room disperses, not as a vague intention afterward.
- Watch for a fix that only addresses this specific release's symptom (a one-off manual double-check) rather than the structural cause; it holds for one cycle and then quietly stops happening.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths