DevOps Engineer (Mid-Level) Interview Preparation Guide
The DevOps engineer interview process typically consists of an initial recruiter screening, followed by technical phone screens, and multiple onsite rounds. These rounds assess your proficiency with CI/CD pipelines, containerization technologies, cloud infrastructure, system design thinking, troubleshooting abilities, and cultural fit. Mid-level candidates are expected to demonstrate strong hands-on experience, the ability to own projects end-to-end, and basic system design understanding.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with the technical recruiter to confirm your background, verify job fit, understand your career motivations, and assess basic communication skills. This round also explains the role, team structure, and interview process.
Tips & Advice
Be clear and concise about your DevOps experience. Highlight specific projects where you improved deployment processes, reduced downtime, or automated infrastructure tasks. Ask thoughtful questions about the team, the current infrastructure challenges, and what success looks like in the first 6 months. Mention your familiarity with the tools mentioned in the job description (Jenkins, Docker, Kubernetes, cloud platforms, monitoring tools). Show enthusiasm for bridging development and operations.
Focus Topics
Cloud Platform Proficiency
Confirm your experience with at least one major cloud platform (AWS, Azure, or GCP). Be ready to discuss VPCs, IAM, load balancers, auto-scaling, and infrastructure provisioning.
Practice Interview
Study Questions
Career Background and Motivation
Articulate your journey into DevOps, key projects that shaped your expertise, and why you're interested in this role at this company.
Practice Interview
Study Questions
CI/CD Pipeline Experience
Discuss hands-on experience building and maintaining continuous integration and deployment pipelines. Mention specific tools (Jenkins, GitLab CI, GitHub Actions) and measurable outcomes.
Practice Interview
Study Questions
Technical Phone Screen - CI/CD and Automation
What to Expect
A technical conversation focused on your hands-on experience with CI/CD pipelines, infrastructure automation, and practical problem-solving. Expect questions about tools, processes, and real-world scenarios you've encountered. You may be asked to explain architecture decisions, troubleshoot deployment issues, or discuss automation strategies.
Tips & Advice
Have specific examples ready from your work: describe a CI/CD pipeline you built, how you automated infrastructure provisioning, or a deployment issue you resolved. Be prepared to explain not just what you did, but why you made those technical choices. Walk through your thought process step-by-step. Discuss trade-offs between different approaches (e.g., containerization vs. VMs, different orchestration strategies). Mid-level expectations: you should own the explanation and demonstrate understanding of the underlying concepts, not just tool usage. Ask clarifying questions to show depth of thinking.
Focus Topics
Deployment Process Automation
Experience automating deployment processes: reducing manual steps, implementing blue-green or canary deployments, handling configuration management across environments.
Practice Interview
Study Questions
Docker and Container Management
Proficiency with Docker: building images, managing containers, Docker Compose, networking, volumes, and best practices for containerization.
Practice Interview
Study Questions
Infrastructure as Code (Terraform, CloudFormation, Ansible)
Hands-on experience with IaC tools to provision and manage cloud infrastructure programmatically. Discuss managing state, modularity, testing IaC, and version control for infrastructure.
Practice Interview
Study Questions
CI/CD Pipeline Design and Implementation
Design and explain CI/CD pipelines: build stages, testing automation, deployment strategies, rollback mechanisms, and integration with version control systems.
Practice Interview
Study Questions
Technical Phone Screen - System Design and Architecture
What to Expect
Focused on basic system design thinking appropriate for mid-level. You'll be asked to design scalable, reliable infrastructure for hypothetical services (e.g., designing a deployment infrastructure, monitoring solution, or high-availability architecture). Emphasis is on trade-offs, scalability considerations, and communication of your design decisions.
Tips & Advice
Start by clarifying requirements and constraints. Discuss your approach systematically: identify components, explain how they interact, and justify your choices. For mid-level, focus on practical, real-world systems rather than enterprise-scale complexity. Be prepared to discuss trade-offs (e.g., cost vs. complexity, ease of management vs. scalability). Talk through how your design handles failure scenarios. Listen to interviewer hints and adjust your design. Show that you understand not just technology choices but why certain approaches are better for specific requirements.
Focus Topics
Disaster Recovery and Backup Strategy
Design backup and recovery strategies: RTO/RPO objectives, data redundancy, backup frequency, testing recovery procedures, and handling data loss scenarios.
Practice Interview
Study Questions
Scalability and Auto-Scaling
Design systems that scale: horizontal vs. vertical scaling, auto-scaling strategies, load distribution, and capacity planning.
Practice Interview
Study Questions
Monitoring, Logging, and Observability
Design monitoring and logging solutions: metrics collection, alerting strategies, log aggregation, dashboards, and identifying key performance indicators.
Practice Interview
Study Questions
High-Availability Architecture Design
Design highly available systems: redundancy, failover mechanisms, load balancing, health checks, and strategies for minimizing downtime.
Practice Interview
Study Questions
Onsite Round 1 - Containerization and Orchestration Deep Dive
What to Expect
In-depth technical discussion focused on Docker and Kubernetes, which are core to modern DevOps. Expect hands-on scenario questions, discussions about containerization best practices, pod management, network policies, resource management, and troubleshooting containerized applications.
Tips & Advice
This is deep technical content. Be comfortable with Kubernetes architecture: master the concepts of pods, services, deployments, StatefulSets, ConfigMaps, and secrets. Discuss real-world troubleshooting: how would you debug a pod that won't start? How do you handle networking issues? Be prepared to discuss Kubernetes best practices for security, resource management, and monitoring. For Docker, understand image layers, multi-stage builds, and registry management. Use kubectl commands confidently. Mid-level expectation: you should be able to design and troubleshoot containerized deployments independently.
Focus Topics
Docker Image Management and Best Practices
Docker image optimization, multi-stage builds, layer caching, security scanning, registry management, and image tagging strategies.
Practice Interview
Study Questions
Kubernetes Resource Management and Scaling
Resource requests/limits, Horizontal Pod Autoscaler (HPA), Vertical Pod Autoscaler (VPA), node affinity, pod disruption budgets, and handling resource constraints.
Practice Interview
Study Questions
Kubernetes Troubleshooting and Debugging
Practical troubleshooting: diagnosing pod failures, debugging networking issues, checking logs, using kubectl describe/logs/exec, and identifying cluster problems.
Practice Interview
Study Questions
Kubernetes Networking and Service Discovery
Kubernetes networking model, service types (ClusterIP, NodePort, LoadBalancer), ingress controllers, network policies, and DNS resolution.
Practice Interview
Study Questions
Kubernetes Architecture and Core Concepts
Deep understanding of Kubernetes: pods, services, deployments, StatefulSets, DaemonSets, ConfigMaps, secrets, namespaces, and RBAC.
Practice Interview
Study Questions
Onsite Round 2 - Cloud Infrastructure and Infrastructure as Code
What to Expect
Technical deep dive into cloud infrastructure management, focusing on one major cloud platform (AWS recommended) and Infrastructure as Code tools like Terraform. Discuss VPCs, IAM, load balancers, auto-scaling, storage solutions, networking, and how you provision and manage infrastructure programmatically. May include scenario-based questions about infrastructure design.
Tips & Advice
Pick one cloud platform (likely AWS based on job description) and know it deeply. Understand VPC architecture, subnetting, routing, IAM policies, security groups, load balancers, and auto-scaling groups. For Terraform, be comfortable with state management, modules, variables, outputs, and infrastructure versioning. Discuss how you've used IaC to manage multiple environments (dev, staging, prod). Be ready to explain trade-offs in infrastructure design (e.g., multi-AZ for reliability vs. cost). Practice writing simple Terraform or CloudFormation in your head. Mid-level expectation: you should design infrastructure changes and explain them clearly, considering both technical and business implications.
Focus Topics
Cost Optimization and Resource Management
Strategies for cost optimization: right-sizing instances, reserved instances, spot instances, monitoring costs, and identifying waste.
Practice Interview
Study Questions
IAM and Security Best Practices
Identity and Access Management: IAM users/roles/policies, least privilege principle, service accounts, credential management, and security auditing.
Practice Interview
Study Questions
AWS Core Services and Architecture
Proficiency with AWS services: EC2, VPC, subnets, security groups, IAM, load balancers, auto-scaling, S3, RDS, CloudWatch, and service integration.
Practice Interview
Study Questions
VPC Design and Networking
VPC architecture, subnetting strategy, routing tables, NAT gateways, VPN, peering, security group configuration, and network troubleshooting.
Practice Interview
Study Questions
Terraform and Infrastructure as Code
Terraform proficiency: modules, state management, backend configuration, variable management, version control for infrastructure, and testing IaC.
Practice Interview
Study Questions
Onsite Round 3 - Production Troubleshooting and Incident Response
What to Expect
Scenario-based technical round focused on real-world troubleshooting and incident response. You'll be presented with production issues (e.g., slow deployments, service outages, resource constraints, security incidents) and asked to diagnose and resolve them. Interviewers assess your systematic troubleshooting approach, technical depth, communication during incidents, and ability to prevent future issues.
Tips & Advice
Use a systematic troubleshooting methodology: gather information, form hypotheses, test them, and iterate. Start by understanding the scope and impact of the issue. Show your thought process as you diagnose: which logs to check, which metrics to examine, what commands to run. For infrastructure issues, walk through: check resource utilization → review logs → verify connectivity → check recent changes → identify root cause → implement fix → monitor. For mid-level, you should demonstrate ownership of incidents, communicate clearly about trade-offs, and think about both immediate fixes and long-term prevention. Be comfortable explaining your reasoning and asking clarifying questions.
Focus Topics
Performance Optimization and Capacity Planning
Identifying performance bottlenecks, optimization strategies, capacity forecasting, and planning for growth.
Practice Interview
Study Questions
Deployment and Rollback Strategies
Managing deployments safely: blue-green deployments, canary releases, health checks, rollback procedures, and handling failed deployments.
Practice Interview
Study Questions
Production Incident Response and Postmortem
Incident response procedures: communication during incidents, severity assessment, escalation, recovery steps, root cause analysis, and preventing recurrence.
Practice Interview
Study Questions
Monitoring, Logging, and Alerting
Collecting metrics, analyzing logs, setting up effective alerts, creating dashboards, and using monitoring tools to identify and diagnose issues.
Practice Interview
Study Questions
Linux System Administration and Troubleshooting
Linux fundamentals: file systems, processes, networking (TCP/IP, DNS), systemd, permissions, log files, and command-line troubleshooting tools.
Practice Interview
Study Questions
Onsite Round 4 - Behavioral and Cultural Fit
What to Expect
Conversation focused on your work style, collaboration, problem-solving approach, learning mindset, and alignment with company values. Expect questions about past experiences, how you handle challenges, cross-functional collaboration with development teams, conflict resolution, and your approach to continuous learning. This round also allows you to learn about the team, culture, and what success looks like.
Tips & Advice
Prepare specific examples using the STAR method (Situation, Task, Action, Result). Focus on stories showing: (1) collaboration with development teams, (2) improving processes or systems, (3) handling production incidents calmly, (4) learning new technologies, (5) mentoring or helping teammates. For mid-level, emphasize ownership, initiative, and your ability to own projects end-to-end. Be genuine about your strengths and areas for growth. Ask meaningful questions about team structure, current challenges, growth opportunities, and how success is measured. Show enthusiasm for the role and company without overselling.
Focus Topics
Learning and Continuous Improvement
Approach to learning new technologies, staying current with industry trends, and improving processes. Examples of skills acquired.
Practice Interview
Study Questions
Handling Pressure and Production Incidents
How you remain calm during critical incidents, prioritize actions, communicate effectively, and learn from mistakes.
Practice Interview
Study Questions
Problem-Solving and Initiative
Taking ownership of problems, proposing solutions, driving improvements, and showing initiative beyond assigned tasks.
Practice Interview
Study Questions
Collaboration and Cross-Functional Teamwork
Working effectively with development teams, operations teams, and other departments. Examples of successful collaborations, handling disagreements, and supporting others.
Practice Interview
Study Questions
Frequently Asked DevOps Engineer Interview Questions
As a security architect, you don't own another team's backlog, but you need your threat-modeling findings built into their design before they start coding. How do you get that prioritized without direct authority over their roadmap?
Sample Answer
Direct answer
As a security architect you rarely have line authority over another team's backlog, so you get findings prioritized by making them cheap to accept and costly to ignore: translate the finding into the other team's own vocabulary (a defect, a customer risk, a compliance control they must attest to) and attach it to a decision they are already about to make, rather than asking them to open a brand-new work item. You lead with a specific, demonstrated risk instead of a policy citation, offer a menu of remediation options at different costs, and use an existing recurring forum, like a design review or architecture council, so the tradeoff is made visible to the team's own stakeholders, not just to you.
Structured elaboration
- Translate, don't mandate: reframe the threat-modeling finding in terms the team already tracks (a customer-facing incident scenario, a compliance control, a defect class QA can reproduce) instead of a generic "security best practice."
- Time it to their planning cycle: bring a written finding before backlog grooming or sprint planning, not after code is merged, so accepting it is a normal prioritization decision instead of a rework request.
- Offer options, not a mandate: propose two or three remediation paths (a quick mitigating control now, a full fix next sprint, an explicit accepted-risk sign-off) so the team's own product owner makes an informed tradeoff instead of feeling overridden.
- Borrow a forum, don't invent one: attach the ask to a ritual the team already respects, like their design review, so it reads as peer-level influence rather than a unilateral security gate.
- Make patterns visible upward: when a team consistently deprioritizes findings, escalate the pattern, not the individual finding, to a shared forum with both engineering and security leadership present, so someone with authority over both sides makes the call.
Worked example (illustrative, adapt to your own experience)
A security architect threat-models a new payments feature two weeks before the product team's sprint planning. Instead of filing a ticket titled "add input validation" into the team's backlog and hoping it gets picked up, they write a one-page finding: the specific attack path, the customer-facing scenario it enables, and three remediation options ranked by effort. They bring it to the team's existing design review, present it alongside the team's own product owner, and let the team choose between a lightweight mitigating control shippable in the current sprint or a fuller fix in the next one. The team picks the lightweight option and schedules the fuller fix on their own board, because the tradeoff was made visible and owned by them, not imposed from outside.
Trade-offs and pitfalls
- Too formal (a mandatory sign-off gate) breeds resentment and workarounds; too informal (a message in passing) gets lost in someone else's priority queue.
- Offering remediation options is powerful but risks a team always choosing the cheapest option indefinitely, so track accepted-risk decisions somewhere durable so a pattern of chronic deferral becomes visible over time.
- Borrowing an existing ritual only works if that ritual has real teeth; if the design review itself gets skipped or ignored, attaching your ask to it just inherits its weakness.
What the interviewer probes next
They typically follow up on how you handle a team that keeps saying "next sprint" indefinitely, whether you would ever reach for a hard gate like a release-blocking scan instead of persuasion, and how this influence model holds up when you are supporting a dozen teams at once instead of just one.
You detect a deployment-induced regression affecting only one user cohort. How do you determine the cause without rolling back globally, and what does a targeted remediation look like?
Sample Answer
Direct answer
Investigating a cohort-specific regression without a global rollback starts by narrowing WHAT'S different about that cohort (a request attribute, a data characteristic, a routing path) and testing that specific hypothesis directly, then remediating narrowly, targeting only the affected cohort, rather than reflexively rolling back everyone because you haven't yet localized the cause.
Structured elaboration
- Characterize the cohort precisely: what defines it (a region, a device type, an account tier, a specific request header combination)? The more precisely you can define the boundary, the faster you can form a real hypothesis about the cause.
- Compare cohort-specific request/response data against the unaffected population: pull a sample of failing requests from the affected cohort and a sample of succeeding requests from elsewhere, diffing what's actually different in the request shape, headers, or data values, not just assuming the deploy is uniformly at fault.
- Test the hypothesis directly: if you suspect the cohort's requests hit a different code path (feature-flagged differently, routed to a specific backend), reproduce it in a controlled environment (staging, or a synthetic request matching the cohort's characteristics) rather than experimenting further in production.
- Targeted remediation options, roughly fastest to slowest: a feature flag scoped specifically to that cohort (if the risky code path can be identified and gated); routing that cohort's traffic specifically back to the old version while everyone else stays on the new one, if your infrastructure supports cohort-targeted routing; a narrow, fast-tracked code fix for the specific edge case, if the root cause is well-understood and the fix is small and low-risk.
- Why not a global rollback: if the new version is working correctly for everyone else, a global rollback discards real value (everyone else's improvement) to fix a problem that's isolated and, once understood, fixable narrowly; it's also slower in a real sense, since you still have to eventually re-ship the fix and re-canary the whole population again from scratch.
Worked example
A regression only affecting users with a specific, unusual account configuration (say, an account with more than a certain number of linked payment methods) traces, via diffing failing versus succeeding request payloads, to an off-by-one array-indexing bug that only triggers past that threshold. The targeted remediation is a feature flag specifically gating the new logic OFF for accounts above that threshold while the fix is developed, leaving the vast majority of users (below the threshold) unaffected and still benefiting from the new version.
Trade-offs and pitfalls
Targeted remediation requires your infrastructure to actually SUPPORT cohort-specific control (a flag with cohort targeting, or routing that can distinguish a specific user segment), which not every system has built in; without it, the only real lever might genuinely be a global rollback, which is itself a useful thing to recognize and invest in fixing for next time. The common mistake is spending too long trying to characterize a rare, hard-to-reproduce cohort-specific bug while it continues affecting real users, when a temporary, narrow mitigation (even an imperfect one, like disabling the specific feature for that cohort entirely) could have limited the damage while the real fix is developed.
Describe secure ways to manage secrets (API keys, database credentials, tokens) used by CI/CD pipelines and ephemeral test environments. Compare approaches like storing environment variables in CI systems, using encrypted files checked into repos, dedicated secrets managers (HashiCorp Vault, AWS/GCP Secrets Manager), and CI-native secret stores. Address rotation, least-privilege access for runners, and how to inject secrets into ephemeral PR environments safely.
Sample Answer
Secrets needed by tests specifically (an API key for a sandbox third-party service, a database credential for an integration test suite) have a distinct wrinkle compared to production secrets: they're often needed across many short-lived, ephemeral, parallel test environments, and the temptation to just check them into a test-config file is strong because 'it's just a test credential'.
Comparing the approaches
Static secrets in the CI provider's own store (a GitHub Actions or GitLab CI secret variable): simplest to set up, but the credential is long-lived and shared across every test run until someone manually rotates it, and access control is only as granular as the CI provider's own permission model for that secret.
Encrypted files checked into the repo: avoids the CI provider dependency but pushes the key-management problem onto the repository itself (where's the decryption key stored, and who can access it), and still leaves a long-lived credential sitting in the repository's history even in encrypted form.
Dedicated secrets managers (Vault, cloud Secrets Manager): the strongest option, since it enables short-lived, dynamically-issued test credentials scoped to exactly the test run that needs them, and every issuance is centrally logged; the added complexity is a real dependency for every ephemeral test environment to authenticate against.
CI-native secret stores: functionally similar to the static-secrets case above, differing mainly in which system holds the value; the same long-lived-credential caveat applies.
Recommendation, and why
Dynamic, short-lived secrets via a dedicated secrets manager is the right target for anything beyond a small team, specifically because test credentials often grant access to a shared sandbox or staging environment that, if leaked, could be abused far beyond just 'a broken test'; the cost is worth it once test-environment access represents real risk, not just inconvenience.
Least-privilege and auditability for ephemeral PR environments
Each ephemeral test environment (spun up per pull request, then torn down) should authenticate with its own scoped, short-lived identity, tied to that specific PR or build, so a test credential issued for one PR's environment cannot be reused once that environment is destroyed; auditability then means every credential issuance is logged against a specific PR and build ID, so an unusual usage pattern (a credential used from an environment it wasn't issued to) is immediately detectable.
Rotation for test secrets specifically
For a dedicated secrets manager, rotation is largely automatic: since credentials are issued dynamically per test run, there is nothing long-lived to rotate on a schedule at all, each run simply gets a fresh, short-lived credential. For the static-secret approaches (CI-provider store, encrypted files, CI-native store), rotation has to be a deliberate, recurring process: rotate on a fixed cadence regardless of whether a leak is suspected, and rotate immediately, out of cadence, the moment a leak is suspected. Either way, the safe sequencing is the same overlap-then-invalidate pattern used for production credential rotation generally: generate the replacement credential first, update every consumer (the CI provider's secret store, the encrypted file, or the CI-native store) to use it, confirm at least one real pipeline run succeeds against the new credential, and only then revoke or invalidate the old one, keeping both valid for a brief overlap window instead of cutting over instantly. Revoking the old credential before the new one is confirmed working risks an outage mid-rotation: every pipeline run failing to authenticate until someone notices and rolls back. Doing it in the opposite order, generate first, verify, then revoke, means a rotation gone wrong just leaves the old credential live a little longer, not the pipeline broken.
Avoiding accidental leakage
Test output is a genuine, often-overlooked leak vector: test frameworks frequently print request/response bodies or environment dumps on failure for debugging purposes, and a test credential embedded in that debug output leaks the same way a production secret would in a build log; masking test secrets in CI log output and being deliberate about what a test failure handler actually prints closes this specific gap.
Trade-offs
The dynamic-secrets approach adds real setup cost (every ephemeral test environment needs to authenticate to the secrets manager, which is more moving parts than just reading an environment variable); for a team running a handful of tests against genuinely low-risk sandbox services, the static-secret approach may be a proportionate choice, but that judgment should be revisited as soon as the test credentials in question could reach anything more sensitive than a disposable sandbox.
Given highly variable per-request CPU and memory usage and unpredictable bursts, propose a hybrid provisioning strategy combining reserved baseline capacity, autoscaling, and burstable instances to minimize cost while meeting SLOs. Quantify the assumptions and demonstrate a break-even analysis for reserving capacity.
Sample Answer
Framing
The goal is to size a reserved baseline just large enough to cover the steady, predictable portion
of demand cheaply, and let autoscaling, backed by on-demand or burstable capacity, absorb the
variable, unpredictable portion, rather than either over-reserving for peak, wasteful, or running
everything on-demand, expensive at steady baseline load.
Strategy design
Reserved baseline: sized to the trough-to-median range of the demand curve, not the peak, since
this capacity runs continuously regardless of instantaneous load, it should only be as big as the
load that's always there. Autoscaling on top: standard on-demand instances that scale in and out to
track demand above the reserved baseline, handling the swings. Burstable instances, a family that
banks CPU credit during low usage and can burst above baseline briefly, are a good fit specifically
for workloads with variable per-request CPU where average utilization is low but occasional short
bursts need more, cheaper than full on-demand for that profile, but risky if bursts are sustained
long enough to exhaust the credit balance, in which case performance drops sharply, throttled to
baseline, right when more capacity is needed, not gracefully.
Quantifying the assumptions
Illustrative hourly demand for one representative day, in a generic capacity unit, on-demand rate
of $0.10 per unit-hour, reserved rate of $0.04 per unit-hour, a 60% discount, a realistic ballpark
for a one-year commitment, though actual discounts vary by provider and commitment length and
should be confirmed against real pricing before committing.
Illustrative 24-hour demand curve, in that same unit, hour 0 (midnight) through hour 23: 9, 7, 6, 5, 5, 6, 9, 14, 19, 24, 27, 29, 30, 30, 31, 32, 33, 34, 36, 35, 34, 30, 25, 18 (sums to 528 unit-hours across the day; peak of 36 falls at hour 18).
Worked example: break-even analysis
cost(b)=b⋅rreserved⋅24+h=1∑24max(0, dh−b)⋅ron_demandPlain-English: for a given reserved baseline size b, you pay the reserved rate for b units across
all 24 hours, plus the on-demand rate for whatever demand exceeds b in each individual hour.
Every extra unit of reserved baseline costs a fixed $0.04 times 24, or $0.96 a day, whether it's
used or not, and saves $0.10 for every hour that unit would otherwise have needed on-demand
capacity. So the break-even rule is: keep raising the reserved baseline as long as the number of
hours per day where demand exceeds the current baseline is greater than $0.96 divided by $0.10,
which is 9.6 hours. Once fewer than roughly 9.6 hours a day would actually use that extra unit of
reserved capacity, it's cheaper to pay on-demand for those occasional hours instead.
Worked example, computed for real against an illustrative 24-hour demand curve: sweeping the
baseline from 0 up to the day's peak of 36 units, total cost is minimized at a baseline of 30
units, giving a total daily cost of $31.30 versus $52.80 for a fully on-demand strategy, a 40.7%
reduction. At that baseline, exactly 7 of the 24 hours have demand above 30, just under the 9.6-hour
break-even threshold, consistent with the marginal rule above. Pushing the baseline higher starts
reserving capacity that fewer than 9.6 hours a day would actually need, and total cost starts
rising again, a baseline set at the full peak of 36 costs $34.56 a day, more than the $31.30
optimum.
Trade-offs
This assumes the demand curve is stable day to day, if it shifts, seasonal growth, a new feature
launch, the optimal baseline shifts too, re-run this analysis periodically, monthly for example,
against updated demand data, not once and forget it. Burstable instances add a third lever between
reserved and full on-demand, appropriate specifically for spiky-but-low-average workloads, but the
credit-exhaustion failure mode means they're a poor fit for anything where a sustained burst is
plausible, validate actual burst duration against the instance family's credit-accumulation and
burst-duration limits before relying on them for anything latency-sensitive.
What are EC2 placement groups (cluster, spread, partition) and when would you use each? What are the trade-offs and limitations?
Sample Answer
Direct answer
Placement groups let you influence how EC2 places instances on underlying hardware, trading network performance against fault isolation. Cluster packs instances close together on the same high-bandwidth network segment inside one Availability Zone (AZ) for the lowest latency and highest throughput. Spread puts a small number of instances on physically distinct hardware to minimize correlated failure. Partition groups instances into logical partitions, each on its own set of racks, so a single rack failure only takes out one partition, which is how large distributed systems like Kafka or Cassandra stay available.
Structured elaboration
| Strategy | What it does | AZ scope | Instance limit | Best for |
|---|---|---|---|---|
| Cluster | Packs instances on one high-bisection-bandwidth network segment | Single AZ only (can span peered VPCs in the same Region, not AZs) | No fixed cap, but mixing instance types or adding instances later increases the chance of an insufficient-capacity error | Tightly-coupled HPC (high performance computing) / MPI (Message Passing Interface) workloads needing low latency and high throughput between nodes |
| Spread | Places each instance on physically distinct hardware (its own rack) | A single rack-level spread group can span multiple AZs in the same Region | Max 7 running instances per AZ per group | A small number of critical instances you want maximally isolated from each other (e.g. a handful of database primaries/replicas) |
| Partition | Divides the group into logical partitions, each on its own racks | A single partition group can span multiple AZs in the same Region | Max 7 partitions per AZ; instance count per partition limited only by account limits (Dedicated Instances cap at 2 partitions) | Large distributed, rack-aware systems (HDFS, Cassandra, Kafka) |
A few rules apply across all three: an instance can belong to only one placement group at a time, placement groups cannot be merged, you can't launch a Dedicated Host into a placement group, and you can't launch a Spot Instance configured to stop or hibernate on interruption into one.
Worked example
For a tightly coupled HPC job needing very low node-to-node latency (say an 8-node MPI simulation), I'd launch a cluster placement group with all 8 instances of the same instance type in a single launch request, in one AZ, using an enhanced-networking-capable instance family; within a cluster placement group, enhanced-networking instances get up to 10 Gbps for single-flow traffic versus 5 Gbps for instances outside one. For a Kafka cluster that needs rack-fault isolation across many more brokers than a spread group's 7-per-AZ cap allows, I'd use a partition placement group instead, with up to 7 partitions in the AZ; Kafka's own rack-awareness can consume the partition topology AWS exposes so a rack failure only affects the brokers in one partition.
Trade-offs & pitfalls
- Cluster placement groups trade fault tolerance for latency: everything sits in one AZ, so an AZ-level event takes the whole group down together. Pair it with cross-AZ checkpointing or keep the control plane elsewhere.
- Adding instances to a cluster group later, or mixing instance types within it, raises the odds of an insufficient-capacity error, since AWS is trying to keep everything on the same network segment.
- You cannot convert a group's strategy after creation and cannot merge two groups; if you outgrow the 7-per-AZ spread limit, the fix is multiple spread groups, which gives no guarantee of spread between the groups themselves.
- You can move a stopped instance into, out of, or between placement groups, but only while it's stopped, so this isn't a live-migration tool.
- Capacity Reservations behave differently per strategy: they don't reserve capacity in spread or partition groups at all, only in cluster groups (via a group-scoped On-Demand Capacity Reservation), which catches people who assume reservations work the same way everywhere.
In a multi-container development environment, one service can't connect to another using its service name on a Docker network, even though both containers are running (for example an API can't reach its database in a docker-compose stack). Walk through how you'd debug DNS resolution and network attachment to find the root cause, and explain the difference between the default bridge network and a user-defined network for this kind of service discovery.
Sample Answer
Direct answer
If an API cannot reach its database by service name inside the same Compose stack, the near-universal cause is that the two containers are not actually on a network with Docker's embedded DNS server enabled, most commonly because they ended up on the default bridge network instead of a user-defined one: Docker's default bridge network intentionally does not do name-based service discovery, while any user-defined network (which is what Compose creates for you automatically, unless you have overridden it) runs an embedded DNS resolver that resolves container and service names to their current IP addresses automatically.
Structured elaboration
Debugging DNS resolution and network attachment, in order
- Confirm both containers are actually attached to the same network:
docker network inspect <network-name>and look for both container names in itsContainerssection. A Compose stack with a mis-scopednetworks:block, or a container started withdocker runoutside the Compose-managed network entirely, is invisible to the other side no matter what DNS is doing. - From inside the calling container, attempt resolution directly:
docker exec <api-container> getent hosts <db-service-name>(ornslookup, if available in that image). A clean IP address back means DNS resolution itself works and the problem is elsewhere (the port, the database not actually listening yet, a firewall rule); a resolution failure means the DNS layer is the actual problem. - Check which network the containers landed on:
docker network lsanddocker inspect <container> --format='{{json .NetworkSettings.Networks}}'. If it shows the network named literallybridge, that is Docker's default network, not a Compose-created user-defined one, and name resolution by container name is not expected to work there at all. - If both containers are confirmed on the same user-defined network and resolution still fails, check for a typo or mismatch between the service name used in code and the actual Compose service name (Compose registers the service name, and, for the container's own hostname, the container name, as DNS entries on the network it creates; a hardcoded old container name or a renamed service is a common, easy-to-miss mismatch).
Default bridge network versus a user-defined network, for service discovery specifically
- The default
bridgenetwork (what you get from plaindocker runwith no--networkflag) does not run an embedded DNS server for container name resolution at all; containers on it can only reach each other by IP address (historically, the deprecated--linkflag patched hostnames into/etc/hostsas a workaround, which is not how modern setups should discover services). - Any user-defined network, including the ones Compose creates automatically for a project, runs Docker's embedded DNS resolver (at the well-known address
127.0.0.11inside each container on that network), which resolves both container names and Compose service names to the correct current IP address, and updates automatically if a container restarts and gets a new IP. - This is precisely why Compose "just works" for service-to-service calls by name out of the box: it always creates a user-defined network for a project unless you explicitly override that, so the embedded DNS behavior is present by default, not something you have to opt into separately.
Worked example
Two containers, api and db, are both started, and api's logs show connection refused or "unknown host" errors when it tries to reach db by name. Running docker exec api getent hosts db returns nothing and exits non-zero, confirming DNS resolution itself is failing, not just the connection. docker network ls shows both containers only ever attached to the default bridge network (the Compose file had accidentally used network_mode: bridge on the api service, which pins it to Docker's default bridge network instead of letting Compose create and use its own project network). Removing that override lets both services join the Compose-created user-defined network instead; getent hosts db then resolves to the container's current IP address, and the application connects successfully on the very next attempt, with no code change at all.
Trade-offs & pitfalls
- It is easy to "fix" this the wrong way by hardcoding the database's current IP address once resolution is confirmed working; container IPs on a bridge network are not stable across restarts, so this reintroduces the same failure the first time the database container restarts and is assigned a different address.
- A resolution failure and a "resolves fine but the connection is refused anyway" failure look similar from the application's point of view (both surface as some flavor of connection error) but have entirely different causes; always separate the DNS step (
getent hosts) from the actual connection step before concluding which layer is broken. - Legacy guidance mentioning the
--linkflag as a way to make one container resolve another by name is outdated; a user-defined network's embedded DNS server does that automatically for every container on the network, with no per-pair linking needed.
Design an experiment to compare two approaches for teaching Kubernetes to engineers: instructor-led workshops versus project-based learning. Define your hypotheses, experimental design (cohorts, duration), metrics (primary and secondary), statistical considerations, and the actions you'd take based on different outcomes.
Sample Answer
Hypotheses
- H0 (null): No difference in Kubernetes proficiency between instructor-led workshops (ILW) and project-based learning (PBL).
- H1 (alternative): PBL yields higher long-term applied proficiency; ILW yields faster short-term knowledge gains.
Experimental design
- Randomized controlled trial with consenting DevOps engineers (n = 80, 40 per arm).
- Stratify by experience (junior/senior) and prior k8s exposure.
- Duration: 6 weeks training + 4 weeks follow-up.
- ILW: 4 half-day instructor sessions + Q&A office hours.
- PBL: 4-week guided team project (build CI/CD on k8s) with mentor checkpoints.
- Pre-test, immediate post-test, and 4-week practical assessment.
Metrics
- Primary: Practical task score — end-to-end deployment & recovery on a cluster (scored rubric).
- Secondary:
- Time-to-complete tasks (efficiency).
- Confidence self-report (Likert).
- Retention: re-run practical at 4 weeks.
- Transfer: ability to debug a seeded failure (mean bugs fixed).
- Behavioral: number of IaC / k8s manifests submitted to repo in 4 weeks.
Statistical considerations
- Power: target 80% power to detect Cohen’s d = 0.6; n ≈ 34 per arm → use 40 to allow attrition.
- Tests: ANCOVA on post-test controlling for pre-test score; nonparametric tests if distributions skewed.
- Multiple comparisons: control FDR with Benjamini-Hochberg for secondary metrics.
- Handle missing data with multiple imputation; report intention-to-treat and per-protocol.
Actions based on outcomes
- PBL significantly better: adopt PBL as primary method; incorporate mini-projects into onboarding; scale with templated project blueprints and mentors.
- ILW better short-term but worse retention: use ILW for crash courses + follow-up projects for retention.
- No significant difference: choose based on cost, scaling, learner preference; consider hybrid (short ILW + capstone project).
- Check subgroup effects (juniors vs seniors) and adjust pathway by experience.
Rationale: emphasis on applied metrics (deployments, debugging) aligns with DevOps objectives — measurable impact on day-to-day operational capability.
Before a Terraform or CloudFormation change ever reaches apply, what automated checks would you want running in the pipeline, and at what stage would each one run? Talk through what kind of mistake each check is actually meant to catch.
Sample Answer
Direct answer
Run checks in a layered pipeline ordered from cheapest to most expensive: format and lint first (seconds, no cloud access), then static policy/security scanning against the plan or template (still no cloud access), then anything that actually stands up real resources (integration tests, plan review with real credentials), and finally post-deploy verification against the live environment. Each layer is designed to catch a different class of mistake, and putting the cheap checks first means a typo never has to wait for an expensive real-resource test to fail.
The two fundamentally different kinds of check
Before mapping tools to stages, it's worth naming the split explicitly: some checks are fast, syntax- or policy-level, and need no cloud access at all; others genuinely stand up real (usually short-lived) infrastructure to prove it actually works. Confusing the two, or skipping straight to the expensive kind, is the most common mistake in a pipeline like this.
| Stage | Check | What mistake it catches | Provisions real resources? |
|---|---|---|---|
| Pre-commit / local | terraform fmt, cfn-lint, ansible-lint | Style drift, malformed syntax, obviously invalid template schema | No |
| CI: fast lint | terraform fmt -check, tflint | Provider-specific misuse, deprecated arguments, obvious logic errors | No |
| CI: static policy/security | checkov, cfn_nag, OPA/Sentinel against the plan JSON | Security misconfiguration (public S3 bucket, overly broad IAM), missing required tags, policy violations | No, reads the plan or template, never calls the cloud provider to create anything |
| CI: unit tests | Module output assertions, template rendering tests | Wrong variable defaults, broken output wiring, logic bugs in the module itself | No |
| CI: plan review and gating | terraform plan posted to the PR, human or automated diff review | Unexpected destroy/replace actions before they ever reach a real environment | No, it's a dry run |
| CI: integration tests | Terratest, Molecule against a container or real cloud instance | Whether the resource actually gets created correctly, whether runtime configuration is right, whether the API actually accepts what you declared | Yes, real (usually ephemeral) resources |
| Post-deploy | Smoke tests, health checks against the deployed environment | Whether the deployed service is actually reachable and healthy in this specific environment | Yes, against the live environment |
Worked example: a pipeline definition
jobs:
fmt-and-lint:
steps:
- run: terraform fmt -check -recursive
- run: tflint
policy-scan:
needs: fmt-and-lint
steps:
- run: checkov -d . --framework terraform
plan:
needs: policy-scan
steps:
- run: terraform plan -out=plan.tfplan
- run: terraform show -json plan.tfplan > plan.json
# a script here fails the job if plan.json contains an unreviewed delete/replace
integration-test:
needs: plan
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
steps:
- run: go test ./test/... -run TestModule -timeout 30m
# this job is the one that actually provisions and tears down real infra
apply:
needs: integration-test
environment: production # requires manual approval
steps:
- run: terraform apply plan.tfplan
The ordering matters: fmt-and-lint and policy-scan run on every commit because they're free and fast; integration-test, which provisions real infrastructure, is scoped to run only on merges to main, not on every push to a feature branch, to bound cost and time.
Trade-offs and pitfalls
- Skipping straight to integration tests (or relying on them to catch what a linter would have caught in seconds) is slow and expensive for no extra safety, put the cheap checks first and let them fail fast.
- Running the real-resource layer on every commit, rather than on merge or nightly, is the single most common way teams accidentally burn cloud spend on a testing pipeline.
- Static policy scanning only catches what it has rules for; it gives false confidence if the rule set isn't kept current with new resource types the team starts using.
- A plan-review gate that only a human reads doesn't scale, encode the "never allow an unreviewed delete on a protected resource" rule as an automated check on the plan JSON, not just a habit.
You are asked to build a capacity trend for disk usage across 200 servers so management can plan storage purchases. What data would you collect, how would you calculate the trend, and what would trigger a purchase order?
Sample Answer
Direct answer
Collect a time series of used capacity per volume, for example daily df samples or whatever your monitoring system already stores, fit a trend line to it, and project forward to when it crosses your action threshold. You need enough history to smooth out noise (weekly patterns, one-off cleanups) but recent enough to reflect current growth, and the output should be a date, not just a percentage, because a date is what triggers a purchase order.
Structured elaboration
Collect at least daily df-style samples of used GB per volume, not just percent, since percent alone hides how many GB a jump actually represents on a large disk. Fit a trend line (ordinary least squares is enough for this) to get a growth rate in GB per day, then project forward to find the day the fit crosses your capacity ceiling.
What triggers the purchase order: pick a lead time longer than your procurement cycle. If buying and provisioning new storage takes 3 weeks, trigger the order when the trend crosses "90% full in 5 weeks," not when the disk is already at 90%.
Worked example
Executed example using 14 days of sampled disk usage on a 500 GB volume, fitting a least-squares line and projecting forward:
measured samples (GB used): [350.5, 350.4, 355.1, 357.4, 355.1, 358.0, 362.2,
363.5, 366.0, 366.7, 371.3, 373.2, 374.1, 377.7]
fitted growth rate: 2.105 GB/day (true underlying rate in this simulation was 2.0 GB/day)
fitted intercept (day 0 est.): 349.3 GB
90% full (450 GB) projected at day 47.9 from day 0
days remaining from today (day 13) to the 90% mark: 34.9 days
The fit (ordinary least squares) recovers the true 2.0 GB/day growth rate closely, 2.105 estimated, even with daily measurement noise. That is the point: a single day's jump or dip should not be read as a trend change, the line across many days is what you act on.
Trade-offs and pitfalls
- A straight-line fit assumes linear growth; a service about to onboard a large new customer or double its retention window will blow past a linear projection, so pair the trend with awareness of planned changes, not just historical data.
- Too little history (a few days) makes the slope noisy and unreliable; too much history (a year) can hide a recent acceleration by averaging it away. Recompute on a rolling window, for example the trailing 30 days, and re-evaluate weekly.
- An aggregate trend across 200 servers hides the one server about to fill up next week; you want both a fleet-wide summary for planning and a per-host projection for the "who pages tonight" question.
What the interviewer probes next
They will usually push on whether the fit is really linear or whether a single unusual day, a big import, a bulk cleanup, is quietly steering it, and on how you would roll a whole fleet of these projections into something a non-technical stakeholder can act on without reading a chart.
You have budget for a 20% infrastructure cost increase and it needs to measurably improve availability. Walk through how you'd decide where that money buys the most reliability, and how you'd justify the spend to someone who isn't an engineer.
Sample Answer
Direct answer
Treat the budget as a portfolio: rank candidate investments by how much measured downtime each one removes per dollar, fund the highest-ROI items first until the budget is spent, then translate the resulting hours of downtime avoided into a revenue or cost-avoidance figure a non-engineer can act on. The common mistake is funding the biggest-sounding project instead of the one with the best marginal return.
Structured elaboration
Step 1: build a downtime-hours ledger from real incidents
Pull the last 12 months of incidents and categorize each by root cause (correlated infrastructure failure, human/process delay, insufficient capacity, and so on). Sum the downtime attributable to each category. This ledger is the input to everything that follows, without it, prioritization is intuition, not evidence.
Step 2: price each candidate fix against the hours it removes
For every candidate investment, estimate its annual cost and the hours of downtime it removes, ideally by fully eliminating a category rather than vaguely helping it. Compute dollars per hour-year removed and rank from cheapest to most expensive.
Step 3: fund down the ranked list until the budget is spent
This is a knapsack-style trade-off: sometimes a slightly worse-ROI item is still funded because a better one does not fit the remaining budget. Treat that as an explicit, stated judgment call rather than a strict greedy pick.
Step 4: translate hours removed into a business number
Multiply hours removed by an estimated revenue-or-cost-per-hour-of-downtime figure, sourced from finance or product and stated as an assumption if not yet confirmed, to produce a dollar risk-reduction figure comparable against the spend.
Worked example
Assume the incident ledger for the last 12 months shows two categories:
| Category | Incidents | Avg duration | Annual downtime |
|---|---|---|---|
| Correlated single-AZ data-layer failures | 3 | 90 min | 4.5 hours |
| Manual failover delay (paging plus investigation before a human triggers failover) | 12 | 25 min | 5.0 hours |
Two candidate investments:
- A: multi-AZ redundancy for the data layer. Removes the correlated-failure category entirely (4.5 hours/year), costs $60k/year.
- B: automated failover tooling. Cuts manual-delay MTTR from 25 minutes to 5 minutes across the same 12 incidents/year, removing (25−5)×12/60=4 hours/year, costs $25k/year.
Dollars per hour-year removed:
A=4.560,000=13,333 per hour-year B=425,000=6,250 per hour-yearB has the better marginal ROI, fund it first. With a $500k baseline infra spend, a 20% increase is a $100k budget: B ($25k) plus A ($60k) totals $85k, funding both, with $15k left for monitoring and alerting improvements that support both fixes.
Total downtime removed: 4.5+4=8.5 hours/year, against a baseline of 9.5 hours/year attributable to these two categories (4.5+5), an 89% reduction in that measured downtime.
Business translation: assume, subject to finance confirmation, $50k of revenue at risk per hour of downtime during business hours:
8.5×50,000=425,000 dollars per year of risk removed 85,000425,000=5.0That is the pitch for the non-engineer conversation: "$85k of this budget removes $425k a year of downtime-driven revenue risk, based on our own incident history."
Trade-offs & pitfalls
- The ROI math only covers failure categories with existing data; it says nothing about the tail risk of a novel failure mode not yet seen, do not present it as a complete risk model.
- Averaging MTTR across incidents of very different severity can hide the one incident that actually mattered; check the distribution, not just the mean, before trusting the "hours removed" estimate.
- A dollar-per-hour-of-downtime figure that is not sourced from finance is a guess dressed as a number; label it explicitly as an assumption when presenting it, and get it validated before it becomes the headline of the pitch.
- Diminishing returns are structural: the third and fourth items on the ranked list will have worse ROI than the first two by construction, do not extrapolate the 5x multiple to the whole budget.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths