Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
When would you pull resources out into a reusable module instead of leaving them in the root configuration? Explain what makes a good module boundary, and what you'd call over-modularization.
Sample Answer
Pull resources into a module once they're reused across more than one stack or environment, or once a single-purpose group of resources, a VPC, an EKS cluster, a standard RDS setup, has grown enough internal complexity that inlining it clutters the root config. A good module boundary maps to one concern with a small, stable interface; over-modularization is wrapping something in a module because it feels tidy, not because anything actually reuses it or needs the extra indirection, and it shows up as one-resource modules, deep nesting, or modules stitched together with a pile of boolean flags to fake reuse across cases that were never really the same.
When to extract, when to leave inline
- Extract when the resource group is reused across environments, teams, or projects, or its complexity, many variables, repeated lifecycle rules, would otherwise be copy-pasted.
- Leave inline when resources are small, unique to one stack, or tightly coupled to that stack's own variables and lifecycle; wrapping them in a module just adds a layer to look through for no reuse benefit.
Layered module hierarchy
A common shape at organization scale: networking, identity, shared-services, platform, and application layers, each owning a narrow slice and exposing only what the next layer needs, networking exposes VPC and subnet IDs, identity exposes role ARNs, shared-services exposes things like a shared logging bucket, platform composes those into a cluster, application deploys onto the platform. The dependency direction only ever points one way, up: a lower layer like networking must never read or depend on a higher layer's outputs like application, or you get a circular dependency that makes independent deploys impossible. Enforce that with separate state per layer and remote-state or published-lookup reads flowing strictly upward, never the other direction.
The module contract, and encapsulation leaks
A module's real contract is its inputs and outputs, nothing else. The failure mode to design against: consumers start depending on a specific resource name or ID inside the module that it never promised to keep stable, referencing an internal resource address directly instead of a proper output, so a refactor inside the module that changes nothing about its behavior still breaks every consumer. Prevent it by never exposing internal resource addresses, only named outputs, and by treating "adding or renaming an output" with the same breaking-change discipline as changing an input.
Choosing a DRY mechanism
| Mechanism | Best for | Downside |
|---|---|---|
| Shared module | Genuine multi-stack reuse of a resource group with real variation between consumers | Overkill for a single repeated value; adds a version to manage |
| Locals / variable maps | Repeated values within one stack, or a small fixed set of environments keyed by name | Doesn't scale across repos; every stack still owns its own copy of the map |
| CI-side templating | Config that's identical except for a few substituted values across many near-identical stacks | Hides the actual config from a plain plan read; harder to review a real diff |
If the same config keeps getting copy-pasted across stacks, start with the cheapest fix, a locals map, before reaching for a full shared module; promote to a module once real behavioral variation, not just value substitution, shows up between consumers.
Trade-offs & pitfalls
Over-modularization is a common overcorrection once a team learns the reuse lesson: they start extracting one-resource modules, or nesting modules three deep, and debugging becomes tracing variable propagation through layers that never needed to exist. A practical rule: start inline, extract when reuse or complexity actually hits, and keep every module's interface small enough that its README fits on one screen.
An internal module is already used by several teams, and you need to add a new capability without breaking existing consumers. How would you evolve the module, version it, and communicate the change so upgrades stay predictable?
Sample Answer
Approach
I treat a shared Terraform module like an API. First I classify the change: additive and backward-compatible, or breaking. If it is additive, I release a new minor version, keep existing variables and outputs unchanged, and make the new capability opt-in with a default that preserves current behavior. If I must rename or remove something, I publish a new major version and keep the old one available for a transition period.
How I keep upgrades predictable
- Use semantic versioning: patch for fixes, minor for new optional features, major for breaking changes.
- Pin module versions in callers, for example
~> 1.4, so teams only receive compatible updates. - Add tests that run old examples and new examples in CI.
- Publish a changelog with migration notes and deprecation dates.
- Announce the change early, then give teams a canary path in one workspace before broad rollout.
Concrete example
If the module currently creates an S3 bucket and I want to add optional access logging, I would add enable_access_logging = false and a new logging block. Existing consumers get the same bucket as before, while teams that want logging can opt in. After a release or two, I can deprecate any old workaround variables without breaking them immediately.
Result
That approach lets teams upgrade on their schedule, keeps state changes predictable, and makes ownership clear.
You need to roll out an infrastructure-level change, say a new machine image or a load balancer/routing config change, with as close to zero downtime as possible. Walk through a blue-green or canary approach: how traffic gets shifted, what health checks and metrics you'd watch before promoting, and what would make you pull the plug and roll back.
Sample Answer
Direct answer
The mechanics are the same regardless of what is being changed: stand up the new version alongside the old one, shift traffic to it gradually (or all at once for blue-green), watch metrics against a defined threshold before promoting further, and have an automated way to shift traffic back if those metrics degrade. What differs by change type is what "the new version" means and which metric actually tells you it is safe, so the mechanics come first below, then how they change for a few different kinds of infrastructure change.
Structured elaboration
Traffic shifting mechanics
- Canary: route a small percentage of traffic to the new version (for example 5%), hold, then ramp (25%, 50%, 100%) with a wait window at each step.
- Blue-green: fully validate the new environment out of band, then cut traffic over in one step (DNS, load balancer listener swap, or weighted routing swap to 100%). Rollback is just swapping back.
What "the new version" and "the signal to watch" are, by change type
| Change type | What actually changes | Primary signal before promoting | What triggers rollback |
|---|---|---|---|
| AMI (Amazon Machine Image) or launch-config replacement | New ASG (Auto Scaling Group, a group of instances that scales automatically) or launch template referencing the new image, registered to the same target group | 5xx rate, app error logs, latency | Error rate or latency breaches threshold for a sustained window |
| Network load balancer or routing config change | Listener rules, target-group weights, or routing paths on the LB itself | Target-group health-check pass rate, connection error rate, latency delta versus baseline | New connection failures or health-check flapping that were not present before |
| Network or security-group rule change | The rule set applied to one canary target group or subnet before the rest of the fleet | Connection-refused and timeout counts, not 5xx, since a bad rule usually blocks the connection before the app ever sees it | Any new connection failures on the canary group that do not exist on the unchanged group |
| ASG instance-type change | Instance type in the launch template, small percentage of the fleet | CPU and memory headroom, p95/p99 latency measured against the SLO's error budget, not just an absolute number | Latency or saturation consumes more of the SLO error budget than the change is worth |
| DB-class change | New instance class, usually validated via a promoted read replica before the primary is touched | Replication lag, query latency, connection-pool saturation | Replication lag or query latency exceeds a defined threshold, or the pool starts queuing |
Stateless versus stateful
For a stateless service, this is close to mechanical: deregister the old instance from the load balancer, let connections drain, terminate it, and the new instance takes over with no coordination needed between old and new.
For a service with state, for example one backed by a database or holding session or connection state, the "new version" cannot simply replace the old one at will:
- Deregistration has to wait for graceful connection draining, not just for the health check to fail, or you drop in-flight work.
- If the state itself is what is changing (the DB-class case above), the new version has to be replication-caught-up before it sees real traffic, and you need a plan for what happens to writes that land during the cutover window.
- The old and new versions may be reading or writing the same underlying data store, so canary steps need to account for that shared state, not just for the request path.
Worked example
A concrete case: shifting 5% of production traffic on an ALB to a target group fronting instances built from a new AMI, before fully cutting over.
resource "aws_lb_target_group" "green" {
name = "app-green"
port = 443
protocol = "HTTPS"
vpc_id = var.vpc_id
health_check {
path = "/health"
healthy_threshold = 3
unhealthy_threshold = 2
timeout = 2
}
}
resource "aws_lb_listener_rule" "canary" {
listener_arn = aws_lb_listener.app.arn
priority = 100
action {
type = "forward"
forward {
target_group {
arn = aws_lb_target_group.green.arn
weight = 5
}
target_group {
arn = aws_lb_target_group.blue.arn
weight = 95
}
}
}
condition {
path_pattern { values = ["/*"] }
}
}
Watch the green target group's 5xx rate and p95 latency for a fixed window (say 15 minutes). If it stays within the same range as blue, bump the weight to 25 and repeat; if 5xx rate or latency breaches the threshold at any step, a second terraform apply (or a CI job wrapping it) sets green's weight back to 0 and blue keeps serving everything.
flowchart TD
A[Provision green environment via Terraform] --> B[Register in target group, run health checks]
B --> C[Shift 5 percent traffic to green]
C --> D{Metrics within SLO?}
D -- Yes --> E[Ramp to 25 percent, then 50 percent]
E --> F{Still within SLO?}
F -- Yes --> G[Shift 100 percent traffic, decommission blue]
F -- No --> H[Roll back: weight to 0 percent, blue keeps serving]
D -- No --> H
Trade-offs & pitfalls
- The biggest mistake is treating every infra change like an app deploy and watching only 5xx rate. A security-group rule change will not show up as a 5xx, it shows up as connections that never complete; a DB-class change will not show up in the app's error rate until replication lag has already gotten bad. Pick the signal that matches what actually breaks for that change type.
- Canary steps for stateful changes need explicit wait-for-caught-up gates, not just a timer; ramping traffic to a replica that has not finished catching up just moves the failure into user-visible latency or stale reads.
- Automating the rollback path, not just the promotion path, is what actually buys "close to zero downtime." If pulling the plug requires someone to remember the right
terraform applyunder pressure, the safety only exists on paper.
A teammate made a change directly in the cloud console, and now terraform plan on that stack shows changes you didn't expect. Walk through how you'd investigate, safely bring things back in line, and stop it from happening again.
Sample Answer
First confirm what actually changed and why, using terraform plan plus the cloud provider's audit log rather than the plan output alone, then decide per change whether to accept it into code or revert it, and only act through the normal pipeline rather than another console edit. The same triage scales up: if a maintenance window leaves plan showing two hundred resources drifted instead of one, the process doesn't change in kind, just in batching, bucket the diffs by risk and category, handle the safe bulk in one deliberate pass, and pull out anything ambiguous for individual review.
Investigate
- Run
terraform plan -detailed-exitcode, orterraform show -jsonpiped through a JSON tool, to get a structured diff of exactly what attributes changed on which resources. - Cross-reference with the cloud audit log filtered to that resource's ID and the plan's time window, to find who or what made the change and when.
- Check whether it was a one-off console edit, an automated actor like an autoscaler, or, in the messier real-world case, a string of manual emergency fixes made during an incident rather than one clean change; the audit log will show a cluster of API calls rather than a single event.
Bring it back in line, safest first
- If the change was legitimate, a genuine fix that should stick: update the Terraform config to match reality, run plan to confirm it now shows no diff, and merge that through the normal review process. Prefer this over reverting a fix someone made under pressure for a real reason.
- If the change was accidental or shouldn't persist: apply the existing config to revert it, but only after confirming the revert itself won't cause an outage by checking what currently depends on the drifted value.
- If the resource was created or modified in a way Terraform doesn't currently track: use
terraform importto bring it under management before deciding accept or revert, rather than reasoning about a resource the state file doesn't know about.
Scaling this to bulk drift
When a maintenance window or incident leaves on the order of two hundred resources showing drift, resource-by-resource triage doesn't scale. Bucket the diff instead: group by change type, tag-only, config-value, security-relevant, resource-created-outside-Terraform, accept the clearly-safe buckets in one batch, updating code to match or applying to revert depending on the bucket, and pull out anything security-relevant or on a stateful resource for individual review. Apply in waves with health checks between waves rather than one big apply, so a bad assumption about one bucket doesn't take down everything drifted in the window at once.
Prevent it from happening again
- Restrict console and API write access on managed resources so a convenient console fix isn't even available to most people outside the pipeline's own role.
- Add scheduled drift detection in CI so the next unauthorized change is caught within a bounded window instead of surfacing as a surprise on the next unrelated plan.
- Give emergency access a proper break-glass path, a documented, time-boxed, logged way to make an urgent manual fix, so the choice isn't "console-edit quietly" versus "wait for a slow pipeline during an outage"; the emergency path should still end with the change immediately back-ported into Terraform config.
Trade-offs & pitfalls
Reverting on reflex is the common mistake: a manual change often exists because something the pipeline didn't handle needed fixing right now, and blindly applying old config over it can reintroduce the original problem or cause a fresh outage. ignore_changes is tempting for silencing the noise but should be temporary and tracked, not a permanent way to stop looking at a field that keeps drifting.
You're advising a team moving from long-lived servers to baked images. What are the trade-offs against staying with configuration management on mutable servers, especially around patching, hotfixes, debugging, and deployment speed?
Sample Answer
Direct answer
Recommend immutable, baked images (images built once with everything preinstalled, then deployed as-is rather than configured after boot) as the default for anything at fleet scale, because they buy reproducibility and safe rollback; keep configuration management for runtime-only concerns (secrets injection, feature toggles, small dynamic config) and as a deliberate bridge during migration, not as the long-term patching strategy. The trade-offs mostly show up in four places: how patches ship, how urgent hotfixes get applied, how debugging works, and how deployment speed scales with fleet size. That default flips for a stateful service, which needs its own answer below.
Structured elaboration
Comparison
| Concern | Baked images (immutable) | Configuration management (mutable) |
|---|---|---|
| Patching | Rebuild the image with the patch, redeploy via rolling replace; consistent across the whole fleet, easy to roll back by pointing at the previous image | Apply in place with Ansible or Chef across the fleet; fast for a handful of hosts, but risks partial application and drift if a run fails partway |
| Hotfixes | Requires a rebuild-and-replace cycle; slower to first fix, but the fix is guaranteed to be in every future image, not just the running instances | Can be applied immediately to running hosts; faster time to fix, but the fix has to be separately folded back into the source of truth or it silently disappears on the next rebuild |
| Debugging | The exact environment is reproducible from the image ID and version, so "what is actually running" is never in doubt; live debugging still needs logging and tracing rather than SSH-and-poke | Easy to SSH in and iterate directly on a running host, which can be faster in the moment, but every such change is an undocumented, untracked snowflake until someone remembers to encode it |
| Deployment speed at scale | Slower for a single one-off change (build, test, replace), faster in aggregate as fleet size grows, because every instance is guaranteed identical and rollouts are automatable | Faster for a handful of hosts, but velocity degrades as fleet size grows because keeping every host consistent becomes a manual, error-prone effort |
When each makes sense
- Prefer immutable images for stateless, autoscaled fleets: the automation investment (image pipeline, CI, registry) pays for itself once more than a handful of instances are being managed.
- Keep configuration management for genuinely dynamic, runtime concerns: injecting secrets at boot, toggling a feature flag, or configuring something that legitimately needs to change without a full rebuild.
- During a migration specifically, a hybrid approach is normal and reasonable: bake the OS and runtime into the image, and use lightweight, idempotent configuration at boot for the parts that still need to vary per environment. That is not indecision, it is how most teams actually get from mutable to immutable without a big-bang cutover.
The stateful-service case
For a stateful service, for example a database node or a broker holding local queue data on disk, naive image-replacement breaks the pattern above. Terminating an old instance and swapping in a fresh one from a new image throws away whatever state lived only on that instance; the deregister-drain-terminate flow that makes stateless replacement safe does not, by itself, move that state anywhere. There is also a handoff problem a stateless fleet does not have: something has to own transferring or re-syncing the state (or explicitly failing it over) before the old instance goes away, not just draining its connections.
Two realistic patterns:
- Separate state from compute where possible. Move the data itself onto something outside the instance's lifecycle, an external volume that reattaches to the new instance, or a managed data store, so the compute layer becomes stateless again and can be image-replaced exactly like the fleet case above. This is the preferred pattern whenever the workload allows it, because it lets the stateful service inherit the same rollback and testing story as everything else.
- In-place with strict drift controls when separation is not possible. If state genuinely has to live on the box, or the migration to external state has not happened yet, stay with configuration management for that host, but treat convergence checks as a required gate after every run rather than an optional audit, since a failed run on a stateful host is far more costly to miss than on a stateless one.
The same three axes from the comparison above still drive the decision, but they cut differently here:
- Rollback complexity: rolling back a baked image is instant for stateless compute, but rolling back a stateful host means rolling back its data too, not just its binaries; an image-based rollback plan for a stateful service has to include a data restore or failover step, not just a launch-template revert.
- Configuration drift: drift is more dangerous to tolerate on a stateful host, because a missed convergence run there does not just leave a config mismatch, it can leave that host holding data no other instance has, which makes it unsafe to simply terminate and replace.
- Testing requirements: testing a stateful change means testing against a realistic copy of the data (a snapshot or a replica), not just a fresh boot; a stateless image can be smoke-tested cold, a stateful one usually cannot.
Worked example
A team migrating a fleet of 40 application servers from Chef-managed mutable hosts to baked AMIs, mid-migration: 15 servers have already moved to immutable images, 25 are still Chef-managed.
- An OS-level CVE needs patching across all 40. For the 15 immutable servers: rebuild the AMI, run it through the standard test and canary gates, and replace those instances via a rolling instance refresh, with no manual steps on any host.
- For the 25 still-mutable servers: run the existing Chef cookbook update across them directly, which is faster to start for that subset but needs a fleet-wide Chef run to confirm every one of the 25 actually converged, since a failed run on any single host leaves it unpatched and undetected until the next audit.
- This is exactly the operational cost mutable configuration carries at scale: the same patch requires two different verification strategies depending on which half of the fleet a given host is in, and the mutable half needs an explicit convergence check the immutable half gets for free from "the image ID changed."
Trade-offs & pitfalls
- The most common mistake is presenting this as all or nothing. A hybrid, migration-aware answer is the stronger answer; a purist "always immutable" answer that ignores legitimate runtime-config needs is not.
- Debugging discipline has to change with immutability: if engineers keep SSHing into production and hand-patching "just this once," the fleet quietly becomes as inconsistent as the mutable model it was supposed to replace, just with extra image-pipeline overhead on top.
- Do not understate the upfront cost of the image pipeline (CI, registry, automated tests, rolling-replace tooling); it is real engineering investment that has to be justified against the ongoing toil of drift remediation and manual patch campaigns it replaces.
- For a stateful service, do not default to "just bake it anyway": image-replacement that ignores where the state lives is not a safer version of the stateless pattern, it is a data-loss risk wearing the same rollout choreography.
Unlock Full Question Bank
Get access to all Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.