Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
A repo has one large Terraform entrypoint that manages several environments and shared pieces of infrastructure. The team is having merge conflicts and accidental cross-environment changes. How would you reorganize the code so engineers can work in parallel with less risk?
Sample Answer
I would separate shared logic from environment-specific entrypoints. Right now one large root module is forcing unrelated changes into the same file set, which creates merge conflicts and increases blast radius.
New structure
modules/networkmodules/appmodules/databaseenvs/devenvs/stagingenvs/prod
Each environment folder has its own backend and only calls the modules it needs. Shared values like tags or account IDs can live in small local files or a common variable file.
Why this helps
- Two engineers can edit dev and prod in parallel with less conflict
- A change to app compute does not force a network review unless inputs change
- Cross-environment mistakes are harder because each root has its own state
Example
Dev can use t3.small and prod can use m6i.large, but both call the same modules/app. That gives consistency without a monolith. I would also keep module outputs explicit so one environment cannot accidentally consume another environment’s state.
Terraform doesn't give you an automatic transactional rollback if an apply goes wrong. What patterns do you rely on instead to safely walk back an infrastructure change, and when is automated rollback the right call versus needing a human to reconcile things by hand?
Sample Answer
Direct answer
Terraform has no built-in transactional rollback because it applies resources one at a time in dependency order, not as a single atomic unit, so "rolling back" is really a set of separate patterns you choose between based on blast radius: reapplying a known-good, version-pinned configuration, restoring a snapshotted state file, or targeted, resource-level remediation. Automated rollback is the right call for reversible, idempotent changes with a clear health signal (a stateless service behind an ASG, a container image swap); a human needs to reconcile by hand whenever the change is destructive, touches stateful data, or the exact post-failure state is ambiguous.
Patterns for walking back a change
Versioned modules and state backups
- Pin every module call to a semantic version tag, so "roll back the infrastructure" can mean "redeploy the previous module version" rather than reconstructing a diff by hand.
- Keep state backend versioning on (S3/GCS object versioning) and snapshot state before every apply, tagged with the pipeline run ID and commit.
Partial rollback versus full rollback
- A full rollback (restoring the entire state file to the last-good snapshot) is the blunt instrument: fast, but it can silently undo any other legitimate change another engineer applied in the meantime, since it doesn't know which parts of the diff are the problem.
- A partial rollback (fixing or re-targeting just the resource(s) that actually broke, using
terraform apply -targetsparingly, or hand-correcting one resource's config) is narrower and safer when you know exactly what failed, but requires that you've actually diagnosed the failure first, guessing wrong under a full state restore is less catastrophic than guessing wrong under a targeted one.
Reapplying configuration management state after an infra rollback
If a configuration management tool (Ansible, Chef, Puppet) runs on top of Terraform-provisioned instances, rolling back the Terraform layer alone (say, reverting an instance's AMI or instance type) does not automatically put that instance's OS-level configuration back in sync. The CM tool's last run was against the instance in its post-change form; after an infra rollback you need to rerun the CM playbook or run-list against the rolled-back instance so both layers agree again, otherwise you end up with infra that matches the old Terraform state but application configuration that still reflects the change you just reverted.
Keeping a multi-resource logical change atomic
A single logical change that spans several resources, for example provisioning a new database instance together with the IAM policies that grant an application access to it, needs to stay consistent even if one part fails. Two practical approaches: sequence the change so a partial failure is safe by construction (grant the IAM policy before the database exists, so a database-creation failure just leaves an unused policy rather than a database nothing can reach), or make each piece individually idempotent so re-running the apply after a partial failure converges to the same end state rather than erroring on "already exists."
Worked example
Say an apply plans five resources in this order: a security group, a DB subnet group, an RDS instance, an IAM policy, and an IAM role-policy attachment. The first four create successfully; the fifth fails because the policy ARN referenced in the attachment has a typo. Terraform writes state incrementally as each resource completes, not only at the end of the run, so at the moment of failure, state already correctly records the security group, subnet group, RDS instance, and IAM policy as created; only the role-policy attachment is missing. Running terraform plan again reflects this accurately: it proposes creating only the missing attachment, not recreating the other four, because state already matches reality for those. The fix is to correct the ARN and reapply; nothing about the first four resources needs to be touched or reasoned about again. This is the diagnostic step that has to happen before deciding between "just reapply" (safe here, since the failure was config-only and everything else succeeded) and "restore from snapshot" (which would be the wrong call here, since it would needlessly discard four correctly-created resources).
Trade-offs and pitfalls
- Automated rollback assumes the operation is reversible and idempotent, true for stateless compute and image swaps, false for anything with an irreversible side effect (a dropped database column, deleted data, a webhook already fired to an external system). Forcing an automated revert onto that class of change is how you turn one incident into two.
- A blanket "restore the old state file" full rollback is dangerous in a team setting precisely because it can undo unrelated, legitimate concurrent changes, prefer targeted remediation once you've actually diagnosed which resource is broken, as in the worked example above.
- Forgetting the configuration-management layer after an infra-level rollback is a common miss: the infra looks reverted, but the OS-level config on that instance can silently stay in the post-change state until the CM tool is rerun.
Before a production apply, what review and automation guardrails would you put around the Terraform workflow so an engineer can catch unexpected destroys or replacements before they reach users?
Sample Answer
I would put both human review and automation around the plan. A Terraform plan is a preview of create, update, destroy, and replace actions. Replace is especially risky because Terraform deletes and recreates a resource.
Guardrails
- Run
terraform fmt -checkandterraform validatein CI - Save the plan as an artifact and require approval before apply
- Fail the pipeline if the plan includes unexpected destroys or replacements on critical resources
- Use policy checks for things like public exposure, open security groups, or deletion of databases
- Add
prevent_destroyto the few resources that should almost never be removed
Example
If the plan shows aws_db_instance.main will be replaced because of an engine change, I would force manual review from an engineer and, ideally, a service owner. If the plan only adds an autoscaling instance, that can follow the normal path.
This catches surprises before users feel them, which is the real goal.
You own a large estate of legacy shell scripts that provision infrastructure imperatively. How would you migrate them to declarative IaC with minimal disruption, and how would you decide when to finally retire the old scripts?
Sample Answer
Direct answer
A strangler-pattern migration: freeze new provisioning through the scripts, pick one resource type at a time, write declarative config that matches what already exists, import the live resource into the new tool's state without recreating it, and only then let the script's write path go dark for that resource type. You retire a script once every resource it used to own is imported, a plan against the new config comes back clean, and nobody still depends on running it manually.
Structured elaboration
What "legacy shell scripts" actually means for the migration
Bash wrapping the cloud CLI and Python wrapping the SDK (boto3, azure-mgmt, etc.) are the two common shapes, and for migration purposes they behave the same way: neither leaves behind a state file, so the IaC tool has no idea these resources exist until you tell it. The first job isn't picking a tool, it's building an inventory.
Phase 1: Inventory and classify
For every script, record:
- What resource(s) it creates or mutates (a name isn't enough, you need the actual resource IDs/ARNs it touches).
- Idempotency: does re-running it error, no-op, or duplicate the resource?
- Side effects beyond provisioning: does it also register the resource somewhere else (DNS, a CMDB, a monitoring config) that the new tool would need to replicate?
- Blast radius: shared resource (a VPC, an IAM role many things depend on) versus isolated (one team's dev bucket).
Order the migration by blast radius, smallest first.
Phase 2: Pilot on one low-risk resource type
Pick something isolated and cheap to redo if wrong (a dev-only S3 bucket, a single security group). Write the declarative config to describe it, then import rather than recreate.
Phase 3: Import without recreating
This is the step teams get wrong: writing config first and letting apply create a duplicate resource next to the one the script made. The safe order is config, import, plan, and only apply if the plan shows zero changes.
Phase 4: Dual-track cutover
Disable the script's write path (comment out the create/update calls, or pull the cron that reruns it) as soon as a resource type is imported, even before the whole estate is converted. Leaving both paths live is how you get drift: someone reruns the old script out of habit and the state file no longer matches reality.
Phase 5: Retirement criteria
Retire a script for a resource type when all of these hold:
- Every resource that script used to own has been imported and a plan against it is clean.
- The script's write path has been disabled for at least a full deploy cycle with no incidents traced back to needing it.
- Nobody on the team can name a workflow that still calls it manually.
- A rollback reference exists: the script itself, kept read-only in version control, in case you need to reconstruct what it used to do.
Worked example
Say a bash script provisioned an S3 bucket with aws s3api create-bucket --bucket acme-prod-assets --region us-east-1. The safe Terraform import sequence:
# 1. Write the resource config that should describe the existing bucket
resource "aws_s3_bucket" "assets" {
bucket = "acme-prod-assets"
}
# 2. Declare the import (Terraform 1.5+ import block, no separate CLI step needed)
import {
to = aws_s3_bucket.assets
id = "acme-prod-assets"
}
Run terraform plan. If the plan shows changes (say, the bucket has versioning enabled in reality but your config didn't set it), that's the config being wrong, not the import: fix the config until plan shows zero diff, then apply. Only after that is the resource genuinely under Terraform's control.
If the resource later gets refactored into a module, use a moved block instead of destroying and recreating it, so state history is preserved:
moved {
from = aws_s3_bucket.assets
to = module.storage.aws_s3_bucket.assets
}
Trade-offs & pitfalls
- Import only checks that the ID exists; it doesn't validate that your config matches every real attribute. A plan that isn't empty after import means your config is wrong, and applying it anyway can modify or, in the worst case, replace the live resource.
- The dual-track period is the highest-risk window: a stray cron rerun of the old script can silently drift the state Terraform thinks it owns. Disable the script's write path per resource type as soon as it's imported, don't wait for the whole estate to finish.
- Not every script maps cleanly to a resource. Ones with side effects (paging a team, writing to an external system) need that side effect re-homed somewhere (a CI step, a webhook) rather than papered over.
- Rewriting from scratch instead of importing is sometimes the right call: when the existing resource's config has drifted so far from any documented baseline that reproducing it faithfully would just be encoding technical debt. In that case, treat it as a planned recreation, with the downtime/cutover that implies, rather than an import.
Walk through init, validate, plan, and apply as they'd run in a typical Terraform workflow. What is each step actually checking, and why does plan specifically belong in your automated PR checks rather than just running at apply time?
Sample Answer
Direct answer
init sets up the working directory (downloads providers and modules, configures the backend), validate checks the configuration is syntactically and internally consistent without touching any real infrastructure, plan computes and previews exactly what would change against real provider state, and apply executes that change. plan belongs in automated PR checks, not just at apply time, because it's the only one of the four that tells a human reviewer what will actually happen before it happens, so the review is of an artifact (a diff) instead of a promise about what the code is supposed to do.
The four steps
terraform init: initializes the backend, downloads the provider plugins and any referenced modules at the versions your config or lock file specify. In PR checks, this is where you'd catch an unexpectedly changed backend configuration, an unpinned provider version, or a module source pointing somewhere it shouldn't.terraform validate: checks HCL syntax and internal consistency (required attributes present, types roughly line up, references resolve) without calling out to any provider API and without needing real credentials. It catches typos and structurally broken config, not semantic errors like an AMI ID that doesn't exist.terraform plan: reads real state and (partially) refreshes against the provider API, then computes the exact set of creates, updates, and destroys needed to reconcile your config with reality, without executing any of them. This is the artifact worth reviewing.terraform apply: executes the plan (ideally a specific saved plan file, not a freshly recomputed one) against real infrastructure.
Why plan belongs in PR checks specifically
- It turns "what will this change do" from something a reviewer has to mentally simulate by reading HCL into something they can read directly: an explicit list of resources and attributes that will be created, updated in place, or destroyed.
- It's the earliest point an unintended destroy on a critical resource becomes visible, well before anyone has run
applyand made it real. - Running
apply -out=tfplanagainst the exact plan file that was reviewed (instead of re-planning at apply time) closes the gap between what was approved and what actually executes; a fresh plan at apply time could differ if something changed underneath in the interim.
Concretely: a PR adding a new subnet triggers CI to run terraform fmt -check, terraform validate, then terraform plan -out=tfplan, and post the plan summary (say, "1 to add, 0 to change, 0 to destroy") as a PR comment. A reviewer approves specifically because the destroy count is zero. Merge triggers terraform apply tfplan using that exact saved plan file, so what gets applied is exactly what was reviewed, not a new plan computed after merge.
Everyday CLI commands beyond the core loop
terraform fmt: canonicalizes HCL whitespace and quoting. Real situation: a PR's diff is noisy because two engineers used different indentation styles; runningterraform fmt -recursive(and wiringfmt -checkinto CI, or a pre-commit hook) keeps formatting out of every plan-review diff.terraform destroy: tears down everything tracked in the current state. Real situation: an ephemeral PR-preview or load-test environment, provisioned nightly against its own isolated state, getsterraform destroy -auto-approverun against that state at the end of the day so it doesn't accrue cost.terraform apply -replace=<address>(the current form of whatterraform taintused to do,taintitself is deprecated as of Terraform 0.15.2): marks one specific resource for recreation on the next apply without changing any configuration. Real situation: an EC2 instance has a corrupted root volume from a bad AMI bake, but everything else about it (security group, subnet, instance profile) is correctly configured; replacing just that one resource avoids touching anything above it.terraform import: brings an existing cloud resource under Terraform management by writing its ID into state, without creating or changing anything. Real situation: someone created an S3 bucket by hand in the console before the team adopted Terraform; import writes it into state, and future applies manage it going forward, once matching configuration for it exists too (older Terraform versions don't generate that configuration for you).
What plan catches, and what it can't
plan is a config-level diff: it recomputes the resource graph against real (partially refreshed) provider data and shows the specific creates, updates, and destroys, and which attributes change. What it reliably catches: unintended destroys, drift between config and last-known state, and the blast radius of a change (how many resources are touched).
What it can miss:
- Runtime, application-level side effects: plan only knows about the infrastructure resource graph, not what happens inside the workload. Rotating an IAM policy might show as a harmless "update in place," but break a running application at runtime because it cached now-invalid credentials, something plan has no visibility into.
- Provider-specific eventual consistency: a cloud API can accept a value and apply it asynchronously (DNS propagation, IAM policy propagation); state can be technically correct right after apply while the live resource hasn't caught up, showing up as spurious drift on a same-day re-plan.
- Values only known after apply: when an attribute depends on a resource that doesn't exist yet, plan shows
(known after apply)as a placeholder, so anything downstream of that value is reviewed incompletely until the change is actually applied. - Out-of-band changes between plan and apply: plan is a point-in-time snapshot; if someone changes the resource manually, or another pipeline applies, in the gap, your apply operates on a stale plan (mitigated by state locking, and by re-planning if the gap between review and apply is long).
Trade-offs and pitfalls
validatepassing tells you nothing about whetherplanwill succeed,validatenever talks to the provider API, so a nonexistent AMI ID or an invalid instance type only surfaces atplan.importwithout matching configuration leaves you with a resource in state that config doesn't fully describe, the next plan may propose changing every attribute config doesn't specify back to a default.-replace(or the oldertaint) forces a full resource replacement, using it on the wrong resource address (for example, a security group instead of the instance) causes far more disruption than intended.destroyrun against the wrong workspace or directory is exactly as irreversible in an ephemeral environment as it is in production, always double check which state you're pointed at before running it, "it's just a dev environment" doesn't help if it was the wrong dev environment.
Unlock Full Question Bank
Get access to all Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.