Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
What's the difference between count and for_each in Terraform? Show a short example using for_each with a map of rule definitions to create several security group rules, so each one keeps a stable identity if the rules change.
Sample Answer
Approach
count indexes resources by position, 0 through n-1, so a resource's identity in state is tied to its numeric slot: remove or reorder an earlier element and everything after it shifts index, which Terraform reads as those resources needing to be destroyed and recreated even though nothing about them actually changed. for_each keys resources by a map key (or a member of a set(string)), so identity is tied to the key itself, independent of what else is added or removed. That's exactly what you want whenever the members have natural, stable names, like a set of named security group rules that get added or removed independently of each other.
Code
variable "rules" {
type = map(object({
from_port = number
to_port = number
protocol = string
cidr_blocks = list(string)
}))
default = {}
}
resource "aws_security_group_rule" "ingress" {
for_each = var.rules
type = "ingress"
from_port = each.value.from_port
to_port = each.value.to_port
protocol = each.value.protocol
cidr_blocks = each.value.cidr_blocks
security_group_id = aws_security_group.example.id
}
Newer AWS provider versions also offer aws_vpc_security_group_ingress_rule / egress_rule as more precise per-rule resources; the same for_each-over-a-map pattern applies identically either way.
Key points
The map's keys become part of the resource address, aws_security_group_rule.ingress["allow_https"], not just an internal label, that's what makes them the unit of identity. Inside the content/resource body, each.key is the current map key and each.value is the corresponding object, use each.value.<field> to reach each rule's attributes. Picking for_each over count here is really a statement that each rule has an independent lifecycle, adding "allow_ssh" shouldn't touch "allow_https" at all.
Complexity
Adding or removing one entry in a for_each map is an O(1) plan diff, only that one resource is affected. Removing element i from a count-based list is O(n - i) in practical terms, every subsequent index shifts, so every resource after position i shows up in the plan as a replacement even though only one rule was actually removed.
Edge cases
for_each accepts a map or a set(string) only, not a list, and not a set of complex objects, if your rules come in as a list you need to key them into a map first (by name, or by a for expression) before they'll work here. Map keys are naturally unique, but if you're deriving them (say, from a "name" field on incoming data) rather than hand-writing them, verify uniqueness explicitly: a for expression that derives the map raises Error: Duplicate object key at plan time on a collision, it does not drop a rule silently, that silent behavior only shows up with zipmap or a map that already arrived pre-collapsed from an external source. And because the key becomes part of the address string, avoid characters in rule names that complicate addressing or CLI targeting (quotes, unescaped special characters).
A teammate's Ansible task changes state every time it runs instead of converging, so re-running the playbook has unwanted side effects. Walk through how you'd diagnose why it isn't idempotent, how you'd fix it, and how the problem could have been caught before it reached production.
Sample Answer
Direct answer
A task that reports "changed" on every run instead of converging is almost always using shell or command (or a module used imperatively) to mutate state unconditionally rather than describing desired state, so Ansible has no way to tell whether the underlying system actually needed to change. Diagnose it by running the playbook twice back to back and finding exactly which task keeps reporting changed=true on the second, unmodified run; fix it by replacing that task with the purpose-built module for the job (or adding an explicit changed_when/creates guard); and catch it before production with a CI check that runs the playbook twice and fails if the second run is not a no-op.
Structured elaboration
Diagnosing
- Run the playbook twice against the same host with no other changes in between, or run it once, then again with
--check --diff, and look at which task or tasks reportchanged=trueboth times. A correctly idempotent task should reportchanged=trueon the first run (it did work) andchanged=falseon the second (nothing left to do). - Use
-vvvto see the actual module invocation and output. Forshell/commandtasks, look for an unconditional write:echo ... >> file,sed -iwithout an idempotent pattern,mkdirwithout a check, anything that acts rather than checks-then-acts. - Remember that
shellandcommandmodules reportchanged=trueby default on every successful run unless you explicitly tell Ansible otherwise viachanged_when, so a task reporting changed every time is often a module-level reporting problem even when the underlying action was harmless to repeat. - Bisect if the play is large: comment out later tasks, or use
--start-at-task, to isolate exactly which task is the offender before touching anything.
Fixing
Replace ad hoc shell with the module built for the job, since these compute their own idempotent diff:
# Non-idempotent: blind append
- name: Append setting (bad)
shell: echo "setting=true" >> /etc/example.conf
# Idempotent: lineinfile checks before writing
- name: Ensure setting present
lineinfile:
path: /etc/example.conf
regexp: '^setting='
line: 'setting=true'
create: yes
The same pattern applies to package installs (apt/yum with state: present instead of shelling out to the package manager) and to file content (template/copy, which compare checksums, instead of cat > or sed -i). Where a module genuinely does not exist for the action, keep shell/command but add changed_when (comparing command output or exit status) and, where applicable, creates/removes so Ansible can determine on its own whether the task actually needs to run.
Catching it before production
- Run
ansible-lintas a baseline check for known non-idempotent patterns, for example rawshell/commandwhere a module already exists (thecommand-instead-of-modulerule) or a command that always reports changed regardless of outcome (theno-changed-whenrule). - Use Molecule's built-in idempotence test: it runs the role, then runs it again, and fails the build if the second run reports any changed tasks. That test only proves the role is a no-op on an unchanged rerun; it will not catch a task like
lineinfilewithoutregexpsilently switching from replace to append the next time the desired value changes, since reproducing that requires a second, different value in CI, not just a second identical run. - Add a review checklist item: any
shell/commandtask must justify why a native module could not be used, and must setchanged_whenexplicitly rather than accepting the default.
Worked example
Take a cron entry managed with a lineinfile task missing its regexp (verified against ansible-core 2.21.2):
# Looks safe at first: no regexp, so lineinfile falls back to an exact-line match
- name: Add cron entry (bad)
lineinfile:
path: /etc/cron.d/myjob
line: '0 2 * * * root /usr/local/bin/backup'
create: yes
Run 1 (initial): changed=1, the file now holds exactly 0 2 * * * root /usr/local/bin/backup.
Run 2, playbook unchanged (identical rerun): changed=0. Without a regexp, lineinfile still does an exact-line match against the file's existing lines, finds 0 2 * * * root /usr/local/bin/backup already present verbatim, and does nothing: the file still holds exactly that 1 line, no duplicate. This is the misleadingly-clean case: it sails through a "run the playbook twice" idempotency check, so the missing regexp survives code review and CI untouched.
The real gotcha shows up when the desired VALUE changes, for example the schedule moves from 2am to 3am and the task is edited to line: '0 3 * * * root /usr/local/bin/backup', still with no regexp:
Run 3 (value changed, still no regexp): changed=1, but the exact-line match is against the OLD line, which no longer equals the new line value, so lineinfile cannot recognize it as "the same entry, updated" and appends a second one instead of replacing the first. The file now contains two conflicting entries:
0 2 * * * root /usr/local/bin/backup
0 3 * * * root /usr/local/bin/backup
Both cron jobs are now live, the backup runs at both 2am and 3am, until someone notices.
The fix is a regexp anchored to the part of the line that stays stable across value changes, here the command path, not the schedule:
- name: Ensure single cron entry (fixed)
lineinfile:
path: /etc/cron.d/myjob
regexp: '/usr/local/bin/backup$'
line: '0 3 * * * root /usr/local/bin/backup'
create: yes
Starting again from a clean file holding just the old-schedule line and running this fixed task with the new schedule: changed=1, and the file now holds exactly 1 line, 0 3 * * * root /usr/local/bin/backup, the old line replaced in place rather than appended alongside. Re-running the same fixed task unchanged: changed=0, still 1 line. Anchoring the regexp to the command path gives lineinfile something stable to match on even after the value it enforces changes, which is what makes the task idempotent across value changes, not merely across identical reruns.
Trade-offs & pitfalls
--checkmode has limits forshell/commandtasks specifically, since Ansible cannot simulate an arbitrary shell command's effect; those tasks needchanged_whenwritten carefully or check mode will either falsely report "no changes" or refuse to run at all.- A task that wrongly reports
changedevery run has a second-order effect beyond noise: any handler that isnotify'd by it (say, a service restart) fires every single run, causing unnecessary restarts on a resource that never actually changed. - A
lineinfiletask withoutregexppasses the identical-rerun idempotency check every single time, so it looks completely safe in CI and code review, right up until the day the desired value itself changes (a new cron time, a new config value); then it silently appends instead of replacing, and two conflicting lines are both live until someone notices. - CI running the playbook twice catches static, deterministic idempotency issues on an unchanged playbook, but not ones that only surface when the declared value changes between commits, and not ones caused by genuinely external state, for example a resource that legitimately looks different every scan because of something outside the playbook's control, like a third-party API's own non-deterministic response.
You are asked to build a reusable Terraform module for a three-tier application that includes networking, application compute, and a managed database. How would you split responsibilities between modules, and what would you expose so another team can compose it safely?
Sample Answer
I would split the solution by responsibility, not by environment. A module should do one job well.
Module layout
network: VPC, subnets, routes, NAT, and network tagscompute: app instances, ECS or ASG, load balancer, security groupsdatabase: managed DB, subnet group, parameter group, DB security grouproot stack: wires the outputs together
Why this works
The network changes slowly, compute changes often, and the database has its own lifecycle and risk. Keeping them separate reduces blast radius and makes reviews easier.
Safe interface
I would expose only what callers need:
- From
network:vpc_id,private_subnet_ids,public_subnet_ids - From
compute:alb_dns_name,app_sg_id - From
database:endpoint,port, and maybe a secret reference, not a password
Example
The root module can pass private_subnet_ids = ["subnet-101", "subnet-202"] into compute and database, while dev and prod use different sizes through variables. That keeps composition flexible without letting one team edit module internals.
What's the practical difference between a resource and a data source in Terraform? Give an example of each in a typical module, and explain how a data source changes what shows up in plan and how it affects the dependency graph.
Sample Answer
Direct answer
A resource block declares something Terraform owns the full lifecycle of: it will create, update, or destroy it to match what's declared. A data block reads information about something that already exists, without ever creating, changing, or destroying it. You reach for a resource when you want Terraform to own an object; you reach for a data source when you need to reference something that already exists, whether that's managed by another team, another Terraform config, or was created outside Terraform entirely.
Definitions and when to use which
- resource: Terraform issues API calls to bring the real object to the declared state, and records that object in state as something it manages.
- data source: Terraform issues read-only API calls to fetch attributes of an existing object, and does not persist it in state as a managed object.
| resource | data source | |
|---|---|---|
| Owns lifecycle | Yes, create/update/destroy | No, read-only |
| Appears in plan as | create / update / destroy action | a "read" during plan/apply |
| Recorded in state as managed | Yes | No (only its fetched values are used) |
| Typical use | New VPC, new RDS instance, new IAM role | An existing shared VPC, an existing AMI, a KMS key owned by another team |
Worked example
# resource: Terraform owns this EC2 instance's full lifecycle
resource "aws_instance" "web" {
ami = data.aws_ami.ubuntu.id
instance_type = "t3.micro"
subnet_id = data.aws_vpc.prod.id
}
# data source: reads an existing, externally-managed AMI
data "aws_ami" "ubuntu" {
most_recent = true
owners = ["099720109477"]
filter {
name = "name"
values = ["ubuntu/images/*"]
}
}
# data source: reads a VPC this config does not manage
data "aws_vpc" "prod" {
filter {
name = "tag:Name"
values = ["prod-vpc"]
}
}
Here aws_instance.web is the only resource, the thing this module actually provisions. The AMI and VPC are read through data sources because some other process owns them, a shared networking module owns the VPC, and AWS itself owns the published Ubuntu AMI.
Effect on plan and the dependency graph
- Plan: resources show up with create/update/destroy actions; data sources show up as reads and never propose a destructive action themselves, but their fetched values can still change what a downstream resource's plan looks like (a new AMI ID means the instance using it may show a forced replacement).
- State: resources are recorded as managed objects; data sources are not, only the values they return get used elsewhere in the config.
- Dependency graph: referencing
data.aws_vpc.prod.idinsideaws_instance.webcreates an implicit dependency, Terraform reads the data source before it can plan the resource that consumes it. If the object a data source reads is itself managed by a different resource in the same configuration, referencing it directly (rather than through a data source) creates an explicit dependency edge instead; mixing the two for the same object, reading viadatasomething also managed viaresourcein the same run, can create ordering ambiguity, since the data source might read a stale value from before that run's own changes land.
Trade-offs and pitfalls
- A data source is re-read on every plan (and refresh), so if the object it points at changes outside Terraform, your plan can shift underneath you with no corresponding resource change in your own config to explain why. This is a common source of "why did my plan suddenly change, I didn't touch anything" surprises.
- Prefer explicit remote-state outputs over a data source when you're referencing something managed by a sibling Terraform config in your own org, since outputs give you an explicit contract and version boundary; reserve data sources for things genuinely outside Terraform's control (cloud-provider-published AMIs, resources owned by another team's tooling entirely).
- Referencing a shared object (a shared VPC, a shared KMS key) via a data source from many independent configs works, but it also means none of those configs can see or coordinate around each other's dependency on it; a breaking change to that shared object has no single blast-radius list to check.
One of your engineers discovers that a database password ended up committed in a Terraform state file that was pushed to a shared repository. Walk me through what you'd do right away, and what you'd change afterward so it can't happen again.
Sample Answer
Direct answer
Rotate and revoke the credential immediately, everywhere it's used, before doing anything else: that's the only step that actually stops the bleeding, and forensics, cleanup, and policy changes all happen after containment, not instead of it. Separately, because this keeps happening across the industry, fix the two structural causes: encrypt the state backend so a leaked state file isn't plaintext-readable, and prefer ephemeral, short-lived credentials so secrets stop landing in state to begin with, rather than relying on catching every leak after the fact.
Structured elaboration
Immediate containment, before root-causing anything
- Revoke or rotate the exposed credential at the source (disable the IAM user/key, rotate the DB password) before investigating how it leaked. Every minute it stays valid is exposure.
- Remove the credential's blast radius: rotating a DB password and updating the app's connection string via the secret manager, not a manual edit, locks out anyone who copied it.
- Pull repository access temporarily if the state file is still sitting in a shared, readable location, so no one can pull a fresh copy while cleanup is in progress.
Investigate blast radius
Was the credential actually used by someone who shouldn't have had it? Check the target system's own access logs (DB connection logs, cloud provider audit logs) for activity from unfamiliar IPs or unusual times in the window since the state file was pushed. This determines whether it's "credential exposed, no evidence of misuse" or "confirmed unauthorized access," which changes the notification obligations.
Two different remediation paths depending on how long it's been exposed
A state file pushed an hour ago and one that's been sitting in git history for months are different problems.
- Fresh push, hours old: rotate, remove the offending commit before it's widely pulled, and a normal coordinated
git push --forceplus history cleanup is enough, because the exposure window is small and you likely know who has pulled it. - Buried in history for months: assume it's been exposed the entire time, not just since discovery, because you can't retroactively know who cloned, forked, or cached the repo, including CI runner caches and backup snapshots, during that window. This needs a full history rewrite (
git filter-repoto strip the file/commit from every ref), coordinated re-clone by every developer and CI system, since a plain force-push doesn't fix clones that already exist, and treating the credential as compromised regardless of what the access logs show, since log retention may not even cover the full exposure window.
Structural fix 1: encrypt the state backend
Even after rotation, the old, now-invalid but still informative, value sitting in a plaintext state file is a problem for anyone auditing what leaked. Use a state backend with encryption at rest by default (S3 with SSE-KMS and a bucket policy denying unencrypted puts, Terraform Cloud/Enterprise's built-in encryption, or a Vault-backed backend) so a copy of the state file alone isn't enough to read secrets even if it leaks again.
Structural fix 2: prefer ephemeral credentials so nothing long-lived is ever in state
The deeper issue is that a static, long-lived credential value, a real DB password, not a reference to one, was ever an attribute in the resource's config to begin with. The fix is architectural, not just procedural: use dynamic, short-TTL credentials issued at apply time (Vault, or a cloud-native equivalent) instead of static passwords, so the value that ever touches Terraform is valid for minutes, not indefinitely. Even that has a gap worth knowing: a value fetched via a data source still gets written into the state file's plaintext by default, because Terraform's state format persists every attribute of every resource and data source it manages, dynamic or not. This is exactly the gap Terraform 1.10+ addressed with ephemeral resources and write-only arguments: values marked ephemeral are used during the apply but are deliberately never written to state or plan files at all, a stronger guarantee than "the credential expires quickly" on its own.
Structural fix 3: a pre-commit/CI scanning gate
Rotation and history rewriting fix one incident; a scanning gate stops the next one from merging in the first place. Run a secret-scanning tool (gitleaks, truffleHog, or similar) both as a pre-commit hook, for fast author feedback, and as a required CI check on every PR, the actual enforcement point since pre-commit hooks can be skipped locally. Scan the Terraform plan/state output specifically, not just source files, since state and plan JSON are exactly where a secret shows up even when the .tf source never had it hardcoded.
Notification and postmortem
If the credential's compromise could have exposed customer data, loop in compliance/legal early, since this determines regulatory notification timelines, often measured in days. Regardless of scale, write a postmortem with a timeline, what was exposed, what was rotated, and concrete owners/dates for the structural fixes above, not just "we rotated it and moved on."
Worked example
Concrete before/after for structural fix 2, showing the actual problem, a static password as a plain resource argument, versus the pattern that avoids it:
# Before: static secret is a plain attribute, gets written to state in plaintext
resource "aws_db_instance" "app" {
identifier = "app-prod"
password = var.db_password # a real, long-lived value, now permanently in state history
}
# After: value is marked ephemeral (Terraform 1.10+), used during apply,
# never persisted to state or plan output
ephemeral "random_password" "db" {
length = 24
}
resource "aws_db_instance" "app" {
identifier = "app-prod"
password_wo = ephemeral.random_password.db.result
password_wo_version = 1
}
The practical difference: if this state file leaks again after the fix, there's no password in it to rotate in a panic, because the value was never written there.
Trade-offs & pitfalls
- Rotating before investigating feels backwards to people who want to preserve evidence first, but a live credential someone else might be actively using is a bigger risk than losing a few minutes of forensic cleanliness; contain first.
- Treating "we rotated it" as the end of the incident misses that the old value is still sitting in every git clone and CI cache that pulled it before rotation; a fresh push and a months-old buried secret need genuinely different remediation, not the same checklist applied at different speeds.
- Ephemeral values and write-only arguments require the specific resource/provider to support them; not every provider has adopted the pattern yet, so this is a direction to move toward for new resources, not a guarantee that already covers the whole estate.
- A pre-commit hook alone isn't enforcement, since
--no-verifyskips it; the CI check on the PR is what actually blocks a merge, and that's the one to treat as required rather than advisory.
Unlock Full Question Bank
Get access to all Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.