Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
When would you pull resources out into a reusable module instead of leaving them in the root configuration? Explain what makes a good module boundary, and what you'd call over-modularization.
Sample Answer
Pull resources into a module once they're reused across more than one stack or environment, or once a single-purpose group of resources, a VPC, an EKS cluster, a standard RDS setup, has grown enough internal complexity that inlining it clutters the root config. A good module boundary maps to one concern with a small, stable interface; over-modularization is wrapping something in a module because it feels tidy, not because anything actually reuses it or needs the extra indirection, and it shows up as one-resource modules, deep nesting, or modules stitched together with a pile of boolean flags to fake reuse across cases that were never really the same.
When to extract, when to leave inline
- Extract when the resource group is reused across environments, teams, or projects, or its complexity, many variables, repeated lifecycle rules, would otherwise be copy-pasted.
- Leave inline when resources are small, unique to one stack, or tightly coupled to that stack's own variables and lifecycle; wrapping them in a module just adds a layer to look through for no reuse benefit.
Layered module hierarchy
A common shape at organization scale: networking, identity, shared-services, platform, and application layers, each owning a narrow slice and exposing only what the next layer needs, networking exposes VPC and subnet IDs, identity exposes role ARNs, shared-services exposes things like a shared logging bucket, platform composes those into a cluster, application deploys onto the platform. The dependency direction only ever points one way, up: a lower layer like networking must never read or depend on a higher layer's outputs like application, or you get a circular dependency that makes independent deploys impossible. Enforce that with separate state per layer and remote-state or published-lookup reads flowing strictly upward, never the other direction.
The module contract, and encapsulation leaks
A module's real contract is its inputs and outputs, nothing else. The failure mode to design against: consumers start depending on a specific resource name or ID inside the module that it never promised to keep stable, referencing an internal resource address directly instead of a proper output, so a refactor inside the module that changes nothing about its behavior still breaks every consumer. Prevent it by never exposing internal resource addresses, only named outputs, and by treating "adding or renaming an output" with the same breaking-change discipline as changing an input.
Choosing a DRY mechanism
| Mechanism | Best for | Downside |
|---|---|---|
| Shared module | Genuine multi-stack reuse of a resource group with real variation between consumers | Overkill for a single repeated value; adds a version to manage |
| Locals / variable maps | Repeated values within one stack, or a small fixed set of environments keyed by name | Doesn't scale across repos; every stack still owns its own copy of the map |
| CI-side templating | Config that's identical except for a few substituted values across many near-identical stacks | Hides the actual config from a plain plan read; harder to review a real diff |
If the same config keeps getting copy-pasted across stacks, start with the cheapest fix, a locals map, before reaching for a full shared module; promote to a module once real behavioral variation, not just value substitution, shows up between consumers.
Trade-offs & pitfalls
Over-modularization is a common overcorrection once a team learns the reuse lesson: they start extracting one-resource modules, or nesting modules three deep, and debugging becomes tracing variable propagation through layers that never needed to exist. A practical rule: start inline, extract when reuse or complexity actually hits, and keep every module's interface small enough that its README fits on one screen.
Tell me about a project where you used Infrastructure as Code. How was it laid out across modules and environments, how did you handle secrets, and what did the approval process look like before a change actually got applied?
Sample Answer
Direct answer
On my last project I codified an AWS microservices platform, a set of backend services running on managed servers with their own database, load balancer, and access controls (VPC, EKS, RDS, IAM, ALB, monitoring) in Terraform, using versioned reusable modules composed per environment, remote state (the file Terraform uses to track what it created, kept separate per environment so a mistake in one can't touch another) isolated per environment, secrets pulled from AWS Secrets Manager and Parameter Store rather than stored in code, and a PR-based workflow where a machine-generated plan had to be reviewed and approved before an apply job with a separate, more privileged role could run.
How I structured it
Module layout and versioning
- Modules lived in a private registry: vpc, eks, rds, iam, alb, monitoring, each with a narrow set of inputs/outputs and no hidden side effects.
- Root configurations per environment composed these modules and pinned each one to a semantic version tag (for example vpc ~> 2.3), so a change to a module's source could not silently change an environment that had not explicitly bumped its pin.
Environment separation and state
- Each environment (dev, staging, prod) had its own remote state file in an S3 backend with a DynamoDB lock table, keyed roughly as s3://infra-state/{env}/{component}.tfstate.
- Isolating state per environment (rather than one shared state with workspaces) meant a mistake in dev could not touch prod's state, and the blast radius of a single apply was limited to one component in one environment.
Secrets handling
- Database credentials lived in AWS Secrets Manager, app configuration lived in Parameter Store encrypted with KMS.
- Terraform read them at apply time through data sources (for example aws_secretsmanager_secret_version) rather than having them typed anywhere in .tf files or CI variables.
- The CI runner itself never held a long-lived key: it assumed a role via OIDC scoped to the minimum permissions needed for that environment's apply.
Change approval and safe apply
- Every change went through a PR. CI ran terraform plan and posted the JSON plan (with sensitive values redacted) as an artifact, plus a readable summary on the PR.
- A change needed sign-off from both the infra owner and the owning service team before the apply stage would even unlock.
- Apply ran as a separate pipeline stage using a more privileged, MFA-gated role for prod, and every apply was logged to CloudTrail.
- Destructive changes (anything showing a resource replacement or delete in the plan) required an explicit manual confirmation step, and we took a fresh snapshot or backup first.
Worked example
Concretely: adding a read replica to an existing RDS instance meant one PR that touched only the rds module's call site, a plan that showed one resource create and zero destroys, review from the database owner since it touched a stateful resource, and an apply that ran with the prod-apply role only after both approvals landed. Because the module was version-pinned and state for that component was isolated, the blast radius of that single PR was exactly one RDS resource in one environment: nothing else in the account could be affected by that apply.
Trade-offs and pitfalls
- Splitting state by component reduces blast radius but adds cross-stack coordination cost (remote state lookups or SSM parameters to pass values between components); too many tiny state files becomes its own operational burden.
- A plan-then-approve workflow is only as safe as the reviewers actually reading the plan. Without a policy-as-code gate (something like OPA/Rego or Sentinel evaluating the plan JSON for guardrails such as "no public security groups" or "no unencrypted volumes"), review can degrade into rubber-stamping on a busy day.
- Reading secrets via data sources at apply time keeps them out of source control, but the values still land in the Terraform state file in plaintext, so state encryption and tightly scoped state-read IAM matter just as much as the CI-side handling.
Why do teams move Terraform state to a remote backend instead of keeping the state file on a laptop or checking it into git? What actually goes wrong once more than one person is working against the same infrastructure?
Sample Answer
Direct answer
Once more than one person, or one person plus CI, can run Terraform against the same infrastructure, a local or git-committed state file has no mechanism to stop two applies from racing against each other. Git specifically makes it worse: state is a large, frequently changing JSON blob that often contains resolved attribute values (including secrets), so committing it turns every apply into a potential merge conflict on a binary-ish file and leaves secrets sitting in permanent git history. A remote backend with locking solves both: one shared source of truth, and a mechanism that serializes concurrent applies instead of letting them collide.
Why local files and git specifically break down
- No locking: two engineers (or an engineer and a CI job) running
applyat the same time both read the same starting state, compute independent plans, and whichever one writes state last silently overwrites the other's recorded changes, even though both sets of underlying cloud resources were actually created. - Merge conflicts on a JSON blob: two people committing an updated
.tfstateproduce a diff Git has no useful way to merge; resolving it by hand risks losing one side's changes entirely. - Secrets in permanent history: state routinely contains resolved attribute values (see the secrets-in-IaC discussion elsewhere in this topic for why that includes values from
sensitive = truevariables and secret-manager lookups alike); committing it to git means those values live in every past commit, not just the latest one, deleting the file later doesn't remove them from history. - No audit trail: no consistent story of who applied what and when, beyond whatever commit messages happen to say.
What remote backends add
| Backend | Locking mechanism | Encryption | Best fit |
|---|---|---|---|
| S3 | Native S3 lockfile locking (use_lockfile): a conditional PutObject with If-None-Match creates a .tflock object, no separate lock table needed | SSE-KMS on the S3 bucket | AWS-native stacks, mature and well-documented. DynamoDB conditional-write locking is the legacy mechanism for this backend: HashiCorp's own docs mark it deprecated and slated for removal, so new setups should use use_lockfile, not stand up a DynamoDB lock table. |
| GCS | Native object generation-based locking | Google-managed or customer-managed (CMEK) encryption | GCP-native stacks |
| Terraform Cloud / Enterprise | Built-in locking per workspace | Managed by HashiCorp, or self-hosted for Enterprise | Teams that also want managed run history, policy checks (Sentinel/OPA), and UI-based approvals without self-hosting a lock table |
Worked example: engineer A runs terraform apply against envs/prod. Moments later, engineer B, unaware, runs terraform apply against the same directory. With S3's use_lockfile locking, B's apply fails immediately with a lock-acquisition error naming the .tflock object holder and start time, and B waits or investigates instead of proceeding (a legacy DynamoDB-locked backend behaves the same way, since it's the same conditional-write mechanism against a different store). Without locking (local state, or a backend configured without it), both applies read the same starting state, compute independent plans, and whichever one writes state last silently discards the other's recorded changes, even though both sets of underlying cloud resources were actually created, leaving state and reality out of sync.
Failure modes once remote state is misconfigured
Moving to a remote backend removes the local/git-specific problems, but introduces its own misconfiguration risks:
- Wrong or reused backend key across environments: if dev and prod point at the same bucket with the same
key, they silently share one state file. The nextapplyagainst "prod" can then try to reconcile prod's real infrastructure against dev's last-applied state, up to and including destroying prod resources dev's config doesn't declare. - Locking disabled, or habitually overridden: some configurations skip locking, or
-lock=falsebecomes a habitual way to "unblock" a stuck pipeline. Two concurrent applies against the same state can then interleave writes and leave the state file with attributes from two different runs, a subtler and harder-to-detect form of corruption than an obvious file truncation. - State bucket versioning or backups disabled: without object versioning on the backend store, there's no rollback point once a bad apply or an accidental
terraform state rmcorrupts state; with versioning on, you can restore the last-good object version. - Overly broad IAM or access on the state bucket/table: since state contains resolved attribute values, including secrets, a state backend with read access broader than the specific CI and deploy identities that need it is itself a secrets-exposure surface, not just a locking concern.
Trade-offs and pitfalls
- Locking prevents concurrent writes; encryption and access control are separate, additional concerns, having one doesn't imply you have the others, all three need to be configured deliberately.
- Treat the state backend with the same access-control rigor as a secrets store, because for practical purposes it often is one.
- Enable object versioning on the backend bucket by default; it's the cheapest possible insurance against a bad apply or an accidental manual state edit.
A client runs workloads both on-premises and in the cloud, and you're deciding between imperative scripting and declarative IaC to automate both. Walk through the trade-offs, including idempotency, debugging and observability, vendor lock-in, and team skills, and what you'd actually recommend.
Sample Answer
Direct answer
For a client running workloads both on-premises and in the cloud, default to declarative IaC (Terraform, OpenTofu, an open source fork of Terraform, or CloudFormation where the target is AWS-only) for anything that represents steady-state infrastructure, because idempotency and drift visibility come largely built in, and reserve imperative scripting for the parts declarative tools genuinely do not express well: multi-step conditional workflows, human approval gates, and one-off data migrations. The two are not mutually exclusive: compose declarative modules for the actual infrastructure state, and wrap them with an imperative pipeline or workflow engine for sequencing across the hybrid environment.
Structured elaboration
Comparing the two approaches on the axes that matter here
| Axis | Declarative (Terraform/OpenTofu, CloudFormation) | Imperative (bash/SDK scripts, ad hoc automation) |
|---|---|---|
| Idempotency | Native: the engine computes a diff against desired state and only changes what needs to change | Not automatic; every script has to be written to check-before-act, and that discipline regresses easily |
| Debugging and observability | A plan shows the intended change before it happens, and the state file is inspectable, but provider-level convergence failures can be opaque | Step-by-step logs make root cause easy to trace, but there is no single artifact you can inspect ahead of time to see everything that will change |
| Vendor lock-in | Provider abstractions can be fairly portable across clouds, but the tool's own licensing and ecosystem is a separate axis of lock-in | Direct SDK/CLI calls are inherently provider-specific unless someone deliberately abstracts them, but there is no tooling-level lock-in |
| Team skills | Lower barrier for infrastructure-focused operators; requires discipline in module design and state hygiene | Needs software-engineering habits, idempotency, retries, testing, that not every operations team has built up |
Vendor lock-in has two separate axes, and one of them changed recently
It is worth separating provider lock-in (which cloud or on-prem APIs your modules target) from tooling lock-in (which organization controls the engine and its licensing terms). HashiCorp moved Terraform off its original open-source MPL license onto the Business Source License in 2023, which is what triggered the OpenTofu fork, maintained under the Linux Foundation as an open-source, largely drop-in alternative. For a hybrid client weighing lock-in, that is a tooling-lock-in consideration distinct from whether their modules are portable across on-prem and cloud providers.
Where the two compose
Use declarative modules as the primary pattern for provisioning and ongoing stateful resources on both on-prem and cloud sides, since that preserves idempotency and auditability everywhere. Wrap the orchestration of a hybrid rollout, sequencing, approvals, cross-environment coordination, in a pipeline (CI/CD) or workflow engine that calls into those declarative modules rather than reimplementing infrastructure logic imperatively.
Worked example
Say the client needs to migrate 40 VMs off on-prem VMware, with some staying on-prem and the rest moving to a cloud provider. The steady-state definitions for both the vSphere-hosted and cloud-hosted VMs live in parameterized Terraform/OpenTofu modules, one per provider, sharing a common interface (name, size, network, tags). The migration sequence itself, drain traffic from a VM, provision its cloud replacement, validate health checks, cut over, decommission the on-prem VM, is not something the declarative modules express well as a single unit, so it runs as a pipeline stage (GitHub Actions or a workflow engine) that calls terraform apply for each phase with an explicit manual approval gate between "provision replacement" and "decommission original," since the decommission step is destructive and on-prem hardware is not trivially recoverable.
Trade-offs & pitfalls
- Embedding imperative logic inside a declarative tool's own provisioners (
local-exec/remote-exec) is a common anti-pattern: it breaks the engine's ability to reason about and safely re-run the plan, since the side effect of a shell command is invisible to the state model. - The opposite mistake also happens: standing up a full workflow engine for orchestration that a plain module dependency graph could already express adds operational surface (another system to run, secure, and understand) with no real benefit.
- Recommending OpenTofu purely on cost or philosophy without checking provider/module compatibility for the client's specific stack can introduce its own migration risk; the fork is close to drop-in but is not automatically identical for every provider version in use.
Write an OPA/Rego policy that rejects any resource missing a tags map with owner and environment keys, and explain how you'd wire that policy into CI so a bad terraform plan can't get merged.
Sample Answer
Direct answer
A tags policy denies any Terraform resource change whose after state has a missing tags attribute, or a tags map without both owner and environment keys, evaluated against Terraform's JSON plan output. Wired into CI as a conftest (a CLI that runs Rego policies against JSON/YAML input) or opa eval step that runs after terraform plan, it fails the pipeline (non-zero exit) before merge, so terraform apply never runs against a non-compliant resource.
Approach
Read Terraform's plan as JSON (terraform show -json), walk resource_changes, skip anything being deleted (a destroy has no future tags to check), and for everything else assert after.tags exists and contains owner and environment. Emit one deny message per violation so CI can print exactly which resource and which key failed, not just "policy failed."
Policy (Rego)
package terraform.tags
import rego.v1
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
not is_map(tags)
msg := sprintf("resource %v missing tags map", [rc.address])
}
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
is_map(tags)
not tags.owner
msg := sprintf("resource %v missing tags.owner", [rc.address])
}
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
is_map(tags)
not tags.environment
msg := sprintf("resource %v missing tags.environment", [rc.address])
}
is_map(x) if {
x != null
type_name(x) == "object"
}
Key points
- Iterate
resource_changes, not the raw HCL, so the policy sees what will actually be created or changed, including values interpolated from variables and modules. - Exclude pure
deleteactions withsome action in rc.change.actions; action != "delete": this is an existential check, so a replace (["delete", "create"]) still gets validated on itscreatehalf, while a plain destroy (["delete"]) is correctly skipped. - Default the tags lookup with
tags := object.get(attrs, "tags", null)before ever callingis_mapon it. Do not callis_map(attrs.tags)directly: whenattrs.tagsdoes not exist, referencing it produces no value at all, and in Rego, negating a call whose argument is itself undefined never resolves to true or false, it just never fires.object.getalways returns a defined value (nullwhen the key is absent), sonot is_map(tags)becomes a normal negation over a normal function result and correctly denies the "no tags key at all" case. - One
denyrule per failure mode (missing map, missing owner, missing environment) so the CI output tells the developer exactly what to fix.
Complexity and edge cases
Evaluation is O(number of resource_changes) with constant work per resource: no recursion, no external calls. Edge cases worth naming explicitly: resource types that do not support tags at all (need an allowlist of resource types, otherwise this policy false-positives on them); for_each/count resources, which appear once per instance in resource_changes and are each checked independently, so a module with ten instances produces ten independent checks; and a tags value that depends on another resource not yet created, which the plan reports as unknown rather than present. A naive check treats "unknown at plan time" the same as "missing," which is a false positive; a stricter version would inspect rc.change.after_unknown.tags and only deny when the value is genuinely absent, not merely unresolved yet.
Wiring into CI
- Generate the plan in CI:
terraform plan -out=plan.binary && terraform show -json plan.binary > plan.json. - Evaluate it:
conftest test --policy ./policy plan.json, oropa eval --input plan.json --data policy.rego "data.terraform.tags.deny". - Treat any non-empty
denyset as a pipeline failure and make that check required on the PR, not advisory, so a merge is mechanically blocked, not just discouraged. - Surface the
denymessages as a PR comment or check annotation so the resource address and missing key are visible without digging into CI logs.
Trade-offs & pitfalls
- A policy this strict on day one breaks every existing untagged module. Roll it out in warn-only mode first, generate a report of current violations, and flip to blocking once the backlog is cleared, not the other way around.
- Checking the JSON plan catches drift-introducing changes before apply, but it cannot catch tags removed later by someone with direct cloud console access. That needs a separate periodic scan against live resources (a detective control), not just this preventive one.
- The
after_unknowncase above is easy to miss in a first pass and is exactly the kind of thing that turns into a noisy false positive once real modules with cross-resource references start hitting the policy. - Rego's "negating a call over an undefined argument never resolves" behavior is a well-known footgun: a
not f(x)guard that looks correct in review can silently never deny anything for the exact input it was written to catch. Always test the "attribute completely absent" case with the realopabinary, not just the "attribute present but wrong" case, since the two paths through a guard function can diverge. - This policy is written in OPA v1 syntax (
deny contains msg if { ... }), which has been the default parser since OPA 1.0 (2024). An older pinned binary that still defaults to v0 needs either an upgrade or the--v0-compatibleflag; don't ship v0-only syntax and call it current.
Unlock Full Question Bank
Get access to all Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.