Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
A service stores database credentials and API keys in Terraform variables today, and the team is worried about exposure. How would you redesign the workflow so Terraform can still provision the stack without making secrets visible to everyone who can read the repo or the state?
Sample Answer
I would redesign this so Terraform manages secret references, not secret values, whenever possible. A secret is the actual password or API key. A secret manager is a service like Vault or AWS Secrets Manager that stores the value and gives apps access at runtime.
Workflow
- Store the real secret in a secret manager outside the repo
- Pass only the secret ARN or path into Terraform, not the plaintext value
- Let the application read the secret at runtime using its IAM role or service identity
- Protect state with encryption, tight backend access, and separate prod permissions
Important caveat
If a provider resource truly needs the raw secret during create, Terraform may still record that value in state. Marking a variable sensitive only hides it from CLI output, not from state. So I would avoid designs that require Terraform to own the secret value unless there is no alternative.
Example
Terraform creates prod/db/password as a managed secret container, and the app reads it later. That keeps repo access from becoming secret access.
Walk me through blue-green versus canary deployment from an infrastructure perspective; when would you reach for each? Think about traffic routing, resource duplication, cost, and what gets harder when a database schema change is part of the rollout.
Sample Answer
Direct answer
Blue-green gives you a full, already-validated environment and an instant, simple rollback, at the cost of running two complete copies of your infrastructure during the change; canary gives you gradual, lower-risk validation at the cost of more routing and monitoring complexity, and it is harder to reason about while it is in flight. Pick blue-green when a clean, all-or-nothing cutover is affordable and rollback speed matters most; pick canary when the blast radius of being wrong needs to stay small and you can invest in the routing and monitoring to support gradual exposure.
Structured elaboration
Comparison
| Dimension | Blue-green | Canary |
|---|---|---|
| Traffic routing | Atomic cutover: DNS, listener swap, or weight flip to 100% | Gradual: weighted routing or service-mesh rules, ramped over time |
| Resource duplication / cost | Full duplicate environment while both are live, higher but short-lived cost | Only a fraction of the fleet runs the new version at any point, lower peak cost |
| Rollback | Instant, swap back to the known-good environment | Ramp weight back down; usually just as fast, but only a subset of traffic was ever exposed |
| Blast radius if wrong | All traffic hits the new version the moment you cut over | Bounded to the canary percentage until you choose to widen it |
| Operational complexity | Lower: mostly "which environment is live" | Higher: routing rules, staged thresholds, automated promote/rollback logic |
| Database schema changes | Blue and green must both work against the same schema at the moment of cutover, since there is no gradual overlap window; this forces backward and forward compatible migrations (expand, migrate, contract) or a maintenance window | Old and new versions coexist against the same schema for longer, so the compatibility window has to hold for however long the canary runs, not just for an instant |
Other strategies worth naming
Blue-green and canary are not the only tools:
- Progressive expansion: widen the set of inputs a change applies to in stages rather than widening a traffic percentage, for example rolling a new IAM policy or network ACL out to one account, then one region, then everywhere, independent of any single request's traffic path.
- Feature flags for infra: gate exposure to a new code or config path behind a flag rather than, or in addition to, a routing change, so the change can be turned off instantly without touching load-balancer configuration at all. This is especially useful when the change is a behavior toggle inside already-deployed infrastructure rather than a new fleet.
Why this gets harder for networking or storage than for a stateless app release
Blue-green and canary are easy to describe for a stateless app release because the unit being duplicated (a fleet of identical, disposable instances) and the unit being routed (an HTTP request) line up cleanly. That stops being true for infrastructure-level changes:
- Networking: a routing or firewall change often affects an entire connection or an entire flow, not a single request, so traffic cannot always be shifted 5% at a time the way it can with an HTTP load balancer. A stateful TCP connection, a VPN tunnel, or a peering change uses the old path or the new path for its whole lifetime; the workaround is to canary at the level of an entire subnet, account, or long-lived connection cohort rather than per request, and to accept a coarser blast radius.
- Storage: there is no cheap duplicate of a stateful data store the way there is for a stateless instance. Standing up a second, fully synced database is expensive and introduces replication lag as a new failure mode. The common workaround is to validate the new storage layer with shadow traffic or dual writes before anything reads from it in production, and to use expand-then-contract schema migrations so both the old and new consumers can run against the same underlying store during the transition instead of trying to canary the store itself.
Worked example
A schema change on a Postgres-backed service being rolled out with blue-green: adding a NOT NULL column safely requires an expand-then-contract sequence rather than a single migration, because both blue and green must work against the same schema during cutover.
- Expand: add the column as nullable, deploy it (both blue and green tolerate a nullable column).
- Backfill: populate the column for existing rows.
- Cut over blue-green as normal, both versions still tolerate the nullable column.
- Contract: once green is fully promoted and blue is decommissioned, add the
NOT NULLconstraint in a separate migration, since only green's code path needs to rely on it.
This is the same three-step compatibility pattern regardless of whether the deployment strategy is blue-green or canary; what changes is only how long the "both versions must tolerate the old and new shape" window needs to hold.
Trade-offs & pitfalls
- Choosing canary by default because it "sounds safer," without the routing and monitoring maturity to support it, is a common mistake; a canary that cannot actually be measured is worse than a blue-green cutover that can be instantly reversed.
- Forgetting that duplication cost for blue-green is temporal, not permanent, leads teams to avoid it for cost reasons when the actual bill is only for the cutover window.
- Assuming the app-release playbook (weighted HTTP routing) transfers directly to networking or storage changes is the single most common failure mode; check whether the thing being changed is even divisible at the granularity being planned for the canary.
What does terraform import actually do, and can you think of a time you'd reach for it instead of just writing a fresh resource block?
Sample Answer
Direct answer
terraform import maps a real-world resource that already exists, created by hand, by a script, or by another tool, into Terraform's state under a specific resource address, so Terraform starts managing something it didn't create. You'd reach for it whenever the resource already exists and deleting-then-recreating it is unacceptable: a production database with real data, a manually created IAM role that other things depend on, or infrastructure inherited from a pre-IaC era or an acquisition, rather than writing a fresh resource block and letting Terraform create it from scratch.
Structured elaboration
What import actually does
It reads the target resource's current attributes from the provider and writes a state entry linking <resource address> to <remote ID>. That's it. The classic terraform import <addr> <id> command does not generate HCL for you and does not reconcile your configuration against reality, so after importing you still need a resource block that matches, and you keep running terraform plan until it shows no diff.
When to reach for it instead of a fresh resource block
Anytime recreation would be destructive or disruptive: a resource holding state or data (a database, an S3 bucket with objects in it), a resource with a live external effect (a DNS record already receiving traffic, a load balancer with active connections), or simply infrastructure someone stood up manually before Terraform existed in the project.
The basic flow
Write a minimal resource block of the correct type first, matching just the arguments you intend to manage. Run terraform import <addr> <id>. Run terraform plan to see what Terraform thinks should change, since your HCL almost never matches the real resource on the first try. Adjust the HCL (and add lifecycle { ignore_changes = [...] } for attributes you don't want to fight with) until plan shows no diff, then only apply if you deliberately want to change something.
Current tooling: declarative import
As of Terraform 1.5, there's also an import block you place directly in configuration, which is plannable (you can see the import as part of a normal plan before committing to it) and can be paired with terraform plan -generate-config-out=generated.tf to scaffold a starting resource block for you instead of hand-writing one. Most day-to-day usage you'll still see is the classic imperative CLI command; the block form is worth knowing for bulk imports or when you want the import itself to go through review.
The related command: taint / -replace
Import and taint solve opposite problems. Import brings an unmanaged resource under management without touching it. terraform taint (deprecated since Terraform 0.15.2 in favor of terraform apply -replace=<address>) does the reverse: it marks a resource Terraform already manages for forced destroy-and-recreate on the next apply, useful when a resource is in a bad state that Terraform's normal attribute diff wouldn't catch on its own, a corrupted disk, or an instance that failed its bootstrap script.
Worked example
# import an existing RDS instance into aws_db_instance.db
terraform import aws_db_instance.db my-db-identifier
# confirm it's tracked
terraform state list | grep aws_db_instance.db
# see what Terraform thinks should change (should shrink to nothing
# as you adjust the HCL to match reality)
terraform plan
If the HCL block for aws_db_instance.db only declares identifier = "my-db-identifier" to start, the first plan will likely propose changes for every other default-valued argument Terraform infers as "should be removed." You reconcile by adding the real values (or ignore_changes for provider-computed ones like endpoint) until plan is clean.
Trade-offs & pitfalls
Import doesn't validate that your HCL is correct, it only validates that the ID exists; if your resource block is wrong, the first plan after import can propose a destructive change on something you specifically imported to avoid touching. Before Terraform 1.5's -generate-config-out, bulk imports of dozens of resources meant either writing every resource block by hand or scripting the CLI, both error-prone. taint/-replace triggers a real destroy-and-recreate, treat it as dangerous on anything stateful even though the command itself is quick to run.
A Terraform run needs to hand a newly provisioned service a database credential or similar application secret. How would you get that credential to the resource without it ever being a long-lived static value baked into your configuration or left sitting in state, and what has to happen when that credential eventually needs to be renewed?
Sample Answer
Direct answer
Don't have Terraform mint or hold the actual database credential at all. Have it provision the plumbing of a secrets engine (Vault, HashiCorp's tool for storing and minting credentials on demand instead of them being fixed values anywhere, has a database secrets engine for this; a cloud-native equivalent like AWS Secrets Manager with rotation configured works too) that mints a short-lived, per-consumer credential on demand, and have both Terraform's own auth to that engine and the application's auth use identity-based, short-lived methods (AppRole, a role-ID-plus-secret-ID login pattern explained below; Kubernetes service-account auth; cloud IAM/OIDC) instead of a static token. Renewal then becomes the secrets engine's job (lease renewal or a scheduled rotation function), not a manual re-apply of a static value.
Structured elaboration
Why baking the secret in is unsafe
A credential written into a resource argument or interpolated into a data source ends up in the Terraform state file (state is plaintext JSON unless you add extra tooling around it) and can appear in plan/apply CLI output and CI logs. It also never expires on its own: a static value is a permanent liability until someone remembers to rotate it by hand.
The pattern: Terraform provisions the engine, not the secret
Terraform's job is to create the database secrets engine mount, the connection, and a role that defines what SQL a dynamic credential is allowed to run and for how long. The application (or the config-management tool bootstrapping it) asks the engine for a credential at the moment it needs one, not before.
Authenticating the secrets engine itself
Two named patterns cover most cases: AppRole-style auth, where a role_id (non-secret, can be baked into config) is paired with a secret_id (short-lived, delivered out of band) and exchanged for a token; and cloud-IAM-based auth, where the caller's existing cloud identity (an EC2 instance profile, a GCP service account, an Azure managed identity) is validated directly by the secrets engine without any bootstrap secret at all, for example Vault's aws auth method validates a signed STS GetCallerIdentity request from the instance's own credentials.
Bootstrapping brand-new hosts and CI runners
A host or CI runner that has never run before has no identity to present yet. Three concrete answers: deliver a short-lived, single-use wrapped AppRole secret_id via cloud-init/user-data at boot; use Kubernetes' projected service-account token with Vault's kubernetes auth method, so a pod authenticates using the token Kubernetes already gives it, no separate secret to distribute; or use cloud-IAM-based auth directly (an EC2 instance profile, or GitHub Actions OIDC federated into a cloud IAM role via AssumeRoleWithWebIdentity) so the runner's platform-issued identity is the credential.
provider "vault" {
address = var.vault_addr
auth_login {
path = "auth/approle/login"
parameters = {
role_id = var.approle_role_id
secret_id = var.approle_secret_id
}
}
}
# Kubernetes-native bootstrap: the pod's projected service-account token
# is the credential, nothing extra to distribute.
provider "vault" {
auth_login {
path = "auth/kubernetes/login"
parameters = {
role = "ci-runner"
jwt = file("/var/run/secrets/kubernetes.io/serviceaccount/token")
}
}
}
resource "vault_database_secret_backend_role" "app_role" {
backend = "database"
name = "my-app-role"
db_name = "my-db"
creation_statements = ["CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; GRANT SELECT ON app_table TO \"{{name}}\";"]
default_ttl = "1h"
max_ttl = "24h"
}
Extending the pattern beyond Terraform
The same idea applies to whatever converges the host after Terraform hands off: Ansible, Chef, or Puppet runs should also fetch secrets at converge time from the same engine using the same identity-based auth, rather than templating a static credential into a playbook variable or a Chef data bag. Terraform's role in that flow stops at "the engine and its access policy exist," it does not hand the config-management tool a credential to pass along.
Leasing and rotation, operationally
Every dynamic credential carries a default_ttl and max_ttl. Short-lived processes just request a fresh credential each run. Long-running processes (an app server holding a DB connection pool, a persistent CI agent) need to renew the lease before it expires, typically via Vault Agent running in auto-auth plus auto-renew mode, or the cloud secrets manager's built-in rotation Lambda on a schedule shorter than the credential's expiry.
The dual-side rotation problem
This is the case that trips people up: a database credential that is used both by a Terraform-managed resource (say, an aws_db_instance master password, or a Vault-issued credential a live application pool is actively holding open connections with) has to be rotated without dropping those connections. You cannot just swap the value in one step. Sequence it:
- Create the new credential alongside the old one, so both are valid simultaneously (a second DB user, or a fresh lease from the secrets engine).
- Push the new credential to the application's config/secret store and let the app pick it up via a rolling restart or a hot-reload of its connection pool, while the old credential still works for any connection that hasn't cycled yet.
- Confirm nothing is still authenticating with the old credential (check the database's connection/audit log for the old username, or the secrets engine's lease usage).
- Only then revoke or rotate out the old credential.
The fallback if it fails partway: if step 2 fails (the app can't reach the new secret, or the rolling restart stalls), do not revoke the old credential. Roll the app config back to the old credential and retry, keeping both valid until a rotation attempt actually completes end to end. Revoking the old credential before confirming cutover is what causes an outage, not the rotation itself.
Keeping Terraform out of the runtime fetch path
Avoid reading dynamic secrets with a Terraform data source at plan or apply time if you can help it; if you truly must, mark the value sensitive = true on any output or local (this only redacts CLI output, it does not encrypt state, so pair it with a remote encrypted backend and restricted ACLs). For long-running CI pipelines, re-authenticate before each terraform invocation with a fresh ephemeral token rather than caching one that could go stale or leak.
Trade-offs & pitfalls
Running a dynamic-secrets engine is real operational overhead (an HA service you now own or pay for) compared to a static secret in a vault-of-passwords tool; that overhead buys you automatic expiry and no long-lived value to leak. sensitive = true is frequently mistaken for encryption, it only hides a value from CLI output, the value is still sitting in state in the clear unless the backend itself encrypts it. The dual-side rotation sequence only works if the application can actually pick up a new credential without a full restart (hot-reload a connection pool); if it can't, budget for a rolling deploy as part of every rotation, not just a config push. Setting default_ttl too aggressively short for a high-traffic service creates a thundering-herd of renewal requests against the secrets engine.
You need one module to create a resource only when a feature flag is enabled, and also create one related object per item in a caller-provided list. How would you keep that configuration maintainable as the list grows or changes order over time?
Sample Answer
I would use count or for_each for the feature flag, but I would prefer for_each for the per-item objects. count is a simple on or off switch. for_each creates one instance per stable key, which is better when the list order changes.
Pattern
- For the feature-flagged singleton, create either one instance or none
- For the repeated objects, convert the caller’s list into a map keyed by a stable ID, such as name
- Avoid indexing directly into a list, because reordering
['api', 'worker']can cause unnecessary replacement
Example
If the caller passes ['api', 'worker'] today and ['worker', 'api'] tomorrow, keys like api and worker still point to the same resources. That keeps Terraform from churning objects just because the order changed.
Rule of thumb
Use count for a single optional resource, and for_each for anything that should survive list reordering. That makes the module much easier to maintain as the list grows.
Unlock Full Question Bank
Get access to all Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.