Infrastructure as Code and Automation Questions
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
A repo has one large Terraform entrypoint that manages several environments and shared pieces of infrastructure. The team is having merge conflicts and accidental cross-environment changes. How would you reorganize the code so engineers can work in parallel with less risk?
Sample Answer
I would separate shared logic from environment-specific entrypoints. Right now one large root module is forcing unrelated changes into the same file set, which creates merge conflicts and increases blast radius.
New structure
modules/networkmodules/appmodules/databaseenvs/devenvs/stagingenvs/prod
Each environment folder has its own backend and only calls the modules it needs. Shared values like tags or account IDs can live in small local files or a common variable file.
Why this helps
- Two engineers can edit dev and prod in parallel with less conflict
- A change to app compute does not force a network review unless inputs change
- Cross-environment mistakes are harder because each root has its own state
Example
Dev can use t3.small and prod can use m6i.large, but both call the same modules/app. That gives consistency without a monolith. I would also keep module outputs explicit so one environment cannot accidentally consume another environment’s state.
You're designing the secrets management layer for a company that runs infrastructure across multiple clouds and is subject to strict compliance requirements, PCI or SOC 2 for example. Walk me through how you'd architect secret storage, access control, and rotation so it satisfies both the engineers who need to use it day to day and the auditors who need to sign off on it.
Sample Answer
Direct answer
Architect it as a zero-trust broker, not a vault of static values: a central secrets engine (Vault or equivalent) issues short-lived, scoped credentials to workloads that authenticate via their platform identity (OIDC/workload identity), so nothing long-lived is stored anywhere and every access is logged at the point of issuance. Layer three things on top of that core for the compliance side: emergency revocation that can kill every active lease in one action, continuous reconciliation so drift gets corrected automatically rather than caught in a quarterly audit, and an audit pipeline that streams every access event to a SIEM so PCI/HIPAA evidence is continuous rather than assembled by hand before an audit.
Structured elaboration
Core architecture: broker, not vault
A central secrets engine (Vault, or a cloud-native equivalent) is the single source of truth for issuing credentials, not storing static ones. Workloads authenticate using their own platform identity, an OIDC token from the CI system, a Kubernetes service account, an IAM role, rather than a shared static token, and the broker exchanges that identity for a short-lived, narrowly-scoped credential: a database role valid for an hour, temporary cloud IAM credentials via STS or its equivalent. Cloud-native secret stores (AWS Secrets Manager, GCP Secret Manager, Azure Key Vault) sit alongside this for the minority of things that genuinely need to be long-lived, root/break-glass credentials, third-party API keys with no dynamic-issuance option, encrypted at rest and access-logged the same way.
Zero-trust framing with emergency revocation procedures
Zero-trust here means specifically: no credential is trusted just because it exists, every issuance is tied to a verified workload identity and a scope, and every lease has a TTL by default rather than an opt-in one. The piece teams often skip is emergency revocation: a single, tested, documented action that invalidates every active lease and credential issued under a given path or role, for the moment something is discovered compromised and can't wait for individual TTLs to expire. This needs to be a rehearsed procedure, a break-glass runbook with a named owner and a tested command, not a theoretical capability, because the first time anyone runs it shouldn't be during a live incident.
Continuous reconciliation via Kubernetes-style reconciler controllers
Rather than injecting a secret once at deploy time and hoping it stays correct, run the secret-delivery layer as a reconciler: a controller (the External Secrets Operator pattern, a Vault Agent sidecar, or an equivalent custom controller) that continuously watches the desired state, what secret a workload should have and its current rotation/expiry status, and reconciles the actual state to match it. That's the same control-loop model Kubernetes itself uses and Crossplane extends to cloud infrastructure. It catches drift, an app still holding a rotated-out credential because a static injection was never refreshed, automatically instead of relying on someone noticing a stale value.
PCI/HIPAA continuous-compliance pipeline with SIEM integration
Whichever regulatory driver applies, PCI DSS, HIPAA, or SOC 2, the controls overlap heavily: access logging, encryption at rest and in transit, rotation, and least privilege. Treat compliance evidence as a continuous pipeline rather than a point-in-time audit exercise:
- Every secret issuance, access, and revocation event streams to a SIEM (Splunk, Datadog, or equivalent) in near real time, not batched nightly.
- Automated policy checks (OPA/Conftest rules, or the secrets engine's own policy layer) continuously verify things auditors ask about anyway: rotation intervals within policy, no credential with a TTL longer than the compliance ceiling, no orphaned access grants.
- The audit evidence, access reports, rotation history, policy-check pass/fail history, is a standing artifact the pipeline produces continuously, so a PCI or HIPAA assessment is "here's months of continuous evidence" instead of a scramble to reconstruct history from raw logs the week before the audit.
What this looks like day to day
Engineers: a workload authenticates once, via its existing platform identity, no new credential to manage, and gets back exactly the scoped, short-lived access it needs; CI/Terraform runs authenticate the same way rather than carrying a long-lived key in a pipeline variable. Auditors: instead of interviewing engineers about their process, they get a standing report, every credential's issuance and access history, proof that TTLs and rotation intervals match policy, and a tested emergency-revocation runbook with an execution log from the last drill.
Worked example
flowchart LR
W[Workload / CI job] -->|OIDC identity| B[Secrets broker: Vault]
B -->|short-lived creds| W
B --> C[Cloud secret stores]
B --> R[Reconciler controller]
R -->|watches and corrects drift| W
B -->|every event| S[SIEM]
B -->|policy checks| P[OPA / Conftest]
O[On-call operator] -->|break-glass revoke| B
Emergency revocation, concretely: an operator with the break-glass role runs a single command that invalidates every lease issued under a compromised path, for example vault lease revoke -prefix database/creds/app-role/, which immediately kills every active credential issued from that role rather than waiting for individual TTLs to lapse.
Trade-offs & pitfalls
- A pure broker model, nothing long-lived, ever, is the right target but not every legacy system or third-party integration supports dynamic credentials; the cloud-native static stores exist precisely for that minority, and pretending everything can be dynamic just pushes the exception handling onto whoever hits the gap first in production.
- Continuous reconciliation adds a new component, the controller, that itself needs to be highly available and monitored; if the reconciler goes down silently, drift stops being caught even though nothing looks obviously broken.
- Streaming every access event to a SIEM in real time is a meaningful volume and cost commitment at scale; sample or aggregate low-value events (routine successful reads) while keeping full fidelity on anything policy-relevant (revocations, failed auth, privilege escalation attempts).
- An untested break-glass procedure is worse than none, since it creates false confidence; schedule the drill on a calendar, not "whenever we get to it."
Your CI pipeline needs to authenticate to your cloud provider to run Terraform, but you don't want a long-lived access key sitting in a repo secret or environment variable. How would you design that authentication so the pipeline gets valid credentials only when it actually needs them, and how would you make sure a pull request from an untrusted fork can't get at anything sensitive?
Sample Answer
Direct answer
Give the pipeline no long-lived key at all: configure your cloud provider's identity federation (AWS IAM OIDC provider, GCP Workload Identity Federation, Azure federated credentials) to trust GitHub's OIDC token issuer (OIDC: a standard that lets GitHub hand over a signed, short-lived proof-of-identity token instead of a stored secret), scope the trust policy to a specific repo/branch/environment claim, and have the pipeline exchange that short-lived, per-run token for temporary cloud credentials only at the moment it needs them; those credentials expire on their own, so there's nothing long-lived sitting in a repo secret to leak. Untrusted fork PRs are handled by GitHub itself before your trust policy even matters: a workflow triggered by the pull_request event on a fork does not receive repository secrets and does not get the id-token: write permission needed to mint an OIDC token, unless you deliberately opt in with pull_request_target, which should be avoided for anything that checks out and runs fork-provided code.
Design
IAM trust policy (OIDC federation)
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
},
"StringLike": {
"token.actions.githubusercontent.com:sub": "repo:your-org/your-repo:ref:refs/heads/main"
}
}
}
]
}
The sub condition pins which repo and branch (or environment) can assume this role, this is the actual boundary that matters, a role trusting repo:your-org/* for any branch is far weaker than one pinned to main or a protected environment.
Workflow configuration
name: terraform
on:
push:
branches: [main]
permissions:
id-token: write
contents: read
jobs:
apply:
runs-on: ubuntu-latest
environment: production
steps:
- uses: actions/checkout@v4
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::123456789012:role/terraform-ci
aws-region: us-east-1
- run: terraform init
- run: terraform apply -auto-approve tfplan
The official aws-actions/configure-aws-credentials action handles requesting the OIDC token and exchanging it for temporary AWS credentials; nothing gets echoed to a secret variable. Triggering only on push to main also means a fork PR can never reach this job in the first place, a separate, unprivileged plan-only workflow can run on pull_request for visibility, but should not carry id-token: write or any cloud credentials at all.
Fork PR isolation
A workflow triggered by pull_request from a fork runs with a read-only GITHUB_TOKEN and does not receive repository secrets or id-token: write, GitHub enforces this at the platform level regardless of what the fork's copy of the workflow file requests. The one way to defeat that protection is switching the trigger to pull_request_target, which runs with the base repository's permissions and secrets against the PR's code, without adding a required manual approval gate (a GitHub Environment with protection rules) that is a well-documented misconfiguration and exactly the setup to avoid here.
Least privilege and log hygiene
- Scope the IAM role's attached policy to only the actions and ARNs that environment's apply actually needs, and use a separate role per environment so a dev-scoped OIDC subject can never assume the prod role.
- Never enable verbose/debug SDK or Terraform logging (
TF_LOG=trace,set -x) in a step that touches plan or apply output, and if any custom script parsesterraform show -jsonto build a summary, apply the same sensitive-field redaction (after_sensitive/before_sensitive) needed anywhere else that output is consumed.
Auditing and rotating the CI identity itself
There's no static credential to rotate with OIDC, the whole point is that nothing long-lived exists, but the trust relationship itself is the credential now and needs periodic review: who can push to the branches or trigger the environments named in the sub condition, does the role's attached policy still match least privilege, and are AssumeRoleWithWebIdentity events in CloudTrail reviewed for anomalies (an unexpected source repo, an unusual time of day).
Worked example
An attacker forks the repo and opens a PR with a modified .tf file that would, if it had credentials, try to exfiltrate values. Because that PR triggers the pull_request event, not push, GitHub does not expose repository secrets to that run and does not grant id-token: write even though the forked copy of the workflow file might ask for it, the platform enforces this regardless of the fork's own YAML. The run gets a short-lived, read-only GITHUB_TOKEN and nothing else; there is no path in this setup for a fork PR to reach the terraform apply job at all, since that job only runs on push to main.
Trade-offs and pitfalls
- OIDC removes the leaked-long-lived-key failure mode but shifts the risk to the trust policy configuration; an overly broad
subcondition effectively grants the role to any branch matching the pattern, including a feature branch anyone can push to. - There's nothing to rotate on a schedule the way you would a static key, so the audit discipline has to shift to reviewing who can satisfy the trust condition and what the role's policy actually grants, that review is easy to let lapse precisely because there's no expiring secret forcing anyone to revisit it.
- Logs still need discipline: OIDC solves the credential-storage problem, not the "did we accidentally print a secret to a log or a PR comment" problem, that's a separate, ongoing concern.
A reusable module has to work across dev, staging, and prod, but each environment needs different sizes, tags, and resource names. How would you design the module interface so callers can customize it without editing the module code?
Sample Answer
I would make the module interface small, typed, and caller-driven. A module should accept inputs through variables and return only the outputs another team needs.
Inputs
envfor environment name likedev,staging,prodname_prefixfor naming consistencytagsas amap(string)so teams can add ownership and cost-center labels- sizing inputs, such as
instance_classorcpu_count - network inputs like
subnet_idsandvpc_id
Design rules
- Use defaults only for safe values, not for important architecture choices
- Add validation so bad values fail early, such as rejecting an empty subnet list
- Use stable names and tags inside the module, but let callers control the prefix
- Output IDs, ARNs, DNS names, or endpoints, not whole resources
Example
Dev can pass instance_class = "t3.small" and prod can pass instance_class = "m6i.large", while the module code stays unchanged. That lets each environment differ in size and naming without forking the module.
Before a production apply, what review and automation guardrails would you put around the Terraform workflow so an engineer can catch unexpected destroys or replacements before they reach users?
Sample Answer
I would put both human review and automation around the plan. A Terraform plan is a preview of create, update, destroy, and replace actions. Replace is especially risky because Terraform deletes and recreates a resource.
Guardrails
- Run
terraform fmt -checkandterraform validatein CI - Save the plan as an artifact and require approval before apply
- Fail the pipeline if the plan includes unexpected destroys or replacements on critical resources
- Use policy checks for things like public exposure, open security groups, or deletion of databases
- Add
prevent_destroyto the few resources that should almost never be removed
Example
If the plan shows aws_db_instance.main will be replaced because of an engine change, I would force manual review from an engineer and, ideally, a service owner. If the plan only adds an autoscaling instance, that can follow the normal path.
This catches surprises before users feel them, which is the real goal.
Unlock Full Question Bank
Get access to all 16 Infrastructure as Code and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.