Infrastructure as Code and GitOps Questions
Defining and managing infrastructure and delivery state declaratively: provisioning as code (Terraform, CloudFormation, Pulumi, Ansible, Puppet), configuration management, idempotency, drift detection and remediation, and version control for infrastructure definitions, extended by GitOps where git is the source of truth for deployment and infrastructure state. Covers keeping environments consistent, treating config as a first-class versioned artifact, pull-based deployment and continuous reconciliation toward the committed state (ArgoCD, Flux, and similar controllers), Kubernetes manifest and configuration delivery via git, secrets handling for IaC and GitOps pipelines, policy-as-code guardrails (OPA, Sentinel), Terraform state management and locking, and auditable change through version control: branching strategy, pull request review, commit conventions, and code review policy for infrastructure code. Distinct from the CI/CD pipeline design topic, which owns generic pipeline structure and platform-scale release orchestration (build, test, artifact publishing, runner mechanics) and the architectural choice between push-based CI/CD and pull-based GitOps, even when the payload is infrastructure code. Distinct from the safe deployment and rollback strategies topic, which owns deployment-strategy mechanics: canary and blue-green traffic shifting, automated rollback triggered by metrics or SLOs, feature-flag progressive delivery, Kubernetes rollout mechanics (maxSurge, maxUnavailable, health-check gating), and database or schema migration safety as it gates a release, even when the delivery mechanism is GitOps. Distinct from the automation and scripting topic, which owns operational-scripting disciplines (retry and backoff logic, CLI tool design, generic file, checksum, or diff utilities) when the task is not specifically about declarative infrastructure or configuration state. This topic keeps the GitOps reconciliation loop itself, drift detection and remediation, and IaC state and module lifecycle management regardless of which adjacent discipline a question also touches.
Explain what 'declarative manifests' means and why storing manifests and environment overlays in version control is considered the 'source of truth' in GitOps. Describe benefits for auditing, rollbacks, CI/CD automation, and one practical limitation or challenge teams face when adopting this approach.
Sample Answer
Direct answer
A declarative manifest specifies the DESIRED END STATE of a resource (how many replicas, which image, what configuration) and leaves it to a controller to figure out HOW to get there, in contrast to an imperative script that specifies the exact SEQUENCE OF STEPS to take. Storing those manifests, plus their environment-specific overlays, in version control makes Git the SOURCE OF TRUTH because Git then holds a complete, ordered, attributable history of every intended state the system has ever been in, and a reconciler can always compare "what Git says should be running" against "what is actually running" to detect and correct any difference, a capability that does not exist if desired state lives only in someone's head or in an imperative script's implicit side effects.
Structured elaboration
Why "declarative" and "source of truth" are linked, not two separate features. A declarative manifest is a snapshot of intended state that can be diffed, reviewed, and reconciled against; an imperative script's "state" is only ever the CUMULATIVE EFFECT of running it, which is not directly comparable to anything without re-running it or inspecting the live system by hand. Git-as-source-of-truth specifically REQUIRES the declarative model: you cannot usefully version-control "the sequence of commands someone ran against production last Tuesday" the way you can version-control "here is the desired Deployment spec," because only the latter is a stable, comparable artifact.
Benefits for auditing. Every change to desired state is a Git commit: who made it, when, what exactly changed (a real diff, not a description of an action), and (with PR-based workflow) who reviewed and approved it. This turns "what changed and why" from a question requiring log archaeology into a question git log and git blame answer directly.
Benefits for rollbacks. Because past states are just earlier commits, rolling back is conceptually git revert (or checking out a prior commit) plus letting the reconciler apply it, not manually reconstructing what the PREVIOUS configuration must have been from memory, monitoring dashboards, or scattered notes.
Benefits for CI/CD automation. A declarative desired-state file is something automation can generate, validate, diff, and apply mechanically (a pipeline can run policy checks against a manifest, produce a human-readable plan of what will change, and gate merges on that), none of which is straightforward for an imperative script whose effects are not knowable without executing it.
One practical limitation. Not everything maps cleanly to "desired state" the way a Deployment's replica count does: operations that are inherently ACTIONS rather than states (rotate this specific secret NOW, run a one-time data migration, manually promote a canary after watching metrics for an hour) resist pure declarative modeling and typically need an imperative escape hatch (a Job, a manually-triggered pipeline step, an operator-specific custom action) layered alongside the declarative core, which means real GitOps systems are rarely 100% declarative in practice, just declarative for the steady-state configuration that dominates day-to-day change.
Trade-offs and pitfalls
- Common mistake: treating "declarative" as meaning "no procedural logic anywhere in the system." The MANIFESTS are declarative; the CONTROLLER reconciling them is free to use arbitrarily complex procedural logic internally (retries, ordering, backoff) to get from current state to desired state. The declarative/imperative distinction is about what gets COMMITTED and DIFFED, not about banning procedural code from existing anywhere in the system.
- A team new to GitOps often struggles most with the "one practical limitation" above, expecting every operational action to have a clean declarative expression and getting frustrated when some genuinely don't; naming this limitation explicitly upfront, rather than discovering it mid-adoption, is worth doing as part of onboarding.
- Git-as-source-of-truth only holds if NOTHING legitimately bypasses it. A team that occasionally "just runs a quick kubectl edit for this one urgent thing" quietly reintroduces exactly the untracked-state problem GitOps exists to solve; the discipline (and the reconciler's drift-correction behavior, which will eventually revert an out-of-band edit) is what makes the source-of-truth claim actually true in practice, not just in principle.
- Auditability benefits assume the Git history itself is trustworthy (protected branches, required reviews, no force-push rewriting history); a Git repo with no branch protection gives the AUDIT-TRAIL APPEARANCE of Git-as-source-of-truth without the actual guarantee, since history itself could be silently altered.
Explain remote Terraform state backends and state locking mechanisms. Discuss best practices for state storage (encryption at rest), state locking (DynamoDB, etc.), access controls, and handling secrets stored in state. Provide an example backend choice for a large organization that spans multiple regions and teams.
Sample Answer
Direct answer
A remote Terraform state backend stores the state file OUTSIDE the local filesystem (object storage like S3, Azure Blob, or GCS, or Terraform Cloud's own managed state) so multiple engineers and CI jobs share one authoritative copy rather than each holding a local file that can silently diverge; state LOCKING (typically a separate mechanism, such as a DynamoDB table, that coordinates with the object-storage backend) prevents two concurrent apply operations from racing and corrupting that shared state. Both together are the minimum bar for ANY team of more than one person running Terraform against the same infrastructure.
Structured elaboration
Why remote state is necessary at all. A local state file is a single point of failure (lost with the laptop it lived on) and a single point of DIVERGENCE (two engineers' local copies drift the moment either one applies without the other pulling first); remote state fixes both by making one location authoritative and fetchable by everyone, including CI runners that have no persistent local disk between runs.
Encryption at rest. The backend's own storage-layer encryption (S3 server-side encryption with a customer-managed key, for example) protects the state file at rest, which matters specifically because Terraform state commonly contains SENSITIVE VALUES in plaintext (database passwords set via a resource argument, generated certificates, sometimes even provider credentials passed through as resource attributes); encryption at rest is a baseline control, not a substitute for also minimizing what sensitive data ends up in state in the first place.
State locking mechanisms. A lock is acquired before plan/apply begins and released after it completes (or on a timeout if the process dies uncleanly); DynamoDB-backed locking (historically the standard pattern for an S3 backend) uses a single table row per state file as the lock record, and a SECOND concurrent apply attempting to acquire the same lock blocks or fails immediately rather than racing the first one's writes. Without locking, two simultaneous applies can each read the same starting state, compute conflicting plans, and the SECOND one to write silently clobbers the first one's recorded changes even though both sets of real infrastructure changes actually happened, a state file that no longer matches reality in either direction.
Access controls. The backend's own IAM (identity and access management) policy should scope WHO can read and write each state file independently of who can run terraform apply against the corresponding infrastructure; state read access alone is a meaningful privilege (state frequently contains secrets, as above), so a CI runner or engineer needing to only READ outputs for a downstream reference should not automatically also have WRITE access to the state file itself.
Handling secrets in state. Because Terraform state routinely captures resource arguments and computed attributes verbatim, including ones a provider marks sensitive at the SCHEMA level but which still land in the raw state file, treat the state file itself as a secret-bearing artifact requiring the same access discipline as a credentials store, not merely as "infrastructure bookkeeping." Where possible, avoid resources that require Terraform to generate or hold a long-lived secret directly (prefer resources that reference an externally-managed secret by ARN or path over ones where Terraform itself sets the secret value).
Worked example
Backend choice for a large organization spanning multiple regions and teams: per-team, per-environment state files (NOT one giant shared state), each in its own S3 key under a consistent naming convention (s3://org-tfstate/<team>/<environment>/<component>/terraform.tfstate), server-side encryption enabled on the bucket with a KMS (key management service) customer-managed key scoped so only the relevant team's CI role and a break-glass admin role can decrypt, DynamoDB table for locking with one row per state key, and bucket-policy-level access control restricting each team's CI role to read/write only under its own prefix. This gives every team an independently locked, independently access-controlled state file while keeping operational overhead centralized (one bucket, one lock table, one encryption key management process) rather than each team standing up its own backend infrastructure from scratch.
Trade-offs and pitfalls
- Common mistake: treating a single monolithic state file as simpler to manage, when it is actually a bigger operational and blast-radius risk as a team grows. Every
applyagainst a shared monolithic state locks out every OTHER team's applies for its duration, and a mistake in one team's change carries the same state-corruption risk for everyone else's resources tracked in that same file; splitting state per team/environment/component trades a bit of cross-references complexity for real isolation. - DynamoDB-based locking is a widely used PATTERN for S3 backends specifically, not a universal Terraform requirement, some backends (Terraform Cloud's managed state, some newer S3-native locking mechanisms) provide locking built in without a separate lock table; the underlying REQUIREMENT (mutual exclusion during writes) is universal, the specific mechanism is backend-dependent.
- Encrypting the state file at rest does nothing to protect it from someone with legitimate READ access to the backend. State access control (who can read the bucket/key) is the control that actually limits secret exposure day to day; encryption at rest primarily protects against a different threat (the underlying storage medium itself being compromised or exfiltrated), and conflating the two leaves the more common risk (an over-broadly-granted read role) unaddressed.
- Common mistake: assuming a resource marked "sensitive" in a Terraform provider schema means the value never appears anywhere in plaintext. The
sensitiveattribute only suppresses the value from CLI OUTPUT and logs; it is still stored in the raw state file exactly as before, sensitive-attribute masking is a display-layer control, not a state-file-encryption or access-control substitute.
Explain how you would use Git for a simple infrastructure change workflow. Cover steps from creating a branch, making changes to infrastructure-as-code (for example Terraform), committing, pushing, opening a pull request, running CI validations, and merging. Include example git commands and explain any differences compared to an application code change.
Example commands you can reference:
git checkout -b feat/update-vpc
git add .
git commit -m "feat(vpc): increase cidr range"
git push origin feat/update-vpc
Sample Answer
Direct answer
The workflow itself (branch, edit, commit, push, open a pull request, wait for CI, merge) is IDENTICAL in mechanics to an application-code change; what differs for infrastructure-as-code (IaC) is what the required CI validation actually checks (a terraform plan showing the real-world effect of the change, not just a test suite) and the weight given to review, since a merged infra change can directly reconfigure live, shared resources the moment it applies, with no separate build-and-deploy step standing between merge and real-world effect the way an application change usually has.
Structured elaboration
Step-by-step, using the question's own example commands as the backbone.
git checkout -b feat/update-vpc
Start from an up-to-date base branch and create a descriptively named feature branch; naming it after the CHANGE ("update-vpc") rather than a ticket number alone helps a reviewer understand intent from the branch list without opening the PR.
Edit the Terraform configuration (for example, widening a VPC's CIDR range in the relevant .tf file).
git add .
git commit -m "feat(vpc): increase cidr range"
A commit message following this project's convention: a concise subject describing WHAT and, ideally, enough context to infer why, referencing a ticket if one exists.
git push origin feat/update-vpc
Push the branch and open a pull request describing the intended change and its expected blast radius.
Running CI validations, the infra-specific step. The pipeline runs terraform fmt -check and terraform validate (syntax and internal consistency), then terraform plan, and posts the PLAN OUTPUT directly on the PR as a comment or check annotation, so the reviewer sees exactly what will change in the real infrastructure (which resources get created, modified in place, or destroyed and recreated) alongside the code diff, not just the code diff alone.
Review. The reviewer reads both the CODE diff and the PLAN output together; particular attention goes to any resource marked for DESTRUCTION or REPLACEMENT in the plan, since those carry real operational risk a pure code-review pass could miss if the reviewer only read the .tf diff without checking what Terraform actually intends to do with it.
Merging. Once approved and CI is green, merge; depending on the pipeline design, the actual apply may run automatically post-merge, or require a SEPARATE manual approval gate before applying, distinct from the PR's own code-review approval.
Worked example
The exact commands from the question, extended with the infra-specific CI step made explicit:
git checkout -b feat/update-vpc
# ... edit network/vpc.tf, widening the CIDR block ...
git add .
git commit -m "feat(vpc): increase cidr range"
git push origin feat/update-vpc
# open PR; CI runs:
# terraform fmt -check
# terraform validate
# terraform plan -out=tfplan (plan output posted to the PR)
# reviewer reads the code diff AND the posted plan output together
# on approval + green CI: merge
# post-merge: terraform apply (auto or gated by a separate approval, per pipeline design)
Trade-offs and pitfalls
- The single biggest practical difference from an application-code change: the CODE DIFF alone does not tell a reviewer what will actually happen. A one-line change to a resource argument can trigger an in-place update OR a destroy-and-recreate depending on that specific argument's
ForceNewbehavior in the provider; reviewing the diff without reviewing the PLAN is reviewing the WRONG artifact for infrastructure changes specifically. - Common mistake: treating the plan-output-on-PR step as informational only, not something the reviewer is expected to actually read. If nobody reads the posted plan, the extra CI step provides no real safety benefit over reviewing application code the normal way; the workflow only earns its infra-specific rigor if the plan output is genuinely part of what gets reviewed, not merely generated.
- Merge and apply being separate steps (rather than merge automatically applying) is itself a real design decision with a trade-off, immediate auto-apply on merge keeps the mental model simple (merged means live) but removes a final safety check between "this was approved" and "this is now actually changing production"; a separate, explicit apply gate adds a step but catches the case where the plan changed between review time and merge time (another change merged in between) before it silently applies.
- Common mistake: naming the branch or commit purely after a ticket number with no description of the actual change ("fix/JIRA-4521"); this workflow's whole value depends on Git history being genuinely READABLE later, and a ticket-only name defeats that the moment the ticket system itself becomes hard to search or is retired.
Design a secure process for rotating secrets (for example database credentials or API keys) that are consumed by GitOps pipelines and CI runners. Consider secret propagation, minimizing downtime, backward compatibility, validation, and auditability. Show how Git workflows and CI/CD jobs coordinate staging, verification, and cutover of rotated credentials.
Sample Answer
Direct answer
Secret rotation for GitOps pipelines and CI runners has to solve one core sequencing problem: the CONSUMERS of a secret (running pods, CI jobs) and the SOURCE of truth for that secret (Git-declared reference, or the secrets backend) are not updated atomically, so a safe rotation is a staged process, provision the NEW credential alongside the old one, verify consumers can use the new one, THEN retire the old one, never a single in-place swap that has no working-backward path if something goes wrong mid-rotation.
Structured elaboration
Secret propagation. For GitOps-managed workloads, the secret reference in Git (an External Secrets Operator path, or a Sealed-Secrets/SOPS-encrypted blob) points at the secrets backend, NOT at a specific credential value baked into a manifest; rotating the underlying value in the backend (Vault, or a cloud key management service) and letting the existing sync mechanism pick it up on its normal interval means most of the propagation work reuses infrastructure that already exists, rather than requiring a special rotation-specific pipeline. For CI runners, credentials are typically injected at job-start time from the same backend via short-lived, scoped tokens rather than a long-lived static secret baked into the runner's own configuration, so "rotation" for CI often means simply updating the backend, the NEXT job invocation picks up the new value automatically, with no separate propagation step needed at all.
Minimizing downtime, the dual-write window. For any credential a running system holds onto (a database password cached in an application's connection pool, for instance), rotation needs a window where BOTH the old and new credentials are simultaneously valid at the source system (the database itself accepts either), so already-running instances using the cached old credential keep working while new connections and restarted instances pick up the new one; only once every consumer has confirmed to be using the new credential does the old one get revoked. Skipping this dual-validity window (revoking the old credential the instant the new one is issued) causes an outage for any consumer that has not yet picked up the change, exactly the disruption the staged approach exists to avoid.
Backward compatibility. The rotation process needs an explicit way to detect "has everything actually cut over yet" before revoking the old credential, connection-attempt metrics or audit logs at the source system showing which credential is actually being used, not just an assumption that enough time has passed. A rotation that revokes on a fixed timer without checking actual cutover status risks revoking while a slow-to-restart consumer (a long-lived batch job, a rarely-redeployed service) is still mid-transition.
Validation. Before the new credential is trusted as the sole valid one, verify it works END TO END, not just "the API call to create it succeeded": a CI job (or a scheduled canary check) that authenticates USING the new credential against the actual target system, confirming both that the value itself is correct and that any downstream permission/scope configuration for it is right, catches a broad class of rotation failures (a typo in a propagated value, a missing grant on the new credential) before any real consumer depends on it.
Auditability. Every step, issuance of the new credential, each system's confirmed cutover, and revocation of the old one, needs its own record: the ISSUANCE and REVOCATION happen at the secrets backend (which should log both natively), while the ROTATION DECISION and its timing should be a Git-tracked event (a commit bumping which secret VERSION or path a manifest references, if the backend supports versioned secrets) so "when did this rotate and who/what triggered it" is answerable from the same Git-commit audit trail that infrastructure changes generally rely on.
Worked example
Rotating a shared database credential consumed by both a GitOps-deployed application and a nightly CI job:
- Provision. Create a new database user/password (or a new version of the existing secret, if the backend supports versioning) in the secrets backend. The OLD credential remains fully valid.
- Propagate to staging first. Update staging's secret reference (a Git commit bumping the referenced secret version, or, for a backend with live sync, updating the backend value for staging's scoped path) and let the existing External Secrets Operator sync pick it up. Run the CI job's canary authentication check against staging using the new credential.
- Propagate to production, staged by consumer. Bump production's secret reference the same way. Running application pods that were already using the OLD credential in an active connection continue working (the database still accepts it); NEWLY started or restarted pods pick up the new value. CI's next scheduled run automatically uses the new value since it fetches fresh from the backend at job start, no separate CI-specific propagation step required.
- Confirm cutover. Check the database's own connection audit log for any connections still authenticating with the OLD credential's identifier; once none remain for a defined observation window (long enough to cover the slowest-restarting consumer, for example one full pod-restart cycle plus one full CI schedule interval), proceed.
- Revoke. Revoke the old credential at the database. The Git commit that bumped the reference in step 3, plus the backend's own issuance/revocation logs, together form the audit trail: who/what triggered the rotation, when each environment cut over, and when the old credential was finally revoked.
Trade-offs and pitfalls
- Common mistake: rotating directly in production first "to save time," skipping the staging validation step. The whole point of validating in staging first is that a broken new credential (wrong permissions, a typo) fails there instead of during a production cutover, where the same mistake means an active incident rather than a caught bug.
- Common mistake: treating "the new value was successfully written to the secrets backend" as equivalent to "rotation complete." Nothing about a successful WRITE confirms any consumer has actually picked it up or that it authenticates correctly against the real target system; the validation and cutover-confirmation steps are not optional formalities, they are what actually closes the loop.
- A fixed rotation-to-revocation timer is a common, reasonable-looking shortcut that breaks for exactly the consumers that matter most: long-lived connections and infrequently-redeployed batch jobs are precisely the ones a timer-based approach is most likely to disrupt, since they are the slowest to naturally pick up the new credential; checking actual cutover status (not just elapsed time) before revoking is the fix.
- CI runners that cache a credential for the DURATION of a single job (rather than re-fetching per step) can still be using a stale value if rotation happens mid-job, even though CI jobs are usually short-lived; a job that's mid-run when rotation occurs should either complete using its already-fetched credential (fine, if the dual-validity window covers it) or, for very long-running jobs, re-fetch periodically rather than caching for the entire run.
Discuss strategies for managing large artifacts and state files that are part of infrastructure workflows (for example Terraform state snapshots, PKI binaries, or large AMI artifacts) without storing them directly in Git. Compare Git LFS, artifact repositories (Artifactory, S3), remote backends, and explain how your chosen approach integrates into PR reviews, CI, and access controls.
Sample Answer
Direct answer
Terraform state snapshots, PKI (public key infrastructure) binaries, and large AMI (Amazon Machine Image) artifacts share one property that makes storing them directly in Git a bad fit: Git's storage model is optimized for TEXT that diffs well and history that stays small, while these are large, frequently-regenerated, mostly-BINARY blobs that Git would store as a full new copy on every change, bloating every clone forever. The right home depends on what the artifact actually needs: Git LFS (Large File Storage) for files that genuinely need Git-native versioning and PR-based review despite being binary; a dedicated artifact repository (Artifactory, or plain object storage like S3) for anything that is really a BUILD OUTPUT with its own release lifecycle; and Terraform's own remote backend, never Git at all, for state specifically, since state's locking, current-value, and sensitive-data properties do not map onto Git's model at all.
Structured elaboration
Git LFS. Stores large files as pointers in Git (small text references) while the actual bytes live in a separate LFS-aware storage backend; Git operations (clone, diff, history) stay fast because the large content is fetched separately and only on demand. This fits files that genuinely benefit from being tied to a specific commit and reviewed via a normal PR flow (a PKI root certificate bundle that changes rarely and where "which commit introduced this exact cert" matters), but LFS storage and bandwidth typically cost more per gigabyte than plain object storage, and LFS is a worse fit for anything that changes FREQUENTLY (every LFS-tracked version is still a distinct stored blob, so frequent large-file churn is expensive regardless of the pointer trick).
Artifact repositories (Artifactory, S3 as a plain artifact store). Purpose-built for exactly this: large binary build outputs, versioned, with lifecycle policies (retention, promotion between repositories for dev/staging/prod), typically far cheaper per gigabyte than Git LFS, and with retrieval APIs designed for CI/CD consumption (pull the specific version a pipeline needs, not the whole history). This is the right home for large AMI artifacts specifically, an AMI has its OWN natural versioning and promotion lifecycle (built once, promoted through environments, eventually deprecated) that maps directly onto an artifact repository's model, and gains nothing from Git's commit-and-diff semantics since an AMI is not something anyone reviews as a text diff.
Remote backends, for Terraform state specifically. State is fundamentally different from the other two categories: it needs LOCKING (to prevent concurrent-write corruption), it represents CURRENT truth (not a historical build artifact with old versions kept around for their own sake), and it routinely contains SENSITIVE VALUES. None of Git, Git LFS, or a generic artifact repository provide locking as a first-class concept; Terraform's own remote backends (S3+DynamoDB, Terraform Cloud) are purpose-built for exactly this and are the only appropriate home, state should never be committed to Git in any form, LFS-tracked or otherwise.
Worked example
A concrete assignment for a mid-size infrastructure org:
| Artifact | Home | Why |
|---|---|---|
| Terraform state files | Remote backend (S3 + DynamoDB, or Terraform Cloud) | Needs locking, represents current truth, contains sensitive values |
| PKI root/intermediate certificate bundles (changed rarely, security-reviewed per change) | Git LFS | Benefits from PR-based review and exact commit-to-cert traceability; low change frequency keeps LFS cost reasonable |
| Large AMI artifacts (built frequently by CI, promoted through environments) | Artifact repository (or a cloud provider's own AMI/image registry) | Has its own build-and-promotion lifecycle; no benefit from Git diff/review semantics; cheaper at this volume and change frequency |
| Terraform provider plugin binaries, module tarballs | Artifact repository / private module registry | Purpose-built versioned distribution already exists (a Terraform module registry) and should be used instead of ad hoc storage |
Integration into PR reviews, CI, and access controls. For the LFS-tracked PKI bundles, a PR reviewer sees the LFS pointer diff and, critically, needs LFS-aware tooling to actually inspect the real content change (a plain git diff on an LFS pointer shows only the pointer hash changing, not the certificate content itself, so the review process needs an explicit step, a CI job that renders and posts the actual diff, to make review meaningful rather than rubber-stamping a hash change). For AMI artifacts in the artifact repository, CI publishes a new version on build and the PROMOTION step (dev to staging to prod) is a separate, reviewed action referencing that specific version by its immutable identifier. Access control differs meaningfully by artifact type: PKI material needs the tightest read scoping of anything on this list (arguably tighter than most application secrets), while AMI artifacts typically need broad READ access (anything provisioning from them) but narrow WRITE access (only the build pipeline that produces them).
Trade-offs and pitfalls
- Common mistake: committing Terraform state, even encrypted, to Git or Git LFS "for convenience," because it is technically just a file. This is a recurring, real mistake distinct from the general large-artifact question; even encrypted, Git provides no locking, so this reintroduces the exact concurrent-write-corruption risk remote backends specifically exist to solve, an LFS-stored state file is still Git-adjacent enough to tempt someone into a normal
git commitworkflow for it, defeating the whole point. - Git LFS's per-gigabyte cost, multiplied by CHANGE FREQUENCY, is the single most common reason a team regrets choosing it for something that turned out to churn more than expected. A file that seemed like a natural LFS candidate at adoption time (infrequent changes) can quietly become expensive if its actual change cadence increases later; periodically reviewing what is tracked in LFS against its actual current churn rate catches this before the cost surprise does.
- A plain
git diffon an LFS pointer is actively misleading if reviewers are not aware LFS is in use, it shows a hash changing, which LOOKS like a trivial change and can get rubber-stamped, while hiding a potentially significant underlying content change; any team adopting LFS needs an explicit review-tooling step (or at minimum, clear reviewer training) to close this gap. - Artifact repository retention/lifecycle policies need to reconcile with any compliance-driven minimum-retention requirement: cost-optimized automatic cleanup can silently violate a longer regulatory retention window if the two policies are not designed together.
Unlock Full Question Bank
Get access to all Infrastructure as Code and GitOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.