Infrastructure as Code and GitOps Questions
Defining and managing infrastructure and delivery state declaratively: provisioning as code (Terraform, CloudFormation, Pulumi, Ansible, Puppet), configuration management, idempotency, drift detection and remediation, and version control for infrastructure definitions, extended by GitOps where git is the source of truth for deployment and infrastructure state. Covers keeping environments consistent, treating config as a first-class versioned artifact, pull-based deployment and continuous reconciliation toward the committed state (ArgoCD, Flux, and similar controllers), Kubernetes manifest and configuration delivery via git, secrets handling for IaC and GitOps pipelines, policy-as-code guardrails (OPA, Sentinel), Terraform state management and locking, and auditable change through version control: branching strategy, pull request review, commit conventions, and code review policy for infrastructure code. Distinct from the CI/CD pipeline design topic, which owns generic pipeline structure and platform-scale release orchestration (build, test, artifact publishing, runner mechanics) and the architectural choice between push-based CI/CD and pull-based GitOps, even when the payload is infrastructure code. Distinct from the safe deployment and rollback strategies topic, which owns deployment-strategy mechanics: canary and blue-green traffic shifting, automated rollback triggered by metrics or SLOs, feature-flag progressive delivery, Kubernetes rollout mechanics (maxSurge, maxUnavailable, health-check gating), and database or schema migration safety as it gates a release, even when the delivery mechanism is GitOps. Distinct from the automation and scripting topic, which owns operational-scripting disciplines (retry and backoff logic, CLI tool design, generic file, checksum, or diff utilities) when the task is not specifically about declarative infrastructure or configuration state. This topic keeps the GitOps reconciliation loop itself, drift detection and remediation, and IaC state and module lifecycle management regardless of which adjacent discipline a question also touches.
Explain how reconciliation loops work in GitOps controllers (for example Argo CD and Flux). Discuss scaling challenges when managing thousands of manifests across many clusters, including informer/watch scaling, reconcile frequency, rate-limiting, caching strategies, and design decisions to reduce API-server and controller load.
Sample Answer
Direct answer
A GitOps controller's reconciliation loop is fundamentally: watch a source (Git, or the cluster via informers) for changes, compute a diff between desired and live state, and apply the difference, repeated continuously. At small scale this is simple; at THOUSANDS of manifests across MANY clusters, the loop's naive form breaks down along four specific axes, informer/watch scaling (how many objects a single controller can efficiently watch), reconcile frequency (how often it re-checks everything), rate-limiting (how fast it can safely apply changes without overwhelming the API server), and caching (how much redundant API-server work each reconcile avoids), and a controller designed for scale has to address all four, not just add more compute.
Structured elaboration
Informer/watch scaling. Kubernetes controllers typically use informers (a local, continuously-updated cache backed by a watch connection to the API server) rather than polling list calls, so each reconcile reads from the LOCAL cache, not the API server directly. At thousands of manifests across many clusters, the watch connections themselves (one per resource type per cluster, potentially) and the memory footprint of the local caches become the bottleneck; a controller architecture built for scale shards this work (per-cluster or per-namespace informer instances, only watching resource kinds actually in use) rather than running one giant informer set watching everything everywhere.
Reconcile frequency. A fixed, short polling interval applied uniformly to every application (checking every 30 seconds regardless of how often that application's config actually changes) wastes API-server and controller CPU on applications that rarely change while still being too slow for applications that need faster drift detection. Event-driven reconciliation (triggering a reconcile on a Git webhook or a Kubernetes watch event, rather than a fixed poll interval) plus EXPONENTIAL BACKOFF for applications that are healthy and unchanged (progressively lengthening the interval between checks when nothing has changed recently, resetting to a fast interval the moment something DOES change) gets fast reaction time where it matters without constant unnecessary work everywhere.
Rate-limiting. Applying many changes simultaneously across a large fleet risks overwhelming the target API server(s) with a burst of writes, especially right after a controller restart when every application needs to be re-evaluated at once (a "thundering herd"). Controllers built for scale use client-side rate limiting (a token-bucket or leaky-bucket limiter on outbound API calls per cluster) and STAGGER the initial reconcile pass after startup (jittered delays rather than firing every application's first reconcile in the same instant) specifically to avoid this.
Caching strategies. Beyond the informer's live-object cache, a scaled controller caches EXPENSIVE derived computations, rendered manifests (re-running a Helm template or Kustomize build on every single reconcile pass is wasteful if the underlying chart/base hasn't changed), diff results, and resolved Git commit SHAs (avoiding a fresh Git fetch on every poll when polling for changes rather than using webhooks), invalidating each cache layer only when its actual input changes.
Trade-offs and pitfalls
- Common mistake: treating "add more controller replicas" as the primary scaling lever. Horizontal scaling helps only if the WORK is genuinely partitionable (sharded by cluster or namespace, with each replica owning a disjoint slice); naively running multiple replicas all reconciling the SAME full set of applications multiplies API-server load without increasing effective throughput, and risks duplicate/conflicting apply operations unless there is a leader-election or sharding scheme.
- Event-driven reconciliation reduces average load but does NOT eliminate the need for a periodic full reconcile. Webhooks and watch events can be missed (a delivery failure, a network partition during the exact window a change happened), so a scaled system still needs an occasional FULL reconciliation pass (hourly, or some longer interval) as a correctness backstop, purely event-driven with no periodic fallback risks silent, permanent drift if even one event is ever lost.
- Aggressive caching trades staleness risk for reduced load, and that trade-off needs an explicit, bounded staleness budget, not an open-ended "cache until something tells us to invalidate," since a caching bug that fails to invalidate correctly can leave a controller confidently reconciling against a manifest render that is silently out of date.
- Rate-limiting protects the API SERVER, but does not by itself protect against a controller being overwhelmed by its OWN internal queue if events arrive faster than they can be rate-limited out; a bounded work queue with backpressure (dropping or coalescing redundant re-reconcile requests for the same object rather than queuing every single trigger separately) is a necessary complement to outbound rate-limiting, not a substitute for it.
Design an infrastructure codebase to support multiple tenants (teams or customers) that require isolation and shared services. Discuss repository layout, use of modules, parameterization, environment management, access controls, onboarding flow for new tenants, and how to use Git constructs (branches, repos, codeowners) to enforce tenancy boundaries.
Sample Answer
Direct answer
The right unit of reuse for a multi-tenant infrastructure codebase is the MODULE, not the repository or the environment: shared platform modules (network, IAM baseline, common services) are authored once and consumed by parameterized per-tenant instantiations, while tenancy BOUNDARIES (who can change what for which tenant) are enforced through Git constructs, CODEOWNERS scoping and branch protection on a per-tenant DIRECTORY structure, rather than through application-layer access control alone. This is a distinct concern from the runtime GitOps question of how a controller reconciles multiple tenant CLUSTERS; this question is about the CODEBASE and REPOSITORY structure that PRODUCES each tenant's infrastructure in the first place.
Structured elaboration
Repository layout. A single repo (shared platform modules genuinely benefit from atomic, coordinated updates) with a clear top-level split: modules/ for shared, versioned, parameterized building blocks, and tenants/<tenant-id>/ for each tenant's specific instantiation of those modules with its own parameter values.
Use of modules and parameterization. Each shared module exposes a deliberately MINIMAL parameter surface (tenant name/ID, environment tier, any genuinely tenant-specific sizing or feature flags), everything else is fixed inside the module so a tenant's own directory cannot accidentally diverge from the platform's baseline security or architecture posture; a tenant directory is then just a short module-invocation file passing that tenant's specific parameter values, not a full copy of the module's internals.
Environment management. Each tenant's directory further splits by environment (tenants/<tenant-id>/prod/, /staging/), following the same overlay-and-promotion pattern used for multi-environment GitOps generally, so a tenant's prod and staging environments are independently reconciled and independently promotable.
Access controls. Two layers, matching how the modules-versus-tenants split above works: shared module changes require PLATFORM-team review (CODEOWNERS on modules/**), since a change here can affect every tenant at once; a tenant's own parameter file changes require review from whoever owns that tenant relationship (could be the platform team for internal tenants, or a delegated team for larger external/enterprise tenants with their own dedicated engineering contact), scoped via CODEOWNERS on tenants/<tenant-id>/** specifically, so tenant A's team cannot approve (or, depending on repo permissions, even SEE) tenant B's parameter changes if isolation requires that level of separation.
Onboarding flow for new tenants. A SCRIPTED or TEMPLATED process, not a hand-copied directory: a generator (a small CLI, or a documented terraform scaffold with placeholder substitution) creates a new tenants/<new-tenant-id>/ directory from a template, pre-populated with sane defaults and the CODEOWNERS entry for the new tenant automatically added, opened as a single PR for platform-team review; this is what keeps onboarding tenant number 50 exactly as consistent and low-error as onboarding tenant number 2, rather than degrading as ad hoc copy-paste drift accumulates.
Using Git constructs to enforce tenancy boundaries. Branch protection with path-scoped CODEOWNERS is the PRIMARY mechanism (as above); for organizations with a genuinely hard isolation requirement between specific tenants (regulatory separation, a competitive-conflict pairing of customers), a full separate REPO per that specific tenant, rather than a directory in the shared repo, is the appropriate exception, accepting the coordination cost of losing atomic, cross-tenant module updates in exchange for repo-level (not merely path-level) access isolation.
Worked example
infra-repo/
modules/
network/ # CODEOWNERS: @org/platform-team
iam-baseline/ # CODEOWNERS: @org/platform-team
tenants/
acme-corp/
prod/main.tf # module "network" { source = "../../../modules/network" tenant_id = "acme-corp" ... }
staging/main.tf
CODEOWNERS-scoped: @org/platform-team, @acme-account-team
globex-inc/
prod/main.tf
staging/main.tf
CODEOWNERS-scoped: @org/platform-team, @globex-account-team
A change to modules/network/ requires platform-team approval and, because the module is shared, triggers a validation pass checking it does not break any CONSUMING tenant's plan. A change to tenants/acme-corp/prod/main.tf requires only the platform team and Acme's own account team, never Globex's, and Globex's team has no approval authority (and, if the repo's permission model supports directory-scoped visibility, potentially no visibility) into Acme's parameters.
Trade-offs and pitfalls
- Common mistake: letting a tenant's directory drift into a full copy of the module's internals "just this once" for a special requirement, rather than extending the shared module's parameter surface (or, if genuinely a one-off, explicitly forking with a documented reason). Once one tenant directory diverges from pure parameterization, every subsequent tenant with a similar special need copies the same pattern, and the module's promised consistency erodes tenant by tenant.
- The onboarding template needs an owner and periodic review, an unmaintained scaffold generates new tenants against an increasingly outdated baseline (missing a security control added to the platform since the template was last updated), silently under-serving every tenant onboarded after that gap opened.
- Path-scoped CODEOWNERS visibility depends on the Git hosting platform's actual capabilities, some platforms only scope APPROVAL rights via CODEOWNERS while still leaving the whole repo readable to anyone with repo access; if genuine READ isolation between tenants is required, this needs to be verified against the specific platform's actual permission model, not assumed from CODEOWNERS alone.
- Full separate repos for hard-isolation tenants should be the deliberate EXCEPTION, not the default, defaulting every tenant to its own repo forfeits the atomic, coordinated shared-module updates that are the main structural advantage of the monorepo-plus-modules design in the first place, for isolation most tenants do not actually require.
Design a secure process for rotating secrets (for example database credentials or API keys) that are consumed by GitOps pipelines and CI runners. Consider secret propagation, minimizing downtime, backward compatibility, validation, and auditability. Show how Git workflows and CI/CD jobs coordinate staging, verification, and cutover of rotated credentials.
Sample Answer
Direct answer
Secret rotation for GitOps pipelines and CI runners has to solve one core sequencing problem: the CONSUMERS of a secret (running pods, CI jobs) and the SOURCE of truth for that secret (Git-declared reference, or the secrets backend) are not updated atomically, so a safe rotation is a staged process, provision the NEW credential alongside the old one, verify consumers can use the new one, THEN retire the old one, never a single in-place swap that has no working-backward path if something goes wrong mid-rotation.
Structured elaboration
Secret propagation. For GitOps-managed workloads, the secret reference in Git (an External Secrets Operator path, or a Sealed-Secrets/SOPS-encrypted blob) points at the secrets backend, NOT at a specific credential value baked into a manifest; rotating the underlying value in the backend (Vault, or a cloud key management service) and letting the existing sync mechanism pick it up on its normal interval means most of the propagation work reuses infrastructure that already exists, rather than requiring a special rotation-specific pipeline. For CI runners, credentials are typically injected at job-start time from the same backend via short-lived, scoped tokens rather than a long-lived static secret baked into the runner's own configuration, so "rotation" for CI often means simply updating the backend, the NEXT job invocation picks up the new value automatically, with no separate propagation step needed at all.
Minimizing downtime, the dual-write window. For any credential a running system holds onto (a database password cached in an application's connection pool, for instance), rotation needs a window where BOTH the old and new credentials are simultaneously valid at the source system (the database itself accepts either), so already-running instances using the cached old credential keep working while new connections and restarted instances pick up the new one; only once every consumer has confirmed to be using the new credential does the old one get revoked. Skipping this dual-validity window (revoking the old credential the instant the new one is issued) causes an outage for any consumer that has not yet picked up the change, exactly the disruption the staged approach exists to avoid.
Backward compatibility. The rotation process needs an explicit way to detect "has everything actually cut over yet" before revoking the old credential, connection-attempt metrics or audit logs at the source system showing which credential is actually being used, not just an assumption that enough time has passed. A rotation that revokes on a fixed timer without checking actual cutover status risks revoking while a slow-to-restart consumer (a long-lived batch job, a rarely-redeployed service) is still mid-transition.
Validation. Before the new credential is trusted as the sole valid one, verify it works END TO END, not just "the API call to create it succeeded": a CI job (or a scheduled canary check) that authenticates USING the new credential against the actual target system, confirming both that the value itself is correct and that any downstream permission/scope configuration for it is right, catches a broad class of rotation failures (a typo in a propagated value, a missing grant on the new credential) before any real consumer depends on it.
Auditability. Every step, issuance of the new credential, each system's confirmed cutover, and revocation of the old one, needs its own record: the ISSUANCE and REVOCATION happen at the secrets backend (which should log both natively), while the ROTATION DECISION and its timing should be a Git-tracked event (a commit bumping which secret VERSION or path a manifest references, if the backend supports versioned secrets) so "when did this rotate and who/what triggered it" is answerable from the same Git-commit audit trail that infrastructure changes generally rely on.
Worked example
Rotating a shared database credential consumed by both a GitOps-deployed application and a nightly CI job:
- Provision. Create a new database user/password (or a new version of the existing secret, if the backend supports versioning) in the secrets backend. The OLD credential remains fully valid.
- Propagate to staging first. Update staging's secret reference (a Git commit bumping the referenced secret version, or, for a backend with live sync, updating the backend value for staging's scoped path) and let the existing External Secrets Operator sync pick it up. Run the CI job's canary authentication check against staging using the new credential.
- Propagate to production, staged by consumer. Bump production's secret reference the same way. Running application pods that were already using the OLD credential in an active connection continue working (the database still accepts it); NEWLY started or restarted pods pick up the new value. CI's next scheduled run automatically uses the new value since it fetches fresh from the backend at job start, no separate CI-specific propagation step required.
- Confirm cutover. Check the database's own connection audit log for any connections still authenticating with the OLD credential's identifier; once none remain for a defined observation window (long enough to cover the slowest-restarting consumer, for example one full pod-restart cycle plus one full CI schedule interval), proceed.
- Revoke. Revoke the old credential at the database. The Git commit that bumped the reference in step 3, plus the backend's own issuance/revocation logs, together form the audit trail: who/what triggered the rotation, when each environment cut over, and when the old credential was finally revoked.
Trade-offs and pitfalls
- Common mistake: rotating directly in production first "to save time," skipping the staging validation step. The whole point of validating in staging first is that a broken new credential (wrong permissions, a typo) fails there instead of during a production cutover, where the same mistake means an active incident rather than a caught bug.
- Common mistake: treating "the new value was successfully written to the secrets backend" as equivalent to "rotation complete." Nothing about a successful WRITE confirms any consumer has actually picked it up or that it authenticates correctly against the real target system; the validation and cutover-confirmation steps are not optional formalities, they are what actually closes the loop.
- A fixed rotation-to-revocation timer is a common, reasonable-looking shortcut that breaks for exactly the consumers that matter most: long-lived connections and infrequently-redeployed batch jobs are precisely the ones a timer-based approach is most likely to disrupt, since they are the slowest to naturally pick up the new credential; checking actual cutover status (not just elapsed time) before revoking is the fix.
- CI runners that cache a credential for the DURATION of a single job (rather than re-fetching per step) can still be using a stale value if rotation happens mid-job, even though CI jobs are usually short-lived; a job that's mid-run when rotation occurs should either complete using its already-fetched credential (fine, if the dual-validity window covers it) or, for very long-running jobs, re-fetch periodically rather than caching for the entire run.
Create a compliance plan that ensures every infrastructure change is auditable. Include signed commits or tags, branch protection policies, enforced code reviews, immutable release artifacts, retention of CI logs, and centralized change records. How would you demonstrate compliance during an external audit using Git history and CI artifacts?
Sample Answer
Direct answer
Auditable infrastructure change rests on five mechanisms working together, not any single one: signed commits or tags (proving WHO authored a change and that history was not altered after the fact), branch protection (guaranteeing every production-bound change actually went through review), enforced code review (the human judgment layer no automated check replaces), immutable release artifacts (guaranteeing what was reviewed is EXACTLY what got deployed), and CI log retention plus centralized change records (the evidence trail an auditor actually reads). During an external audit, the whole story is told by pointing directly at Git history and CI artifacts rather than a separately maintained change log: git log with signed commits and required-approval metadata IS the change record, and the artifact registry IS the proof of what shipped when.
Structured elaboration
Signed commits or tags. Require GPG- or SSH-signed commits on protected branches (enforced as a required status check, unsigned commits rejected at merge time), so git log --show-signature proves both authorship and that the commit content has not been altered since signing, closing the gap a plain author-name field leaves open (anyone can set their local Git username to anything).
Branch protection policies. Protected branches require at minimum: N approving reviews from someone other than the author, passing status checks (including the policy-as-code and signature checks above), and no force-push or history rewriting. This is what makes "every production-bound change went through review" a structurally enforced fact rather than a documented expectation someone could route around.
Enforced code reviews. Beyond the branch-protection MECHANISM (which enforces that SOME review happened), the POLICY needs to specify who counts as a qualified reviewer for a given path (CODEOWNERS scoping production-infrastructure paths to a smaller, accountable group) and what a reviewer is actually expected to check (a documented review checklist), since an unscoped "any teammate can approve" rule satisfies the branch-protection mechanism without satisfying most compliance frameworks' actual segregation-of-duties intent.
Immutable release artifacts. The artifact that gets deployed (a container image referenced by digest, a Terraform plan file) must be built ONCE, from the reviewed commit, and referenced by an immutable identifier from that point forward; if the artifact could be rebuilt or modified after review, the reviewed diff and the deployed bytes are no longer provably the same thing, which breaks the entire audit chain at its weakest link even if every other control here is correctly implemented.
Retention of CI logs. CI run logs (which checks ran, their results, who triggered the run, exact timestamps) need a retention policy that matches the audit window the compliance framework requires (often one to seven years depending on the regulation), stored somewhere the CI system's own default log rotation will not silently delete before that window closes.
Centralized change records. For most infrastructure-audit purposes, the merged PR itself (title, description, diff, approvals, linked ticket, CI results) IS the centralized change record; a SEPARATE change-management system is only needed if the compliance framework requires change records to exist independently of the engineering tooling itself (some regulated environments do), in which case that system should be populated automatically from the PR/CI data, not maintained as a manually duplicated parallel record that can drift out of sync with what Git actually shows.
Worked example
Demonstrating compliance for a specific past change during an external audit, walking the evidence trail an auditor would actually be shown:
- Find the change.
git log --all --grep="<ticket-id>"or the PR search UI locates the specific merged PR by its linked ticket. - Show the review happened. The PR's approval history (who approved, when, against which commit SHA) demonstrates the required review occurred, and the branch-protection RULE configuration (screenshotted or exported from the Git platform's API) demonstrates that review was structurally REQUIRED, not merely present by coincidence on this one PR.
- Show the commit's integrity.
git log --show-signature <commit-sha>shows the signature verifying both authorship and content integrity. - Show what actually deployed matched what was reviewed. The CI run tied to that exact commit SHA produced an artifact with a specific digest; the deployment record (Argo CD's sync history, or the CI deploy job's log) shows that SAME digest was what actually got applied to production, closing the loop from "this diff was reviewed" to "this exact artifact ran."
- Show the CI checks that gated the merge. The CI run's log (retained per the policy above) shows every required check (policy-as-code, security scan, plan validation) passed before merge was even possible.
Trade-offs and pitfalls
- Common mistake: treating signed commits as sufficient proof of review, when they only prove authorship and integrity, not that review happened. Signing and reviewing are two independent guarantees; both need their own enforcement mechanism (commit signing verification, and branch-protection-required-approvals), and conflating them leaves a real gap an auditor will find.
- A maintained-in-parallel change-management system is a common, avoidable compliance liability. If it is populated by hand rather than automatically from PR/CI events, it WILL eventually drift from what Git actually shows, and an auditor comparing the two systems and finding a discrepancy is a worse outcome than having only the Git-native record in the first place.
- Immutable artifacts are the load-bearing link auditors most often find missing. Even with perfect commit signing and review enforcement, if the deployed artifact could theoretically have been rebuilt from a slightly different commit after review (a mutable tag, a non-reproducible build), the entire chain of evidence has a gap at exactly the step that matters most, "is what I reviewed what actually ran."
- CI log retention needs an explicit owner and a periodic verification that it is actually working, a retention policy that exists only in a document, with nobody having confirmed logs from six months ago are still retrievable, is a compliance gap waiting to be discovered mid-audit rather than beforehand.
Design release gates and approval workflows for a GitOps pipeline that must satisfy compliance requirements (auditable approvals for production). Describe where gates are enforced (CI, PR, Argo CD), what metadata should be stored in Git, and how to implement human approvals without breaking declarative principles.
Sample Answer
Direct answer
Enforce release gates at TWO complementary layers: the pull request itself (peer/approver review, required status checks, branch protection rules) as the primary compliance-auditable approval, and Argo CD's sync policy (manual sync for production, or SyncWindows restricting when auto-sync can fire) as a technical backstop that prevents a merged-but-not-yet-approved-for-deployment change from auto-applying. The key design constraint: human approval happens on the PR (a Git-native, naturally auditable event with reviewer identity and timestamp already captured), never as an out-of-band manual step performed against the live cluster, which is what would actually break the declarative, Git-is-source-of-truth principle.
Structured elaboration
Where gates are enforced. CI enforces AUTOMATED gates (linting, terraform plan/manifest-diff generation, policy-as-code checks, security scanning) as required status checks that must pass before merge is even possible. The PR itself enforces the HUMAN approval gate via required reviewers (branch protection rules requiring N approvals from a CODEOWNERS-defined group for production-targeting paths). Argo CD enforces a FINAL technical gate: production Application objects configured with syncPolicy.automated OMITTED (manual sync only) or restricted via SyncWindows, so even a merged commit doesn't immediately roll out until an authorized operator (or an automated promotion step gated on its own separate approval) triggers the sync.
What metadata should live in Git. The PR itself is the metadata record: title/description linking to the change ticket, the diff itself (what actually changed), the list of approvers and their approval timestamps (captured natively by the Git hosting platform, exportable via API for audit), and any required-context status check results (policy-as-code pass/fail, security scan results) attached to the commit. For compliance frameworks requiring a documented "why" beyond the diff, a structured commit message or PR template field (change ticket ID, risk assessment, rollback plan) keeps that context IN the same auditable artifact rather than in a separate system that can drift out of sync with what was actually deployed.
Implementing human approval without breaking declarative principles. The subtlety: "declarative" means the DESIRED STATE is fully specified in Git and reconciliation is automatic given that state, not that every state transition must be fully automatic with no human in the loop. A human approval gate is compatible with GitOps as long as the approval ITSELF is captured declaratively (the merged PR, an approved state) and the reconciler then acts deterministically on whatever is in Git; what would BREAK the principle is a human directly editing the live cluster to "approve" a change out of band, since that reintroduces exactly the drift-inducing manual intervention GitOps exists to eliminate. A common, clean pattern: a promotion PR (bumping an image tag or config value in the production overlay) requires its own separate, higher-bar approval from the PR that merged the change into staging, so "promote to production" is itself a reviewed, auditable Git event, not a live cluster action.
Trade-offs and pitfalls
- Common mistake: implementing the "human approval" step as a manual
argocd app syncperformed by an operator outside of any recorded process. This satisfies "a human looked at it" but produces NO durable audit record of WHO approved WHAT and WHY beyond whatever is (or isn't) in a chat log; the PR-based approval model is specifically what makes the approval auditable without extra tooling, since the Git platform already records it. - Common mistake: conflating CI's automated checks with the compliance-required human approval. A passing lint/policy-as-code check is necessary but not sufficient for most compliance frameworks, which specifically require a human decision-maker's sign-off; treating "all checks green" as equivalent to "approved" removes the human judgment step the compliance requirement exists to capture.
- SyncWindows and manual-sync settings are a technical backstop, not a substitute for PR-level review. They prevent an ALREADY-MERGED change from auto-deploying at the wrong time, but by the time a change is merged, the PR-level approval should already have happened; relying on the sync gate as the PRIMARY approval mechanism (skipping rigorous PR review because "someone will catch it before syncing") inverts the intended order and risks a rushed, under-reviewed sync-time approval replacing a properly reviewed PR.
- CODEOWNERS-based required-reviewer rules need to be scoped to the RIGHT paths. A blanket "any 2 approvers" rule does not satisfy most compliance requirements for segregation of duties if the two approvers can be the same person who wrote the change plus a rubber-stamping teammate; production-path changes typically need a reviewer who is NOT the author and, in stricter regimes, a reviewer from a DIFFERENT team or a designated compliance role.
Unlock Full Question Bank
Get access to all Infrastructure as Code and GitOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.