Infrastructure as Code and GitOps Questions
Defining and managing infrastructure and delivery state declaratively: provisioning as code (Terraform, CloudFormation, Pulumi, Ansible, Puppet), configuration management, idempotency, drift detection and remediation, and version control for infrastructure definitions, extended by GitOps where git is the source of truth for deployment and infrastructure state. Covers keeping environments consistent, treating config as a first-class versioned artifact, pull-based deployment and continuous reconciliation toward the committed state (ArgoCD, Flux, and similar controllers), Kubernetes manifest and configuration delivery via git, secrets handling for IaC and GitOps pipelines, policy-as-code guardrails (OPA, Sentinel), Terraform state management and locking, and auditable change through version control: branching strategy, pull request review, commit conventions, and code review policy for infrastructure code. Distinct from the CI/CD pipeline design topic, which owns generic pipeline structure and platform-scale release orchestration (build, test, artifact publishing, runner mechanics) and the architectural choice between push-based CI/CD and pull-based GitOps, even when the payload is infrastructure code. Distinct from the safe deployment and rollback strategies topic, which owns deployment-strategy mechanics: canary and blue-green traffic shifting, automated rollback triggered by metrics or SLOs, feature-flag progressive delivery, Kubernetes rollout mechanics (maxSurge, maxUnavailable, health-check gating), and database or schema migration safety as it gates a release, even when the delivery mechanism is GitOps. Distinct from the automation and scripting topic, which owns operational-scripting disciplines (retry and backoff logic, CLI tool design, generic file, checksum, or diff utilities) when the task is not specifically about declarative infrastructure or configuration state. This topic keeps the GitOps reconciliation loop itself, drift detection and remediation, and IaC state and module lifecycle management regardless of which adjacent discipline a question also touches.
Explain how reconciliation loops work in GitOps controllers (for example Argo CD and Flux). Discuss scaling challenges when managing thousands of manifests across many clusters, including informer/watch scaling, reconcile frequency, rate-limiting, caching strategies, and design decisions to reduce API-server and controller load.
Sample Answer
Direct answer
A GitOps controller's reconciliation loop is fundamentally: watch a source (Git, or the cluster via informers) for changes, compute a diff between desired and live state, and apply the difference, repeated continuously. At small scale this is simple; at THOUSANDS of manifests across MANY clusters, the loop's naive form breaks down along four specific axes, informer/watch scaling (how many objects a single controller can efficiently watch), reconcile frequency (how often it re-checks everything), rate-limiting (how fast it can safely apply changes without overwhelming the API server), and caching (how much redundant API-server work each reconcile avoids), and a controller designed for scale has to address all four, not just add more compute.
Structured elaboration
Informer/watch scaling. Kubernetes controllers typically use informers (a local, continuously-updated cache backed by a watch connection to the API server) rather than polling list calls, so each reconcile reads from the LOCAL cache, not the API server directly. At thousands of manifests across many clusters, the watch connections themselves (one per resource type per cluster, potentially) and the memory footprint of the local caches become the bottleneck; a controller architecture built for scale shards this work (per-cluster or per-namespace informer instances, only watching resource kinds actually in use) rather than running one giant informer set watching everything everywhere.
Reconcile frequency. A fixed, short polling interval applied uniformly to every application (checking every 30 seconds regardless of how often that application's config actually changes) wastes API-server and controller CPU on applications that rarely change while still being too slow for applications that need faster drift detection. Event-driven reconciliation (triggering a reconcile on a Git webhook or a Kubernetes watch event, rather than a fixed poll interval) plus EXPONENTIAL BACKOFF for applications that are healthy and unchanged (progressively lengthening the interval between checks when nothing has changed recently, resetting to a fast interval the moment something DOES change) gets fast reaction time where it matters without constant unnecessary work everywhere.
Rate-limiting. Applying many changes simultaneously across a large fleet risks overwhelming the target API server(s) with a burst of writes, especially right after a controller restart when every application needs to be re-evaluated at once (a "thundering herd"). Controllers built for scale use client-side rate limiting (a token-bucket or leaky-bucket limiter on outbound API calls per cluster) and STAGGER the initial reconcile pass after startup (jittered delays rather than firing every application's first reconcile in the same instant) specifically to avoid this.
Caching strategies. Beyond the informer's live-object cache, a scaled controller caches EXPENSIVE derived computations, rendered manifests (re-running a Helm template or Kustomize build on every single reconcile pass is wasteful if the underlying chart/base hasn't changed), diff results, and resolved Git commit SHAs (avoiding a fresh Git fetch on every poll when polling for changes rather than using webhooks), invalidating each cache layer only when its actual input changes.
Trade-offs and pitfalls
- Common mistake: treating "add more controller replicas" as the primary scaling lever. Horizontal scaling helps only if the WORK is genuinely partitionable (sharded by cluster or namespace, with each replica owning a disjoint slice); naively running multiple replicas all reconciling the SAME full set of applications multiplies API-server load without increasing effective throughput, and risks duplicate/conflicting apply operations unless there is a leader-election or sharding scheme.
- Event-driven reconciliation reduces average load but does NOT eliminate the need for a periodic full reconcile. Webhooks and watch events can be missed (a delivery failure, a network partition during the exact window a change happened), so a scaled system still needs an occasional FULL reconciliation pass (hourly, or some longer interval) as a correctness backstop, purely event-driven with no periodic fallback risks silent, permanent drift if even one event is ever lost.
- Aggressive caching trades staleness risk for reduced load, and that trade-off needs an explicit, bounded staleness budget, not an open-ended "cache until something tells us to invalidate," since a caching bug that fails to invalidate correctly can leave a controller confidently reconciling against a manifest render that is silently out of date.
- Rate-limiting protects the API SERVER, but does not by itself protect against a controller being overwhelmed by its OWN internal queue if events arrive faster than they can be rate-limited out; a bounded work queue with backpressure (dropping or coalescing redundant re-reconcile requests for the same object rather than queuing every single trigger separately) is a necessary complement to outbound rate-limiting, not a substitute for it.
Propose how Service Level Objectives (SLOs) and error budgets should be represented, versioned, and deployed via Git workflows alongside infrastructure code. Describe validation steps, who should review SLO changes, how SLO changes propagate to alerting and automation, and how to roll back SLO changes if they cause undesired automation behavior.
Sample Answer
Direct answer
Service-level objectives (SLOs) and error budgets should be represented as PLAIN, VERSION-CONTROLLED DECLARATIVE ARTIFACTS (a YAML or similar spec defining the objective, the underlying service-level indicator query, and the burn-rate alerting thresholds) living in the SAME repository, and going through the SAME PR-and-review workflow, as the infrastructure they govern, deployed to the alerting/automation system by the SAME reconciliation mechanism, not a separate, dashboard-configured, un-versioned side channel. Because an SLO artifact directly controls automated behavior (paging thresholds, and in some organizations automated rollback triggers), its review bar and rollback path need to be at least as rigorous as the infrastructure changes it governs, arguably more so, since a wrong SLO can either cause alert fatigue (too tight) or silently miss real degradation (too loose).
Structured elaboration
Representation. A structured spec per service: the SLI (service-level indicator) query defining what is actually measured (a specific latency percentile, a specific error-rate calculation), the objective threshold and measurement window, and the burn-rate alert thresholds derived from it (fast-burn and slow-burn windows, following the widely-used multi-window burn-rate alerting pattern). This spec is data, not code, so it is reviewable as a clean diff the same way any other declarative configuration is.
Versioning. Committed to Git alongside (or in a clearly cross-referenced sibling path to) the service's own infrastructure definition, so git log on the SLO spec answers "when did this objective change and why" the same way it would for any other infra change, and a service's SLO history is auditable independent of whoever happens to remember the reasoning.
Deployment via Git workflows. A GitOps controller (or a dedicated SLO-config sync job, if the alerting platform is not itself Kubernetes-native) reconciles the deployed alerting-platform configuration to match the Git-declared spec, the same pull-based, continuously-reconciled model already used for application infrastructure generally, so a manually-edited alert threshold in the alerting platform's own UI gets treated as drift, not silently tolerated as a parallel source of truth.
Validation steps. Before merge: schema validation (is the spec well-formed), a SANITY check on the objective itself (is the target achievable given recent actual measured performance, catching an SLO set so tight it would page constantly from day one), and, where the alerting platform supports it, a DRY-RUN evaluation against recent historical data showing what the proposed burn-rate thresholds WOULD have alerted on over, say, the last 30 days, so a reviewer sees concretely whether the change would have caused excessive noise or a dangerous blind spot before it goes live.
Who should review SLO changes. The OWNING team proposes (they have the domain context for what "acceptable" performance means for their service), but a change LOOSENING an existing objective, or changing what a downstream automated system does in response to burn (see below), should require a second reviewer from OUTSIDE the owning team, typically the SRE or platform function with cross-service context, specifically because a team under pressure has an incentive to loosen its own SLO to escape being paged, which is exactly the review-conflict this second-reviewer requirement exists to catch.
How changes propagate to alerting and automation. The reconciler updates the ALERTING PLATFORM's configuration (the actual burn-rate alert rules) directly from the Git-declared spec; any DOWNSTREAM automation keyed off SLO/error-budget state (an automated deployment-freeze trigger when a budget is exhausted, for instance) reads the SAME reconciled objective, so there is exactly one source of truth for "what counts as within budget" feeding every consumer, rather than each consumer holding its own copy that could drift out of sync.
Rolling back an SLO change that causes undesired automation behavior. Because the spec is just a Git-tracked file, rollback is the same git revert plus reconciliation pattern as any other infra change; the practical difference is URGENCY, an SLO change causing a pager storm or, worse, an incorrectly-triggered automated rollback/freeze needs the same fast-path emergency-change mechanism as any infrastructure change: a short-lived, audited bypass with mandatory reconciliation back into Git afterward, since waiting for a normal-cadence PR review while pages fire is not acceptable.
Trade-offs and pitfalls
- Common mistake: configuring SLO thresholds directly in the alerting platform's UI "to iterate quickly," planning to formalize into Git later. This almost never gets formalized in practice, and it recreates exactly the un-auditable, un-reviewable side channel this whole answer exists to prevent; the dry-run-against-historical-data validation step exists specifically to make the Git-first path fast enough that UI-first iteration stops being tempting.
- The second-reviewer requirement for loosening an SLO is the single most-skipped control on this list under deadline pressure, precisely because it exists to catch the exact situation (a team wanting relief from its own pages) where the proposing team has the least incentive to enforce it on themselves.
- A DRY-RUN validation is only as good as the historical window it evaluates against. A recent quiet period can make an objectively too-loose threshold look safe; validating against a window that includes at least one known past incident, where the team can check "would this new threshold have caught it," is a meaningfully stronger check than an arbitrary recent-N-days window.
- Downstream automation consuming SLO state (deployment freezes, automated rollback triggers) needs to fail SAFE if the SLO reconciliation itself is stale or broken, an automation system that silently treats "no recent SLO data" as "budget is fine" can mask exactly the kind of problem it exists to catch; explicit staleness detection on the SLO data feed itself is a real, easy-to-omit requirement.
Discuss the trade-offs between using immutable image tags (digests) versus mutable tags (like 'latest' or 'v1') in a GitOps workflow. Explain how immutable tagging affects reconciliation, security (CVE remediation), reproducibility, and developer iteration. Propose a recommended tagging policy for production and for developer environments.
Sample Answer
Direct answer
Immutable digests (image@sha256:...) guarantee that the manifest committed to Git and the bytes actually pulled and run are the SAME image forever, which is exactly what reconciliation, reproducibility, and rollback correctness all depend on; mutable tags (:latest, :v1 reused across builds) let the SAME manifest silently resolve to DIFFERENT bytes at different times, breaking the core GitOps guarantee that Git fully describes what is running. The practical policy: digests (or at minimum strictly immutable, never-reused tags) in every environment a GitOps controller actually reconciles, with a separate, explicit "update the digest" commit as the mechanism for shipping a new version, never implicit re-resolution of a mutable tag; mutable tags are acceptable ONLY in inner-loop developer environments that are explicitly outside the GitOps-reconciled path.
Structured elaboration
Effect on reconciliation. A GitOps controller's reconciliation loop compares the manifest's declared image reference against the cluster's running state. With a digest, "does the running Pod match desired state" is an unambiguous, stable comparison, the digest either matches or it doesn't, and once it matches, reconciliation correctly does nothing further. With a mutable tag, the CONTROLLER sees no drift (the tag string in the manifest hasn't changed), but the ACTUAL running bytes can differ from what was originally deployed if a Pod restarts and re-pulls a tag that has since been overwritten upstream, a form of drift the reconciliation loop is structurally blind to because it only compares the STRING in the manifest, not the resolved content.
Effect on security (CVE, Common Vulnerabilities and Exposures, remediation). A patched image (fixing a CVE) pushed under the SAME mutable tag changes what NEW pods pull without any Git commit recording that change, so there is no audit trail of when the fix actually rolled out, and existing running pods do not automatically pick it up (they keep running the old, pulled bytes until they restart), creating an unpredictable, unrecorded window of mixed-version exposure. A digest-pinned deployment makes CVE remediation an explicit, auditable act: bump the digest in Git, and reconciliation deploys exactly that patched image, deterministically, everywhere, with a Git commit as the permanent record of when the fix went out.
Effect on reproducibility. "Reproducible" means the SAME Git commit always produces the SAME running system. Digests give this property unconditionally. Mutable tags give it only until someone pushes a new image under that tag, at which point the SAME Git commit now resolves differently than it did before, which specifically breaks rollback (reverting to an OLD commit that references :v1 does not guarantee you get back the ORIGINAL :v1 bytes if :v1 has since been overwritten) and breaks any forensic "what was actually running at time T" investigation.
Effect on developer iteration. For an INNER-LOOP dev environment (a developer's own sandbox, rapidly rebuilding and testing), requiring a fresh digest and a Git commit for every single iteration is genuine friction that does not serve any of the guarantees above, since nobody is trying to prove reproducibility or audit a dev sandbox's history the way they would a production deployment.
Worked example
A recommended tiered policy:
| Environment | Tag policy | Rationale |
|---|---|---|
| Developer/sandbox | Mutable tag (:dev) OK, often OUTSIDE GitOps reconciliation entirely (direct kubectl/local tooling) | Fast iteration matters more than audit trail; not a GitOps-reconciled environment in the first place |
| Staging/CI-integration | Immutable, unique tag per build (:sha-<commit> or :build-<n>), digest-equivalent in practice since each is never reused | Needs traceability back to the exact commit/build without full production rigor |
| Production | Digest (@sha256:...) required, enforced by policy (admission-controller rule rejecting mutable tags or bare tags without a resolved digest) | Full reconciliation correctness, CVE-remediation auditability, and rollback guarantees required |
The transition from a build artifact to a production digest reference should be an EXPLICIT step in the pipeline (CI resolves the newly built image's digest and commits a manifest update referencing it), not something a human copies by hand, since a manually-typed digest is exactly the kind of error-prone step this policy exists to eliminate.
Trade-offs and pitfalls
- Common mistake: using a "unique-looking" tag (a commit SHA or build number) and assuming it is equivalent to a digest. It usually is, in practice, AS LONG AS the registry and CI pipeline genuinely never reuse it; but this is a PROCESS guarantee (nobody re-pushes that tag), not a CRYPTOGRAPHIC one the way a digest is (the digest IS the content hash, so it cannot silently point to different bytes by definition). A digest removes the reliance on that process discipline entirely.
- Common mistake: enforcing digest-only policy in production manifests but leaving the ADMISSION path open to mutable tags, so a manifest applied outside the normal GitOps flow (a manual
kubectl apply, a different pipeline) can still introduce a mutable-tag image; an admission-controller policy (Kyverno/Gatekeeper) rejecting any Pod spec without a resolved digest closes this gap regardless of HOW the manifest arrived. - Digests make manifests harder for a human to read at a glance (
sha256:a1b2c3...conveys no version information the way:v2.3.1does); a common, workable mitigation is keeping a human-readable tag in the commit message or a companion annotation/label while the actualimage:field uses the digest, giving both machine-verifiable immutability and human-readable context. - Rollback via Git revert only fully works if EVERY commit in history referenced digests, not mutable tags. A rollback to a historical commit that referenced
:v1is not a true rollback if:v1has since been overwritten; digest pinning is what makes "revert the commit" and "actually get back the old running state" the same operation.
You must lead adoption of infrastructure-as-code (IaC) and CI/CD across four product teams with different maturity. Draft a six-month change management plan covering pilot selection, training, templates, governance, rollback strategy, and success metrics to show value quickly.
Sample Answer
Direct answer
A six-month adoption plan across four teams of different maturity needs to sequence by DEMONSTRATED VALUE, not by uniform rollout: start with the team most ready to succeed quickly (proving the templates and process work before asking a skeptical or lower-maturity team to trust them), expand using what that first team's success actually showed rather than a fixed timeline alone, and treat rollback and governance as CONSTANT infrastructure present from day one, not phase-4 additions bolted on after adoption is already underway.
Structured elaboration
Pilot selection. Choose the team with the HIGHEST readiness (existing scripting discipline, an engineer with some Terraform/CI exposure already) and a project with GENUINE but bounded risk (not a toy project nobody would notice failing, but also not the highest-stakes system in the company); a pilot chosen purely for safety produces weak evidence, a pilot chosen purely for visibility risks a highly public stumble before the process is proven.
Training. Differentiated by team maturity, not one-size-fits-all: the pilot team gets HANDS-ON, project-embedded training (learning by doing the actual pilot, with close support); teams 2 through 4 get progressively more SELF-SERVICE training material (documentation, recorded walkthroughs, office hours) as the templates and common patterns mature from the pilot's own experience, since by the time team 4 onboards, the training material itself should be measurably better than what team 1 had, built from real lessons rather than speculation.
Templates. Built FROM the pilot, not designed in the abstract beforehand and handed down; the pilot's real Terraform modules, CI pipeline configuration, and PR checklist become the STARTING templates for the next team, refined incrementally as each subsequent team's own real usage surfaces gaps the pilot did not happen to hit.
Governance. Present from the START, not deferred to a later phase: even the pilot's very first PR goes through review, plan validation, and the same policy-as-code baseline the whole program will eventually require everywhere; introducing governance late, after teams have already gotten used to a looser process, is far harder than establishing it as the norm from the pilot onward.
Rollback strategy. A concrete, TESTED path back to each team's PREVIOUS process (their prior deployment mechanism) for at least the first month of their own adoption, not just a theoretical claim that "we can always go back"; a team asked to adopt new tooling without a genuinely tested fallback is taking on risk the plan should not be asking of them, especially the lower-maturity teams with less slack to absorb a rocky transition.
Success metrics to show value quickly. Metrics that are MEANINGFUL within weeks, not only visible after six months: time from PR-merge to deployed (should shrink measurably even for the pilot alone), count of manual/undocumented emergency changes needed (should trend down), and a simple, honest survey of the ADOPTING team's own confidence/friction, since a technically-successful pilot that the team itself experienced as painful is weak evidence for teams 2 through 4, who will be watching how team 1 actually felt about it, not just the dashboard numbers.
Worked example
A concrete six-month timeline:
| Month | Activity |
|---|---|
| 1 | Pilot team selected; hands-on adoption begins on a bounded, real project; governance baseline established from day one |
| 2 | Pilot's first metrics available (deploy latency, emergency-change count); templates refined from real pilot friction points |
| 3 | Team 2 (next-highest readiness) onboards using refined templates and more self-service training material; pilot team's success metrics presented org-wide |
| 4 | Team 3 onboards; team 2's own feedback further refines templates and training |
| 5 | Team 4 (lowest initial maturity) onboards with the MOST mature templates, training, and by now, three teams' worth of internal champions/support available |
| 6 | Org-wide retrospective; governance and rollback discipline reviewed for what to keep permanently versus what was pilot-phase-specific |
Trade-offs and pitfalls
- Common mistake: rolling out to all four teams simultaneously on a fixed calendar, rather than sequencing by demonstrated pilot success. This forfeits the single biggest advantage a phased plan has, letting the templates, training, and even the ROLLBACK plan itself improve with each successive team's real experience; a simultaneous rollout means every team hits the SAME rough edges independently, with no team benefiting from a predecessor's lessons.
- Choosing the pilot team purely for its lower risk (an unimportant, low-visibility project) rather than for genuine readiness produces weak, unconvincing evidence for the skeptical teams still to come, the same pilot-selection principle that applies to any staged technical migration applies here to organizational adoption too.
- Deferring governance to "once adoption is further along" is a common, costly sequencing mistake, a team that spends month 1 in a looser process and only encounters real review/policy discipline in month 3 experiences that discipline as a regression, actively worse for change management than establishing the bar from day one and never lowering it.
- A success metric measured only at the six-month mark gives the LATER teams (3 and 4) no early evidence to build confidence from, exactly the metrics that matter most for keeping the whole six-month plan credible are the ones visible within the first four to six weeks of the PILOT alone, not the ones that only show up at the very end.
Architect a GitOps deployment platform that enforces policy-as-code (for example OPA/Gatekeeper) across clusters, automatically rejects non-compliant manifests, and supports human approval workflows for regulated workloads. Describe repo layout, admission flows, and policy testing.
Sample Answer
Direct answer
Enforcing policy-as-code across clusters while supporting human approval for REGULATED workloads specifically means the platform needs TWO distinct enforcement modes operating side by side, not one uniform gate: hard REJECTION for policies that are always non-negotiable (no exceptions, ever), and a SOFT-BLOCK-PENDING-APPROVAL mode for regulated-workload policies where a human can, through an explicit, audited process, approve a specific exception; conflating the two either makes hard rules bypassable (unacceptable) or makes every regulated workload's legitimate exceptions impossible without weakening policy for everyone (equally unacceptable).
Structured elaboration
Repo layout. A policies/ directory (platform-owned) holding the Rego/Gatekeeper ConstraintTemplate definitions, separate from apps/ (per-service manifests); each Constraint (the specific parameterized instance of a policy, e.g., "no privileged containers, applied to these namespaces") declares which clusters/namespaces it applies to and, critically, whether it is HARD (always enforced) or SOFT (enforceable but exception-eligible for regulated workloads with approval).
Admission flows. Gatekeeper's admission webhook evaluates every manifest against BOTH categories at admission time (not just at CI/PR time): a HARD violation is rejected outright by the webhook, no path around it. A SOFT violation on a workload NOT flagged as regulated is also rejected (soft policies are not "optional" for everyone, only exception-eligible for the specific regulated category they are scoped to). A SOFT violation on a workload flagged as REGULATED checks for a recorded, approved exception (the same exception-tag Rego pattern shown below) and admits ONLY if a valid, current exception is present, rejecting otherwise.
Human approval workflows for regulated workloads. An exception request is its OWN reviewed artifact (a PR adding an exception annotation, referencing a ticket and an expiry date), requiring sign-off from whoever owns policy exceptions for regulated workloads specifically (a compliance or security function, distinct from the requesting team's own approvers), so a regulated exception cannot be self-approved by the team that wants it.
Policy testing. As with any Rego policy (executed, not merely asserted): every policy needs both a VIOLATING test case (confirms the policy actually catches what it is meant to catch) and a COMPLIANT test case including adjacent, similar-but-legitimate configuration (confirms it does not false-positive), run in CI on every change to the policies/ directory itself, since a broken or overly-broad policy change deployed cluster-wide is a platform-level incident, not a routine bug.
Worked example
A concrete constraint distinguishing HARD from SOFT enforcement, and its exception mechanism:
package cluster.security
import rego.v1
# HARD: no privileged containers, ever, no exception path at all.
deny contains msg if {
some c in input.review.object.spec.containers
c.securityContext.privileged == true
msg := sprintf("container %v: privileged containers are never permitted", [c.name])
}
# SOFT: no host network access, exception-eligible for workloads
# explicitly labeled 'regulated-workload-exception-eligible' AND carrying
# a current, approved exception annotation.
deny contains msg if {
input.review.object.spec.hostNetwork == true
not has_approved_host_network_exception(input.review.object)
msg := "host network access requires an approved regulated-workload exception"
}
has_approved_host_network_exception(obj) if {
obj.metadata.labels["regulated-workload-exception-eligible"] == "true"
obj.metadata.annotations["policy-exception-ticket"]
startswith(obj.metadata.annotations["policy-exception-ticket"], "COMP-")
}
A privileged container is rejected unconditionally, no annotation can ever admit it. A hostNetwork: true pod is rejected UNLESS it carries both the eligibility label AND a valid, ticket-referencing exception annotation, giving regulated workloads a real, audited path while every other workload is held to the same bar as the hard rule.
Output (actually executed with opa eval --format pretty -d regulated_policy.rego -i <input> "data.cluster.security.deny", OPA 1.18.2)
A privileged container (the hard rule, no exception path exists for it at all):
[
"container app: privileged containers are never permitted"
]
A hostNetwork: true pod carrying the eligibility label and a valid COMP--prefixed exception ticket:
[]
The same hostNetwork: true pod with no exception annotation:
[
"host network access requires an approved regulated-workload exception"
]
The three runs confirm the two-tier design works as intended: the hard rule denies unconditionally regardless of any annotation (no exception input was even tested against it, since none exists), the soft rule correctly ADMITS a workload carrying a genuine, well-formed regulated exception, and correctly DENIES an otherwise-identical workload with no exception present.
Trade-offs and pitfalls
- Common mistake: implementing exceptions as a single, uniform mechanism applied to every policy, hard and soft alike. This either makes truly non-negotiable rules (privileged containers, for instance) technically bypassable by anyone who can construct the right annotation, or forces every genuinely exception-eligible regulated case through the SAME rigor as a rule that should never bend at all; the two-tier design exists specifically to avoid both failure modes.
- Admission-time enforcement alone, without the SAME policy also evaluated at PR/CI time, means a violation is only discovered when someone tries to deploy it; running the identical policy at BOTH points (not two independently-drifting implementations) gives fast PR-time feedback while admission remains the actual, unbypassable backstop.
- Exception approval authority needs to sit OUTSIDE the requesting team specifically for regulated workloads, a self-approved regulated exception defeats the entire point of having a distinct, audited approval workflow for this category; this is a deliberate, not incidental, separation-of-duties requirement.
- Policy testing that only checks the violating case, never the compliant/adjacent case, is a common, dangerous gap: an over-broad policy that also blocks LEGITIMATE regulated workloads with valid exceptions would only be caught by a compliant-case test explicitly exercising the exception path, not by testing the violation alone.
Unlock Full Question Bank
Get access to all Infrastructure as Code and GitOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.