Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
Design a CI workflow to build and publish multi-architecture Docker images (amd64 and arm64) for a microservice using docker buildx: building for multiple platforms, reusing a shared build cache across architectures, and publishing a manifest list. What would you test before promoting the image, and what's your rollback plan if one architecture's build is broken?
Sample Answer
Direct answer
I would build each target architecture as its own job sharing one registry-backed cache reference, gate the manifest-list publish step (a manifest list is a single registry reference that points to multiple per-architecture images, so one tag resolves to the correct image for whichever architecture is pulling it) on every architecture independently passing a real runtime smoke test (not just a successful build), and treat rollback as retagging the previous known-good manifest list rather than rebuilding anything.
Caching across architectures
Multi-architecture builds go through docker buildx (Docker's CLI plugin for building images for multiple target architectures in one invocation), which is powered by BuildKit (Docker's build engine, the component that actually executes build steps and manages caching): docker buildx build --platform linux/amd64,linux/arm64 --cache-to type=registry,ref=myregistry/myservice-cache,mode=max --cache-from type=registry,ref=myregistry/myservice-cache ... uses one shared cache reference for both architectures; BuildKit stores platform-specific cache entries under that same reference internally, so both architectures benefit from a single, shared cache location without their genuinely different compiled layers colliding with each other. It is worth being precise about what actually gets shared here: early, architecture-independent steps (copying source files, for instance) can dedupe across platforms since their content is identical regardless of target architecture, but compiled or platform-specific layers do not and should not share cache entries across architectures, since they are legitimately different bytes.
What to test before promoting the image
A successful build proves the image was assembled, not that it runs correctly on its target architecture; a native dependency can compile without error but behave subtly differently on a different architecture (a different default in a base image's OpenSSL or C library version, for instance), so each architecture's image needs to actually run before the pipeline treats it as promotable. Concretely: run each platform's image with docker run --rm --platform linux/arm64 <staging-tag> <smoke-test-command> (using the registered QEMU emulation if no native arm64 runner is available) hitting the application's own health endpoint and a small set of representative requests, before that architecture's result is allowed to feed into the final manifest list. Pushing first to a staging tag, testing against it, and only then promoting (or directly assembling the manifest list from the tested per-platform images) keeps a bad build from ever reaching the production tag in the first place.
Rollback plan if one architecture's build is broken
The design constraint that prevents most of the pain here is refusing to ever publish a partial manifest list: if the amd64 build passes its smoke test but the arm64 build fails, the pipeline does not publish a manifest list under the production tag with only amd64 present, since a client resolving that tag on arm64 would either fail to pull anything runnable or, worse on some registry configurations, get served stale or unexpected content. The promotion step is gated on every target architecture succeeding and passing its smoke test together; a single failing architecture blocks the whole promotion, which is a deliberately conservative trade (slower to ship a fix for one architecture) in exchange for never shipping an inconsistent multi-architecture release. If a regression is discovered only after promotion, rollback does not mean rebuilding: it means retagging the previous manifest list's digest back onto the mutable production tag (docker buildx imagetools create -t myregistry/myservice:prod myregistry/myservice:prod@sha256:<previous-known-good-digest>, or the registry's equivalent retag operation), which only moves a pointer and completes immediately rather than waiting on a fresh build.
Trade-offs and pitfalls
Running each architecture as a separate parallel job is faster in wall-clock terms when native runners are available for each target, but needs either that native-runner infrastructure or QEMU emulation set up per job, adding CI complexity; a single buildx invocation covering all platforms is simpler to maintain but ties up one job for the combined duration of every architecture, and any QEMU-emulated architecture in that single job is slower than it would be on a native runner. The most consequential design decision is the promotion gate itself: treating "the build succeeded" as sufficient to promote, without an actual per-architecture runtime smoke test, is the gap that lets an architecture-specific runtime bug reach production behind a green build.
Design a safe image-promotion pattern so the same artifact moves from development to staging to production with minimal drift: how tags, digests, and registry permissions ensure the exact image that was tested is the one that gets deployed, and how you'd prevent accidental rebuilds or tag-drift surprises along the way.
Sample Answer
Direct answer
The core principle is build once, promote everywhere: an artifact is built and given an immutable identity exactly once, and every later environment, staging and then production, references that same identity rather than triggering a second build from source. Drift, and the class of bug where something works in staging but not in production, overwhelmingly comes from an environment quietly building its own slightly different copy of "the same" release instead of promoting the one that was already tested.
The design
Build once. CI builds the image a single time per release candidate, tags it immutably and uniquely (a commit hash or a semantic release version), and records its digest (a digest is a cryptographic hash of the image's exact bytes; two images with the same digest are guaranteed to be byte-for-byte identical, unlike a tag, which can later be moved to point at different bytes) at that moment. No later pipeline stage is allowed to rebuild this artifact; staging and production only ever pull or reference it.
Promotion is a metadata operation, not a build. Promoting an artifact from staging to production re-tags the same digest for the next environment, for example adding app:v1.4.2-prod-approved alongside the existing app:v1.4.2, or a promotion tool simply records that this digest is now the production candidate. In a digest-pinned deploy configuration, promotion can be as simple as updating production's deploy manifest to reference the already-tested digest directly.
Registry permissions enforce the guardrail. CI's build identity gets write access to push new images and tags. Production's deploy identity gets read-only access, able to pull, never to push. If the registry supports tag immutability or protected tags, lock release tags so nothing, accidental or malicious, can silently repoint a tag after staging has already tested it. A human or an approval step gets permission to move a "promoted" pointer or write a deployment record, never to trigger a rebuild.
Preventing accidental rebuilds. The biggest practical source of drift is a Dockerfile or CI configuration that is not fully deterministic: a base image referenced by a floating tag like python:3.12-slim instead of a pinned digest, or an install step with no version pins, meaning two builds from "the same" source a day apart can legitimately produce different bytes. Pin base images by digest, pin dependency versions, and treat anything environment-specific, a database URL, a feature flag, as runtime configuration injected into the same image, never as a build-time difference that forces separate, environment-specific images.
flowchart LR
Commit["source commit"] --> Build["CI build (once)"]
Build --> Registry["registry: immutable digest"]
Registry --> Staging["staging deploy (pull-only)"]
Staging -->|"tests pass, promote"| Prod["production deploy (pull-only, same digest)"]
Worked example
A release pipeline for a checkout service: CI builds on merge to the main branch, tags checkout-service:2026.09.23-abc1234, and records the digest in a release manifest. Staging's deploy configuration is templated to always deploy checkout-service@<digest from the latest release manifest>. After staging's smoke tests pass, a promotion job, not a rebuild, copies that exact digest reference into the production deploy configuration and opens the standard change-approval process. Production's deploy identity has pull-only registry permissions, so even a compromised production deploy credential cannot push a different image under that same tag.
Trade-offs and pitfalls
Locking down write access this tightly adds real friction for a legitimate emergency change: a hotfix now has to go through the same build-once pipeline rather than a quick local build pushed straight to production, a deliberate trade of speed for guaranteed provenance. "Build once" is easy to violate by accident the moment any Dockerfile step reaches for something mutable, an unpinned base image tag, an unpinned package install, or a dependency resolved with no lockfile, so this whole design only holds if the Dockerfile itself is actually reproducible; otherwise the promotion discipline is being built on top of a build that was never deterministic to begin with.
Why is tagging an image latest risky in production, and how would you eliminate tag drift in a pipeline where developers currently deploy app:latest? Describe how you'd publish, promote, and reference images so that, months later, an audit can reconstruct exactly what was running in production at a given time.
Sample Answer
Direct answer
latest is not "the newest version"; it is simply the tag Docker applies by default when none is specified, and many teams separately choose to also push it as a rolling "current" alias by convention, not by anything Docker enforces. The real risk is that a tag, including latest, is a mutable pointer: deploying app:latest on two different days can pull two genuinely different images, and nothing about the tag itself records which one actually ran. That gap is exactly what an audit hits months later. The fix is to make every deploy reference an image by its immutable content digest, and to record that reference durably outside the registry, rather than trusting that latest happened to point at the right thing at deploy time.
Why tags drift
A tag is just a mutable pointer in the registry to a manifest digest. Re-pushing to the same tag repoints it, with no built-in history of what it used to point to, and no guarantee that two docker pull app:latest calls a minute apart even resolve to the same bytes if a push landed in between.
The fix, layered
- Immutable, unique tags at build time. Every CI build tags the image uniquely, most commonly by version control commit hash (
app:sha-abc1234) or a monotonically increasing build number, and never re-pushes over an existing unique tag. - Deploys reference the digest, not the tag. Wherever deploy tooling supports it, reference
app@sha256:...directly. A digest is cryptographically bound to the exact bytes of that manifest; it cannot silently repoint the way a tag can. - Promotion re-tags the same digest; it never rebuilds. Moving an artifact from development to staging to production repoints or re-references the same digest, guaranteeing that the artifact tested in staging is bit-for-bit the one that reaches production.
- The deployment record is durable and independent of the registry. The pipeline's own deployment log stores which digest went to which environment at which time. If the registry later garbage-collects an old tag, or someone force-pushes over one, the audit trail still exists elsewhere.
Worked example
A service currently deployed as app:latest migrates as follows: CI starts also pushing app:sha-<commit> alongside latest (keeping latest purely as a convenience for local development, never referenced by any deploy pipeline). The deploy pipeline changes to resolve the digest for the just-built SHA tag, for example via docker inspect --format='{{index .RepoDigests 0}}' right after the push, and passes that digest to the deploy step. The deployment system logs a record such as {service: app, digest: sha256:..., deployed_at: ..., deployed_by: pipeline_run_id} to a durable store, even a simple append-only table is enough. Months later, an audit queries that log for the timestamp in question, gets the digest, and pulls the exact same image bytes that ran that day, regardless of whatever app:latest currently points to.
Trade-offs and pitfalls
A unique tag per build is an improvement, but not a complete fix on its own, unless the registry also enforces tag immutability, since a careless or compromised process could still repoint even a "unique" tag if nothing prevents it. Pure digest references are correct but unreadable to a human scanning a deploy manifest, so most real setups keep both: a readable unique tag for the log line, resolved to its digest for the actual deploy reference and audit record. Retention matters too: if the registry garbage-collects layers behind a tag that later gets deleted, an audit six months out needs both the digest and a registry retention policy long enough to still have it, or the audit trail points at bytes that no longer exist anywhere.
You need to integrate automated image vulnerability scanning (a tool like Trivy or Clair) into the CI pipeline. Where in the build/test/publish stages would you run it, what severity threshold would block a release, and how would you handle false positives without training the team to ignore every alert? In a regulated environment, what would you specifically look for in a scan result, and how would you prioritize findings?
Sample Answer
Direct answer
Run image vulnerability scanning right after the image is built but before it is pushed anywhere a deploy step can reach it, so a failing scan blocks the artifact from ever becoming pullable, not just from ever reaching production. Gate on severity combined with whether a fix actually exists, rather than a blanket "zero vulnerabilities," and treat false positives as a reviewed, time-boxed exception, not a blanket ignore list, so the team is never trained to click past every alert.
Where in the pipeline
Build stage produces the image, the scan stage runs immediately against that exact image by digest, not a floating tag, because a digest names exactly one immutable image while a tag can be repointed after the scan, and only a passing scan proceeds to push or publish. Scanning after the build but before the push, rather than after the push, means a failed scan never creates a pullable artifact at all, a stronger guarantee than scanning post-deploy and opening a ticket afterward.
Severity threshold
A defensible baseline: block on CRITICAL and HIGH severity findings that have a known available fix. Allow, with a tracked, time-boxed exception, HIGH or CRITICAL findings with no fix available yet, since blocking on those permanently reds the pipeline for something the team cannot act on. MEDIUM and LOW findings typically report to a dashboard rather than gate the release.
Handling false positives without training the team to ignore alerts
Use an explicit, reviewed suppression naming the specific vulnerability identifier and the specific package, with a required justification and an expiry date, never a permanent blanket ignore. A suppression with no expiry quietly becomes technical debt nobody revisits when the "false positive" later becomes exploitable, for example after a refactor makes a previously unreachable code path reachable. Track the suppression count as its own metric: a gate that always passes because everything got suppressed has stopped doing its job.
In a regulated environment
What changes is less the scanning mechanism and more what has to be provable afterward. Retain scan results per released artifact, tied to its digest, for the required audit retention period. Be able to show who approved any exception and when. Be able to answer quickly whether a specific vulnerable component is running anywhere in production right now, which is exactly why linking a scan result to the exact deployed digest matters here specifically: an auditor's question is almost always what is running right now, not what was scanned once. Prioritize findings by a combination of severity, whether the vulnerable component is actually reachable or exercised by the running application, not every vulnerability in an unused code path is equally urgent, and whether the component sits in a path that touches regulated data, for example anything handling payment data under a standard like PCI-DSS (the Payment Card Industry Data Security Standard) or health data under HIPAA (the Health Insurance Portability and Accountability Act), over one buried in a build-only tool that never ships in the runtime image.
Worked example
A CI stage added right after the build:
docker build -t registry.example.com/app:${SHA} .
trivy image --exit-code 1 --severity CRITICAL,HIGH --ignore-unfixed registry.example.com/app:${SHA}
# only on success, push by tag and capture the digest the registry actually assigned:
docker push registry.example.com/app:${SHA} | tee /tmp/push.log
DIGEST=$(grep -oE 'sha256:[0-9a-f]{64}' /tmp/push.log | tail -1)
echo "published as registry.example.com/app@${DIGEST}"
--exit-code 1 fails the CI step, and therefore blocks the push, on any matching finding. --ignore-unfixed implements the "do not block on findings with no available fix" policy directly at the tool level, while those findings still get reported elsewhere for tracking. Note the push step: docker inspect --format='{{index .RepoDigests 0}}' already returns a full repo@sha256:... reference, not a bare digest, so wrapping it inside another registry.example.com/app@$(...) template produces a garbled, invalid reference (confirmed directly: Docker rejects that construction with "invalid reference format"). Capturing the digest from the docker push output itself, which always prints its own digest: sha256:... line, avoids that entirely and is the reliable way to record exactly what was published.
Trade-offs and pitfalls
Scanning too early, source-only or dependency-manifest-only, misses vulnerabilities introduced by the base image itself, so image-level scanning after the build is necessary even when a separate dependency scan also runs earlier in the pipeline. Over-blocking on every MEDIUM finding is the single most common way teams end up routing around the gate entirely, a merge-and-suppress culture that defeats the point far more thoroughly than a slightly looser threshold would.
Your organization builds dozens of services from a monorepo on distributed CI runners with poor cache locality, so Docker layer caches rarely hit and builds are slow. Design a build-cache strategy: shared base images, BuildKit remote cache export/import (--cache-from/--cache-to), cache namespacing per service, and how you'd secure and store the remote cache.
Sample Answer
Direct answer
On distributed CI (continuous integration) runners, the local build cache that normally makes instruction-ordering discipline work does not help, because a given job can land on a different, cold runner every time with no prior build history at all. The fix is to stop relying on any one runner's local cache and make the cache itself a portable artifact, using the remote cache export and import built into BuildKit (Docker's modern build engine, driven from the command line through the docker buildx plugin) against a registry, so a fresh runner can pull down another runner's cache before it builds anything.
The design, piece by piece
Shared base images. Build and publish one pinned internal base image, with the OS, language runtime, and common system packages already installed, that every service's Dockerfile starts FROM, instead of forty services each independently installing the same packages. This collapses the expensive, rarely-changing part of every build into one artifact that only needs rebuilding when the base itself changes.
BuildKit remote cache export and import. docker buildx build --cache-to=type=registry,ref=<registry>/<service>-cache,mode=max ... pushes the build's layer cache to a registry as its own artifact after every build. --cache-from=type=registry,ref=<registry>/<service>-cache on the next build, on any runner, cold or warm, pulls that cache down before building starts, so a fresh runner behaves as if it had just built this service minutes ago. mode=max, versus the default mode=min, exports cache for every intermediate layer, including ones that never end up in the final image, which matters for multi-stage Dockerfiles where a builder stage's cache would otherwise be discarded.
Cache namespacing per service. One shared cache reference for all forty services in the monorepo causes constant cross-service cache thrashing, where one service's build evicts or overwrites cache entries another service needed. Namespace the cache reference per service, and typically per branch too (<registry>/cache/<service>:<branch>, with a fallback import from :main when a feature branch has no cache of its own yet), so each service's cache import only ever competes with its own history, never the whole monorepo's.
Securing and storing the remote cache. The cache registry needs the same access control as the image registry, because a poisoned cache blob is later reused as literal layer content in a real build, making it a genuine supply-chain risk, not a throwaway artifact. Use a private repository with the CI system's own short-lived, narrowly scoped credentials rather than a long-lived shared token. Plan for retention: cache blobs accumulate fast under mode=max across many services and branches, so a time-to-live or a periodic prune job on the cache repository is needed, since unlike the production image registry, cache artifacts have no long-term retention requirement.
flowchart LR
RunnerA["cold CI runner A"] -->|"pull"| Base["shared base image"]
RunnerA -->|"--cache-from"| CacheReg["cache registry (per-service namespace)"]
RunnerA -->|"build"| Image["service image"]
Image -->|"--cache-to"| CacheReg
RunnerB["cold CI runner B"] -->|"--cache-from"| CacheReg
Worked example
A Node.js service in the monorepo starts FROM internal/node-base:20.3-slim (the shared base), then COPY package*.json . followed by RUN npm ci (the dependency layer, ordered so the most rarely changing inputs come first), then COPY . .. The CI job runs docker buildx build --cache-from=type=registry,ref=registry.internal/cache/checkout-service:main --cache-to=type=registry,ref=registry.internal/cache/checkout-service:main,mode=max -t checkout-service:$SHA .. The first build on a brand-new runner still pulls the npm ci layer from the registry cache blob instead of recomputing it, because the cache import happens before the build starts, entirely independent of anything that specific runner has ever built before.
Trade-offs and pitfalls
mode=max cache exports are noticeably larger to push and pull than mode=min, a bandwidth-versus-hit-rate trade-off that gets worse the more services share one CI cluster's egress (its outbound network capacity to the registry). Per-branch namespacing needs a sane fallback chain, branch cache miss falling back to the main branch's cache, or every new feature branch pays a full cold build. A shared base image becomes a coordination point in its own right: bumping it now needs a rollout plan across every dependent service, rather than each service upgrading on its own schedule, a real cost worth naming rather than treating this design as a pure win.
Unlock Full Question Bank
Get access to all 7 Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.