Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
Design a retention and cleanup policy for a private registry that's accumulated hundreds of tags from experiments, feature branches, and old releases. Balance storage cost against developer convenience, reproducibility, and rollback capability: lifecycle rules (keep latest N per branch pattern), tag immutability, and how you'd schedule garbage collection.
Sample Answer
Direct answer
I design retention as tiers keyed to what a tag actually represents, not a single blanket rule: release and production-facing tags are protected and effectively kept forever, branch and experiment tags decay automatically on a short clock, and garbage collection runs as a scheduled, low-traffic maintenance job rather than an ad hoc cleanup, so nobody has to manually decide what is safe to delete.
The design
| Tag class | Retention rule | Immutability | Rationale |
|---|---|---|---|
Release / semantic-version tags (v1.4.2, prod-*) | Keep indefinitely (or until an explicit, reviewed deprecation) | Immutable: once pushed, the tag can never be overwritten | These are the rollback targets; losing one or letting it silently change content is a production incident waiting to happen |
main/default-branch build tags | Keep last N (for example, last 20) plus anything referenced by a currently-deployed environment | Mutable is acceptable, but each build gets a unique tag (commit SHA or build number) so nothing is overwritten in place | Bounds storage growth from continuous integration while preserving enough history for a same-day rollback |
Feature-branch tags (feature/*, pr-*) | Keep the last 3 to 5 pushes per branch pattern, and expire automatically after a fixed idle window (for example, 14 days with no new push and no pull) | Mutable | These exist for review and short-lived testing; nobody rolls back to a feature branch's image in production |
Experiment / scratch tags (exp-*, ad hoc developer pushes) | Keep for a short fixed window (for example, 3 to 7 days) regardless of activity | Mutable | Genuinely disposable; the cost of keeping hundreds of these indefinitely outweighs the rare case someone wants one back |
Garbage collection (GC, the process of reclaiming storage for image layers no longer referenced by any remaining tag or manifest) is scheduled during low-traffic hours rather than triggered manually, because a naive GC run that walks the whole blob store while pushes are actively happening can race with an in-flight push and, on registries that require it, needs the registry put into a read-only or maintenance mode for the duration. This differs sharply by registry implementation and matters when choosing one: a self-hosted registry (the open-source Docker Registry, or Harbor, a registry that adds retention policies and tag immutability rules on top of it) needs an explicit, scheduled GC job and, on the open-source registry specifically, historically needs the registry to be read-only while GC runs to avoid corrupting in-flight uploads. A managed registry like Amazon ECR (Elastic Container Registry, AWS's managed container registry) has no manual GC step at all: you configure lifecycle policies (rules like "expire untagged images after 14 days" or "keep only the last 10 images matching feature/*") and the service reclaims storage automatically and continuously.
Worked example
For a registry that has accumulated hundreds of tags from months of unmanaged experimentation, the rollout is staged rather than a one-shot deletion: first, tag every existing image against the classification above (a one-time audit script that reads push metadata and branch-name patterns), then apply the new lifecycle rules going FORWARD only (new pushes obey the tiers immediately), then run a single manual cleanup pass on the pre-existing backlog using the same rules with a dry-run flag first, so the retention design is validated against real data before anything is actually deleted. Concretely, an ECR lifecycle policy rule expressing the feature-branch tier looks like: match tag prefix feature/, keep the most recently pushed 5 images matching that prefix, expire the rest. A second rule expresses the experiment tier: match tag prefix exp/, expire any image pushed more than 7 days ago, with no count-based exception.
Trade-offs and pitfalls
The main cost-versus-convenience trade-off is the idle-window length on feature and experiment tags: a shorter window saves more storage but increases the chance someone needs an image that already expired (usually recoverable by rebuilding from source, since the tag was never a durable artifact, but still a real interruption). The most common design mistake is applying one retention rule uniformly across all tag classes, which either keeps release tags too loosely (a rollback target could theoretically be garbage-collected) or expires experimental tags too slowly (storage cost creeps up for images nobody was ever going to reuse). Tag immutability on release tags specifically closes a separate but related risk: without it, someone can accidentally push new content to an existing release tag, which silently breaks reproducibility for anyone who redeploys that "same" version later, independent of whatever the retention window is set to.
You inspect a running production container and notice ps aux inside it shows a growing number of <defunct> zombie processes that never clear, even though the main application appears to be working fine. Explain what PID 1 is responsible for inside a container's PID namespace that it wouldn't have to worry about as an ordinary process on a normal host, why an application binary running directly as PID 1 often fails at that responsibility, and what concrete change you'd make to the container's entrypoint to fix it. What would you gain and lose if you instead solved this by running a full init system inside the container, or by having the container share the host's PID namespace?
Sample Answer
Direct answer
PID 1, the process the kernel assigns process ID 1 to, either at boot on a normal host, or as the first process started inside a container's own PID namespace (an isolated view of process IDs separate from the host's), has two duties an ordinary process never has to think about: it inherits any process whose original parent has already died (an orphan) and must call wait()/waitpid() on it to collect its exit status, and the kernel gives it special signal-handling behavior, a signal being the kernel's way of asynchronously telling a process about an event such as a shutdown request, where most signals simply aren't delivered to PID 1 unless it explicitly installs a handler for them (only SIGKILL and SIGSTOP, which force-terminate or pause a process, always work regardless). An ordinary application binary run directly as PID 1 usually does neither: it never calls wait() on children it didn't fork itself, so those orphans exit but are never reaped and pile up as <defunct> zombies, and it never registered a SIGTERM handler, so it may not even shut down cleanly on docker stop. The concrete fix is to put a tiny, purpose-built init process in front of the application: run docker run --init (which uses tini, bundled into the Docker Engine since version 1.13) or add tini/dumb-init as the image's ENTRYPOINT, so the real application runs as PID 2 and the init process handles reaping and signal forwarding on its behalf.
Structured elaboration
Why zombies accumulate specifically because of PID 1, not just "processes exiting"
A zombie is a process that has already exited but whose exit status hasn't been collected yet; the kernel keeps a small process-table entry for it until something calls wait(). On a normal host, the system init, the first process the kernel starts at boot and the one responsible for starting and supervising every other service on the machine (systemd is the most common example on modern Linux distributions), reaps any orphan automatically, all the time, as a routine part of its job. Inside a container, whatever process the container engine starts is that namespace's PID 1, whether or not it was written with that responsibility in mind. Most application binaries (a Python script, a Node server, a compiled Go binary) were never written to reap other processes; they assume some init layer is already doing that for them, true on a host, false inside a bare container. The practical trigger is usually an application that shells out to short-lived helper processes, each one that exits without being reaped adds one more zombie, and because a single zombie barely consumes resources, this can run for a long time before it becomes visible as process-table pressure or confusing monitoring output.
Comparing the three fixes
| Approach | What it does | Gains | Costs |
|---|---|---|---|
Lightweight init (tini, dumb-init, or docker run --init) | A few-kilobyte process becomes PID 1, forwards signals to the real app and reaps everything else | Minimal image and runtime overhead, predictable signal forwarding, a drop-in ENTRYPOINT change, preserves the single-main-process-per-container model orchestrators expect | Doesn't manage multiple services, no dependency ordering, no logging or cgroup management, it does exactly one job |
Full init system (systemd) inside the container | The container runs a complete OS-style init, with service units, its own logging, and cgroup management | Genuine multi-service management inside one container, familiar operational model if the team already runs systemd everywhere | Needs privileged mode or specific extra mounts to manage cgroups, a much larger image and attack surface, and it breaks the single-process-per-container assumption most orchestrators, health checks, and log collectors are built around |
Share the host's PID namespace (docker run --pid=host) | The container's processes live in the host's real PID namespace, so orphans reparent to the host's real PID 1 and get reaped there | No init process needed inside the container at all | The container can see, and in some configurations signal, every process on the host, a serious isolation break that's unacceptable on any multi-tenant or security-sensitive host, and it doesn't fit typical orchestrator scheduling models either |
sequenceDiagram
participant App as PID 1 process
participant Child as Worker process
App->>Child: fork
Child-->>Child: exits
Child->>App: SIGCHLD delivered
alt PID 1 calls waitpid
App->>Child: reap exit status
else PID 1 never calls waitpid
Child->>Child: stays a zombie, shown as defunct
end
The right default for almost every case is the lightweight init: a one-line ENTRYPOINT or flag change that doesn't meaningfully alter the container's resource footprint, and it fixes both halves of the problem, reaping and signal forwarding, without touching the container's isolation model. systemd-in-a-container earns its cost only when deliberately running something that behaves like a small VM with several cooperating services, a rare, specific need, not a general containerization pattern. Sharing the host PID namespace is a debugging tool, letting you inspect host processes from inside a container, not a production fix, because the isolation it gives up is exactly the isolation a container exists to provide.
Worked example
An engineer notices ps aux inside a long-running container shows dozens of <defunct> entries, and the container's memory climbs slowly over days despite the application's own metrics looking flat. The application is a Python service that shells out to a short-lived helper process per request, and one of those helper processes occasionally forks a further child that gets orphaned before it exits, a grandchild the application's own subprocess-handling code never sees and so never reaps. The Dockerfile's ENTRYPOINT currently runs the Python app directly as PID 1: ENTRYPOINT ["python", "app.py"]. Adding docker run --init (or, equivalently, ENTRYPOINT ["tini", "--", "python", "app.py"]) puts tini at PID 1; the Python app becomes PID 2, and any orphaned grandchild now reparents to tini, which reaps it immediately. No application code changes, no measurable image size increase, and docker stop now delivers SIGTERM to tini, which forwards it to the app, instead of the kernel silently declining to deliver it to an application that never asked to handle it.
Trade-offs & pitfalls
- A common wrong turn is treating this as "the app should just add a
SIGTERMhandler and move on," which fixes shutdown but not reaping; an app can handle its own termination signal correctly and still accumulate zombies from subprocesses it spawns and neverwait()s on. docker run --initonly takes effect atdocker runtime; if the container is launched by an orchestrator that constructs its own run configuration, the equivalent has to be configured there, or baked into the image'sENTRYPOINTinstead, which works regardless of how the container is launched and is usually the more portable choice.--pid=hostlooks like it "just works" in a quick test because reaping does start happening; that's exactly the trap, the fix appears to work while quietly removing process isolation, which won't surface as a problem until it's exploited or audited.- Running
systemdin a container without the specific cgroup and mount configuration it expects tends to fail in ways that look unrelated to init at all (units failing to start, logging silently going nowhere), which sends people debugging the wrong layer if they don't already knowsystemdneeds that container-specific setup.
A container is showing unhealthy in production according to its HEALTHCHECK, and it keeps restarting before it ever receives real traffic. Walk through a step-by-step troubleshooting process: what you'd check in the Dockerfile and application startup path, and how you'd replicate the health probe locally to isolate the cause.
Sample Answer
Direct answer
A container that is marked unhealthy and keeps getting restarted before it ever serves traffic is, on plain Docker or Compose, almost always the application process itself exiting (crashing), not the health status "causing" a restart. Docker's health status is informational: a restart policy reacts to the container's process actually exiting, not to HEALTHCHECK reporting unhealthy by itself. So the fastest path is to confirm what actually killed the process first, then work outward through the Dockerfile and the startup path, and finally reproduce the exact health probe locally to separate "the app never came up" from "the app came up but fails this specific check."
Structured elaboration
1. Confirm what is actually happening, before touching any code
docker inspect --format='{{json .State}}' <container>forExitCode,OOMKilled,Status, and.State.Health.Log(the last several probe attempts, each with its own exit code and output).docker logs --tail=200 --timestamps <container>to see stdout and stderr right up to the crash.- If
docker ps -akeeps showing new container IDs (or the same container's restart count keeps climbing under a--restartpolicy), this is a genuine crash loop: the process is exiting. If instead one container ID just sits atUp ... (unhealthy)and never gets recreated, the process is alive and only the probe itself is failing, which is a narrower problem than "restarting."
2. Dockerfile-side checks
- Read the
HEALTHCHECKinstruction exactly as built:docker inspect --format='{{json .Config.Healthcheck}}' <image>. Confirm the tool it invokes (curl,wget, a custom binary) is actually present in the final image. A hardened or distroless (a minimal base image with no shell, package manager, or other OS tooling beyond what the application itself needs) final stage that droppedcurlmakes the probe itself fail with a shell or exec error, which is indistinguishable from "the app is broken" until you look closely at.State.Health.Log. - Confirm the checked port and path match what the app actually binds, and that the app listens on
0.0.0.0, not only127.0.0.1. - Check whether
CMD/ENTRYPOINTuse shell form (CMD myapp, which runs under/bin/sh -c, making the shell PID 1, the first and special process in the container's own process namespace, responsible for reaping child processes and receiving signals) or exec form (CMD ["myapp"], which makes the application itself PID 1). If the app is not written to handleSIGTERMand reap children as PID 1, shutdowns get slow and messy, which is worth ruling out even if it is not this incident's root cause.
3. Application startup-path checks
- Does startup do a single, no-retry dependency check (a database ping, a config-service fetch) that calls
panic(),os.Exit(), orSystem.exit()on the very first failure, instead of retrying with backoff? This is the most common real cause: a dependency is briefly not ready, the app gives up immediately, the process exits, the restart policy relaunches it into the same race, and it loops forever. - Are all required environment variables and secrets actually present in this environment? A value that only exists in a local
.envfile is a common gap between a workstation and production. - Is there a legitimately slow one-time step (schema migration, cache warm-up, model load) that takes longer than the health check's timing allows before the app can answer the probe at all?
4. Replicate the health probe locally
- Run the same image in the foreground, without
-d, so stdout is visible directly:docker run --rm -p 8080:8080 myimage. - Extract the literal
HEALTHCHECK CMDand run it by hand against a live container:docker exec <container> <the exact command>, and read its real exit code and output instead of trusting the summarizedunhealthylabel. - If the container dies before you can attach, override the entrypoint to keep it alive for inspection:
docker run --rm -it --entrypoint sh myimage, then execute the startup steps one at a time to isolate the exact failing line.
Worked example
A team's Dockerfile ends with HEALTHCHECK CMD curl -f http://localhost:8080/health || exit 1, and the container flaps every few seconds. docker inspect shows ExitCode: 1 and a new container ID on each cycle, a genuine crash loop. docker logs shows Starting application... followed immediately by panic: dial tcp db:5432: connect: connection refused. The startup code opens a database connection and calls a single ping, panicking on any error instead of retrying. Running docker run --rm myimage locally against a database that is not yet reachable reproduces the exact panic in under a second, confirming this has nothing to do with the health check or the Dockerfile. The fix is to make the startup path retry the database connection with backoff and only begin listening (so HEALTHCHECK can start passing) once the dependency is actually reachable, and separately to widen HEALTHCHECK's start_period so a slow but legitimate first connection is not flagged before the app has had a fair chance.
Trade-offs & pitfalls
- Do not "fix" this by only loosening the health check (longer timeout, fewer retries, more lenient
start_period). That hides a real startup-ordering bug and defers discovering it to a worse moment, such as a live outage. - A probe that fails only because its own tool was stripped out of a hardened final image gets misread as "the app is broken," wasting time on the wrong layer. Check the probe command's exit code and output specifically before assuming the application is at fault.
depends_on: condition: service_healthyin Compose reduces this exact class of bug at startup by not starting a dependent service until the thing it needs is already healthy, but it does not replace retry logic inside the app itself: a dependency can still disappear later, after the container is already running and serving traffic.
What role do image registries (Docker Hub, AWS ECR, GCR, Azure ACR) play in the container workflow, and why might an organization run a private registry instead of only using a public one? Cover tagging conventions (semantic versions vs. commit SHAs), and authentication/access control for pushing and pulling in production.
Sample Answer
Direct answer
An image registry stores, versions, and serves container images the way a package repository serves libraries: docker push uploads layers and a manifest to it, docker pull (or a deployment step) downloads them. Docker Hub is the default public registry; AWS ECR (Elastic Container Registry), Azure Container Registry (ACR), and Google Artifact Registry are the major cloud-managed alternatives. An organization runs a private registry, rather than relying only on a public one, mainly for access control, reliability guarantees it does not get from a shared public service, and to keep proprietary images out of a public namespace entirely.
Naming note: Google Container Registry is retired
The question's third option, GCR (Google Container Registry, gcr.io), is retired for new use: Google shut down write access on March 18, 2025, so docker push to a Container Registry-hosted repository no longer works. Read access was not cut off on a fixed date the way write access was; Google migrated existing gcr.io image references to be served through Artifact Registry under that same hostname, so images already published there keep pulling successfully rather than disappearing. Google's current offering for new work is Artifact Registry, which also supports legacy gcr.io-style hostnames for images that were already there. A candidate who has used the older name should recognize it, but new work should target Artifact Registry directly, and the operational risk to plan for is push/CI credentials configured against Container Registry breaking, not already-published images becoming unreachable.
Why run a private registry
A public registry like Docker Hub is a fine default for open-source or non-sensitive images, but production teams usually want: access control (who can push a new version of a production image, and who can pull it, tied into the same identity system as everything else); rate-limit and reliability guarantees a shared public service does not promise for free-tier pulls at production scale; and simply not publishing proprietary application code as an image anyone on the internet can pull. Cloud-managed private registries (ECR, ACR, Artifact Registry) add tight integration with each cloud's identity and access management (IAM) system, so a compute service can be granted pull access without embedding long-lived credentials.
Tagging conventions
Two conventions dominate, and most mature pipelines use both for different purposes: semantic version tags (v1.4.2) are meant for humans and for deliberate promotion between environments (staging tests v1.4.2, then production is told to run v1.4.2), while commit-SHA tags (myapp:a1b2c3d) are generated automatically on every build and give perfect traceability back to the exact source commit, which a semantic version alone does not (two builds of "the same" v1.4.2 from a hotfixed branch would otherwise be indistinguishable). Many pipelines push both tags to the same underlying image so a human-readable release name and an exact build provenance trail both exist.
Authentication and access control in production
Pushing and pulling from a private registry requires authenticating first (docker login, or a cloud CLI's registry-login helper that exchanges cloud credentials for a short-lived registry token). In production, the strong pattern is to avoid long-lived, shared credentials entirely: a compute service (a CI runner, a Kubernetes node, a serverless container platform) is granted an IAM role or workload identity scoped to pull (and, for CI, push) specific repositories, and the registry issues short-lived tokens against that identity rather than a static username and password baked into a config file.
What's the difference between COPY and ADD in a Dockerfile? Give two concrete examples where ADD's extra behavior (remote URL fetching, auto-extracting local tar archives) matters, and explain why COPY is generally preferred in production Dockerfiles.
Sample Answer
Direct answer
COPY does exactly one thing: it copies files or directories from the build context into the image. ADD does that too, plus two extra behaviors: it can fetch a file from a remote URL directly into the image, and it will automatically extract a recognized local archive (like a .tar.gz) into the destination directory. COPY is preferred in production Dockerfiles precisely because it has no hidden behavior: what you see in the instruction is exactly what happens to the filesystem.
ADD's extra behavior, made concrete
Both behaviors were verified directly in the same Dockerfile:
FROM alpine:3.20
WORKDIR /app
ADD archive.tar.gz /app/extracted/
ADD https://www.google.com/robots.txt /app/remote/robots.txt
RUN ls -la /app/extracted && head -c 100 /app/remote/robots.txt
Building this produced the extracted contents of archive.tar.gz under /app/extracted/ with no separate tar command needed, and the fetched contents of robots.txt readable under /app/remote/, confirming both auto-extraction of a local archive and remote URL fetching happen from a single ADD line with no other instruction involved. (A remote URL source is the one exception to auto-extraction: since a newer Dockerfile syntax release, a tar archive fetched from a URL is left compressed unless you pass ADD --unpack <url> <dest>; only a local archive source auto-extracts with no flag, which is what this example uses.)
Example 1, auto-extraction: shipping a vendored dependency or dataset that arrives as a .tar.gz; ADD dataset.tar.gz /data/ extracts it in one step, where COPY would leave the compressed archive sitting in the image unextracted, requiring an additional RUN tar -xzf ... (which then also leaves the original archive file taking up space in that layer unless you clean it up in the same instruction).
Example 2, remote URL fetching: pulling a fixed, versioned release artifact directly by URL during the build. Two assumptions worth checking here, both tested directly rather than assumed: first, ADD for a URL is not simply cached on the URL and Dockerfile text staying the same. Serving a file locally, building ADD http://.../file.txt /app/file.txt, then changing the file's contents on the server without touching the URL or Dockerfile and rebuilding with no --no-cache flag produced the new content, not a stale cached copy; a further rebuild with the content unchanged then did hit the cache. BuildKit revalidates the remote resource on every build and only reuses the cached layer when the content is actually unchanged. Second, ADD does support integrity verification: ADD --checksum=sha256:<hash> <url> <dest> fails the build outright on a mismatch, confirmed directly by supplying a deliberately wrong digest against a real server and getting ERROR: ... digest mismatch, so "no way to verify integrity" is only true for a bare ADD <url> with no --checksum supplied, not for ADD categorically.
What still makes many teams prefer an explicit RUN curl -fsSL ... -o file && sha256sum -c ... over ADD for this case: --checksum is opt-in and easy to forget (a plain ADD <url> with no flag verifies nothing, and the Dockerfile still builds fine without it, so there is no forcing function to add it), the network fetch happens as an opaque step in the build graph rather than a scriptable command you can retry, mirror, or attach custom auth headers to inline, and either way the build still depends on a live, reachable URL: neither approach makes the build reproducible without network access.
Why COPY is generally preferred in production
COPY's behavior is fully predictable from reading the instruction: no network access happens during the build, no archive is silently unpacked when you only meant to copy a file that happened to be named something.tar.gz, and there is no ambiguity about what ended up in the image. ADD's extra behaviors are occasionally exactly what you want (vendoring a local archive, or a URL fetch with --checksum pinned), but reaching for ADD as a habit means every reader of the Dockerfile has to remember to check whether a given source path might trigger auto-extraction or a network fetch, and whether that fetch is actually checksummed or just a bare unverified download, which is exactly the kind of hidden behavior a production build should not depend on implicit knowledge to get right.
Unlock Full Question Bank
Get access to all Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.