Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
What is a Docker image layer, and how does Docker's build cache decide whether a layer from a previous build can be reused? Given a project where dependencies rarely change but source code changes often, how would you order your COPY and RUN instructions to maximize cache reuse, and what causes a cache miss?
Sample Answer
Direct answer
Each instruction in a Dockerfile that touches the filesystem (RUN, COPY, ADD) produces one layer, and Docker's build cache reuses a previous layer instead of re-executing the instruction when two conditions hold: the instruction's own inputs are unchanged (the exact command text for RUN, or the exact file contents and metadata for COPY/ADD), and every layer before it in the build was also a cache hit. That second condition is the one people forget: a single cache miss invalidates every layer after it, even if those later instructions did not change at all.
How the cache actually decides reuse
For a RUN instruction, Docker compares the instruction string itself (and the parent layer's cache key) against previous builds; if they match, it reuses the cached layer without re-executing the command. For COPY/ADD, Docker additionally checksums the source files being copied; if any file's contents or metadata changed, that layer is a cache miss, no matter how the instruction was worded. Once a layer misses, Docker cannot know whether the filesystem state going into the next instruction is still what it was before, so it must invalidate and rebuild every instruction after it too.
Worked example
Given requirements.txt (rarely changes) and app.py (changes often), compare two orderings.
Bad order (dependencies installed after the whole context is copied):
FROM python:3.12-slim
WORKDIR /app
COPY . .
RUN pip install --no-cache-dir -r requirements.txt
CMD ["python", "app.py"]
Good order (dependency files copied and installed first, source copied last):
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD ["python", "app.py"]
After an initial build of both, changing only app.py and rebuilding gives, for the bad order:
=> CACHED [2/4] WORKDIR /app
=> [3/4] COPY . . # cache miss: app.py is part of this COPY, so its checksum changed
=> [4/4] RUN pip install --no-cache-dir -r requirements.txt # re-runs, no CACHED tag
and for the good order:
=> CACHED [2/5] WORKDIR /app
=> CACHED [3/5] COPY requirements.txt .
=> CACHED [4/5] RUN pip install --no-cache-dir -r requirements.txt # CACHED
=> [5/5] COPY . .
The pip install step is a re-run every time in the bad ordering (that was measured directly: the same one-line change to app.py forced RUN pip install to execute in the bad Dockerfile and stay CACHED in the good one), because COPY . . includes app.py, so any source change invalidates that layer and every layer after it, including the expensive dependency install. In the good ordering, requirements.txt is copied and installed before the source code is ever touched, so a source-only change cannot invalidate the install step at all.
What causes a cache miss
A RUN line's text changing (even whitespace or a reordered flag), any file changing that a COPY/ADD references (including files you did not mean to include, if your .dockerignore is missing or incomplete), a different base image digest, or a build argument (ARG) used earlier in the file changing value. The single most common unintentional miss is copying the entire build context (COPY . .) before installing dependencies, which ties the expensive install step's cache key to files that change on every commit.
What's the difference between a process being alive inside a container and the application actually being healthy? How would you use a Docker HEALTHCHECK to reflect that distinction rather than just checking that the process hasn't crashed?
Sample Answer
Direct answer
A process being "alive" only means the operating system's process table still has an entry for PID 1 (the container's first and main process); it says nothing about whether that process is doing useful work. "Healthy" means the application can actually serve a request correctly right now: its dependencies (database, cache, downstream services) are reachable, and its own internal state (thread pool, connection pool, event loop) is not stuck. Docker's HEALTHCHECK instruction exists precisely to close that gap: it lets you define an active, application-level probe instead of relying on the passive fact that the process has not crashed.
Structured elaboration
- Why "alive" is not enough. A process can be alive and still be useless: deadlocked on a lock it will never release, stuck in an infinite retry loop against a database that no longer exists, or simply never having finished its startup sequence.
docker psreportingUp 12 minutesreflects only that the process has not exited; it makes no claim about correctness. - What
HEALTHCHECKadds. It runs a command on a schedule (--interval), gives it a--timeout, allows a--start-periodduring which failures do not count against the container (used for legitimately slow startup), and needs--retriesconsecutive failures before flipping the reported status fromstartingtounhealthy. Crucially, the command should exercise the same code path a real client would use, not just check that a socket accepts a TCP connection. - Writing a check that reflects health, not just aliveness. A weak check hits
/and expects any HTTP response. A meaningful check hits a dedicated endpoint (commonly/healthor/healthz) that the application itself computes: does it have an open, working connection to its database, is its background worker queue processing, has startup fully completed. The check should return a definite pass or fail, cheaply and quickly, without doing expensive work (do not run a full business transaction inside a health probe that fires every few seconds). - A concrete instruction:
HEALTHCHECK --interval=10s --timeout=3s --start-period=15s --retries=3 \
CMD curl -f http://localhost:8080/health || exit 1
Here /health is application code, for example: return 200 only if a lightweight SELECT 1 against the database succeeds within a short timeout, and 503 otherwise. That single distinction, "did I check a real dependency" versus "did I just answer any request", is what separates a liveness-flavored check from one that reflects actual health.
Worked example
An API's process stays alive but its database connection pool has silently exhausted (every connection is checked out and never returned, due to a bug that forgets to close a result set). Requests hang until they time out, but the process itself never crashes, so a plain "is the process running" check would report the container as fine indefinitely. A HEALTHCHECK calling /health, where the handler tries to borrow a connection from the pool with a one-second timeout and fails fast if none is available, flips to unhealthy within a few probe cycles. That gives an operator (or an external supervisor watching the status) a concrete, actionable signal well before every user-facing request starts failing.
Trade-offs & pitfalls
- Too shallow a check (any HTTP
200, or just a TCP connect) never catches internal degradation like the connection-pool example above; it only proves the listener thread is alive. - Too deep a check (a full end-to-end transaction against every downstream dependency) makes the health check itself slow, expensive, and a new source of false failures when an optional, non-critical dependency has a bad moment.
- On plain Docker,
HEALTHCHECKonly changes the status Docker reports (visible viadocker psanddocker inspect); it does not by itself restart the container. Acting on that status (restarting, rerouting traffic, alerting) is something you or an orchestrator has to build on top of it.
A faulty release tagged prod just got deployed. Walk through how you'd perform an immediate rollback to the previous known-good image using image digests with a Docker Compose stack, and how automation (recording the last-known-good digest, a one-command rollback script) could make this safer and faster next time.
Sample Answer
Direct answer
A rollback under a faulty prod tag has one job: get back to the exact previous image bytes as fast as possible, without relying on a rebuild, since even a rebuild from "the same" source tag is not guaranteed to reproduce identical bytes if a base image or a dependency resolves differently since then. With Docker Compose, that means pointing the affected service at the previous image's content digest directly and recreating just that one service.
The immediate rollback
- Find the last known-good digest. If digests were already being recorded at deploy time, this is a lookup in the deploy log, not a hunt. Without that record,
docker images --digestsor the registry's own tag history is the fallback, which is exactly the gap worth closing before an incident, not during one. - Point the compose file at that digest, not the floating tag:
services:
app:
image: registry.example.com/app@sha256:d6cd6f0583314b6592c3e0b0ae2f6778a6b2373e9cf283c3d979c65c7ebf47d7
- Recreate only that service:
docker compose up -d --no-deps app(adding--force-recreateif Compose believes nothing changed). This pulls the pinned digest, guaranteed to be the same bytes as before since digests are content-addressed, and swaps only that one container, leaving sibling services and the network untouched. - Verify before declaring it done. Confirm the running container's actual image digest matches what was intended (
docker inspect --format='{{.Image}}'cross-checked against the pulled digest), and that health checks or smoke tests pass. A rollback that silently landed on the wrong digest is worse than a slower, correct one.
Making this safer and faster with automation
- Record last-known-good automatically. Every successful, health-checked deploy writes its digest to a small, durable record, a file, a row in a deploy-tracking table, or a tag like
app:last-goodthat only the automated pipeline is ever allowed to move. This turns step one above from a stressful hunt through logs into a lookup that already exists before the incident happens. - A one-command rollback script. Wraps the steps above into a single invocation: it reads the last-known-good record, updates the compose file's image reference for that one service, runs the targeted
docker compose up -d --no-deps, then runs the same smoke test the normal deploy pipeline runs before declaring success. The goal is that an on-call engineer under pressure runs one command instead of reconstructing this whole sequence from memory. - Guardrails. The rollback script should refuse to "roll back" to the digest that is already running, protecting against accidentally rolling back twice, and should log its own action to the same deploy-tracking record a normal deploy would, so the audit trail stays accurate even during an incident.
Trade-offs and pitfalls
Rolling back the image alone does not roll back anything that shipped alongside it: a database migration that ran with the bad release does not reverse itself, so a rollback plan needs a documented answer for what happens when the bad release also changed schema, not just code. --no-deps matters specifically because a full docker compose up -d can recreate dependent services too if their configuration appears to have changed, a much bigger blast radius (a bigger set of things a mistake could break) than the incident called for. Automating "last known good" purely as "the previous deploy" is subtly wrong if that previous deploy was also bad; the record should track the last deploy that actually passed a real health check, not merely the one immediately before this one chronologically.
A latency-sensitive service shows occasional CPU throttling, memory pressure, and noisy-neighbor effects when several containers share a host. How would you tune container runtime settings (cgroups, CPU shares, --cpuset-cpus) and application behavior to improve tail latency and node stability, including how you'd share GPU resources fairly if the workload uses them?
Sample Answer
Direct answer
On a shared host, protecting tail latency means giving the latency-sensitive container a guaranteed floor rather than trusting best-effort scheduling to sort itself out: dedicate specific cores to it and cap what noisy neighbors can burst to, rather than only weighting priorities and hoping. GPU sharing needs the same discipline but a different mechanism entirely, because a GPU has no built-in kernel-level time-slicing fairness the way CPU cores do under cgroups (control groups, the Linux kernel mechanism containers use to enforce per-process CPU, memory, and I/O limits).
Diagnosing which kind of contention you actually have
Before tuning anything, distinguish a container being throttled by its own configured quota (checkable per-container via /sys/fs/cgroup/cpu.stat's nr_throttled and throttled_usec fields, which count how often and for how long the cgroup's CPU quota was exhausted) from the host's CPUs genuinely being oversubscribed across containers despite each individual container having headroom in its own quota (checkable via the host's overall load average and run-queue length). The fixes for these two are different: the first is a quota problem for one container, the second is a host-level capacity or scheduling-weight problem across all of them.
Tuning levers
--cpuset-cpusfor hard isolation. Pinning the latency-sensitive container to specific physical cores it does not share with best-effort workloads eliminates cross-core migration and cache-thrashing entirely for that container, at the cost of those cores being reserved even when idle (no automatic work-stealing across the pin boundary).--cpu-sharesfor prioritization under contention. Shares are a relative weight (default 1024) that only changes behavior when the host is actually CPU-contended; giving the sensitive container a higher share (say 4096 against everything else's default) means it wins a larger fraction of CPU time specifically when there is competition, with zero effect when there is not. This alone does not stop a noisy neighbor from bursting to 100% of a shared core for the length of one scheduling period before the weighting evens out, which is enough to cause a visible latency spike even if the long-run average looks fine.- A hard
--cpusquota on the noisy batch containers, not just low shares, caps their absolute ceiling so they cannot burst into full core usage at all, directly bounding the size of the disruption they can cause to a co-located latency-sensitive workload. - Memory:
--memoryplus--memory-swappiness=0for the latency-sensitive container, so a neighbor's memory pressure does not push it into swap, which is catastrophic for tail latency.--memory-reservationsets a soft floor the kernel tries to protect during host-wide memory pressure, which is the safer lever than--oom-kill-disable(removing the kernel's out-of-memory safety valve entirely is not a fix, it just changes what kind of failure happens next). - Application behavior: connection pooling and backpressure so the application sheds load explicitly under contention rather than queuing silently, since an unbounded queue is what actually turns a brief CPU stall into a multi-second latency spike downstream.
GPU fairness
Unlike CPU cores, a GPU is not preemptible the same way: once a compute kernel is running, it generally cannot be paused mid-execution to let another process's work through, so naive sharing means one workload's launch can block the GPU's execution queue for as long as it runs. Two real hardware/software mechanisms exist for fairness on newer NVIDIA data-center GPUs: MIG (Multi-Instance GPU), which physically partitions one GPU into isolated hardware instances with their own memory and compute slice, giving hard capacity isolation at the cost of fixed, coarse-grained partition sizes (you cannot carve out an arbitrary 10% slice); and MPS (Multi-Process Service), which time- and space-shares a single GPU context across processes with a configurable per-client limit (the CUDA_MPS_ACTIVE_THREAD_PERCENTAGE environment variable), giving finer-grained sharing but no fault isolation between clients, since a client that crashes can potentially disrupt the shared MPS server context for everyone co-scheduled on it. Where neither is available, the practical fallback on a single host is coarser: assign whole GPUs per workload (--gpus device=<id>) rather than attempting to share one GPU between a latency-sensitive and a best-effort workload at all.
Trade-offs and pitfalls
--cpuset-cpus pinning is a static partition: it wastes capacity when the pinned cores sit idle while other cores are saturated, since nothing can migrate across the boundary automatically. MIG's hard isolation can under-utilize a large GPU for a small job because partition sizes are fixed, not arbitrary. MPS's finer granularity comes at the cost of losing the fault isolation MIG provides. The most common mistake on a single shared host is reaching only for --cpu-shares and assuming it solves burst-driven tail-latency spikes, when a hard quota on the noisy workload, or physical core separation, is what actually bounds the worst case rather than just improving the average.
You inspect a running production container and notice ps aux inside it shows a growing number of <defunct> zombie processes that never clear, even though the main application appears to be working fine. Explain what PID 1 is responsible for inside a container's PID namespace that it wouldn't have to worry about as an ordinary process on a normal host, why an application binary running directly as PID 1 often fails at that responsibility, and what concrete change you'd make to the container's entrypoint to fix it. What would you gain and lose if you instead solved this by running a full init system inside the container, or by having the container share the host's PID namespace?
Sample Answer
Direct answer
PID 1, the process the kernel assigns process ID 1 to, either at boot on a normal host, or as the first process started inside a container's own PID namespace (an isolated view of process IDs separate from the host's), has two duties an ordinary process never has to think about: it inherits any process whose original parent has already died (an orphan) and must call wait()/waitpid() on it to collect its exit status, and the kernel gives it special signal-handling behavior, a signal being the kernel's way of asynchronously telling a process about an event such as a shutdown request, where most signals simply aren't delivered to PID 1 unless it explicitly installs a handler for them (only SIGKILL and SIGSTOP, which force-terminate or pause a process, always work regardless). An ordinary application binary run directly as PID 1 usually does neither: it never calls wait() on children it didn't fork itself, so those orphans exit but are never reaped and pile up as <defunct> zombies, and it never registered a SIGTERM handler, so it may not even shut down cleanly on docker stop. The concrete fix is to put a tiny, purpose-built init process in front of the application: run docker run --init (which uses tini, bundled into the Docker Engine since version 1.13) or add tini/dumb-init as the image's ENTRYPOINT, so the real application runs as PID 2 and the init process handles reaping and signal forwarding on its behalf.
Structured elaboration
Why zombies accumulate specifically because of PID 1, not just "processes exiting"
A zombie is a process that has already exited but whose exit status hasn't been collected yet; the kernel keeps a small process-table entry for it until something calls wait(). On a normal host, the system init, the first process the kernel starts at boot and the one responsible for starting and supervising every other service on the machine (systemd is the most common example on modern Linux distributions), reaps any orphan automatically, all the time, as a routine part of its job. Inside a container, whatever process the container engine starts is that namespace's PID 1, whether or not it was written with that responsibility in mind. Most application binaries (a Python script, a Node server, a compiled Go binary) were never written to reap other processes; they assume some init layer is already doing that for them, true on a host, false inside a bare container. The practical trigger is usually an application that shells out to short-lived helper processes, each one that exits without being reaped adds one more zombie, and because a single zombie barely consumes resources, this can run for a long time before it becomes visible as process-table pressure or confusing monitoring output.
Comparing the three fixes
| Approach | What it does | Gains | Costs |
|---|---|---|---|
Lightweight init (tini, dumb-init, or docker run --init) | A few-kilobyte process becomes PID 1, forwards signals to the real app and reaps everything else | Minimal image and runtime overhead, predictable signal forwarding, a drop-in ENTRYPOINT change, preserves the single-main-process-per-container model orchestrators expect | Doesn't manage multiple services, no dependency ordering, no logging or cgroup management, it does exactly one job |
Full init system (systemd) inside the container | The container runs a complete OS-style init, with service units, its own logging, and cgroup management | Genuine multi-service management inside one container, familiar operational model if the team already runs systemd everywhere | Needs privileged mode or specific extra mounts to manage cgroups, a much larger image and attack surface, and it breaks the single-process-per-container assumption most orchestrators, health checks, and log collectors are built around |
Share the host's PID namespace (docker run --pid=host) | The container's processes live in the host's real PID namespace, so orphans reparent to the host's real PID 1 and get reaped there | No init process needed inside the container at all | The container can see, and in some configurations signal, every process on the host, a serious isolation break that's unacceptable on any multi-tenant or security-sensitive host, and it doesn't fit typical orchestrator scheduling models either |
sequenceDiagram
participant App as PID 1 process
participant Child as Worker process
App->>Child: fork
Child-->>Child: exits
Child->>App: SIGCHLD delivered
alt PID 1 calls waitpid
App->>Child: reap exit status
else PID 1 never calls waitpid
Child->>Child: stays a zombie, shown as defunct
end
The right default for almost every case is the lightweight init: a one-line ENTRYPOINT or flag change that doesn't meaningfully alter the container's resource footprint, and it fixes both halves of the problem, reaping and signal forwarding, without touching the container's isolation model. systemd-in-a-container earns its cost only when deliberately running something that behaves like a small VM with several cooperating services, a rare, specific need, not a general containerization pattern. Sharing the host PID namespace is a debugging tool, letting you inspect host processes from inside a container, not a production fix, because the isolation it gives up is exactly the isolation a container exists to provide.
Worked example
An engineer notices ps aux inside a long-running container shows dozens of <defunct> entries, and the container's memory climbs slowly over days despite the application's own metrics looking flat. The application is a Python service that shells out to a short-lived helper process per request, and one of those helper processes occasionally forks a further child that gets orphaned before it exits, a grandchild the application's own subprocess-handling code never sees and so never reaps. The Dockerfile's ENTRYPOINT currently runs the Python app directly as PID 1: ENTRYPOINT ["python", "app.py"]. Adding docker run --init (or, equivalently, ENTRYPOINT ["tini", "--", "python", "app.py"]) puts tini at PID 1; the Python app becomes PID 2, and any orphaned grandchild now reparents to tini, which reaps it immediately. No application code changes, no measurable image size increase, and docker stop now delivers SIGTERM to tini, which forwards it to the app, instead of the kernel silently declining to deliver it to an application that never asked to handle it.
Trade-offs & pitfalls
- A common wrong turn is treating this as "the app should just add a
SIGTERMhandler and move on," which fixes shutdown but not reaping; an app can handle its own termination signal correctly and still accumulate zombies from subprocesses it spawns and neverwait()s on. docker run --initonly takes effect atdocker runtime; if the container is launched by an orchestrator that constructs its own run configuration, the equivalent has to be configured there, or baked into the image'sENTRYPOINTinstead, which works regardless of how the container is launched and is usually the more portable choice.--pid=hostlooks like it "just works" in a quick test because reaping does start happening; that's exactly the trap, the fix appears to work while quietly removing process isolation, which won't surface as a problem until it's exploited or audited.- Running
systemdin a container without the specific cgroup and mount configuration it expects tends to fail in ways that look unrelated to init at all (units failing to start, logging silently going nowhere), which sends people debugging the wrong layer if they don't already knowsystemdneeds that container-specific setup.
Unlock Full Question Bank
Get access to all Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.