Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
You built an image using a multi-stage build, but at runtime the container fails with an ImportError or a missing shared library. Walk through a step-by-step debugging plan (using tools like docker history, an interactive docker run --rm -it, or dive) to find which stage or layer omitted the dependency, and how you'd fix it without bloating the final image back up.
Sample Answer
Direct answer
A multi-stage build failing with a missing shared library at runtime almost always means a runtime-only dependency (a shared library that a compiled extension needs, but that was only ever installed as a side effect of the builder stage's toolchain) never made it across the COPY --from=builder boundary; the fix is to find exactly which file or package is missing and add that one thing to the final stage, not to widen what gets copied or fall back to a single fat stage.
Step-by-step debugging plan
- Reproduce interactively against the final image, not the builder.
docker run --rm -it --entrypoint sh <final-image>and attempt the failing import or command by hand, to get the exact error and confirm it happens in the shipped image, not just in some other environment. - Identify precisely what is missing.
ldd(a standard Linux utility that lists the shared libraries a binary depends on) run against the specific binary or shared object (ldd $(which python3), orldd path/to/the_compiled_extension.so) prints=> not foundnext to any dependency the runtime image cannot resolve. This turns a vagueImportErrorinto a concrete missing filename. - Compare against the discarded builder stage.
docker historyonly shows the layers of the image you run it against, which for a multi-stage build is the FINAL stage; the intermediate builder stage's layers are discarded once the build finishes and are not visible this way. To inspect what the builder stage actually had, build up to it explicitly and tag it:docker build --target <builder-stage-name> -t debug-builder ., then rundocker history debug-builderanddocker run --rm -it debug-builder shto confirm the missing file genuinely exists there. - Use
dive(a terminal tool for browsing a Docker image's layers and file contents) on the taggeddebug-builderimage to search for the specific missing path or package by name, confirming both that it exists in the builder and which layer introduced it, then check theCOPY --from=builderline in the final stage to see whether that path was actually included in what got copied over. - Fix with the smallest addition that closes the gap, either a targeted
COPY --from=builder /exact/path/to/lib.so /usr/lib/for a single file, or installing the specific runtime package via the distribution's package manager (for exampleapt-get install -y --no-install-recommends libpq5), never the corresponding-dev/full build-toolchain package, which drags headers, compilers, and static libraries back into the image the multi-stage build was written to keep out. - Verify without reintroducing bloat. Re-run the same reproduction from step 1 to confirm the import now succeeds, and check
docker imagesfor the size delta: a targeted runtime library addition should move the image size by kilobytes, not the hundreds of megabytes a reintroduced build toolchain would add.
Worked example
I verified this exact failure shape directly: python:3.12-slim, a common runtime base for a multi-stage Python build, does not have the PostgreSQL client library installed by default (ldconfig -p | grep -i libpq inside a fresh python:3.12-slim container returns nothing). A realistic version of this bug: a python:3.12 builder stage compiles psycopg2 from source, which links against libpq because the builder stage has libpq-dev installed (it needs the development headers to compile against); the final stage COPY --from=builder /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packagess the compiled package across, but never installs the runtime libpq5 package in the slim final stage. The result at runtime is ImportError: libpq.so.5: cannot open shared object file: No such file or directory, exactly the symptom ldd against the compiled psycopg2 .so file would surface directly. The fix is a single added line in the runtime stage: RUN apt-get update && apt-get install -y --no-install-recommends libpq5 && rm -rf /var/lib/apt/lists/*, which adds the shared library alone, not the development headers or compiler that libpq-dev would bring back.
Trade-offs and pitfalls
The tempting shortcut, when a missing dependency shows up, is to just install the same -dev package in the final stage that the builder stage used, since it is guaranteed to fix the immediate error; this reliably reintroduces the exact bloat multi-stage builds exist to avoid, because dev packages typically pull in a compiler toolchain and headers that a running application never needs. A second pitfall is fixing the symptom for one missing library and stopping, without checking whether the same class of gap exists for other compiled extensions in the same image; ldd against every compiled .so file the application actually loads, not just the first one that errored, catches the rest before they surface one at a time in later incidents.
You're asked to harden a container image that currently runs as root, ships shell utilities, and is built from a large general-purpose base image, exposing an API to untrusted input. Propose a runtime hardening policy: seccomp profiles, dropping Linux capabilities, a read-only root filesystem, rootless containers, and how you'd enforce and audit these at deploy time, while keeping the service actually operable.
Sample Answer
Direct answer
For a container that runs as root, ships general shell utilities, is built from a large general-purpose base, and is directly exposed to untrusted input, the goal is defense in depth: assume the application itself will eventually be exploited, since it is the one thing an attacker directly controls, and make sure a successful exploit lands in a process with almost nothing useful available to it. No root privileges, no unnecessary Linux capabilities, no writable filesystem, and no shell or general-purpose tooling to pivot with.
The policy, layer by layer
- Run as a non-root user. Add a dedicated user in the Dockerfile and set
USER, so a process compromise does not automatically hand the attacker root inside the container. This alone does not stop a container-escape-class exploit, but it removes the most common privilege-escalation shortcut and is a prerequisite for the other controls below being fully effective: a read-only root filesystem matters far less if the compromised process is still root and able to remount things. - Drop Linux capabilities. Linux capabilities are the finer-grained privileges historically bundled together into "root": binding low network ports, changing file ownership, loading kernel modules, and dozens more. Docker grants a small default set to every container, and a typical API service needs almost none of it.
--cap-drop=ALL, then adding back only what is strictly required (commonly nothing at all for a plain HTTP service;CAP_NET_BIND_SERVICEonly if it must bind a port below 1024), shrinks what a compromised process can even ask the kernel to do, regardless of its user ID. - Read-only root filesystem.
--read-onlyat runtime makes the entire root filesystem immutable, combined with atmpfsmount for the one or two paths the application genuinely needs to write to, a cache directory,/tmp. Even a successfully injected payload cannot persist itself to disk inside the container or tamper with the application's own binaries. - Rootless containers, a different control from a non-root USER. The
USERinstruction changes which user ID the process runs as inside the container. Rootless mode goes a layer deeper: it runs the container engine itself without root privileges on the host, using Linux user namespaces so that even a process that is user ID 0 inside the container maps to an unprivileged, ordinary account on the host. A container-escape exploit that would normally hand an attacker host root instead hands them an unprivileged host account, a meaningfully smaller blast radius if the layers above somehow fail. - Seccomp. A seccomp (secure computing mode) profile is a kernel-enforced allowlist or denylist of which system calls a process may make at all. The engine's default profile already blocks a range of dangerous, rarely needed syscalls, including loading kernel modules; a custom profile generated from tracing what a specific workload actually uses, denying everything else, closes off entire exploitation techniques, since many container-breakout and privilege-escalation chains rely on a syscall the application itself never legitimately needs.
Enforcing and auditing at deploy time
None of the above helps if it is only a suggestion in a document. Enforce it at the deploy gate, a policy engine or admission controller that rejects a deployment specification missing a non-root user requirement, a dropped-capabilities list, or a read-only filesystem setting, rather than trusting every team to remember. Audit continuously by periodically diffing running container configurations against the required policy and alerting on drift, since a policy checked only at initial deploy time misses a manual run or an emergency change that bypassed the pipeline.
Worked example, measured
This exact hardening stack was built and run against a small HTTP service: a dedicated non-root user baked into the image, started with --read-only --cap-drop=ALL --security-opt no-new-privileges --tmpfs /tmp. Inside the running container, whoami returned the unprivileged application user, not root, and id showed no supplementary groups beyond that user's own group. The service still answered a real request over its published port with the expected response body, confirming the hardening did not break operability. Attempting to write a file outside the tmpfs mount (touch /app/newfile) failed with "Read-only file system," confirming the control was actually enforced rather than merely configured.
Trade-offs and pitfalls
A read-only root filesystem breaks any application that writes logs, temp files, or cache data to an unexpected path by default; audit what the application actually touches, or test it once with --read-only in staging, before enforcing it in production, and provide tmpfs or volume mounts for the legitimate write paths rather than disabling the control entirely. Dropping every capability can break something subtle that quietly relied on one, binding a low port, or a health check needing a specific network privilege, so validate functionally, not just that the image built successfully. A custom seccomp profile generated from one narrow test run can be too restrictive if it never exercised every code path, an error-handling branch that needs a syscall the happy path never calls, so generate it from realistic, broad testing, or accept a slightly looser default profile over an untested custom one that breaks in production on the first edge case.
A faulty release tagged prod just got deployed. Walk through how you'd perform an immediate rollback to the previous known-good image using image digests with a Docker Compose stack, and how automation (recording the last-known-good digest, a one-command rollback script) could make this safer and faster next time.
Sample Answer
Direct answer
A rollback under a faulty prod tag has one job: get back to the exact previous image bytes as fast as possible, without relying on a rebuild, since even a rebuild from "the same" source tag is not guaranteed to reproduce identical bytes if a base image or a dependency resolves differently since then. With Docker Compose, that means pointing the affected service at the previous image's content digest directly and recreating just that one service.
The immediate rollback
- Find the last known-good digest. If digests were already being recorded at deploy time, this is a lookup in the deploy log, not a hunt. Without that record,
docker images --digestsor the registry's own tag history is the fallback, which is exactly the gap worth closing before an incident, not during one. - Point the compose file at that digest, not the floating tag:
services:
app:
image: registry.example.com/app@sha256:d6cd6f0583314b6592c3e0b0ae2f6778a6b2373e9cf283c3d979c65c7ebf47d7
- Recreate only that service:
docker compose up -d --no-deps app(adding--force-recreateif Compose believes nothing changed). This pulls the pinned digest, guaranteed to be the same bytes as before since digests are content-addressed, and swaps only that one container, leaving sibling services and the network untouched. - Verify before declaring it done. Confirm the running container's actual image digest matches what was intended (
docker inspect --format='{{.Image}}'cross-checked against the pulled digest), and that health checks or smoke tests pass. A rollback that silently landed on the wrong digest is worse than a slower, correct one.
Making this safer and faster with automation
- Record last-known-good automatically. Every successful, health-checked deploy writes its digest to a small, durable record, a file, a row in a deploy-tracking table, or a tag like
app:last-goodthat only the automated pipeline is ever allowed to move. This turns step one above from a stressful hunt through logs into a lookup that already exists before the incident happens. - A one-command rollback script. Wraps the steps above into a single invocation: it reads the last-known-good record, updates the compose file's image reference for that one service, runs the targeted
docker compose up -d --no-deps, then runs the same smoke test the normal deploy pipeline runs before declaring success. The goal is that an on-call engineer under pressure runs one command instead of reconstructing this whole sequence from memory. - Guardrails. The rollback script should refuse to "roll back" to the digest that is already running, protecting against accidentally rolling back twice, and should log its own action to the same deploy-tracking record a normal deploy would, so the audit trail stays accurate even during an incident.
Trade-offs and pitfalls
Rolling back the image alone does not roll back anything that shipped alongside it: a database migration that ran with the bad release does not reverse itself, so a rollback plan needs a documented answer for what happens when the bad release also changed schema, not just code. --no-deps matters specifically because a full docker compose up -d can recreate dependent services too if their configuration appears to have changed, a much bigger blast radius (a bigger set of things a mistake could break) than the incident called for. Automating "last known good" purely as "the previous deploy" is subtly wrong if that previous deploy was also bad; the record should track the last deploy that actually passed a real health check, not merely the one immediately before this one chronologically.
What's the difference between CMD and ENTRYPOINT in a Dockerfile, and between their shell form and exec form? Give one example where ENTRYPOINT is preferred and one where CMD is preferred, and explain why exec form matters for how the container's main process receives signals like SIGTERM.
Sample Answer
Direct answer
CMD sets the default command a container runs, and it can be fully overridden by whatever is passed after the image name on docker run. ENTRYPOINT sets the fixed executable that always runs; anything from CMD, or arguments passed to docker run, is appended to it as arguments instead of replacing it. Shell form (CMD myapp --flag) runs the command through /bin/sh -c; exec form (CMD ["myapp", "--flag"]) runs the binary directly, without a shell in between, which matters specifically because of how Linux delivers signals like SIGTERM.
When to prefer each
Prefer ENTRYPOINT when the image should behave like a fixed, single-purpose binary and you want docker run arguments to feel like command-line flags to that binary; the classic example is a CLI tool image like docker run mytool --version, where ENTRYPOINT ["mytool"] plus no CMD means any arguments the user supplies are passed straight through to mytool. Prefer plain CMD when the image is meant to run one sensible default command that a user might reasonably want to replace entirely, such as a base development image where CMD ["python", "app.py"] is the default but someone might run docker run myimage bash to get a shell instead; with only CMD set, that fully replaces the default command as intended.
Shell form, exec form, and why exec form matters for signals
docker stop sends SIGTERM to the container's PID 1 process, waits a grace period (10 seconds by default), and then sends SIGKILL if the process has not exited. In shell form, the actual application is often started as a child of /bin/sh -c "..."; whether the shell then forwards SIGTERM to that child, or the shell instead sits as PID 1 and the application never sees the signal at all, is implementation detail you should not rely on, and it varies by exactly what the command string does (a single simple final command is sometimes optimized away by the shell into a direct exec, but anything with multiple statements, pipes, or && may leave the shell resident as PID 1). Exec form removes that ambiguity entirely by making your application the literal PID 1 process, guaranteed, with no shell in between.
That said, there is a second, separate wrinkle worth knowing: being PID 1 inside a container's process namespace changes how the Linux kernel applies default signal dispositions. A process that has not explicitly installed a SIGTERM handler normally terminates on SIGTERM by default, but the kernel suppresses that default action specifically for PID 1, so an application with no explicit signal handler can ignore SIGTERM even in exec form. This was verified directly: a plain sleep 100 running as exec-form PID 1 with no init process ignored SIGTERM completely and rode out the full grace period before being SIGKILLed (docker stop -t 3 took the full 3.2 seconds), while the same binary run under docker run --init (where a tiny init process, not the application, is PID 1, and sleep runs as a normal child at PID 2) received SIGTERM's default action immediately and exited in about 0.1 seconds.
What this means in practice
Exec form is necessary but not sufficient for graceful shutdown: it guarantees your application (not a shell) receives the signal Docker sends, but if your application does not install its own SIGTERM handler, running it as PID 1 without an init process can still mean it never actually reacts to the signal and just waits out the timeout before being force-killed. The two-part fix is exec form plus either an explicit signal handler in the application, or docker run --init (or an init process like tini baked into the image) so the application is not PID 1 at all and gets ordinary signal-handling semantics.
What's the difference between a process being alive inside a container and the application actually being healthy? How would you use a Docker HEALTHCHECK to reflect that distinction rather than just checking that the process hasn't crashed?
Sample Answer
Direct answer
A process being "alive" only means the operating system's process table still has an entry for PID 1 (the container's first and main process); it says nothing about whether that process is doing useful work. "Healthy" means the application can actually serve a request correctly right now: its dependencies (database, cache, downstream services) are reachable, and its own internal state (thread pool, connection pool, event loop) is not stuck. Docker's HEALTHCHECK instruction exists precisely to close that gap: it lets you define an active, application-level probe instead of relying on the passive fact that the process has not crashed.
Structured elaboration
- Why "alive" is not enough. A process can be alive and still be useless: deadlocked on a lock it will never release, stuck in an infinite retry loop against a database that no longer exists, or simply never having finished its startup sequence.
docker psreportingUp 12 minutesreflects only that the process has not exited; it makes no claim about correctness. - What
HEALTHCHECKadds. It runs a command on a schedule (--interval), gives it a--timeout, allows a--start-periodduring which failures do not count against the container (used for legitimately slow startup), and needs--retriesconsecutive failures before flipping the reported status fromstartingtounhealthy. Crucially, the command should exercise the same code path a real client would use, not just check that a socket accepts a TCP connection. - Writing a check that reflects health, not just aliveness. A weak check hits
/and expects any HTTP response. A meaningful check hits a dedicated endpoint (commonly/healthor/healthz) that the application itself computes: does it have an open, working connection to its database, is its background worker queue processing, has startup fully completed. The check should return a definite pass or fail, cheaply and quickly, without doing expensive work (do not run a full business transaction inside a health probe that fires every few seconds). - A concrete instruction:
HEALTHCHECK --interval=10s --timeout=3s --start-period=15s --retries=3 \
CMD curl -f http://localhost:8080/health || exit 1
Here /health is application code, for example: return 200 only if a lightweight SELECT 1 against the database succeeds within a short timeout, and 503 otherwise. That single distinction, "did I check a real dependency" versus "did I just answer any request", is what separates a liveness-flavored check from one that reflects actual health.
Worked example
An API's process stays alive but its database connection pool has silently exhausted (every connection is checked out and never returned, due to a bug that forgets to close a result set). Requests hang until they time out, but the process itself never crashes, so a plain "is the process running" check would report the container as fine indefinitely. A HEALTHCHECK calling /health, where the handler tries to borrow a connection from the pool with a one-second timeout and fails fast if none is available, flips to unhealthy within a few probe cycles. That gives an operator (or an external supervisor watching the status) a concrete, actionable signal well before every user-facing request starts failing.
Trade-offs & pitfalls
- Too shallow a check (any HTTP
200, or just a TCP connect) never catches internal degradation like the connection-pool example above; it only proves the listener thread is alive. - Too deep a check (a full end-to-end transaction against every downstream dependency) makes the health check itself slow, expensive, and a new source of false failures when an optional, non-critical dependency has a bad moment.
- On plain Docker,
HEALTHCHECKonly changes the status Docker reports (visible viadocker psanddocker inspect); it does not by itself restart the container. Acting on that status (restarting, rerouting traffic, alerting) is something you or an orchestrator has to build on top of it.
Unlock Full Question Bank
Get access to all Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.