Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
How do you set CPU and memory limits for a Docker container, both at runtime (--memory, --cpus, --memory-swap) and in Compose? Why do these limits matter for reliability when many containers share a host, and what's the practical difference between limiting resources in Docker versus relying on the host OS alone?
Sample Answer
Direct answer
At runtime you cap a container's resource use with --memory (a hard cap on RAM the container's cgroup, the Linux kernel mechanism that accounts and limits a group of processes' resource use, may use), --cpus (a fractional cap on CPU cores, for example 1.5 for one and a half cores' worth of CPU time), and --memory-swap (the combined cap on RAM plus swap; setting it equal to --memory disables additional swap for that container entirely). Compose expresses the same caps declaratively under a service's deploy.resources.limits (and, for guaranteed minimums, reservations). These limits matter because a host is shared: without them, one noisy or leaking container can starve every other container's CPU or trigger the kernel's out-of-memory killer against processes that had nothing to do with the problem. Setting limits in Docker, rather than trusting the host OS alone, means each container's ceiling is explicit, enforced per-container by the kernel's cgroup mechanism, and visible in the container's own configuration, instead of being an emergent, unpredictable property of whatever else happens to be running alongside it.
Structured elaboration
The three runtime flags
--memory=256m: a hard ceiling on the container's RAM. Exceeding it does not throttle the container gracefully; the kernel's out-of-memory killer terminates a process inside it, which is why containers hitting this limit exit with code 137 (128 plus signal 9,SIGKILL).--cpus=1.5: a fractional CPU cap enforced by the kernel's CPU scheduler quota mechanism (internally, a quota and period pair, for example 150 milliseconds of CPU time allowed per 100-millisecond period, which works out to one and a half cores' worth of time). The container can still use more than one core briefly if the host has spare capacity within that period, but its average use over time is capped.--memory-swap: the ceiling on RAM plus swap combined. Set it equal to--memoryto give the container no extra swap at all (predictable, fails fast under memory pressure instead of degrading into slow, swapped-out behavior); leave it unset or higher to allow some swap as a cushion, at the cost of unpredictable latency when the container is actually swapping.
The same caps in Compose
services:
app:
image: myapp:latest
deploy:
resources:
limits:
cpus: "1.5"
memory: 256M
reservations:
memory: 128M
limits is the hard ceiling, equivalent to the runtime flags above. reservations is a soft guarantee used for scheduling decisions (more relevant in a multi-node Swarm context, but harmless to declare locally). As of current Compose versions, plain docker compose up applies deploy.resources.limits directly to the container, the same as if you had passed --memory and --cpus on the command line; this is worth knowing explicitly because older documentation associated deploy.resources with Swarm mode only.
Why this matters for reliability with many containers on one host
- CPU is a compressible resource: a container that goes over its fair share slows its neighbors down (added latency) but does not crash them. Memory is not compressible: a container that overshoots gets killed, and without a per-container ceiling, the kernel's OOM killer picks a victim using its own heuristics across the whole host, which can land on an unrelated, otherwise healthy container instead of the actual offender.
- A hard memory limit turns "some process on this host eventually gets killed, chosen by a heuristic" into "this specific container gets killed, predictably, when it exceeds the budget you gave it," which is a much easier failure mode to detect, alert on, and attribute correctly.
Docker limits versus relying on the host OS alone
Without container-level limits, isolation depends entirely on however many containers happen to be co-located and how they individually behave that day; a single runaway process can degrade or take down everything sharing the host, and there is no per-workload accounting to tell you which container caused it. Explicit per-container limits make capacity planning a per-service decision (this service gets this much CPU and RAM, always) rather than an emergent property of the day's neighbor mix, and they make a container's own docker inspect output an honest record of what it is entitled to.
Worked example
A host runs four API containers with no memory limits set. A memory leak in one of them slowly consumes RAM until the host is under pressure; the kernel's OOM killer fires and terminates a process in a different, healthy container, because that container's process happened to score worse on the kernel's heuristics at that moment. After adding --memory=512m --memory-swap=512m to every container, the leaking one instead gets its own process killed once it crosses 512 MB, its docker inspect shows OOMKilled: true against that specific container, and the other three keep running untouched. The failure is now attributable and comes with a hard, checkable number instead of a shared, host-wide guess.
Trade-offs & pitfalls
- Setting
--memorytoo tight causes healthy workloads to be killed under normal, non-leaky peak load (a burst in traffic, a large batch job); size it from measured peak usage with headroom, not from idle-state observation. - Relying only on
--cpuswithout ever measuring actual usage under load can silently throttle a service during real traffic spikes, showing up as added latency that is easy to misdiagnose as a code regression. --memory-swapequal to--memory(no swap) is usually the right default for latency-sensitive services, since swapping trades a hard, visible failure for a slow, hard-to-diagnose one; leaving swap available can be appropriate for batch or background work where an occasional slowdown is preferable to being killed outright.
How do you make Docker image builds reproducible, so the same source and dependencies produce an equivalent image weeks or months later? Discuss lockfiles, pinning base images by digest instead of a tag, avoiding latest in build args, and how you'd verify reproducibility (hashes, binary diffs) and respond when a previously-reproducible build breaks.
Sample Answer
Direct answer
Reproducibility means the same source and dependency inputs produce an equivalent image weeks later, and every input that can silently change without your knowledge has to be pinned: the base image by content digest instead of a mutable tag, the application dependencies by a lockfile instead of a version range, and any build-time tooling by an explicit version instead of whatever latest happens to resolve to that day.
What actually breaks reproducibility, and how to close each gap
- Base image tags are mutable pointers, digests are not. A tag like
node:20-alpinecan point at a different set of bytes tomorrow (a security patch rebuild, a base OS bump). A digest,sha256:<hash of the image manifest>, is content-addressed: pulling that exact digest always returns the same bytes. As of this writing,docker inspect --format='{{index .RepoDigests 0}}'on a freshly pulledalpine:3.20resolves toalpine@sha256:d9e853e87e55526f6b2917df91a2115c36dd7c696a35be12163d44e6e2a4b6bcandnode:20-alpineresolves tonode@sha256:fb4cd12c85ee03686f6af5362a0b0d56d50c58a04632e6c0fb8363f609372293(verified locally against the live registry; those specific digests will eventually be superseded as the tags are rebuilt, which is exactly the point: pin the digest that was true for the build you shipped). A reproducibleFROMline looks likeFROM node@sha256:fb4cd12c85ee...rather thanFROM node:20-alpine. - Lockfiles pin the dependency graph, not just the top-level versions.
package-lock.json,poetry.lock,Pipfile.lock,go.sumrecord the exact resolved version of every transitive dependency.COPYthe lockfile before the manifest-driven install step, and install FROM the lockfile in strict mode (npm ciinstead ofnpm install,pip install --require-hashes -r requirements.txt,poetry install --sync) so the build fails loudly instead of silently re-resolving if the lockfile and manifest disagree. - Never let
latestleak in through a build argument. ADockerfilethat is otherwise pinned but takesARG TOOL_VERSION=latestas a default, or shells out topip install some-toolwith no version pin inside aRUNstep, reintroduces exactly the nondeterminism the digest pin was meant to remove.
Worked example: verifying and responding to a break
To verify reproducibility in practice, rebuild the image from the same pinned inputs and compare content, not just "it built successfully": docker build twice from a clean cache, then compare either the final image ID (a hash of the image config, which BuildKit, Docker's build engine and the default backend behind docker build, computes deterministically from identical layer content) or export both images with docker save and diff the resulting tarballs. Two images built minutes apart from the same digest-pinned base and the same lockfile should produce identical layer digests for every layer whose inputs did not change.
If a previously-reproducible build breaks, the pinning itself narrows the search space to a short list, because you have already eliminated the two most common culprits (base image drift and dependency resolution drift):
- Check whether the pinned digest still resolves. Rare on a primary registry, but a pull-through mirror or proxy cache can occasionally serve stale or re-tagged content; re-pull directly from the origin registry to rule this out.
- Check for an unlocked transitive dependency. Some ecosystems allow optional or platform-specific dependencies (npm's
optionalDependencies, Python extras) that a lockfile does not fully pin across platforms; a build on a different host architecture can resolve a different optional package even with an otherwise-locked file. - Check the build toolchain itself. If the base image includes a compiler or interpreter that was not separately pinned (only the OS layer was digest-pinned, not the language runtime inside it), a rebuild against a "same" digest but a different day can still differ if the image was rebuilt with an updated toolchain before you re-pulled it.
- Only then suspect true non-determinism in the build process (embedded timestamps, non-deterministic file ordering when creating archives), which is a real but deeper problem: fully byte-identical output additionally requires controlling for timestamps embedded in file metadata (commonly addressed with a fixed
SOURCE_DATE_EPOCHenvironment variable that build tools read instead of the real clock) and deterministic ordering when packing files into layers. Most teams do not need this level: matching content digests for the layers that should not have changed is the practical bar; true byte-identical output tarballs is a specialized requirement (regulated software supply chains, formal reproducible-builds programs) worth calling out as a deeper tier rather than the default target.
Trade-offs and pitfalls
Digest pinning trades automatic security patching for determinism: a digest never moves, so you stop getting base-image CVE fixes for free and need an active process (a dependency bot that opens a pull request when the upstream tag's digest changes, checked in CI before merging) to intentionally bump the pin on a schedule. Pinning and then never revisiting the pin is the common failure mode: it looks like "we solved reproducibility" while quietly accumulating unpatched vulnerabilities in a base layer nobody is watching.
For a stateful database service, would you use a named Docker volume or a host bind mount in development versus in production, and why? Discuss portability, performance, backup strategy, and permission issues you might hit (including SELinux on some Linux hosts).
Sample Answer
Direct answer
In development, a bind mount (a direct mapping of a path on the host's own file system into the container) is usually the better fit for a stateful database, because it puts the data somewhere a developer can see, back up, and delete with ordinary host tools, at the cost of portability and occasional permission friction. In production, a named Docker volume (storage that Docker itself creates and manages, identified by a name rather than a host path) is the better fit, because it is portable across hosts, does not depend on a specific host directory existing with the right permissions, and integrates with Docker's own backup and volume-driver tooling rather than whatever ad hoc scripts a bind mount setup would need.
Structured elaboration
Portability
A bind mount hardcodes a host path into the container's configuration; moving the workload to a different host, or a different developer's machine with a different directory layout, means updating that path everywhere it is referenced. A named volume is referenced only by name; Docker manages where its data actually lives on disk, and that mapping does not need to change when you move the container definition to a new host, or attach the same volume to a replacement container after an upgrade.
Performance
On Linux hosts, both a bind mount and a named volume are ordinary file system access with negligible overhead. The gap shows up specifically on Docker Desktop for macOS or Windows, where a bind mount crosses into a virtualized Linux environment and back, and can be noticeably slower for I/O-heavy workloads (a database doing many small reads and writes is exactly this case) than a named volume, which lives natively inside that virtualized environment and avoids the cross-boundary file-sharing layer entirely.
Backup strategy
A bind mount's data is just files at a known host path, so host-level backup tools (a cron job running rsync or a snapshotting file system) work without any Docker-specific knowledge, which is convenient for a single developer's machine. A named volume's data lives inside Docker's own storage area, so backing it up properly means either running a helper container that mounts the volume and streams a tar archive out (docker run --rm -v myvolume:/data -v $(pwd):/backup alpine tar czf /backup/myvolume.tar.gz -C /data .), or using a volume driver that has its own backup integration. This is a small amount of extra ceremony that is worth it in production for the portability and permission benefits above.
Permission issues, including SELinux
A bind mount exposes the container to the host's actual file ownership and permission model directly: a container process running as a different user ID (UID) than whatever owns the host directory can get permission-denied errors immediately, since the container is not exempt from ordinary Unix file permissions just because it is a container. On a host with SELinux (Security-Enhanced Linux, a mandatory access control system that restricts what a process may do to specific files beyond ordinary Unix permissions, common on Red Hat family distributions) enforcing, a bind-mounted directory additionally needs the right SELinux label before a container process can access it at all, even if the Unix permission bits look correct; Docker's :z and :Z mount flags request that label automatically (:z shares the label across multiple containers that need to read the same host directory; :Z marks it private to just this one container). A named volume mostly avoids this class of problem, since Docker creates and manages its storage directly and typically already applies a compatible label, rather than depending on whatever an existing host directory happened to be labeled for some other purpose.
Worked example
A developer runs Postgres locally with -v ./pgdata:/var/lib/postgresql/data (a bind mount) so they can browse the raw data files, delete the directory to reset state instantly, and back it up by just copying the folder; on a Fedora workstation with SELinux enforcing, the container fails to start with a permission error until the mount is changed to -v ./pgdata:/var/lib/postgresql/data:Z, which applies the correct private SELinux label to that directory for this container. In production, the same service instead uses -v pgdata:/var/lib/postgresql/data (a named volume): the same container definition now works unchanged on any host in the fleet, regardless of that host's specific directory layout or SELinux configuration, and backups run through a scheduled helper container rather than a host-specific script tied to one machine's file paths.
Trade-offs & pitfalls
- Using a bind mount in production because "it worked in development" reintroduces exactly the portability and permission problems a named volume exists to avoid, the first time the container needs to run on a different host.
- Using a named volume in development removes the convenience of browsing and editing the data directly with ordinary host tools, which is often exactly what a developer wants while debugging.
- Forgetting the SELinux label on a bind mount on an enforcing host produces a permission error that looks identical to a plain Unix permissions problem; check
getenforceon the host before assuming the fix is a UID orchmodchange.
A containerized service doesn't shut down cleanly during deployments and leaves requests hanging for several seconds. Walk through how you'd investigate PID 1 behavior, signal handling, and graceful termination in Docker to find the cause.
Sample Answer
Direct answer
When a container hangs for several seconds during shutdown, the first thing I check is what process is actually running as PID 1 (the first process in the container's process namespace, which the Linux kernel treats specially: it does not apply default signal behavior the way it does for every other process) and whether that process, or whatever is standing between it and the real application, actually reacts to SIGTERM (the signal Docker sends to ask a process to shut down gracefully). Most "hangs on shutdown" bugs are the application never receiving or never acting on that signal, and Docker eventually giving up and force-killing it once its grace period expires, which is exactly the multi-second delay the symptom describes.
Investigation steps
- Confirm what PID 1 actually is.
docker exec <container> ps auxand look at the process with PID 1. If it is the application binary directly, it is receiving signals directly. If it is a shell (/bin/sh -c "...", which is what Docker uses when aCMDis written in shell form rather than exec form), the shell is PID 1 and the real application is a child process; many shells do not automatically forward a receivedSIGTERMto their children, so the child, the process actually serving requests, may never see the signal at all. - Reproduce and time the shutdown. I verified this directly: running a container whose entrypoint script traps and ignores
SIGTERM(trap 'echo ignoring' TERM; while true; do sleep 1; done),docker stoptook 10.177 seconds to return, anddocker inspect --format='{{.State.ExitCode}}'showed exit code 137 (128 + 9, meaning the process was terminated bySIGKILL, not by exiting on its own). That 10-second figure is not incidental: it is Docker's documented default grace period between sendingSIGTERMand escalating toSIGKILL(configurable viadocker stop -t <seconds>, or aSTOPSIGNAL/stop_grace_periodsetting), which is exactly why a process that ignoresSIGTERMproduces a delay in that neighborhood rather than an immediate or indefinite hang. - Distinguish "ignoring the signal" from "not receiving it." If PID 1 is a shell wrapping the real application, the fix target is different than if the application itself is PID 1 and simply has no signal handler: in the shell case, the shell needs to
execthe application (replacing itself rather than forking a child) so the application becomes PID 1 directly, or a minimal init process needs to sit in front of it to forward signals correctly. - Fix by giving the application a real handler, not just fixing signal delivery. Docker's
--initflag (or an explicit minimal init binary such astini) solves signal forwarding and reaps zombie processes, but it does not by itself make the application shut down gracefully; the application still needs to catchSIGTERM, stop accepting new connections, finish in-flight requests, and exit on its own. Fixing only the forwarding without adding an actual handler still ends in aSIGKILLafter the same grace period, just with the signal correctly reaching the right process. - Verify the fix the same way you reproduced the bug. Re-running the same
docker stoptest after adding a handler should return in well under the grace period, with a clean exit code (typically 0, or 128 + the signal number the application chose to exit on) instead of 137.
Worked example
A Dockerfile using CMD ["node", "server.js"] (exec form) makes the Node process PID 1 directly, so SIGTERM reaches it immediately; if that process has no listener for the SIGTERM event, Node's default behavior does terminate the process, but without ever draining in-flight HTTP requests first, which is why "leaves requests hanging" points more precisely at either a shell-form CMD swallowing the signal, or a caught-but-mishandled signal (a handler registered that does something other than start a graceful drain, such as the trap in my reproduction above, which explicitly ignored the signal and did nothing). The fix in the application is to register a SIGTERM handler that stops the HTTP server from accepting new connections, waits for existing requests to complete (with its own internal timeout shorter than Docker's grace period), and then exits, rather than relying on Docker's grace period as the mechanism that ends in-flight work.
Trade-offs and pitfalls
Reducing docker stop -t to make deployments feel faster without fixing the application's signal handling just changes the symptom from "hangs for several seconds" to "drops in-flight requests sooner," which is worse, not better, for exactly the reliability reason the question is asking about. Adding tini/--init is close to free (negligible overhead, one extra thin process in the tree) and fixes real classes of bugs (signal forwarding through a shell, zombie process accumulation), but it is not a substitute for the application actually implementing a graceful drain, and treating it as one is the most common half-fix I see.
You're asked to harden a container image that currently runs as root, ships shell utilities, and is built from a large general-purpose base image, exposing an API to untrusted input. Propose a runtime hardening policy: seccomp profiles, dropping Linux capabilities, a read-only root filesystem, rootless containers, and how you'd enforce and audit these at deploy time, while keeping the service actually operable.
Sample Answer
Direct answer
For a container that runs as root, ships general shell utilities, is built from a large general-purpose base, and is directly exposed to untrusted input, the goal is defense in depth: assume the application itself will eventually be exploited, since it is the one thing an attacker directly controls, and make sure a successful exploit lands in a process with almost nothing useful available to it. No root privileges, no unnecessary Linux capabilities, no writable filesystem, and no shell or general-purpose tooling to pivot with.
The policy, layer by layer
- Run as a non-root user. Add a dedicated user in the Dockerfile and set
USER, so a process compromise does not automatically hand the attacker root inside the container. This alone does not stop a container-escape-class exploit, but it removes the most common privilege-escalation shortcut and is a prerequisite for the other controls below being fully effective: a read-only root filesystem matters far less if the compromised process is still root and able to remount things. - Drop Linux capabilities. Linux capabilities are the finer-grained privileges historically bundled together into "root": binding low network ports, changing file ownership, loading kernel modules, and dozens more. Docker grants a small default set to every container, and a typical API service needs almost none of it.
--cap-drop=ALL, then adding back only what is strictly required (commonly nothing at all for a plain HTTP service;CAP_NET_BIND_SERVICEonly if it must bind a port below 1024), shrinks what a compromised process can even ask the kernel to do, regardless of its user ID. - Read-only root filesystem.
--read-onlyat runtime makes the entire root filesystem immutable, combined with atmpfsmount for the one or two paths the application genuinely needs to write to, a cache directory,/tmp. Even a successfully injected payload cannot persist itself to disk inside the container or tamper with the application's own binaries. - Rootless containers, a different control from a non-root USER. The
USERinstruction changes which user ID the process runs as inside the container. Rootless mode goes a layer deeper: it runs the container engine itself without root privileges on the host, using Linux user namespaces so that even a process that is user ID 0 inside the container maps to an unprivileged, ordinary account on the host. A container-escape exploit that would normally hand an attacker host root instead hands them an unprivileged host account, a meaningfully smaller blast radius if the layers above somehow fail. - Seccomp. A seccomp (secure computing mode) profile is a kernel-enforced allowlist or denylist of which system calls a process may make at all. The engine's default profile already blocks a range of dangerous, rarely needed syscalls, including loading kernel modules; a custom profile generated from tracing what a specific workload actually uses, denying everything else, closes off entire exploitation techniques, since many container-breakout and privilege-escalation chains rely on a syscall the application itself never legitimately needs.
Enforcing and auditing at deploy time
None of the above helps if it is only a suggestion in a document. Enforce it at the deploy gate, a policy engine or admission controller that rejects a deployment specification missing a non-root user requirement, a dropped-capabilities list, or a read-only filesystem setting, rather than trusting every team to remember. Audit continuously by periodically diffing running container configurations against the required policy and alerting on drift, since a policy checked only at initial deploy time misses a manual run or an emergency change that bypassed the pipeline.
Worked example, measured
This exact hardening stack was built and run against a small HTTP service: a dedicated non-root user baked into the image, started with --read-only --cap-drop=ALL --security-opt no-new-privileges --tmpfs /tmp. Inside the running container, whoami returned the unprivileged application user, not root, and id showed no supplementary groups beyond that user's own group. The service still answered a real request over its published port with the expected response body, confirming the hardening did not break operability. Attempting to write a file outside the tmpfs mount (touch /app/newfile) failed with "Read-only file system," confirming the control was actually enforced rather than merely configured.
Trade-offs and pitfalls
A read-only root filesystem breaks any application that writes logs, temp files, or cache data to an unexpected path by default; audit what the application actually touches, or test it once with --read-only in staging, before enforcing it in production, and provide tmpfs or volume mounts for the legitimate write paths rather than disabling the control entirely. Dropping every capability can break something subtle that quietly relied on one, binding a low port, or a health check needing a specific network privilege, so validate functionally, not just that the image built successfully. A custom seccomp profile generated from one narrow test run can be too restrictive if it never exercised every code path, an error-handling branch that needs a syscall the happy path never calls, so generate it from realistic, broad testing, or accept a slightly looser default profile over an untested custom one that breaks in production on the first edge case.
Unlock Full Question Bank
Get access to all Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.