Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
You built an image using a multi-stage build, but at runtime the container fails with an ImportError or a missing shared library. Walk through a step-by-step debugging plan (using tools like docker history, an interactive docker run --rm -it, or dive) to find which stage or layer omitted the dependency, and how you'd fix it without bloating the final image back up.
Sample Answer
Direct answer
A multi-stage build failing with a missing shared library at runtime almost always means a runtime-only dependency (a shared library that a compiled extension needs, but that was only ever installed as a side effect of the builder stage's toolchain) never made it across the COPY --from=builder boundary; the fix is to find exactly which file or package is missing and add that one thing to the final stage, not to widen what gets copied or fall back to a single fat stage.
Step-by-step debugging plan
- Reproduce interactively against the final image, not the builder.
docker run --rm -it --entrypoint sh <final-image>and attempt the failing import or command by hand, to get the exact error and confirm it happens in the shipped image, not just in some other environment. - Identify precisely what is missing.
ldd(a standard Linux utility that lists the shared libraries a binary depends on) run against the specific binary or shared object (ldd $(which python3), orldd path/to/the_compiled_extension.so) prints=> not foundnext to any dependency the runtime image cannot resolve. This turns a vagueImportErrorinto a concrete missing filename. - Compare against the discarded builder stage.
docker historyonly shows the layers of the image you run it against, which for a multi-stage build is the FINAL stage; the intermediate builder stage's layers are discarded once the build finishes and are not visible this way. To inspect what the builder stage actually had, build up to it explicitly and tag it:docker build --target <builder-stage-name> -t debug-builder ., then rundocker history debug-builderanddocker run --rm -it debug-builder shto confirm the missing file genuinely exists there. - Use
dive(a terminal tool for browsing a Docker image's layers and file contents) on the taggeddebug-builderimage to search for the specific missing path or package by name, confirming both that it exists in the builder and which layer introduced it, then check theCOPY --from=builderline in the final stage to see whether that path was actually included in what got copied over. - Fix with the smallest addition that closes the gap, either a targeted
COPY --from=builder /exact/path/to/lib.so /usr/lib/for a single file, or installing the specific runtime package via the distribution's package manager (for exampleapt-get install -y --no-install-recommends libpq5), never the corresponding-dev/full build-toolchain package, which drags headers, compilers, and static libraries back into the image the multi-stage build was written to keep out. - Verify without reintroducing bloat. Re-run the same reproduction from step 1 to confirm the import now succeeds, and check
docker imagesfor the size delta: a targeted runtime library addition should move the image size by kilobytes, not the hundreds of megabytes a reintroduced build toolchain would add.
Worked example
I verified this exact failure shape directly: python:3.12-slim, a common runtime base for a multi-stage Python build, does not have the PostgreSQL client library installed by default (ldconfig -p | grep -i libpq inside a fresh python:3.12-slim container returns nothing). A realistic version of this bug: a python:3.12 builder stage compiles psycopg2 from source, which links against libpq because the builder stage has libpq-dev installed (it needs the development headers to compile against); the final stage COPY --from=builder /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packagess the compiled package across, but never installs the runtime libpq5 package in the slim final stage. The result at runtime is ImportError: libpq.so.5: cannot open shared object file: No such file or directory, exactly the symptom ldd against the compiled psycopg2 .so file would surface directly. The fix is a single added line in the runtime stage: RUN apt-get update && apt-get install -y --no-install-recommends libpq5 && rm -rf /var/lib/apt/lists/*, which adds the shared library alone, not the development headers or compiler that libpq-dev would bring back.
Trade-offs and pitfalls
The tempting shortcut, when a missing dependency shows up, is to just install the same -dev package in the final stage that the builder stage used, since it is guaranteed to fix the immediate error; this reliably reintroduces the exact bloat multi-stage builds exist to avoid, because dev packages typically pull in a compiler toolchain and headers that a running application never needs. A second pitfall is fixing the symptom for one missing library and stopping, without checking whether the same class of gap exists for other compiled extensions in the same image; ldd against every compiled .so file the application actually loads, not just the first one that errored, catches the rest before they surface one at a time in later incidents.
An image that used to be around 900 MB is now over 4 GB after recent changes. Walk through how you'd investigate layer by layer to find the source of the bloat, and name the most common Dockerfile or package-manager mistakes that cause images to balloon like this.
Sample Answer
Direct answer
Start with docker history, which prints every layer in the image along with the size it added and the instruction that created it; sort mentally (or pipe through sort -h on the size column) to find the one or two layers responsible for most of the growth, then go fix whatever Dockerfile instruction produced them. Growth from 900 MB to 4+ GB is almost always one of a handful of well-known mistakes, not death by a thousand small cuts.
A layer-by-layer investigation method
- Run
docker history --no-trunc <image>and look at theSIZEcolumn for each layer; the--no-truncflag matters because it shows the full command instead of a truncated one, which you need to tell two similarRUNlines apart. - For any suspiciously large layer, look at what
CreatedBycommand produced it; aRUN apt-get installorRUN pip installline that jumps by hundreds of megabytes is your first suspect. - Check the base image itself:
docker images <base-image>for the tag you are using; a fullpython:3.12ornode:20image can be over a gigabyte on its own, versus low hundreds of megabytes for the-slimvariant. - Check whether the build context has a
.dockerignore; without one, aCOPY . .can silently pull in.git, local virtual environments, ornode_modulesfrom the host, none of which belong in the image.
Worked example
Building a deliberately unoptimized image and inspecting it:
FROM python:3.12
RUN apt-get update && apt-get install -y build-essential
RUN pip install --no-cache-dir pandas==2.2.3 scikit-learn==1.5.2
COPY . .
$ docker history --no-trunc --format "{{.Size}}\t{{.CreatedBy}}" bloated-image
393MB RUN /bin/sh -c pip install --no-cache-dir pandas==2.2.3 scikit-learn==1.5.2
22.6MB RUN /bin/sh -c apt-get update && apt-get install -y build-essential
...
73.5MB RUN ... (compiling Python from source inside the base image)
$ docker images bloated-image --format '{{.Size}}'
2.14GB
Two things jump out immediately from that history: the pip install layer alone is 393 MB, and the base image is python:3.12 (the full, non-slim variant), which compiles CPython from source and measures roughly 1.6 GB to 2 GB on its own before a single application dependency is added (the exact figure moves a little with CPU architecture and the current Debian point release). For comparison, python:3.12-slim pulled fresh is 204 MB, roughly an eighth of that (measured directly: 204 MB against a 1.62 GB python:3.12 pull), not the mere rounding difference a smaller gap would suggest. In a real "900 MB became 4 GB" case, docker history on the new image would show the same signature: either a newly-added large RUN layer, or a base image tag that quietly changed from a slim variant to a full one (for example, someone bumped FROM python:3.12-slim to FROM python:3.12 to "fix" a missing build tool, and never reverted it).
The usual suspects
Using a full, non-slim base image instead of a slim or distroless one; installing build tools (compilers, -dev headers) in the same layer as the final application instead of a multi-stage build, so the toolchain ships in production; not cleaning package-manager caches in the same RUN layer they were populated in (rm -rf /var/lib/apt/lists/* has to happen in the same RUN as the apt-get install, or the cache is already baked into an earlier layer and cannot be removed by a later one); and a missing or incomplete .dockerignore letting COPY . . pull in .git, log files, or dependency directories from the host.
Discuss the trade-offs of squashing image layers (--squash or tools like docker-slim) versus keeping many granular layers: rebuild performance, caching, image size, transparency for security scanners, and the ability to patch incrementally. Give two scenarios where squashing helps and two where it hurts.
Sample Answer
Direct answer
Squashing collapses an image's layer history into one (or few) layers, which can shrink the download for that specific image and hide files that were later deleted, but it also throws away the layer-level cache sharing, incremental patchability, and per-command scan attribution that make Docker's layered model useful in the first place. My default is to not squash, and instead architect the Dockerfile (multi-stage builds, combining cleanup into the same RUN that created the mess) so squashing is unnecessary; I reach for it only in the narrow cases below.
What each dimension actually costs
| Dimension | Many granular layers | Squashed image |
|---|---|---|
| Cross-build cache reuse | Each unchanged layer is reused on rebuild, and can be shared across different images built from the same base, saving real registry storage and pull bandwidth fleet-wide | Destroyed: a squashed image is one monolithic layer with nothing to share, even with an unsquashed image built from the identical Dockerfile |
| This image's own size | Can carry real dead weight: a RUN apt-get install build-essential && ... && apt-get remove build-essential still leaves the installed bytes in an earlier layer, with the removal only adding a "whiteout" marker in a later layer, not shrinking the tar | Removes dead weight for real: deleted-then-recreated files collapse away, often producing a genuinely smaller push and pull for that one image |
| Security scanner attribution | Scanners can map a package or vulnerability back to the specific instruction/layer that introduced it, which matters for knowing what to patch and where | That per-instruction provenance is lost; the scanner can still see final file contents, but "which Dockerfile line added this" is gone |
| Incremental patching / rollback | A base image bump only requires pulling the changed layer; downstream nodes reuse everything else they already have | Any change requires re-pulling and re-pushing the entire squashed blob, since there is no smaller unit to diff against |
Worth noting on current tooling: the legacy docker build --squash flag required enabling the (non-default) experimental daemon feature and only worked with the classic, non-BuildKit builder. On a current install (verified here: Docker 29.4.0, buildx 0.33.0, daemon experimental features off), neither docker build --help nor docker buildx build --help expose a --squash flag at all, since BuildKit, today's default builder, does not carry that flag forward. The practical modern equivalents are a multi-stage build that COPYs only the final artifacts into a fresh minimal stage (getting squash's size benefit by construction, with none of its downsides), or a third-party post-processing tool like docker-slim, which analyzes and strips an image after the fact.
Two scenarios where squashing helps
- A one-off, ad hoc image with no shared base and no iteration. A nightly batch or experimentation image cobbled together from a long chain of manual
apt/pipinstalls and cleanup commands, built once, run once, and discarded, never used as a base for anything else. There is no caching benefit to lose because the image is never rebuilt incrementally, so squashing's dead-weight removal is pure upside. - A public, single-purpose CLI tool image where the goal is the smallest possible download for end users who pull it once and do not care whether its layers match anything else in the world.
Two scenarios where squashing hurts
- A shared base image consumed by many downstream services. If thirty services build
FROMthe same well-ordered base, granular layers let every one of those builds and every node pulling them reuse the shared, cached layers. Squashing the base multiplies registry storage and pull bandwidth across the whole organization for no compensating benefit, since the base was never the bottleneck. - A pipeline with per-layer security attestation or audit requirements, where a reviewer needs to prove which build step introduced a specific package or file for compliance or incident-response purposes. Squashing destroys exactly the evidence trail that requirement depends on.
Trade-offs and pitfalls
The core pitfall is treating squashing as a general-purpose size-reduction technique rather than a narrow tool for a specific shape of image (one-off, unshared, un-iterated). The available fix that most senior engineers reach for instead is designing the Dockerfile so the dead weight never exists in the first place (installing, using, and cleaning up build tools in the same RUN layer, and separating the true final artifact into its own minimal stage), which gets the size benefit without sacrificing cache reuse or scanner provenance for anyone.
You need to ship a container for a Python model that depends on NumPy, SciPy, and a custom C++ or CUDA extension compiled from source. Explain how you'd structure a multi-stage build to compile the native dependencies in one stage and copy only the runtime artifacts into a minimal final stage, and how you'd avoid runtime linker or ABI mismatches between the two stages.
Sample Answer
Direct answer
Put the compiler toolchain, headers, and native build (NumPy, SciPy, and the custom C++ or CUDA, NVIDIA's GPU programming toolkit, extension) in a builder stage, then copy only the compiled artifacts, the extension's shared object file and whatever Python packages depend on it, into a separate, minimal final stage that never installs a compiler. The one rule that prevents runtime linker and ABI (application binary interface, the low-level contract, such as function calling conventions and symbol layout, that a compiled extension and the interpreter or C library it links against must agree on) failures is that both stages must share the same base distribution and C library, so a .so built in one loads cleanly in the other; mixing an Alpine (musl C library) builder with a Debian or Ubuntu (glibc) final stage is the single most common way this breaks, because the two C libraries are not binary-compatible.
Structured elaboration
Structuring the multi-stage build
FROM python:3.12-slim AS builder
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc python3-dev && rm -rf /var/lib/apt/lists/*
# For a CUDA-enabled extension, use an NVIDIA CUDA "devel" base here instead
# (it ships the compiler and headers); the final stage below would then use
# the matching CUDA "runtime" base, not a plain python:slim image.
WORKDIR /build
COPY setup.py fastmath.c ./
RUN pip install --no-cache-dir setuptools \
&& python3 setup.py build_ext --inplace
FROM python:3.12-slim AS final
WORKDIR /app
RUN pip install --no-cache-dir numpy scipy # plain runtime deps, prebuilt
# wheels, nothing to compile here
COPY --from=builder /build/fastmath*.so ./
COPY app.py .
CMD ["python3", "app.py"]
The final stage never installs gcc or python3-dev; it only ever receives the already-compiled .so file via COPY --from=builder. This is what actually shrinks the image and reduces its attack surface: the compiler, headers, and any build-only tooling simply never exist in the artifact that ships to production. NumPy and SciPy themselves need no build stage at all here: on a standard glibc-based image they install from prebuilt wheels with no compilation, so only the genuinely custom, source-compiled extension needs the multi-stage treatment.
Why the ABI mismatch happens, specifically
A compiled Python extension is tagged with the exact platform it was built for, encoded right into its filename, for example fastmath.cpython-312-aarch64-linux-musl.so when built on Alpine (musl) versus ...-linux-gnu.so (or a bare .so) when built on Debian or Ubuntu (glibc). If the builder stage is Alpine and the final stage is Debian-based, the compiled extension links against musl's C library at build time; the glibc-based final stage does not have that library at all, so loading the extension fails at import time with an error naming the missing shared library, not a helpful "wrong platform" message. The fix is not a build flag; it is choosing the same base distribution family for both stages, or, if a smaller final image than the builder's own base is wanted, choosing a final base from the same family (for example both Debian-derived) rather than crossing between musl and glibc.
Avoiding the mismatch in practice
- Pin both stages to the same base image family and, ideally, the same major version, so their C libraries match exactly.
- If the final image genuinely needs to be smaller than the builder's base, use a slimmer image from the same family (
python:3.12-slimas both builder and final, rather thanpython:3.12-alpineas builder andpython:3.12-slimas final) instead of crossing library implementations. - For CUDA specifically, match the CUDA toolkit's minor version between the builder's "devel" image (compiler and headers) and the final stage's "runtime" image (no compiler, just the shared libraries the compiled code links against at run time); a CUDA extension compiled against one minor version can fail to load, or silently misbehave, against a runtime image pinned to a different one.
- Any Python-level runtime dependency the extension itself imports (NumPy, SciPy) still needs to be installed in the final stage too; copying the compiled
.sofile alone does not bring its Python-level dependencies with it.
Worked example
This exact scenario was built and run directly. A team builds a custom extension in an Alpine-based builder stage; the compiled artifact lands as fastmath.cpython-312-aarch64-linux-musl.so, its platform tag baked into the filename exactly as the build tool reports it. Copying that file as-is into a python:3.12-slim (Debian, glibc) final stage and importing it does not even get as far as the dynamic linker: Python's import system matches a compiled extension's filename against the interpreter's own list of recognized platform suffixes before ever trying to load it, and -linux-musl is not one of the suffixes a glibc-based interpreter recognizes, so the import fails with ModuleNotFoundError: No module named 'fastmath', the same error you would get if the file were simply missing entirely. If that filename check is bypassed, for example by renaming the file to a bare fastmath.so so the interpreter actually attempts to load it, the real, deeper failure surfaces: ImportError: libc.musl-aarch64.so.1: cannot open shared object file: No such file or directory, because the extension was in fact linked against musl and the glibc-based image has no such library at any path. Either way the fix is the same: switching the builder stage to python:3.12-slim as well (installing gcc and python3-dev via apt-get instead of apk) produces fastmath.cpython-312-aarch64-linux-gnu.so, which loads cleanly in the Debian-based final stage and runs correctly, computing the same result it would on any other correctly matched glibc setup. The only Dockerfile change that mattered was making both stages share the same base family; no code changed.
Trade-offs & pitfalls
- Chasing the smallest possible final image by reaching for Alpine is a common instinct, but for anything with compiled native extensions it trades a real, hard-to-diagnose runtime failure for a modest size saving; measure whether the size difference actually matters before accepting that risk.
- A
ModuleNotFoundErroror a missing-shared-library error at import time, for an extension that clearly built successfully in the builder stage, should immediately raise the question of a base-image mismatch between stages, before looking anywhere else. - Forgetting to install the extension's own Python-level dependencies (NumPy, SciPy) in the final stage is a separate, equally common mistake: the compiled
.sofile loads, but importing the module still fails because the pure-Python packages it depends on at import time were never installed there.
What is a multi-stage Docker build, and when would you use one? Walk through a scenario where it substantially reduces final image size, explain how COPY --from transfers artifacts between stages, and name one trade-off multi-stage builds introduce during local development.
Sample Answer
Direct answer
A multi-stage build lets a Dockerfile use one image to build an application (with compilers, build tools, and full source available) and a separate, minimal image to run it, copying across only the finished artifact between them. You reach for one whenever your build tools are not needed at runtime, which is most compiled or bundled applications: the final image ships the binary or bundled assets, not the toolchain that produced them.
How COPY --from transfers artifacts between stages
Each FROM line in a Dockerfile starts a new, independent stage; naming a stage (FROM golang:1.23 AS build) lets a later stage reference it by name in COPY --from=build <path> <path>, which copies files directly from that earlier stage's filesystem, not from the build context on disk. Nothing else about the build stage (its installed compilers, its intermediate object files, its full source tree) carries over; only the specific files named in the COPY --from line do.
Worked example, verified end to end
A Go binary built in one stage and shipped from scratch in the next:
FROM golang:1.23-bookworm AS build
WORKDIR /src
COPY main.go .
RUN CGO_ENABLED=0 GOOS=linux go build -o /out/app main.go
FROM scratch
COPY --from=build /out/app /app
ENTRYPOINT ["/app"]
Building and running this produced a 3.4 MB final image (measured directly) and printed the program's output correctly:
$ docker build -t scratch-demo .
... (build output omitted for length; ends with "naming to ... scratch-demo")
$ docker images scratch-demo --format '{{.Size}}'
3.4MB
$ docker run --rm scratch-demo
hello from scratch
The golang:1.23-bookworm build stage itself is several hundred megabytes (the Go toolchain, standard library, build cache), none of which appears in the final image at all; only the compiled 3.4 MB binary was copied across by COPY --from=build. Compare that to a naive single-stage Dockerfile that built and ran the same binary from golang:1.23-bookworm directly: the final image would carry the entire Go toolchain into production for no runtime benefit.
The trade-off during local development
Multi-stage builds add a layer of indirection that costs you at inner-loop iteration speed: every local rebuild re-runs the build stage (even with caching, verifying the cache is warm and correctly ordered takes some care), and debugging a build failure means figuring out which stage failed and inspecting that stage's intermediate filesystem specifically, rather than just looking at one flat set of layers. Some teams keep a separate, simpler single-stage Dockerfile (or a Compose override) purely for local development with hot-reload, and reserve the full multi-stage build for CI and production images, precisely to avoid paying that iteration cost on every local change.
Unlock Full Question Bank
Get access to all Containerization and Docker Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.