Containerization and Docker Fundamentals Questions
Packaging applications into containers: images and layers, Dockerfiles, registries, image optimization and security, container networking and storage, and the container runtime model. Covers how containers differ from virtual machines, image build and management, and the fundamentals that underpin any orchestration platform. The container primitive before orchestration.
You're asked to harden a container image that currently runs as root, ships shell utilities, and is built from a large general-purpose base image, exposing an API to untrusted input. Propose a runtime hardening policy: seccomp profiles, dropping Linux capabilities, a read-only root filesystem, rootless containers, and how you'd enforce and audit these at deploy time, while keeping the service actually operable.
Sample Answer
Direct answer
For a container that runs as root, ships general shell utilities, is built from a large general-purpose base, and is directly exposed to untrusted input, the goal is defense in depth: assume the application itself will eventually be exploited, since it is the one thing an attacker directly controls, and make sure a successful exploit lands in a process with almost nothing useful available to it. No root privileges, no unnecessary Linux capabilities, no writable filesystem, and no shell or general-purpose tooling to pivot with.
The policy, layer by layer
- Run as a non-root user. Add a dedicated user in the Dockerfile and set
USER, so a process compromise does not automatically hand the attacker root inside the container. This alone does not stop a container-escape-class exploit, but it removes the most common privilege-escalation shortcut and is a prerequisite for the other controls below being fully effective: a read-only root filesystem matters far less if the compromised process is still root and able to remount things. - Drop Linux capabilities. Linux capabilities are the finer-grained privileges historically bundled together into "root": binding low network ports, changing file ownership, loading kernel modules, and dozens more. Docker grants a small default set to every container, and a typical API service needs almost none of it.
--cap-drop=ALL, then adding back only what is strictly required (commonly nothing at all for a plain HTTP service;CAP_NET_BIND_SERVICEonly if it must bind a port below 1024), shrinks what a compromised process can even ask the kernel to do, regardless of its user ID. - Read-only root filesystem.
--read-onlyat runtime makes the entire root filesystem immutable, combined with atmpfsmount for the one or two paths the application genuinely needs to write to, a cache directory,/tmp. Even a successfully injected payload cannot persist itself to disk inside the container or tamper with the application's own binaries. - Rootless containers, a different control from a non-root USER. The
USERinstruction changes which user ID the process runs as inside the container. Rootless mode goes a layer deeper: it runs the container engine itself without root privileges on the host, using Linux user namespaces so that even a process that is user ID 0 inside the container maps to an unprivileged, ordinary account on the host. A container-escape exploit that would normally hand an attacker host root instead hands them an unprivileged host account, a meaningfully smaller blast radius if the layers above somehow fail. - Seccomp. A seccomp (secure computing mode) profile is a kernel-enforced allowlist or denylist of which system calls a process may make at all. The engine's default profile already blocks a range of dangerous, rarely needed syscalls, including loading kernel modules; a custom profile generated from tracing what a specific workload actually uses, denying everything else, closes off entire exploitation techniques, since many container-breakout and privilege-escalation chains rely on a syscall the application itself never legitimately needs.
Enforcing and auditing at deploy time
None of the above helps if it is only a suggestion in a document. Enforce it at the deploy gate, a policy engine or admission controller that rejects a deployment specification missing a non-root user requirement, a dropped-capabilities list, or a read-only filesystem setting, rather than trusting every team to remember. Audit continuously by periodically diffing running container configurations against the required policy and alerting on drift, since a policy checked only at initial deploy time misses a manual run or an emergency change that bypassed the pipeline.
Worked example, measured
This exact hardening stack was built and run against a small HTTP service: a dedicated non-root user baked into the image, started with --read-only --cap-drop=ALL --security-opt no-new-privileges --tmpfs /tmp. Inside the running container, whoami returned the unprivileged application user, not root, and id showed no supplementary groups beyond that user's own group. The service still answered a real request over its published port with the expected response body, confirming the hardening did not break operability. Attempting to write a file outside the tmpfs mount (touch /app/newfile) failed with "Read-only file system," confirming the control was actually enforced rather than merely configured.
Trade-offs and pitfalls
A read-only root filesystem breaks any application that writes logs, temp files, or cache data to an unexpected path by default; audit what the application actually touches, or test it once with --read-only in staging, before enforcing it in production, and provide tmpfs or volume mounts for the legitimate write paths rather than disabling the control entirely. Dropping every capability can break something subtle that quietly relied on one, binding a low port, or a health check needing a specific network privilege, so validate functionally, not just that the image built successfully. A custom seccomp profile generated from one narrow test run can be too restrictive if it never exercised every code path, an error-handling branch that needs a syscall the happy path never calls, so generate it from realistic, broad testing, or accept a slightly looser default profile over an untested custom one that breaks in production on the first edge case.
Why should a container process run as non-root in production, especially one that processes sensitive data or is exposed to untrusted input? Show the Dockerfile techniques to create a non-root user and switch to it, and explain how to still allow something like binding to a privileged port below 1024 without running as root.
Sample Answer
Direct answer
A container process should run as non-root because the container's process namespace boundary is not a security boundary by itself: if an attacker escapes the container (through a kernel vulnerability, a misconfigured mount, or a vulnerable dependency processing untrusted input), running as root inside the container often means they land as root on the host's view of that process too, and root inside a container frequently has more host-visible capability than people expect. For a service that parses untrusted input or handles sensitive data, that escape path is exactly the one you are trying to make as narrow as possible.
Dockerfile technique
Create a dedicated, low-privilege system user and group, then switch to it before the final CMD/ENTRYPOINT:
FROM python:3.12-slim
RUN addgroup --system app && adduser --system --ingroup app app
WORKDIR /app
COPY --chown=app:app . .
USER app
CMD ["python", "app.py"]
Everything before USER app (installing packages, writing files as part of the build) still runs as root, which is fine since that happens at build time, not against untrusted input; USER app switches the actual runtime process. --chown=app:app on the COPY avoids a separate chown layer and ensures the application files are actually readable by the user that will run them.
Binding a privileged port without running as root
Ports below 1024 are privileged on Linux by default, and a non-root process cannot bind one directly. The clean fix that does not require root at all is a Linux capability: grant the specific CAP_NET_BIND_SERVICE capability to the binary itself with setcap, rather than granting the whole process root:
RUN setcap 'cap_net_bind_service=+ep' /app
USER app
Verified directly: a small Go HTTP server built with this exact technique, running as the non-root app user, still bound to port 80 successfully and served a request:
$ docker exec nonroot-demo whoami
app
$ curl http://127.0.0.1:18080/
ok
(port 18080 on the host was mapped to port 80 in the container). The far more common real-world alternative, when you do not strictly need the container's internal process to bind a low port, is simply to have the application listen on an unprivileged port (like 8080) and let Docker's own port publishing do the remapping (-p 80:8080), since the privileged-port restriction only applies inside the container's own network namespace, not to the host-side port a container is published on.
Trade-offs and gotchas
setcap requires the capability tool (libcap2-bin on Debian-based images) to be present at build time, and the capability has to be reapplied if the binary is rebuilt or replaced later in the Dockerfile, since it is a filesystem attribute on that specific binary, not a property of the user. If you only need the container to appear on port 80 externally and are not otherwise constrained (no host networking, no orchestrator requirement to bind low ports internally), prefer the simpler port-remapping approach; reach for setcap when the application specifically must bind a privileged port inside its own namespace.
A container orchestrator (the system that decides where and how containers run in production, such as Kubernetes) usually layers its own version of this same control on top of the Dockerfile: a Kubernetes pod's securityContext.runAsNonRoot: true tells the orchestrator to refuse to start the container at all if the image's own configuration would otherwise run it as root, enforcing the requirement from outside the image in addition to, not instead of, the USER instruction inside it.
As a new microservice is onboarded to production, what concrete Dockerfile and runtime hardening changes would you require before sign-off? Cover user permissions, filesystem settings, dropped capabilities, network exposure, and secrets handling, with example commands or Dockerfile snippets where relevant.
Sample Answer
Direct answer
Before sign-off, a new microservice's Dockerfile and its deploy manifest both need to demonstrate a small, concrete checklist, not a claim of following best practices in general: a non-root user, a locked-down filesystem, no unnecessary Linux capabilities, no unnecessary network exposure, and no secret material baked into the image or checked into its configuration. Each item below has a specific artifact a reviewer can point to and verify directly, not just a policy to trust.
The checklist
1. User permissions. The Dockerfile creates a dedicated non-root user and sets USER before the final CMD or ENTRYPOINT:
RUN groupadd -r app && useradd -r -g app -d /app -s /usr/sbin/nologin app
WORKDIR /app
COPY --chown=app:app . .
USER app
Reviewer check: docker run <image> whoami must not print root.
2. Filesystem settings. The deploy manifest, or the equivalent runtime flags, sets the root filesystem read-only, with an explicit, named exception for any path the application genuinely writes to:
docker run --read-only --tmpfs /tmp <image>
Reviewer check: attempt a write outside the declared exception path and confirm it fails, for example touch /app/x returning "Read-only file system."
3. Dropped capabilities. Start from zero and add back only what is justified:
docker run --cap-drop=ALL --security-opt no-new-privileges <image>
no-new-privileges additionally blocks the process, or anything it executes, from gaining more privileges than it started with through a setuid binary (an executable that runs with its owner's privileges, often root, rather than the privileges of whoever launched it, one of the most common ways a process escalates to root on Linux), closing a common local privilege-escalation path even for an already non-root process. Reviewer check: docker inspect shows an empty or minimal capability-add list; if the service genuinely needs one, for example CAP_NET_BIND_SERVICE to bind a privileged port, that specific exception is documented with a reason, not silently granted alongside everything else.
4. Network exposure. Only the ports the service actually serves are exposed and published, and the service sits on a network scoped to what it actually needs to reach, its own database, its own cache, rather than a flat network shared with every other service in the organization. Reviewer check: the deploy manifest's port list and network membership match the service's actual documented dependencies, not a copied default.
5. Secrets handling. No secret, an API key, a database credential, a registry token, appears as a Dockerfile ARG or ENV default, in a committed configuration file, or hardcoded in source. Build-time secrets use BuildKit (Docker's own build engine, the default for docker build since Docker Engine 23) and its --mount=type=secret flag, which leaves no trace in image history or the final filesystem. Runtime secrets are injected through the platform's own secret store or a mounted file, never a plaintext environment variable baked into an image layer. Reviewer check: docker history --no-trunc <image> and a full filesystem export contain no credential-shaped strings.
Worked example, measured
A minimal image implementing items one through three together was built and run: a dedicated non-root USER, --read-only with a --tmpfs /tmp exception, --cap-drop=ALL, and --security-opt no-new-privileges. The container's own whoami returned the dedicated application user (a non-root system user ID assigned by useradd -r on this base image), the service still answered a real HTTP request correctly through its published port, and an attempted write outside the tmpfs exception failed with "Read-only file system," confirming these controls compose together without breaking the service.
Trade-offs and pitfalls
This is an onboarding gate, not a one-time audit, so it should be automatable and re-checked, a linter against the Dockerfile plus a runtime assertion in a smoke test, because a checklist a human re-reads by eye for every new service degrades the moment there are ten services a week instead of one. The most common real failure is not skipping an item outright; it is an exception, a capability, a writable path, an open port, that was justified for the first service and then copy-pasted into every service afterward without being re-justified against that service's actual needs.
That is every published Containerization and Docker Fundamentals question for Cybersecurity Engineer so far. Browse the other topics in this category, or practice this one interactively.