Andrew Mercer
on this page

Bare-Bones Linux Containers: How They Actually Work

Docker, containerd, and Kubernetes' CRI are orchestration layered on top of about four Linux kernel primitives. Nothing more is required for a "container" in the technical sense — no magic, no special binary format, no hypervisor. None of it is tied to any language: it's syscalls and files. This guide shows exactly which ones, and proves it by building a working container using nothing but bash and util-linux — then maps each step to the raw syscall so you can implement it in C, Rust, Go, Python, whatever you want.

The Four Pillars

Pillar What it does Kernel mechanism
Namespaces Isolate what a process can see clone(2) / unshare(2) / setns(2) with CLONE_NEW* flags
Filesystem root change Isolate what a process can access on disk chroot(2) / pivot_root(2)
cgroups Limit what a process can consume cgroupfs v2 — plain file reads/writes, no syscall at all
Capabilities Limit what a process can do as root capset(2) / prctl(2)

A "container" is a regular Linux process that has had all four of these applied to it before its real program is exec'd. Every container runtime — runc, crun, LXC, gVisor's runsc — is built on exactly this list. The only thing that differs between languages is how you invoke these syscalls (libc wrapper, raw syscall(2), a language's FFI), not what the syscalls do.

Namespace types

Flag (C constant) Isolates
CLONE_NEWUTS hostname/domainname
CLONE_NEWPID process ID tree (child becomes PID 1 inside)
CLONE_NEWNS mount table
CLONE_NEWNET network stack (interfaces, routes, ports)
CLONE_NEWIPC System V IPC / POSIX message queues
CLONE_NEWUSER UID/GID mapping (needed for rootless containers)
CLONE_NEWCGROUP cgroup root view

These are the same integer flags whether you reach them via clone(), unshare(), or a language binding — they're passed straight to the kernel.


Step 1 — Get a root filesystem

Language-independent; it's just a directory tree:

sudo debootstrap --variant=minbase jammy /containers/rootfs

or hand-roll one from a static busybox:

mkdir -p rootfs/{bin,proc,sys,dev,etc}
cp $(which busybox) rootfs/bin/
for cmd in sh ls cat mount; do ln -s busybox rootfs/bin/$cmd; done

Step 2 — Do it with zero code at all (unshare + chroot)

Before writing a single line in any language, prove the mechanism with tools already on your system. unshare(1) is a thin CLI wrapper around unshare(2):

sudo unshare --uts --pid --mount --net --ipc --fork \
    chroot /containers/rootfs /bin/sh

Inside that shell: hostname is independent of the host, echo $$ shows PID 1, ip addr shows only a down lo, and mount shows a fresh (if empty) mount table. That's a container. Everything from here is just re-implementing unshare --fork + chroot yourself, in a language of your choice, so you can control it programmatically.


Step 3 — The syscalls, in order

Whatever language you use, this is the exact sequence:

  1. clone(2) with flags CLONE_NEWUTS | CLONE_NEWPID | CLONE_NEWNS | CLONE_NEWNET | CLONE_NEWIPC — creates a child process that starts life inside new namespaces. (Alternative: fork(2) then unshare(2) with the same flags from inside the child — simpler API, subtly different semantics for the PID namespace, see note below.)
  2. sethostname(2) — only affects the new UTS namespace.
  3. mount(2) to bind-mount the new root onto itself: mount(new_root, new_root, NULL, MS_BIND|MS_REC, NULL) — required because pivot_root demands its target be a mount point.
  4. pivot_root(2) — swaps / to the new root, moving the old root to a subdirectory you specify.
  5. chdir(2) to /.
  6. umount2(2) with MNT_DETACH on the old-root path, then delete that directory — makes the old root fully unreachable (this is what makes pivot_root safe where chroot alone is not).
  7. mount(2) to mount a fresh proc filesystem at /proc — only shows meaningful data because you're inside CLONE_NEWPID + CLONE_NEWNS.
  8. capset(2) (or prctl(2, PR_CAPBSET_DROP, ...) per-capability) — drop everything you don't explicitly need.
  9. execve(2) — replace the process image with the real program.

Separately, from the parent, after clone() returns a PID:

  1. Write that PID into a cgroup's cgroup.procs file, having already written limits into memory.max, pids.max, cpu.max in the same directory. No syscall — open()/write()/close() on ordinary files under /sys/fs/cgroup/.

That's the entire list. Any language that can call these nine syscalls plus do basic file I/O can build a container:

  • C: direct libc calls (clone, pivot_root isn't in glibc — use syscall(SYS_pivot_root, ...)).
  • Rust: the nix crate wraps all of these directly.
  • Go: syscall/golang.org/x/sys/unix, or set SysProcAttr.Cloneflags on exec.Cmd (this is literally how Docker's own Go code does it).
  • Python: ctypes.CDLL(None).unshare(flags) / os.chroot (no pivot_root binding, call via ctypes with syscall(SYS_pivot_root, ...)).
  • Zig, Nim, whatever: same story — FFI to the same four syscalls.

Step 4 — cgroups v2: plain files, any language

No API, no library — cgroups v2 is just a filesystem:

mkdir /sys/fs/cgroup/mycontainer
echo "+memory +cpu +pids" > /sys/fs/cgroup/cgroup.subtree_control
echo 268435456   > /sys/fs/cgroup/mycontainer/memory.max   # 256 MiB
echo 64          > /sys/fs/cgroup/mycontainer/pids.max
echo "50000 100000" > /sys/fs/cgroup/mycontainer/cpu.max   # 50% of one core
echo $CHILD_PID  > /sys/fs/cgroup/mycontainer/cgroup.procs

The kernel enforces limits the instant the PID lands in cgroup.procs. Any language's standard "write a string to a file" function is sufficient — there is no lower-level API to reach for.


Step 5 — Capabilities: drop before execve

Two equivalent approaches, available from any language via prctl(2) or capset(2):

# quickest to verify from the shell before you exec your real program:
capsh --drop=cap_sys_admin,cap_net_admin,cap_sys_module -- -c /bin/sh

Programmatically: iterate the full capability set, and for anything not in your explicit keep-list, call prctl(PR_CAPBSET_DROP, cap) and clear it from the effective/permitted sets via capset(2). Do this as the very last step before execve, once no further privileged operation (mounting, chroot, hostname) remains.


Full sequence, restated

  1. Parent: clone()/fork()+unshare() with namespace flags → child PID
  2. Parent: write child PID + limits into cgroupfs
  3. Child: sethostname()
  4. Child: bind-mount new root onto itself, pivot_root(), chdir("/"), unmount+delete old root
  5. Child: mount fresh /proc, /sys, /dev as needed
  6. Child: drop capabilities
  7. Child: execve() into the target program
  8. Parent: waitpid()

This is, step for step, what runc does before handing off to libcontainer. The difference between runc and a weekend project is OCI spec parsing, veth/bridge network wiring, seccomp-bpf profiles, and rootless UID mapping — all bookkeeping layered on the same nine syscalls, not new mechanism.


A note on clone() vs fork()+unshare()

clone(2) with CLONE_NEWPID set creates the child already as PID 1 in the new namespace. If you instead fork() normally and then call unshare(CLONE_NEWPID) from inside the child, the calling process itself does not move into the new PID namespace — only its next forked child will. So unshare-based approaches typically need an extra fork after the unshare() call. This is why unshare(1)'s CLI needs --fork when you pass --pid, and it's the one place where the two syscalls aren't drop-in equivalents.


What's deliberately left out (and where to go next)

  • Networking — a veth pair, one end moved into the new CLONE_NEWNET namespace via ip link set veth1 netns <pid> (or setns(2) if scripting it), bridged on the host side.
  • Rootless containers — add CLONE_NEWUSER and write /proc/<pid>/uid_map and /proc/<pid>/gid_map before the child does anything else.
  • Seccomp — restrict callable syscalls via prctl(PR_SET_SECCOMP, ...) or libseccomp.
  • Image layering / overlayfs — how Docker builds a rootfs from layers; that's overlay mount options, unrelated to namespaces.

Everything above namespaces + pivot_root + cgroups + capabilities is bookkeeping, not new kernel mechanism — and none of it cares what language issues the calls.