Bare-Bones Linux Containers: How They Actually Work¶
Docker, containerd, and Kubernetes' CRI are orchestration layered on top of about four Linux kernel primitives. Nothing more is required for a "container" in the technical sense — no magic, no special binary format, no hypervisor. None of it is tied to any language: it's syscalls and files. This guide shows exactly which ones, and proves it by building a working container using nothing but bash and util-linux — then maps each step to the raw syscall so you can implement it in C, Rust, Go, Python, whatever you want.
The Four Pillars¶
| Pillar | What it does | Kernel mechanism |
|---|---|---|
| Namespaces | Isolate what a process can see | clone(2) / unshare(2) / setns(2) with CLONE_NEW* flags |
| Filesystem root change | Isolate what a process can access on disk | chroot(2) / pivot_root(2) |
| cgroups | Limit what a process can consume | cgroupfs v2 — plain file reads/writes, no syscall at all |
| Capabilities | Limit what a process can do as root | capset(2) / prctl(2) |
A "container" is a regular Linux process that has had all four of these applied to it before its real program is exec'd. Every container runtime — runc, crun, LXC, gVisor's runsc — is built on exactly this list. The only thing that differs between languages is how you invoke these syscalls (libc wrapper, raw syscall(2), a language's FFI), not what the syscalls do.
Namespace types¶
| Flag (C constant) | Isolates |
|---|---|
CLONE_NEWUTS |
hostname/domainname |
CLONE_NEWPID |
process ID tree (child becomes PID 1 inside) |
CLONE_NEWNS |
mount table |
CLONE_NEWNET |
network stack (interfaces, routes, ports) |
CLONE_NEWIPC |
System V IPC / POSIX message queues |
CLONE_NEWUSER |
UID/GID mapping (needed for rootless containers) |
CLONE_NEWCGROUP |
cgroup root view |
These are the same integer flags whether you reach them via clone(), unshare(), or a language binding — they're passed straight to the kernel.
Step 1 — Get a root filesystem¶
Language-independent; it's just a directory tree:
sudo debootstrap --variant=minbase jammy /containers/rootfs
or hand-roll one from a static busybox:
mkdir -p rootfs/{bin,proc,sys,dev,etc}
cp $(which busybox) rootfs/bin/
for cmd in sh ls cat mount; do ln -s busybox rootfs/bin/$cmd; done
Step 2 — Do it with zero code at all (unshare + chroot)¶
Before writing a single line in any language, prove the mechanism with tools already on your system. unshare(1) is a thin CLI wrapper around unshare(2):
sudo unshare --uts --pid --mount --net --ipc --fork \
chroot /containers/rootfs /bin/sh
Inside that shell: hostname is independent of the host, echo $$ shows PID 1, ip addr shows only a down lo, and mount shows a fresh (if empty) mount table. That's a container. Everything from here is just re-implementing unshare --fork + chroot yourself, in a language of your choice, so you can control it programmatically.
Step 3 — The syscalls, in order¶
Whatever language you use, this is the exact sequence:
clone(2)with flagsCLONE_NEWUTS | CLONE_NEWPID | CLONE_NEWNS | CLONE_NEWNET | CLONE_NEWIPC— creates a child process that starts life inside new namespaces. (Alternative:fork(2)thenunshare(2)with the same flags from inside the child — simpler API, subtly different semantics for the PID namespace, see note below.)sethostname(2)— only affects the new UTS namespace.mount(2)to bind-mount the new root onto itself:mount(new_root, new_root, NULL, MS_BIND|MS_REC, NULL)— required becausepivot_rootdemands its target be a mount point.pivot_root(2)— swaps/to the new root, moving the old root to a subdirectory you specify.chdir(2)to/.umount2(2)withMNT_DETACHon the old-root path, then delete that directory — makes the old root fully unreachable (this is what makespivot_rootsafe wherechrootalone is not).mount(2)to mount a freshprocfilesystem at/proc— only shows meaningful data because you're insideCLONE_NEWPID+CLONE_NEWNS.capset(2)(orprctl(2, PR_CAPBSET_DROP, ...)per-capability) — drop everything you don't explicitly need.execve(2)— replace the process image with the real program.
Separately, from the parent, after clone() returns a PID:
- Write that PID into a cgroup's
cgroup.procsfile, having already written limits intomemory.max,pids.max,cpu.maxin the same directory. No syscall —open()/write()/close()on ordinary files under/sys/fs/cgroup/.
That's the entire list. Any language that can call these nine syscalls plus do basic file I/O can build a container:
- C: direct libc calls (
clone,pivot_rootisn't in glibc — usesyscall(SYS_pivot_root, ...)). - Rust: the
nixcrate wraps all of these directly. - Go:
syscall/golang.org/x/sys/unix, or setSysProcAttr.Cloneflagsonexec.Cmd(this is literally how Docker's own Go code does it). - Python:
ctypes.CDLL(None).unshare(flags)/os.chroot(nopivot_rootbinding, call viactypeswithsyscall(SYS_pivot_root, ...)). - Zig, Nim, whatever: same story — FFI to the same four syscalls.
Step 4 — cgroups v2: plain files, any language¶
No API, no library — cgroups v2 is just a filesystem:
mkdir /sys/fs/cgroup/mycontainer
echo "+memory +cpu +pids" > /sys/fs/cgroup/cgroup.subtree_control
echo 268435456 > /sys/fs/cgroup/mycontainer/memory.max # 256 MiB
echo 64 > /sys/fs/cgroup/mycontainer/pids.max
echo "50000 100000" > /sys/fs/cgroup/mycontainer/cpu.max # 50% of one core
echo $CHILD_PID > /sys/fs/cgroup/mycontainer/cgroup.procs
The kernel enforces limits the instant the PID lands in cgroup.procs. Any language's standard "write a string to a file" function is sufficient — there is no lower-level API to reach for.
Step 5 — Capabilities: drop before execve¶
Two equivalent approaches, available from any language via prctl(2) or capset(2):
# quickest to verify from the shell before you exec your real program:
capsh --drop=cap_sys_admin,cap_net_admin,cap_sys_module -- -c /bin/sh
Programmatically: iterate the full capability set, and for anything not in your explicit keep-list, call prctl(PR_CAPBSET_DROP, cap) and clear it from the effective/permitted sets via capset(2). Do this as the very last step before execve, once no further privileged operation (mounting, chroot, hostname) remains.
Full sequence, restated¶
- Parent:
clone()/fork()+unshare()with namespace flags → child PID - Parent: write child PID + limits into cgroupfs
- Child:
sethostname() - Child: bind-mount new root onto itself,
pivot_root(),chdir("/"), unmount+delete old root - Child: mount fresh
/proc,/sys,/devas needed - Child: drop capabilities
- Child:
execve()into the target program - Parent:
waitpid()
This is, step for step, what runc does before handing off to libcontainer. The difference between runc and a weekend project is OCI spec parsing, veth/bridge network wiring, seccomp-bpf profiles, and rootless UID mapping — all bookkeeping layered on the same nine syscalls, not new mechanism.
A note on clone() vs fork()+unshare()¶
clone(2) with CLONE_NEWPID set creates the child already as PID 1 in the new namespace. If you instead fork() normally and then call unshare(CLONE_NEWPID) from inside the child, the calling process itself does not move into the new PID namespace — only its next forked child will. So unshare-based approaches typically need an extra fork after the unshare() call. This is why unshare(1)'s CLI needs --fork when you pass --pid, and it's the one place where the two syscalls aren't drop-in equivalents.
What's deliberately left out (and where to go next)¶
- Networking — a veth pair, one end moved into the new
CLONE_NEWNETnamespace viaip link set veth1 netns <pid>(orsetns(2)if scripting it), bridged on the host side. - Rootless containers — add
CLONE_NEWUSERand write/proc/<pid>/uid_mapand/proc/<pid>/gid_mapbefore the child does anything else. - Seccomp — restrict callable syscalls via
prctl(PR_SET_SECCOMP, ...)orlibseccomp. - Image layering / overlayfs — how Docker builds a rootfs from layers; that's
overlaymount options, unrelated to namespaces.
Everything above namespaces + pivot_root + cgroups + capabilities is bookkeeping, not new kernel mechanism — and none of it cares what language issues the calls.