Andrew Mercer
on this page

systemd: A Comprehensive Overview

1. What systemd Actually Is

systemd is a system and service manager for Linux that replaced SysV init (and largely displaced Upstart) as PID 1 on nearly every major distribution — RHEL/CentOS/Fedora since RHEL 7, Debian/Ubuntu since ~15.04/16.04, SUSE, Arch, and others. It's not just an init system; it's a suite of ~70 integrated components covering service supervision, logging, device management, login sessions, network configuration, time sync, and more, all built around a small set of shared design principles: aggressive parallelization, dependency-based ordering (rather than fixed runlevel sequencing), socket/D-Bus activation, cgroups-based process tracking, and declarative unit files instead of imperative shell scripts.

Core binaries/daemons worth knowing: - systemd — PID 1 - systemd-journald — structured/binary logging - systemd-logind — session/seat management - systemd-networkd — optional network configuration - systemd-resolved — DNS resolution/caching - systemd-timesyncd — SNTP client - systemd-udevd — device management (absorbed the old standalone udev) - systemd-tmpfiles — temp file/directory lifecycle management - systemd-cgtop, systemd-analyze, systemctl, journalctl, loginctl, timedatectl, hostnamectl, busctl — tooling

2. Unit Types

Everything systemd manages is a unit, defined declaratively in an INI-style file. The major types:

Unit type Purpose
.service A managed process/daemon
.socket A socket (inet, unix) for socket activation
.target Grouping/synchronization point (analogous to old runlevels)
.mount / .automount Filesystem mounts, lazy-mount on access
.timer Cron-like scheduled activation
.path Trigger on filesystem path changes (inotify-based)
.device Kernel device units, auto-generated from udev
.swap Swap devices
.slice cgroup hierarchy grouping for resource control
.scope Externally-created process groups (e.g. from systemd-run, containers, user sessions)

Unit files live in a three-tier precedence hierarchy: - /etc/systemd/system/ — local admin overrides (highest priority) - /run/systemd/system/ — runtime, volatile - /usr/lib/systemd/system/ — vendor/package-supplied (lowest priority)

systemctl edit foo.service creates a drop-in at /etc/systemd/system/foo.service.d/override.conf rather than editing the vendor file directly — this is the correct way to customize a packaged unit and survives package upgrades. systemctl edit --full copies the whole file for a full override instead.

3. Anatomy of a Service Unit

[Unit]
Description=Example service
After=network-online.target
Wants=network-online.target
StartLimitIntervalSec=60
StartLimitBurst=3

[Service]
Type=notify
ExecStartPre=/usr/local/bin/preflight.sh
ExecStart=/usr/local/bin/myapp --config /etc/myapp/config.yml
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5
User=myapp
Group=myapp
DynamicUser=no
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
MemoryMax=512M
CPUQuota=50%
TimeoutStartSec=30
WatchdogSec=10

[Install]
WantedBy=multi-user.target

Key [Service] semantics

Type= controls when systemd considers the unit "started" — this is one of the most misunderstood settings: - simple (default) — considered started as soon as the process forks (ExecStart returns immediately, process runs in foreground) - exec — like simple, but waits for the actual execve() to succeed before considering it started (better failure detection than simple) - forking — expects the process to fork and the parent to exit; systemd tracks the child via PIDFile= or process tracking - oneshot — for scripts that run and exit; commonly paired with RemainAfterExit=yes - notify / notify-reload — process calls sd_notify() to signal readiness (READY=1); the gold standard for correctness since systemd knows exactly when startup logic has actually finished - dbus — considered started when it acquires a specified D-Bus name - idle — like simple but delays execution until other jobs finish (mostly cosmetic, for TTY output ordering)

Restart= — no, on-success, on-failure, on-abnormal, on-watchdog, on-abort, always. Combine with RestartSec= and StartLimitIntervalSec/StartLimitBurst (in [Unit]) to avoid restart storms.

Ordering vs. dependency are separate axes — this trips people up constantly: - Wants= / Requires= / Requisite= / BindsTo= control dependency (does starting this pull in that, and what happens on failure) - Before= / After= control ordering only - Declaring Requires= without After= means both units start in parallel with no ordering guarantee — you almost always want both together (Wants=+After= is the common safe default; Requires= is stricter and will take your unit down if the dependency fails/stops, which is often not what you want for soft dependencies) - network-online.target is worth calling out specifically: network.target only means networking is configured, not that connectivity exists — daemons that need actual network reachability should order After=network-online.target and pull in Wants=network-online.target, and this requires the network management stack (NetworkManager-wait-online, systemd-networkd-wait-online) to actually be enabled or the target returns instantly without waiting.

4. Targets Replace Runlevels

Targets are synchronization points, not sequential states. Rough SysV mapping:

SysV runlevel systemd target
0 poweroff.target
1 rescue.target
3 multi-user.target
5 graphical.target
6 reboot.target

systemctl get-default / set-default replace /etc/inittab editing. systemctl isolate <target> switches active target set immediately (like changing runlevel live). Targets can Wants=/Requires= other units, forming a dependency graph you can inspect with systemctl list-dependencies <target> or visualize with systemd-analyze plot > boot.svg.

5. Socket & Bus Activation

A .socket unit can create a listening socket before the corresponding service exists, and systemd hands off the fd on first connection (Type=notify services get this via LISTEN_FDS/sd_listen_fds()). Benefits: - Faster parallel boot — dependents can connect immediately even if the service behind the socket hasn't started yet - On-demand starting — the service doesn't need to run until something actually connects - Socket survives service restarts (no connection drops during a crash/restart cycle) since systemd, not the service, owns the socket

D-Bus activation works similarly for bus-name-based services.

6. cgroups and Resource Control

systemd uses cgroups (v2 unified hierarchy on modern distros) as its primary mechanism for process tracking — not PID files, not process groups. Every unit gets its own cgroup, which is why systemd can reliably kill an entire service tree (KillMode=) even if the process double-forks and daemonizes to escape traditional tracking.

Resource control directives map directly onto cgroup controllers: - CPUQuota=, CPUWeight= → cpu controller - MemoryMax=, MemoryHigh=, MemorySwapMax= → memory controller - IOWeight=, IOReadBandwidthMax= → io controller - TasksMax= → pids controller

.slice units let you group services into a resource-controlled hierarchy (e.g. all containers under one CPU budget). systemd-cgtop gives a live view analogous to top but scoped by cgroup/unit. systemctl status <unit> shows the cgroup tree of live processes under CGroup:, which is often the fastest way to find a rogue child process a "simple" daemon spawned.

7. Sandboxing / Hardening Directives

This is where systemd earns its keep for security-conscious deployments — no need for a separate MAC policy or container just to sandbox a single daemon:

  • ProtectSystem=strict|full|true — mounts /usr, /boot, etc. read-only
  • ProtectHome=yes|read-only|tmpfs
  • PrivateTmp=yes — service gets its own /tmp and /var/tmp namespace
  • PrivateNetwork=yes — isolated network namespace
  • PrivateDevices=yes — minimal /dev
  • NoNewPrivileges=yes — blocks setuid/setgid privilege escalation via execve
  • ProtectKernelTunables=, ProtectKernelModules=, ProtectControlGroups=
  • RestrictAddressFamilies=, RestrictNamespaces=, SystemCallFilter= (seccomp allow/deny lists)
  • CapabilityBoundingSet= / AmbientCapabilities= — fine-grained capability control instead of full root
  • DynamicUser=yes — allocates a transient UID/GID for the unit's lifetime, no persistent service account needed

systemd-analyze security <unit> scores a unit's hardening and tells you exactly which directives you're missing — genuinely useful for an audit pass across a fleet of unit files.

8. Logging: journald

journald captures stdout/stderr from every unit plus kernel ring buffer, audit records, and structured fields (_SYSTEMD_UNIT, _PID, _UID, custom fields via sd_journal_send()) into a binary, indexed, optionally-signed (FSS/forward secure sealing) log store.

Common journalctl patterns:

journalctl -u myapp.service -f              # follow a unit
journalctl -u myapp.service --since "1h ago"
journalctl -k                               # kernel messages only
journalctl -p err                           # priority filter
journalctl --disk-usage
journalctl -o json-pretty                   # structured output for parsing
journalctl -b -1                            # previous boot
journalctl _PID=1234                        # arbitrary field match

Persistence is controlled by Storage= in /etc/systemd/journald.conf (volatile = /run only, lost on reboot; persistent = /var/log/journal, created automatically if the directory exists). Given your ELK background: journald can forward to syslog via ForwardToSyslog=, or you can skip that hop entirely and have Logstash/Vector/Filebeat read directly from the journal (journalbeat/Vector's journald source, or journalctl -o json --follow piped in) — usually cleaner than double-handling through rsyslog.

9. Boot Analysis & Performance

systemd-analyze                    # total boot time, kernel vs userspace split
systemd-analyze blame              # per-unit startup time, slowest first
systemd-analyze critical-chain     # the actual critical path — often more useful than blame
systemd-analyze plot > boot.svg    # visual timeline
systemd-analyze verify foo.service # lint a unit file before deploying it
systemd-analyze dot 'myapp.*' | dot -Tsvg > deps.svg  # dependency graph via graphviz

critical-chain matters more than blame in practice — a unit can be "slow" per blame but off the critical path entirely (running in parallel with something slower), so optimizing it buys you nothing.

10. Timers vs. Cron

.timer units are the systemd-native replacement for cron, and they're a meaningfully better fit for anything that needs observability or resilience:

# backup.timer
[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=300

[Install]
WantedBy=timers.target

Advantages over cron: Persistent=true catches up missed runs if the machine was off at trigger time; jobs get full journald logging and cgroup accounting like any other service; you can add OnFailure= to a paired .service for alerting; systemctl list-timers shows next/last run for everything at a glance; calendar syntax (OnCalendar=) is more expressive than cron syntax for things like "last day of month" (*-*~1) or "every 15 min" (*:0/15). Monotonic timers (OnBootSec=, OnUnitActiveSec=) cover the "N seconds after boot/after last run" cases cron can't express cleanly at all.

11. Other Notable Components

  • systemd-networkd — declarative .network/.netdev/.link files as an alternative to NetworkManager or ifupdown; common on servers/containers where a GUI-oriented network manager is overkill
  • systemd-nspawn — lightweight container/chroot tool, good for quick namespace isolation without a full container runtime
  • systemd-run — ad-hoc transient unit creation from the CLI (systemd-run --scope --uid=nobody stress-ng --cpu 4) — handy for one-off resource-constrained or isolated command execution without writing a unit file
  • systemd-tmpfiles — declarative rules (/etc/tmpfiles.d/*.conf) for creating/cleaning /tmp, /run, PID files, etc. on boot or on a timer
  • systemd-sysusers — declarative system user/group creation, commonly used by RPM/deb packaging instead of useradd in post-install scriptlets
  • systemd-resolved — caching stub DNS resolver; worth knowing because it can conflict with containers/VPNs that expect to manage /etc/resolv.conf directly (the classic symptom is resolv.conf pointing at 127.0.0.53 and confusing tools that expect real nameservers)
  • systemd-boot — a minimal UEFI boot manager (distinct from GRUB), used by default on some distros/appliance images

12. Practical Debugging Toolkit

systemctl status <unit>              # state, recent logs, cgroup tree
systemctl list-units --failed        # what's currently broken
systemctl list-unit-files --state=enabled
systemctl show <unit> -p <Property>  # dump any single resolved property
systemctl cat <unit>                 # show the fully-merged unit file incl. drop-ins
systemd-delta                        # show which vendor files are overridden and how
journalctl -xe                       # verbose recent errors with explanatory text
systemctl mask <unit>                # hard-disable (symlinks to /dev/null; stops accidental re-enable)

systemctl show is underused — when a unit isn't behaving as the unit file "reads," it's often because a drop-in, a default, or a vendor preset is overriding what you'd expect, and show -p <Property> gives you the actual resolved value systemd is using rather than what's in the file you happen to be looking at.

13. Common Criticisms (for balance)

Worth knowing the counter-arguments since systemd remains genuinely contentious in parts of the Linux community: - Scope creep — critics argue it's absorbed too much (networking, DNS, login management, boot) for what was originally "just an init system," violating Unix's do-one-thing philosophy - Binary logs — journald's binary format is less trivially greppable/shippable than plain text, though journalctl mitigates this and structured querying is arguably an upgrade once you're used to it - Tight coupling — components are harder to run independently of each other than the SysV-era daemons were, making alternative init systems (runit, OpenRC, s6) harder to substitute in piecemeal - PID 1 complexity/attack surface — more code running as PID 1 is more attack surface and more that can go wrong catastrophically if it crashes

Alternatives that reject systemd entirely — Devuan (Debian without systemd), Void Linux (runit), Artix (OpenRC/runit/s6/dinit) — exist specifically along these fault lines, though systemd remains the default on effectively every major enterprise distro (RHEL, SUSE, Ubuntu LTS).


Happy to go deeper on any one section — sandboxing directives for a specific hardening pass, networkd config for a homelab box, or journald-to-OpenSearch pipeline design given your Vector/RabbitMQ setup, for instance.