Linux Memory Management: The Complete Guide¶
Memory is usually the resource that decides whether a server feels fast or slow. Linux manages it aggressively: it hands out more address space than physically exists, uses nearly all free RAM as disk cache, and moves cold pages out to swap when it needs room. Much of what looks alarming in free or top is normal behavior, and much of what is actually a problem (swap thrashing, fragmentation, runaway processes) can be spotted quickly once you know what to look at.
This guide covers the theory, the tools, the tuning knobs, and the troubleshooting workflow. Sections marked Field notes hold practical notes and lab exercises kept separate from the general material.
Conventions
- Commands need root unless noted.
- Kernel parameters are shown as
sysctlnames (vm.swappiness), which map to files under/proc/sys/vm/.- Change one parameter at a time and measure. No setting is right for every workload.
Table of Contents¶
- Core concepts
- Reading memory usage
- The page cache (and why "Linux ate my RAM" is a myth)
- Swap
- Overcommit and virtual vs resident memory
- The OOM killer
- Huge pages and Transparent Huge Pages
- Memory fragmentation
- NUMA
- Tuning with sysctl
- Limiting memory with cgroups and systemd
- Inspecting hardware memory
- Testing memory (memtest86+)
- Clearing caches
- Finding and debugging memory leaks
- Field notes: leak lab exercise
- Troubleshooting workflow
- Cheat sheet
- Further reading
1. Core concepts¶
1.1 Virtual memory¶
Linux uses virtual memory: every process sees its own private, contiguous address space, and the kernel plus the CPU's memory management unit (MMU) translate those virtual addresses to physical RAM. This gives:
- Isolation: one process cannot read or write another's memory. The kernel builds a fresh address space when a process is created (
fork()), and copy-on-write keeps forks cheap because parent and child share pages until one writes. - Flexibility: memory can be allocated lazily, shared between processes (libraries, shared memory), memory-mapped from files, or pushed out to swap.
- Overcommit: the total virtual memory handed out may exceed physical RAM.
From an administrator's viewpoint, "virtual memory" also means the total usable pool: RAM plus swap. The knobs that control how it behaves live under /proc/sys/vm/.
1.2 Pages¶
Memory is managed in fixed-size pages. Everything the kernel does with memory (allocating, mapping, reading from disk, swapping) happens in page-sized chunks.
getconf PAGESIZE # base page size, in bytes
- x86-64: 4 KiB base pages (fixed), with 2 MiB and 1 GiB "huge" pages available on top.
- ARM64: 4, 16, or 64 KiB base pages, chosen when the kernel is built (Red Hat and Apple-silicon-era kernels differ from each other).
- 32-bit x86: 4 KiB.
Small pages are efficient for many small files and allocations. For very large working sets (databases, virtual machines), tracking millions of 4 KiB pages is costly, which is what huge pages solve (section 7).
1.3 Memory kinds¶
| Kind | What it is | Reclaimable? |
|---|---|---|
| Anonymous memory | Heap, stack, malloced data. Has no file behind it. |
Only by writing to swap |
| File-backed memory | Page cache: file contents, executables, mmapped files | Yes: clean pages are simply dropped; dirty pages are written back first |
| Shared memory / tmpfs | /dev/shm, tmpfs mounts, SysV/POSIX shared memory |
Only via swap (it is anonymous in nature) |
| Kernel memory | Slab caches (dentries, inodes, network buffers), page tables, kernel stacks | Some of it (SReclaimable), not all |
| Huge pages (hugetlb) | Pages explicitly reserved at boot/runtime | No; not swappable, pre-reserved |
1.4 RAM vs disk speed¶
Memory is roughly 1,000 times faster than a spinning disk and still orders of magnitude faster than SSDs when you count latency. A CPU only runs smoothly when its data is in RAM (or even better, in the CPU's own caches). When the working set does not fit, the system falls back to slower I/O and everything queues behind it. That is why memory often matters more than CPU on a server, and why swap activity is the first thing to check on a slow machine.
1.5 Zones, watermarks, and reclaim¶
The kernel groups physical memory into zones (for example DMA, DMA32, Normal) and per NUMA node. It keeps free memory above watermarks (min/low/high). When free memory dips below the low watermark, the kswapd thread wakes up and reclaims pages in the background: it drops clean cache, writes back dirty pages, and swaps out cold anonymous pages. If allocations outrun kswapd and free memory reaches the min watermark, processes must reclaim memory themselves ("direct reclaim"), which causes latency spikes. If reclaim fails entirely, the OOM killer steps in (section 6).
2. Reading memory usage¶
2.1 free¶
free -h # human-readable
free -m # MiB
free -h -s 2 # refresh every 2 seconds
total used free shared buff/cache available
Mem: 15Gi 3.1Gi 1.2Gi 210Mi 11Gi 12Gi
Swap: 4.0Gi 0B 4.0Gi
| Column | Meaning |
|---|---|
total |
Installed usable RAM |
used |
Memory in use by programs and the kernel (excluding reclaimable cache) |
free |
Completely unused memory. Low is normal. |
shared |
tmpfs and shared memory |
buff/cache |
Page cache, buffers, and reclaimable slab |
available |
Estimate of memory available to start new work without swapping. This is the number to watch. |
2.2 /proc/meminfo¶
cat /proc/meminfo
grep -E 'MemTotal|MemAvailable|Cached|Dirty|Swap|Commit|Huge|Slab|Shmem' /proc/meminfo
Key fields:
| Field | Meaning |
|---|---|
MemAvailable |
Best estimate of allocatable memory without swapping |
Buffers / Cached |
Block-device buffers / page cache |
Active(anon), Inactive(anon), Active(file), Inactive(file) |
LRU lists; inactive pages are reclaimed first |
Dirty, Writeback |
Data waiting to be / being written to disk |
AnonPages |
Anonymous (non-file) memory |
Mapped |
File pages mapped into process address spaces |
Shmem |
Shared memory and tmpfs |
Slab, SReclaimable, SUnreclaim |
Kernel object caches |
PageTables |
Memory used for address-translation tables |
SwapTotal, SwapFree, SwapCached |
Swap capacity, free, and pages present in both RAM and swap |
CommitLimit, Committed_AS |
Overcommit accounting (section 5) |
HugePages_Total/Free/Rsvd, Hugepagesize, AnonHugePages |
Huge page pools and THP usage |
2.3 Per-system activity: vmstat, sar, top¶
vmstat 1 # one line per second
vmstat -SM 1 # in MiB
sar -r 1 5 # memory utilization (sysstat)
sar -B 1 5 # paging statistics
sar -S 1 5 # swap utilization
Important vmstat columns:
| Column | Meaning | Concern when... |
|---|---|---|
si / so |
Swap-in / swap-out (KiB/s) | Sustained non-zero values |
bi / bo |
Blocks read/written from disk | Very high with low CPU |
free, buff, cache |
Memory pools | (see available instead) |
wa |
CPU time waiting for I/O | High, together with swapping |
r, b |
Runnable / blocked processes | b consistently > 0 |
In top/htop, the header shows memory and swap. Per-process columns:
| Column | Meaning |
|---|---|
VIRT |
Total virtual address space the process has mapped (includes untouched and file-mapped memory) |
RES |
Resident memory: physical RAM the process is using now |
SHR |
Portion of RES that is shared (libraries, shared memory) |
%MEM |
RES as a percentage of RAM |
SWAP/VmSwap |
Memory of the process currently in swap |
In top, press M to sort by memory. In htop, F6 selects the sort column.
2.4 Per-process detail¶
ps aux --sort=-rss | head # biggest resident users
ps -eo pid,comm,rss,vsz,%mem --sort=-rss | head
pmap -x <pid> # mapping-by-mapping breakdown
cat /proc/<pid>/status | grep -E 'Vm|Rss'
cat /proc/<pid>/smaps_rollup # PSS (proportional set size) totals
smem -tk -s pss # accurate accounting for shared memory (package: smem)
RSS counts shared pages fully in every process that maps them, so summing RSS overstates usage. PSS divides shared pages proportionally and adds up correctly.
2.5 Kernel memory and pressure¶
sudo slabtop -o # kernel slab caches, sorted by size
cat /proc/pressure/memory # pressure stall information (PSI), kernel 4.20+
PSI output:
some avg10=0.00 avg60=0.12 avg300=0.05 total=123456
full avg10=0.00 avg60=0.00 avg300=0.00 total=6789
some is the share of time at least one task was stalled waiting on memory; full is the share of time all non-idle tasks were stalled. Sustained non-zero full is real trouble. PSI is a more honest signal than "free memory" for deciding whether a system is short of RAM.
3. The page cache¶
Whenever a file is read, Linux keeps its contents in RAM (the read cache) so the next read is served from memory instead of the disk. Writes go into RAM first and are flushed later (the write cache, dirty pages). Free RAM is wasted RAM, so the kernel uses almost all otherwise unused memory as cache.
Consequences:
- A busy file server or database host will show very little
freememory and a largebuff/cache. That is healthy. Clean cache is released instantly whenever a program needs the memory. - Judge memory health by
available, swap activity, and PSI, not byfree. - Cache is not swap. Cache is RAM used to accelerate disk access; swap is disk space used to extend RAM.
- For a read-heavy server, a healthy amount of cache directly translates to speed. If the working set of hot files no longer fits and the cache keeps shrinking and refilling (rising
biinvmstat, falling cache hit ratios), the machine needs more RAM. An old rule of thumb suggests trouble when the cache falls under about 40% of RAM on a read-mostly server, but treat that as a rough hint rather than a threshold: measure disk reads and latency instead.
Dirty pages and writeback¶
Writes accumulate as dirty pages and are flushed in the background by kernel flusher threads.
grep -E 'Dirty|Writeback' /proc/meminfo
sysctl vm.dirty_ratio vm.dirty_background_ratio vm.dirty_expire_centisecs vm.dirty_writeback_centisecs
See section 10 for tuning. In short: a low dirty_background_ratio starts flushing early (smoother I/O, less cache benefit); a high dirty_ratio allows large bursts to absorb but risks long stalls when the limit is hit and writers get throttled.
4. Swap¶
Swap is disk space used as overflow for anonymous memory (and to hold shared memory pages). Because disks are far slower than RAM, heavy swapping ("thrashing") can cripple a server. But "swap usage" is not automatically bad:
- A small amount of swap used by long-idle processes is healthy. The kernel moves cold pages out so RAM can hold hot data and cache.
- The signal for trouble is ongoing swap-in/swap-out activity (
si/soinvmstat), especially with high I/O wait, not the raw amount used. - Some applications (traditionally Oracle databases) are designed around swap usage. For everything else, be suspicious if more than a modest amount of swap is actively in use.
- If a server is slow, look at swapping first.
4.1 Inspect swap¶
swapon --show # devices/files, size, used, priority
free -h
cat /proc/swaps
vmstat 1 # watch si/so
grep -E 'VmSwap|Name' /proc/*/status 2>/dev/null | paste - - | awk '$4>0 {print $4, $2}' | sort -nr | head # biggest swap users
4.2 Create a swap partition or LV¶
mkswap /dev/vg_root/lv_swap
swapon /dev/vg_root/lv_swap
echo '/dev/vg_root/lv_swap none swap sw 0 0' >> /etc/fstab
4.3 Create a swap file¶
fallocate -l 4G /swapfile # or: dd if=/dev/zero of=/swapfile bs=1M count=4096
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab
Notes:
- The file must be contiguous with no holes.
fallocateworks on ext4 and XFS; on some filesystems (older kernels, or copy-on-write ones) useddinstead. - Btrfs: swap files need copy-on-write disabled and must not be compressed or snapshotted. Use
btrfs filesystem mkswapfile --size 4g /swapfile(btrfs-progs 6.1+) or setchattr +Con an empty file before writing to it. - ZFS: swap on a ZFS file is not supported; use a zvol.
- Swap on an LVM logical volume is fine and easy to grow (see the LVM guide). Encrypted swap should use a random key per boot (
swapoption incrypttab), or sit on an encrypted volume, so secrets in memory are not left readable on disk.
4.4 Remove or resize swap¶
swapoff /swapfile # pages are moved back into RAM/other swap first; needs enough free memory
rm /swapfile # and delete its fstab line
# To resize: swapoff, recreate at the new size, mkswap, swapon
swapoff -a && swapon -a # cycle all swap, e.g. to flush swapped pages back into RAM
4.5 Swap priority¶
Multiple swap areas can be used in a priority order. Equal priority means round-robin (striping, which helps throughput across separate disks):
swapon --priority 10 /dev/nvme0n1p3 # higher number is used first
# fstab: /swapfile none swap sw,pri=5 0 0
4.6 How much swap?¶
Traditional guidance, still used by Red Hat for sizing:
| RAM | Recommended swap | With hibernation |
|---|---|---|
| 2 GB or less | 2 × RAM | 3 × RAM |
| more than 2 GB up to 8 GB | equal to RAM | 2 × RAM |
| more than 8 GB up to 64 GB | at least 4 GB | 1.5 × RAM |
| more than 64 GB | at least 4 GB | hibernation not recommended |
Older editions of this table recommended 0.5 × RAM for the 8-64 GB range and "4 GB" above 64 GB; current guidance simply sets a floor. Hibernation needs swap at least as large as the RAM in use because the whole memory image is written there.
In practice, swap sizing depends on the workload: memory-hungry batch jobs want more; latency-sensitive services often want a small amount so the kernel can evict cold pages but never has to thrash. Kubernetes nodes have historically disabled swap entirely (the kubelet refuses to start with swap on unless configured), though recent versions support it.
Field notes: personal rule of thumb. For servers with 2 GB of RAM or more, cap swap at around 4 GB or less. Large swap on a big server only delays the point at which a runaway process is noticed and killed.
4.7 Compressed swap alternatives: zram and zswap¶
- zram creates a compressed RAM-backed block device used as swap. Great for small-memory machines and desktops. Install
systemd-zram-generator(Fedora enables it by default) and configure/etc/systemd/zram-generator.conf, for examplezram-size = ram / 2. - zswap is a compressed cache in front of a real swap device. Pages are compressed into a RAM pool first and only spill to disk when the pool fills. Enable with the kernel parameter
zswap.enabled=1, or at runtimeecho 1 > /sys/module/zswap/parameters/enabled. It requires a backing swap device.
4.8 swappiness¶
vm.swappiness (default 60; range 0-200 on kernels 5.8+, 0-100 before) tunes how eagerly the kernel swaps anonymous pages out versus reclaiming file cache.
sysctl vm.swappiness
sudo sysctl vm.swappiness=10
Lower values favor keeping application memory in RAM and dropping cache; higher values swap more readily. 0 does not disable swap; it only makes the kernel reluctant to use it until memory is very tight. To truly avoid swap, don't configure any (swapoff -a), or constrain it per workload with cgroups (memory.swap.max).
5. Overcommit and virtual vs resident memory¶
5.1 Overcommit¶
A program may ask for far more memory than exists (malloc(16 GiB) on an 8 GiB machine) because most requested memory is never touched. The kernel usually agrees and only allocates physical pages when they are first written. This is called overcommit, and it is controlled by:
sysctl vm.overcommit_memory vm.overcommit_ratio vm.overcommit_kbytes
grep -E 'CommitLimit|Committed_AS' /proc/meminfo
vm.overcommit_memory |
Behavior |
|---|---|
0 (default) |
Heuristic: refuse only obviously excessive single allocations |
1 |
Always allow (used by some workloads such as Redis fork-based persistence and certain HPC codes) |
2 |
Strict: total commit may not exceed CommitLimit = swap + RAM × overcommit_ratio% (default 50%) |
Committed_AS is the total address space currently promised to processes; compare it with CommitLimit. Mode 2 makes allocation failures deterministic (malloc returns NULL) instead of the OOM killer firing later, but it also requires that software handle failed allocations gracefully, and it demands more swap or RAM than a default system needs.
5.2 Virtual vs resident, and what a "leak" looks like¶
Two things can grow in a leaking program, and they matter very differently:
- Virtual size (
VIRT,Committed_AS): memory the program has requested. Leaking virtual memory is untidy, but if the pages are never touched no physical memory is consumed. - Resident size (
RSS,RES): physical RAM actually used. Growth here causes a true memory shortage, swapping, and eventually the OOM killer.
When investigating, always watch both (ps -o pid,vsz,rss, or watch -d -n1 'free -m; grep -i commit /proc/meminfo'). Hands-on demonstration: section 16.
6. The OOM killer¶
When the kernel cannot reclaim enough memory to satisfy an allocation, the out-of-memory (OOM) killer picks a process and kills it to free memory.
6.1 Detecting it¶
dmesg -T | grep -iE 'out of memory|oom-kill|killed process'
journalctl -k | grep -i oom
The log shows the victim, its oom_score, memory statistics, and a task table, which is invaluable for finding who was using memory at the time.
6.2 Influencing the choice¶
Every process has a score (/proc/<pid>/oom_score), roughly proportional to its memory use. Adjust it with oom_score_adj (-1000 to +1000):
cat /proc/<pid>/oom_score
echo -500 > /proc/<pid>/oom_score_adj # less likely to be chosen
echo -1000 > /proc/<pid>/oom_score_adj # exempt (use very sparingly: protects one process by endangering everything else)
For services, prefer systemd: OOMScoreAdjust=-500 in the unit file.
6.3 Tuning and alternatives¶
| Setting | Effect |
|---|---|
vm.panic_on_oom |
1 panics instead of killing; useful for clusters that fail over |
vm.oom_kill_allocating_task |
Kill the task that triggered the OOM rather than scanning for the "best" victim |
systemd-oomd, earlyoom |
User-space daemons that act on PSI or free-memory thresholds before the kernel gets desperate, avoiding long stalls |
cgroup memory.oom.group |
Kill every process in a cgroup together (avoids leaving a half-dead service) |
Containers (including Kubernetes pods) are OOM-killed when they exceed their cgroup memory limit even if the host has plenty of RAM; see section 11.
7. Huge pages and Transparent Huge Pages¶
With 4 KiB pages, a process using 64 GiB needs 16 million page-table entries, and the CPU's translation cache (TLB) can only cover a tiny fraction of them. Huge pages (2 MiB, or 1 GiB on x86-64) shrink page-table overhead and TLB misses. They benefit large-memory workloads: databases, JVMs with big heaps, virtual machines, DPDK.
There are two mechanisms.
7.1 Explicit huge pages (hugetlbfs)¶
You reserve a pool of huge pages; applications must be configured to use them.
grep -i huge /proc/meminfo
sysctl vm.nr_hugepages # number of 2 MiB pages reserved
sudo sysctl vm.nr_hugepages=1024 # reserve 2 GiB
Characteristics:
- Reserved memory is not used for anything else and cannot be swapped, so do not over-allocate. Reserve early (at boot) or memory fragmentation may prevent allocation later.
- Use the sizing formula:
nr_hugepages = shared memory needed / Hugepagesize. - For 1 GiB pages, use kernel parameters:
default_hugepagesz=1G hugepagesz=1G hugepages=8. - Persist with
/etc/sysctl.d/:
vm.nr_hugepages = 1024
- Applications need explicit support (for example PostgreSQL huge_pages = try, Oracle USE_LARGE_PAGES, Java -XX:+UseLargePages, QEMU/KVM -mem-path / hugepage backing).
- Control group access with vm.hugetlb_shm_group for SysV shared memory.
7.2 Transparent Huge Pages (THP)¶
THP lets the kernel automatically back suitable anonymous memory with 2 MiB pages, without application changes.
cat /sys/kernel/mm/transparent_hugepage/enabled # always | madvise | never
cat /sys/kernel/mm/transparent_hugepage/defrag
grep -i AnonHugePages /proc/meminfo
always: the kernel uses huge pages whenever it can.madvise: only for memory regions where the application asked (madvise(MADV_HUGEPAGE)). A common, safe choice.never: disabled.
Caveat: THP's background compaction and page-splitting can cause latency spikes and memory bloat. Databases such as Redis, MongoDB, Oracle, and some others explicitly recommend disabling THP or setting it to madvise. Set it persistently on the kernel command line (transparent_hugepage=madvise) or via a systemd unit / tuned profile. Test with your real workload.
8. Memory fragmentation¶
Over time, free memory can become scattered as many small free chunks. Total free memory may be plentiful, yet a request for a large contiguous block (for a huge page, a big DMA buffer, or a high-order kernel allocation) can fail or stall.
8.1 Inspecting fragmentation¶
cat /proc/buddyinfo
Node 0, zone Normal 4210 1820 903 311 87 12 1 0 0 0 0
Each column is the count of free blocks of order 0, 1, 2 ... (block size = 4 KiB × 2^order). Many blocks in the low columns and zeros on the right means memory is fragmented: plenty of free pages, but not in big contiguous runs.
cat /proc/pagetypeinfo # per-migration-type breakdown
cat /sys/kernel/debug/extfrag/extfrag_index # debugfs: how fragmented for each order
8.2 Symptoms¶
- Failed or slow huge page allocation (
nr_hugepagescannot reach the target; THP rarely used). page allocation failure: order:Nmessages indmesg.- High system CPU in
kcompactdor in direct compaction. - Latency spikes when a process needs a large allocation.
8.3 Mitigation¶
echo 1 > /proc/sys/vm/compact_memory # trigger manual compaction (can stall briefly)
- Reserve huge pages at boot, before memory fragments.
- Raise
vm.min_free_kbytesmodestly so reclaim starts earlier and keeps some headroom (careful: too high wastes RAM or triggers OOM). - Set THP to
madviseto avoid constant compaction (section 7.2). - Fix the root cause: restart long-running services that fragment memory, or avoid workloads allocating and freeing many odd-sized kernel objects.
See the reading list for detailed articles on how kernel memory fragmentation works and what compaction does.
9. NUMA¶
On multi-socket servers (and some large single-socket ones), memory is attached to particular CPU sockets: Non-Uniform Memory Access. Accessing remote memory is slower.
numactl --hardware # nodes, sizes, distances
numastat -m # per-node memory usage
numastat -p <pid> # per-node usage for a process
lscpu | grep -i numa
Control placement:
numactl --cpunodebind=0 --membind=0 ./app # pin CPU and memory to node 0
numactl --interleave=all ./database # spread allocations across nodes
Symptoms of NUMA trouble: one node exhausted while another has plenty of free memory, causing swapping or reclaim on a machine that appears to have free RAM. vm.zone_reclaim_mode (default 0) controls whether the kernel prefers reclaiming locally over using remote memory; leave it at 0 unless you have measured a benefit. Modern kernels can also migrate pages automatically (kernel.numa_balancing).
10. Tuning with sysctl¶
Parameters live in /proc/sys/vm/, readable and writable via sysctl.
sysctl -a | grep '^vm\.' # list all
sudo sysctl vm.swappiness=10 # set temporarily
Persist in a drop-in file, then reload:
sudo tee /etc/sysctl.d/90-memory.conf <<'EOF'
vm.swappiness = 10
EOF
sudo sysctl --system
Method: change one parameter, measure under a realistic load, and keep only what demonstrably helps. There are no universally right values.
| Parameter | Default | What it does | Typical tuning |
|---|---|---|---|
vm.swappiness |
60 | Preference for swapping anon pages vs reclaiming cache | 10 for databases/latency-sensitive; 60+ for cache-heavy file servers |
vm.vfs_cache_pressure |
100 | Tendency to reclaim dentry/inode caches; >100 reclaims faster, <100 keeps them | 50 for servers with millions of files; avoid 0 |
vm.dirty_background_ratio |
10 | % of memory of dirty data at which background writeback starts | 5 or use dirty_background_bytes on large-RAM systems |
vm.dirty_ratio |
20 | % of dirty data at which writers are forced to write synchronously | 10-15; use dirty_bytes on large-RAM systems (percentages get huge) |
vm.dirty_expire_centisecs |
3000 | Age at which dirty data must be written | Lower for quicker durability |
vm.dirty_writeback_centisecs |
500 | How often flusher threads wake | Rarely changed |
vm.min_free_kbytes |
auto | Free memory floor kept for the kernel/atomic allocations | Raise slightly for network-heavy or fragmented systems |
vm.watermark_scale_factor |
10 | Gap between watermarks; higher wakes kswapd earlier |
Raise (e.g., 100-200) to reduce direct reclaim stalls |
vm.overcommit_memory / _ratio |
0 / 50 | Overcommit policy (section 5) | 1 for Redis-like fork workloads; 2 for strict environments |
vm.max_map_count |
65530 | Max memory-mapped regions per process | Raise to 262144 for Elasticsearch/OpenSearch and similar |
vm.nr_hugepages |
0 | Reserved 2 MiB huge pages | See section 7 |
vm.panic_on_oom / oom_kill_allocating_task |
0 / 0 | OOM behavior | Cluster-specific |
vm.zone_reclaim_mode |
0 | NUMA local reclaim preference | Leave at 0 |
vm.compact_memory |
write-only | Trigger compaction | Ad-hoc |
vm.drop_caches |
write-only | Drop caches (section 14) | Diagnostics only |
Tip: on machines with very large RAM, dirty ratios translate into enormous byte amounts (20% of 256 GiB is over 50 GiB), so writers can stall for a long time when the limit is finally hit. Prefer vm.dirty_background_bytes and vm.dirty_bytes, which override the ratios when set.
11. Limiting memory with cgroups and systemd¶
Control groups cap and account memory per group of processes. They underpin containers (Docker, Podman, Kubernetes) and systemd services.
stat -fc %T /sys/fs/cgroup # "cgroup2fs" = unified v2 hierarchy (default on current distros)
11.1 cgroup v2 memory files¶
| File | Meaning |
|---|---|
memory.current |
Current usage |
memory.min / memory.low |
Protected memory (guaranteed / best-effort) |
memory.high |
Soft limit: exceeding it throttles the group with reclaim (no kill) |
memory.max |
Hard limit: exceeding it triggers the cgroup OOM killer |
memory.swap.max |
Swap allowance (0 disables swap for this group) |
memory.stat |
Detailed breakdown (anon, file, slab, ...) |
memory.events |
Counters for high, max, oom, oom_kill events |
memory.pressure |
PSI for the group |
11.2 systemd¶
# in a service unit or drop-in (systemctl edit myservice)
[Service]
MemoryHigh=1500M
MemoryMax=2G
MemorySwapMax=0
OOMScoreAdjust=200
Ad-hoc test:
systemd-run --scope -p MemoryMax=512M -p MemorySwapMax=0 ./memory-hungry-program
systemctl status myservice # shows current memory and limits
systemd-cgtop -m # top-like view of memory per cgroup
11.3 Containers and Kubernetes¶
- A container's
memory.maxcomes fromdocker run --memory=or the Kubernetesresources.limits.memory. Going over it triggers an OOMKilled (exit code 137) even if the node has RAM to spare. - Kubernetes
requests.memoryinfluences scheduling and OOM-kill priority (QoS classes: Guaranteed, Burstable, BestEffort), but is not enforced as a hard cap. - Page cache counts toward the cgroup's usage but is reclaimed under pressure, so
memory.currentcan look near the limit without being in danger. Look atmemory.statand the working set (working_set_bytesin metrics), not just raw usage.
12. Inspecting hardware memory¶
12.1 Maximum supported memory¶
The motherboard's limit comes from the firmware's DMI/SMBIOS tables:
sudo dmidecode -t 16
Physical Memory Array
Location: System Board Or Motherboard
Use: System Memory
Error Correction Type: None
Maximum Capacity: 32 GB
Number Of Devices: 4
Maximum Capacity is the total the board supports and Number Of Devices is the number of DIMM slots. Error Correction Type shows whether ECC is supported.
12.2 Memory installed¶
sudo lshw -short -C memory
H/W path Device Class Description
=============================================
/0/1 memory 64KiB BIOS
/0/26 memory 256KiB L1 cache
/0/28 memory 1MiB L2 cache
/0/29 memory 6MiB L3 cache
/0/2a memory 20GiB System Memory
/0/2a/0 memory 2GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/1 memory 2GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/2 memory 8GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/3 memory 8GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
More detail per slot (size, speed, type, manufacturer, part number, which slots are empty):
sudo dmidecode -t memory
sudo dmidecode -t 17 # memory devices only
Compare the mix of DIMMs against what the board supports; mismatched sizes or slot layouts can cause reduced channel interleaving.
12.3 ECC errors¶
On ECC systems, the kernel's EDAC subsystem reports corrected and uncorrected errors:
edac-util -v
ras-mc-ctl --summary
ras-mc-ctl --errors
dmesg | grep -i -E 'edac|mce|ecc'
Corrected errors that keep repeating on one DIMM indicate a module heading for failure.
12.4 Quick facts¶
free -h
lsmem # memory blocks and online state
cat /proc/meminfo | head -3
getconf PAGESIZE
13. Testing memory (memtest86+)¶
Faulty RAM causes random crashes, segfaults, and data corruption. memtest86+ boots outside the OS and hammers every address with test patterns.
13.1 Installing on RHEL-family systems (BIOS boot)¶
sudo yum install memtest86+ memtest-setup # dnf on newer releases
sudo memtest-setup # adds a GRUB entry
sudo grub2-mkconfig -o /boot/grub2/grub.cfg # regenerate the GRUB config
On UEFI systems the GRUB config lives elsewhere (/boot/efi/EFI/<distro>/grub.cfg). Many current packages of memtest86+ (version 6+) support UEFI directly.
13.2 Debian/Ubuntu¶
sudo apt install memtest86+
sudo update-grub
Reboot and choose Memtest86+ from the GRUB menu. If the OS package is awkward, boot the official memtest86+ image from a USB stick instead.
13.3 Practical advice¶
- Run at least one full pass; for suspected problems, run several passes or overnight. Some faults show up only after hours.
- Any error is a failure. Reseat DIMMs, then test each module alone in a known-good slot to identify the bad one.
- If the machine can't be taken offline,
memtester(user-space,memtester 1G 3) can test the portion of RAM it can allocate, though it cannot cover memory used by the kernel or other programs.
14. Clearing caches¶
The kernel can be asked to discard clean cache. This is a diagnostic or benchmarking tool, not a routine "memory cleanup". Dropping caches forces the system to re-read data from disk afterward, so performance is worse until the cache warms up again. Cache is normally released automatically whenever programs need the memory.
# Flush dirty pages to disk first, so more cache is clean and droppable
sync
# Drop the page cache only
echo 1 | sudo tee /proc/sys/vm/drop_caches
# Drop reclaimable slab objects (dentries and inodes)
echo 2 | sudo tee /proc/sys/vm/drop_caches
# Drop both page cache and slab objects
echo 3 | sudo tee /proc/sys/vm/drop_caches
(sudo echo 3 > /proc/... fails because the redirection is done by your shell, not by sudo; use tee or a root shell.)
Legitimate uses: benchmarking cold-cache performance, reproducing a problem that depends on a cold cache, or measuring how much of a process's I/O really hits the disk. Illegitimate use: putting it in cron to make free look nicer. That hurts performance and hides real problems.
To free swapped-out pages, cycle swap instead: swapoff -a && swapon -a (only when there is enough free RAM to absorb them).
15. Finding and debugging memory leaks¶
A leak is memory a program allocates but never releases, so usage grows over time. First confirm the leak is real, then find where.
15.1 Confirming growth¶
# Track a process over time
while sleep 60; do ps -o pid,rss,vsz,cmd -p <pid>; done
pidstat -r -p <pid> 60 # sysstat
watch -d -n1 'free -m; grep -i commit /proc/meminfo'
- RSS steadily rising: a real shortage is developing (resident leak).
- VSZ/Committed_AS rising, RSS flat: virtual-only leak (allocated, never touched); less harmful but still worth reporting.
- Rising then plateauing: probably a cache, not a leak.
- Also consider fragmentation and allocator behavior: some allocators (glibc
malloc) retain freed memory in arenas and never return it to the OS, which looks like a leak. SettingMALLOC_ARENA_MAX=2or usingjemalloc/tcmalloccan change the picture.
15.2 valgrind¶
valgrind --tool=memcheck ./program [program options]
valgrind --tool=memcheck /usr/bin/httpd # any binary (note: very slow, a 10-50x slowdown)
# Detailed leak report with allocation stack traces
valgrind --tool=memcheck --leak-check=full --show-leak-kinds=all --track-origins=yes ./program
Leak categories in the summary: definitely lost (real leaks), indirectly lost (reachable only through a lost block), possibly lost, and still reachable (memory alive at exit, usually harmless). Compile with -g (and without heavy optimization) to get file and line numbers.
Related valgrind tools:
valgrind --tool=massif ./program # heap profiler: who allocates memory over time
ms_print massif.out.<pid> # view the profile
15.3 Other tools¶
| Tool | Use |
|---|---|
AddressSanitizer / LeakSanitizer (-fsanitize=address) |
Compile-time instrumentation, much faster than valgrind; reports leaks at exit |
| heaptrack | Low-overhead heap profiler with a GUI |
memleak-bpfcc (BCC/eBPF) |
Attach to a running process and list outstanding allocations by stack: sudo memleak-bpfcc -p <pid> |
pmap -x, /proc/<pid>/smaps |
See which mappings (heap, anon, libraries) grew |
| jemalloc/tcmalloc heap profiling | Production-friendly profiling for C/C++ services |
| Language runtime tools | Java (jcmd GC.heap_dump, MAT), Go (pprof heap), Python (tracemalloc, objgraph), Node (--inspect heap snapshots) |
15.4 What to do next¶
A leak found in third-party software goes to its maintainers, with the valgrind or profiler output, the version, and steps to reproduce. Fixing it requires changing the source, rebuilding, and re-testing. Until then, contain it: restart on a schedule, or set a cgroup memory limit (MemoryMax=) and let systemd restart the service, so a leak becomes a controlled restart instead of taking down the host.
16. Field notes: leak lab exercise¶
This exercise makes the two kinds of leak (section 5.2) visible. The original notes used a training tool called bigmem; here is an equivalent you can build anywhere.
16.1 The test program¶
/* bigmem.c: allocate N MiB and either touch it (resident) or not (virtual only) */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
int main(int argc, char **argv) {
int virtual_only = 0, arg = 1;
if (argc > 1 && strcmp(argv[1], "-v") == 0) { virtual_only = 1; arg++; }
if (arg >= argc) { fprintf(stderr, "usage: %s [-v] MiB\n", argv[0]); return 1; }
size_t bytes = (size_t)atol(argv[arg]) * 1024 * 1024;
char *p = malloc(bytes);
if (!p) { perror("malloc"); return 1; }
if (!virtual_only) memset(p, 1, bytes); /* touching the pages makes them resident */
printf("allocated %zu MiB (%s); sleeping, Ctrl-C to exit\n",
bytes >> 20, virtual_only ? "virtual only" : "resident");
pause(); /* never free(p): this is the "leak" */
return 0;
}
gcc -g -o bigmem bigmem.c
16.2 Step 1: run under valgrind¶
valgrind --tool=memcheck ./bigmem 256 # resident: allocates and touches 256 MiB
valgrind --tool=memcheck ./bigmem -v 256 # virtual only: allocates 256 MiB but never touches it
Valgrind flags a leak in both cases. It cannot tell you which kind of leak it is.
16.3 Step 2: watch the system from a second terminal¶
watch -d -n1 'free -m; grep -i commit /proc/meminfo'
This shows physical memory and the committed (promised) totals side by side.
16.4 Step 3: run without valgrind and compare¶
./bigmem 256 # resident
./bigmem -v 256 # virtual only
Expected observations:
| Run | Committed_AS |
used / RSS |
|---|---|---|
./bigmem 256 |
+256 MiB | +256 MiB (physical memory really consumed) |
./bigmem -v 256 |
+256 MiB | Only a few MiB (program image, page tables); the untouched pages never became resident |
Committed memory rises by the requested amount in both cases; only the resident case consumes real RAM. That distinction is exactly what separates "a leak worth reporting" from "a leak that will take the server down".
17. Troubleshooting workflow¶
"The server is slow." Work through this list from top to bottom:
- Is it swapping?
vmstat 1: sustainedsi/sovalues, highwa. If yes, memory is short (or something is bloated). - How much is truly available?
free -h: look atavailable, notfree. Checkcat /proc/pressure/memoryfor stalls. - Who is using it?
ps aux --sort=-rss | head,smem -tk -s pss,systemd-cgtop -m. Includetmpfs(df -h /dev/shm /tmp), since tmpfs files consume RAM and swap. - Kernel memory?
slabtop, and compareSlab/SUnreclaimandPageTablesin/proc/meminfo. A growing slab points to a kernel or driver leak (or a huge dentry cache). - Is anything being OOM-killed?
dmesg -T | grep -i oom;journalctl -k -b | grep -i 'killed process'. - Is the cache being squeezed? High
biinvmstatwith low cache and rising disk latency means the hot working set no longer fits. - Fragmentation?
cat /proc/buddyinfo; failed huge-page allocations;kcompactdCPU use. - NUMA imbalance?
numastat -mshows one node exhausted. - Limits reached? Container/cgroup
memory.eventsandmemory.max; overcommit (Committed_ASvsCommitLimitin strict mode). - Hardware? ECC counters in
dmesg/ras-mc-ctl; run memtest86+ if crashes look random.
| Symptom | Likely cause | First action |
|---|---|---|
free shows almost no free memory |
Normal page-cache usage | Check available instead |
| Swap in use but system fast | Cold pages parked in swap | Nothing; only worry if si/so are active |
Constant swap in/out, high wa |
Working set exceeds RAM | Find the memory hog, add RAM, add limits, tune swappiness |
| Process killed unexpectedly | OOM killer or cgroup limit | dmesg, check memory.events and limits |
| RSS grows without bound | Leak | valgrind, heaptrack, memleak-bpfcc |
| Periodic latency stalls | THP compaction, direct reclaim, or dirty-page throttling | Set THP to madvise, raise watermark_scale_factor, use dirty_bytes |
page allocation failure: order:N |
Fragmentation | Compaction; reserve huge pages at boot; raise min_free_kbytes |
Container OOMKilled (137) with host RAM to spare |
cgroup limit too low | Raise the limit; check working set vs limit |
| Random crashes or corruption | Bad RAM | memtest86+, check ECC logs |
18. Cheat sheet¶
Look¶
free -h # summary (watch "available")
vmstat 1 # swap, IO, CPU activity
cat /proc/meminfo # everything
cat /proc/pressure/memory # PSI
ps aux --sort=-rss | head # top memory users
smem -tk -s pss # shared-memory-aware accounting
slabtop -o # kernel caches
swapon --show # swap devices
cat /proc/buddyinfo # fragmentation
numastat -m # NUMA usage
dmesg -T | grep -i oom # OOM events
Hardware¶
sudo dmidecode -t 16 # max capacity and slot count
sudo dmidecode -t 17 # each DIMM
sudo lshw -short -C memory # installed memory summary
getconf PAGESIZE # base page size
Tune¶
sudo sysctl vm.swappiness=10
sudo sysctl vm.vfs_cache_pressure=50
sudo sysctl vm.dirty_background_ratio=5 vm.dirty_ratio=15
sudo sysctl vm.nr_hugepages=1024
echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
sudo tee /etc/sysctl.d/90-memory.conf <<< 'vm.swappiness = 10' && sudo sysctl --system
Swap¶
fallocate -l 4G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
swapoff /swapfile
swapon --show
Limit¶
systemd-run --scope -p MemoryMax=512M -p MemorySwapMax=0 ./cmd
echo -500 | sudo tee /proc/<pid>/oom_score_adj
Debug leaks¶
valgrind --tool=memcheck --leak-check=full --show-leak-kinds=all ./prog
valgrind --tool=massif ./prog && ms_print massif.out.*
sudo memleak-bpfcc -p <pid>
watch -d -n1 'free -m; grep -i commit /proc/meminfo'
Caches (diagnostics only)¶
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches
19. Further reading¶
- Man pages:
proc(5)(the/proc/meminfoand/proc/sys/vmsections),free(1),vmstat(8),swapon(8),mkswap(8),sysctl(8),numactl(8),valgrind(1),systemd.resource-control(5) - Kernel documentation:
Documentation/admin-guide/sysctl/vm.rst, Transparent Hugepage Support, cgroup v2, PSI - Linux ate my RAM!: why free memory is low on a healthy system
- Debian Wiki: Hugepages
- Thomas-Krenn: Linux Page Cache Basics
- PingCAP: Linux kernel vs memory fragmentation, part 1 and part 2
- Red Hat knowledge base: 641323
- Tecmint: Clear RAM memory cache, buffer and swap space on Linux
- Red Hat: swap space sizing guidance, Server Fault: how much swap on a high-memory system, cyberciti.biz: Linux swap space
- Pro Ubuntu Server Administration, chapters on memory and performance (pages 49, 57, and 91 covered the top output, the cache versus swap distinction, and virtual memory tuning)