Andrew Mercer
on this page

Linux Memory Management: The Complete Guide

Memory is usually the resource that decides whether a server feels fast or slow. Linux manages it aggressively: it hands out more address space than physically exists, uses nearly all free RAM as disk cache, and moves cold pages out to swap when it needs room. Much of what looks alarming in free or top is normal behavior, and much of what is actually a problem (swap thrashing, fragmentation, runaway processes) can be spotted quickly once you know what to look at.

This guide covers the theory, the tools, the tuning knobs, and the troubleshooting workflow. Sections marked Field notes hold practical notes and lab exercises kept separate from the general material.

Conventions

  • Commands need root unless noted.
  • Kernel parameters are shown as sysctl names (vm.swappiness), which map to files under /proc/sys/vm/.
  • Change one parameter at a time and measure. No setting is right for every workload.

Table of Contents

  1. Core concepts
  2. Reading memory usage
  3. The page cache (and why "Linux ate my RAM" is a myth)
  4. Swap
  5. Overcommit and virtual vs resident memory
  6. The OOM killer
  7. Huge pages and Transparent Huge Pages
  8. Memory fragmentation
  9. NUMA
  10. Tuning with sysctl
  11. Limiting memory with cgroups and systemd
  12. Inspecting hardware memory
  13. Testing memory (memtest86+)
  14. Clearing caches
  15. Finding and debugging memory leaks
  16. Field notes: leak lab exercise
  17. Troubleshooting workflow
  18. Cheat sheet
  19. Further reading

1. Core concepts

1.1 Virtual memory

Linux uses virtual memory: every process sees its own private, contiguous address space, and the kernel plus the CPU's memory management unit (MMU) translate those virtual addresses to physical RAM. This gives:

  • Isolation: one process cannot read or write another's memory. The kernel builds a fresh address space when a process is created (fork()), and copy-on-write keeps forks cheap because parent and child share pages until one writes.
  • Flexibility: memory can be allocated lazily, shared between processes (libraries, shared memory), memory-mapped from files, or pushed out to swap.
  • Overcommit: the total virtual memory handed out may exceed physical RAM.

From an administrator's viewpoint, "virtual memory" also means the total usable pool: RAM plus swap. The knobs that control how it behaves live under /proc/sys/vm/.

1.2 Pages

Memory is managed in fixed-size pages. Everything the kernel does with memory (allocating, mapping, reading from disk, swapping) happens in page-sized chunks.

getconf PAGESIZE          # base page size, in bytes
  • x86-64: 4 KiB base pages (fixed), with 2 MiB and 1 GiB "huge" pages available on top.
  • ARM64: 4, 16, or 64 KiB base pages, chosen when the kernel is built (Red Hat and Apple-silicon-era kernels differ from each other).
  • 32-bit x86: 4 KiB.

Small pages are efficient for many small files and allocations. For very large working sets (databases, virtual machines), tracking millions of 4 KiB pages is costly, which is what huge pages solve (section 7).

1.3 Memory kinds

Kind What it is Reclaimable?
Anonymous memory Heap, stack, malloced data. Has no file behind it. Only by writing to swap
File-backed memory Page cache: file contents, executables, mmapped files Yes: clean pages are simply dropped; dirty pages are written back first
Shared memory / tmpfs /dev/shm, tmpfs mounts, SysV/POSIX shared memory Only via swap (it is anonymous in nature)
Kernel memory Slab caches (dentries, inodes, network buffers), page tables, kernel stacks Some of it (SReclaimable), not all
Huge pages (hugetlb) Pages explicitly reserved at boot/runtime No; not swappable, pre-reserved

1.4 RAM vs disk speed

Memory is roughly 1,000 times faster than a spinning disk and still orders of magnitude faster than SSDs when you count latency. A CPU only runs smoothly when its data is in RAM (or even better, in the CPU's own caches). When the working set does not fit, the system falls back to slower I/O and everything queues behind it. That is why memory often matters more than CPU on a server, and why swap activity is the first thing to check on a slow machine.

1.5 Zones, watermarks, and reclaim

The kernel groups physical memory into zones (for example DMA, DMA32, Normal) and per NUMA node. It keeps free memory above watermarks (min/low/high). When free memory dips below the low watermark, the kswapd thread wakes up and reclaims pages in the background: it drops clean cache, writes back dirty pages, and swaps out cold anonymous pages. If allocations outrun kswapd and free memory reaches the min watermark, processes must reclaim memory themselves ("direct reclaim"), which causes latency spikes. If reclaim fails entirely, the OOM killer steps in (section 6).


2. Reading memory usage

2.1 free

free -h            # human-readable
free -m            # MiB
free -h -s 2       # refresh every 2 seconds
               total        used        free      shared  buff/cache   available
Mem:            15Gi       3.1Gi       1.2Gi       210Mi        11Gi        12Gi
Swap:          4.0Gi          0B       4.0Gi
Column Meaning
total Installed usable RAM
used Memory in use by programs and the kernel (excluding reclaimable cache)
free Completely unused memory. Low is normal.
shared tmpfs and shared memory
buff/cache Page cache, buffers, and reclaimable slab
available Estimate of memory available to start new work without swapping. This is the number to watch.

2.2 /proc/meminfo

cat /proc/meminfo
grep -E 'MemTotal|MemAvailable|Cached|Dirty|Swap|Commit|Huge|Slab|Shmem' /proc/meminfo

Key fields:

Field Meaning
MemAvailable Best estimate of allocatable memory without swapping
Buffers / Cached Block-device buffers / page cache
Active(anon), Inactive(anon), Active(file), Inactive(file) LRU lists; inactive pages are reclaimed first
Dirty, Writeback Data waiting to be / being written to disk
AnonPages Anonymous (non-file) memory
Mapped File pages mapped into process address spaces
Shmem Shared memory and tmpfs
Slab, SReclaimable, SUnreclaim Kernel object caches
PageTables Memory used for address-translation tables
SwapTotal, SwapFree, SwapCached Swap capacity, free, and pages present in both RAM and swap
CommitLimit, Committed_AS Overcommit accounting (section 5)
HugePages_Total/Free/Rsvd, Hugepagesize, AnonHugePages Huge page pools and THP usage

2.3 Per-system activity: vmstat, sar, top

vmstat 1                  # one line per second
vmstat -SM 1              # in MiB
sar -r 1 5                # memory utilization (sysstat)
sar -B 1 5                # paging statistics
sar -S 1 5                # swap utilization

Important vmstat columns:

Column Meaning Concern when...
si / so Swap-in / swap-out (KiB/s) Sustained non-zero values
bi / bo Blocks read/written from disk Very high with low CPU
free, buff, cache Memory pools (see available instead)
wa CPU time waiting for I/O High, together with swapping
r, b Runnable / blocked processes b consistently > 0

In top/htop, the header shows memory and swap. Per-process columns:

Column Meaning
VIRT Total virtual address space the process has mapped (includes untouched and file-mapped memory)
RES Resident memory: physical RAM the process is using now
SHR Portion of RES that is shared (libraries, shared memory)
%MEM RES as a percentage of RAM
SWAP/VmSwap Memory of the process currently in swap

In top, press M to sort by memory. In htop, F6 selects the sort column.

2.4 Per-process detail

ps aux --sort=-rss | head                 # biggest resident users
ps -eo pid,comm,rss,vsz,%mem --sort=-rss | head
pmap -x <pid>                             # mapping-by-mapping breakdown
cat /proc/<pid>/status | grep -E 'Vm|Rss'
cat /proc/<pid>/smaps_rollup              # PSS (proportional set size) totals
smem -tk -s pss                           # accurate accounting for shared memory (package: smem)

RSS counts shared pages fully in every process that maps them, so summing RSS overstates usage. PSS divides shared pages proportionally and adds up correctly.

2.5 Kernel memory and pressure

sudo slabtop -o                # kernel slab caches, sorted by size
cat /proc/pressure/memory      # pressure stall information (PSI), kernel 4.20+

PSI output:

some avg10=0.00 avg60=0.12 avg300=0.05 total=123456
full avg10=0.00 avg60=0.00 avg300=0.00 total=6789

some is the share of time at least one task was stalled waiting on memory; full is the share of time all non-idle tasks were stalled. Sustained non-zero full is real trouble. PSI is a more honest signal than "free memory" for deciding whether a system is short of RAM.


3. The page cache

Whenever a file is read, Linux keeps its contents in RAM (the read cache) so the next read is served from memory instead of the disk. Writes go into RAM first and are flushed later (the write cache, dirty pages). Free RAM is wasted RAM, so the kernel uses almost all otherwise unused memory as cache.

Consequences:

  • A busy file server or database host will show very little free memory and a large buff/cache. That is healthy. Clean cache is released instantly whenever a program needs the memory.
  • Judge memory health by available, swap activity, and PSI, not by free.
  • Cache is not swap. Cache is RAM used to accelerate disk access; swap is disk space used to extend RAM.
  • For a read-heavy server, a healthy amount of cache directly translates to speed. If the working set of hot files no longer fits and the cache keeps shrinking and refilling (rising bi in vmstat, falling cache hit ratios), the machine needs more RAM. An old rule of thumb suggests trouble when the cache falls under about 40% of RAM on a read-mostly server, but treat that as a rough hint rather than a threshold: measure disk reads and latency instead.

Dirty pages and writeback

Writes accumulate as dirty pages and are flushed in the background by kernel flusher threads.

grep -E 'Dirty|Writeback' /proc/meminfo
sysctl vm.dirty_ratio vm.dirty_background_ratio vm.dirty_expire_centisecs vm.dirty_writeback_centisecs

See section 10 for tuning. In short: a low dirty_background_ratio starts flushing early (smoother I/O, less cache benefit); a high dirty_ratio allows large bursts to absorb but risks long stalls when the limit is hit and writers get throttled.


4. Swap

Swap is disk space used as overflow for anonymous memory (and to hold shared memory pages). Because disks are far slower than RAM, heavy swapping ("thrashing") can cripple a server. But "swap usage" is not automatically bad:

  • A small amount of swap used by long-idle processes is healthy. The kernel moves cold pages out so RAM can hold hot data and cache.
  • The signal for trouble is ongoing swap-in/swap-out activity (si/so in vmstat), especially with high I/O wait, not the raw amount used.
  • Some applications (traditionally Oracle databases) are designed around swap usage. For everything else, be suspicious if more than a modest amount of swap is actively in use.
  • If a server is slow, look at swapping first.

4.1 Inspect swap

swapon --show            # devices/files, size, used, priority
free -h
cat /proc/swaps
vmstat 1                 # watch si/so
grep -E 'VmSwap|Name' /proc/*/status 2>/dev/null | paste - - | awk '$4>0 {print $4, $2}' | sort -nr | head   # biggest swap users

4.2 Create a swap partition or LV

mkswap /dev/vg_root/lv_swap
swapon /dev/vg_root/lv_swap
echo '/dev/vg_root/lv_swap none swap sw 0 0' >> /etc/fstab

4.3 Create a swap file

fallocate -l 4G /swapfile            # or: dd if=/dev/zero of=/swapfile bs=1M count=4096
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab

Notes:

  • The file must be contiguous with no holes. fallocate works on ext4 and XFS; on some filesystems (older kernels, or copy-on-write ones) use dd instead.
  • Btrfs: swap files need copy-on-write disabled and must not be compressed or snapshotted. Use btrfs filesystem mkswapfile --size 4g /swapfile (btrfs-progs 6.1+) or set chattr +C on an empty file before writing to it.
  • ZFS: swap on a ZFS file is not supported; use a zvol.
  • Swap on an LVM logical volume is fine and easy to grow (see the LVM guide). Encrypted swap should use a random key per boot (swap option in crypttab), or sit on an encrypted volume, so secrets in memory are not left readable on disk.

4.4 Remove or resize swap

swapoff /swapfile            # pages are moved back into RAM/other swap first; needs enough free memory
rm /swapfile                 # and delete its fstab line
# To resize: swapoff, recreate at the new size, mkswap, swapon
swapoff -a && swapon -a      # cycle all swap, e.g. to flush swapped pages back into RAM

4.5 Swap priority

Multiple swap areas can be used in a priority order. Equal priority means round-robin (striping, which helps throughput across separate disks):

swapon --priority 10 /dev/nvme0n1p3          # higher number is used first
# fstab: /swapfile none swap sw,pri=5 0 0

4.6 How much swap?

Traditional guidance, still used by Red Hat for sizing:

RAM Recommended swap With hibernation
2 GB or less 2 × RAM 3 × RAM
more than 2 GB up to 8 GB equal to RAM 2 × RAM
more than 8 GB up to 64 GB at least 4 GB 1.5 × RAM
more than 64 GB at least 4 GB hibernation not recommended

Older editions of this table recommended 0.5 × RAM for the 8-64 GB range and "4 GB" above 64 GB; current guidance simply sets a floor. Hibernation needs swap at least as large as the RAM in use because the whole memory image is written there.

In practice, swap sizing depends on the workload: memory-hungry batch jobs want more; latency-sensitive services often want a small amount so the kernel can evict cold pages but never has to thrash. Kubernetes nodes have historically disabled swap entirely (the kubelet refuses to start with swap on unless configured), though recent versions support it.

Field notes: personal rule of thumb. For servers with 2 GB of RAM or more, cap swap at around 4 GB or less. Large swap on a big server only delays the point at which a runaway process is noticed and killed.

4.7 Compressed swap alternatives: zram and zswap

  • zram creates a compressed RAM-backed block device used as swap. Great for small-memory machines and desktops. Install systemd-zram-generator (Fedora enables it by default) and configure /etc/systemd/zram-generator.conf, for example zram-size = ram / 2.
  • zswap is a compressed cache in front of a real swap device. Pages are compressed into a RAM pool first and only spill to disk when the pool fills. Enable with the kernel parameter zswap.enabled=1, or at runtime echo 1 > /sys/module/zswap/parameters/enabled. It requires a backing swap device.

4.8 swappiness

vm.swappiness (default 60; range 0-200 on kernels 5.8+, 0-100 before) tunes how eagerly the kernel swaps anonymous pages out versus reclaiming file cache.

sysctl vm.swappiness
sudo sysctl vm.swappiness=10

Lower values favor keeping application memory in RAM and dropping cache; higher values swap more readily. 0 does not disable swap; it only makes the kernel reluctant to use it until memory is very tight. To truly avoid swap, don't configure any (swapoff -a), or constrain it per workload with cgroups (memory.swap.max).


5. Overcommit and virtual vs resident memory

5.1 Overcommit

A program may ask for far more memory than exists (malloc(16 GiB) on an 8 GiB machine) because most requested memory is never touched. The kernel usually agrees and only allocates physical pages when they are first written. This is called overcommit, and it is controlled by:

sysctl vm.overcommit_memory vm.overcommit_ratio vm.overcommit_kbytes
grep -E 'CommitLimit|Committed_AS' /proc/meminfo
vm.overcommit_memory Behavior
0 (default) Heuristic: refuse only obviously excessive single allocations
1 Always allow (used by some workloads such as Redis fork-based persistence and certain HPC codes)
2 Strict: total commit may not exceed CommitLimit = swap + RAM × overcommit_ratio% (default 50%)

Committed_AS is the total address space currently promised to processes; compare it with CommitLimit. Mode 2 makes allocation failures deterministic (malloc returns NULL) instead of the OOM killer firing later, but it also requires that software handle failed allocations gracefully, and it demands more swap or RAM than a default system needs.

5.2 Virtual vs resident, and what a "leak" looks like

Two things can grow in a leaking program, and they matter very differently:

  • Virtual size (VIRT, Committed_AS): memory the program has requested. Leaking virtual memory is untidy, but if the pages are never touched no physical memory is consumed.
  • Resident size (RSS, RES): physical RAM actually used. Growth here causes a true memory shortage, swapping, and eventually the OOM killer.

When investigating, always watch both (ps -o pid,vsz,rss, or watch -d -n1 'free -m; grep -i commit /proc/meminfo'). Hands-on demonstration: section 16.


6. The OOM killer

When the kernel cannot reclaim enough memory to satisfy an allocation, the out-of-memory (OOM) killer picks a process and kills it to free memory.

6.1 Detecting it

dmesg -T | grep -iE 'out of memory|oom-kill|killed process'
journalctl -k | grep -i oom

The log shows the victim, its oom_score, memory statistics, and a task table, which is invaluable for finding who was using memory at the time.

6.2 Influencing the choice

Every process has a score (/proc/<pid>/oom_score), roughly proportional to its memory use. Adjust it with oom_score_adj (-1000 to +1000):

cat /proc/<pid>/oom_score
echo -500 > /proc/<pid>/oom_score_adj      # less likely to be chosen
echo -1000 > /proc/<pid>/oom_score_adj     # exempt (use very sparingly: protects one process by endangering everything else)

For services, prefer systemd: OOMScoreAdjust=-500 in the unit file.

6.3 Tuning and alternatives

Setting Effect
vm.panic_on_oom 1 panics instead of killing; useful for clusters that fail over
vm.oom_kill_allocating_task Kill the task that triggered the OOM rather than scanning for the "best" victim
systemd-oomd, earlyoom User-space daemons that act on PSI or free-memory thresholds before the kernel gets desperate, avoiding long stalls
cgroup memory.oom.group Kill every process in a cgroup together (avoids leaving a half-dead service)

Containers (including Kubernetes pods) are OOM-killed when they exceed their cgroup memory limit even if the host has plenty of RAM; see section 11.


7. Huge pages and Transparent Huge Pages

With 4 KiB pages, a process using 64 GiB needs 16 million page-table entries, and the CPU's translation cache (TLB) can only cover a tiny fraction of them. Huge pages (2 MiB, or 1 GiB on x86-64) shrink page-table overhead and TLB misses. They benefit large-memory workloads: databases, JVMs with big heaps, virtual machines, DPDK.

There are two mechanisms.

7.1 Explicit huge pages (hugetlbfs)

You reserve a pool of huge pages; applications must be configured to use them.

grep -i huge /proc/meminfo
sysctl vm.nr_hugepages                        # number of 2 MiB pages reserved
sudo sysctl vm.nr_hugepages=1024              # reserve 2 GiB

Characteristics:

  • Reserved memory is not used for anything else and cannot be swapped, so do not over-allocate. Reserve early (at boot) or memory fragmentation may prevent allocation later.
  • Use the sizing formula: nr_hugepages = shared memory needed / Hugepagesize.
  • For 1 GiB pages, use kernel parameters: default_hugepagesz=1G hugepagesz=1G hugepages=8.
  • Persist with /etc/sysctl.d/:

vm.nr_hugepages = 1024 - Applications need explicit support (for example PostgreSQL huge_pages = try, Oracle USE_LARGE_PAGES, Java -XX:+UseLargePages, QEMU/KVM -mem-path / hugepage backing). - Control group access with vm.hugetlb_shm_group for SysV shared memory.

7.2 Transparent Huge Pages (THP)

THP lets the kernel automatically back suitable anonymous memory with 2 MiB pages, without application changes.

cat /sys/kernel/mm/transparent_hugepage/enabled     # always | madvise | never
cat /sys/kernel/mm/transparent_hugepage/defrag
grep -i AnonHugePages /proc/meminfo
  • always: the kernel uses huge pages whenever it can.
  • madvise: only for memory regions where the application asked (madvise(MADV_HUGEPAGE)). A common, safe choice.
  • never: disabled.

Caveat: THP's background compaction and page-splitting can cause latency spikes and memory bloat. Databases such as Redis, MongoDB, Oracle, and some others explicitly recommend disabling THP or setting it to madvise. Set it persistently on the kernel command line (transparent_hugepage=madvise) or via a systemd unit / tuned profile. Test with your real workload.


8. Memory fragmentation

Over time, free memory can become scattered as many small free chunks. Total free memory may be plentiful, yet a request for a large contiguous block (for a huge page, a big DMA buffer, or a high-order kernel allocation) can fail or stall.

8.1 Inspecting fragmentation

cat /proc/buddyinfo
Node 0, zone   Normal   4210   1820    903    311     87     12      1      0      0      0      0

Each column is the count of free blocks of order 0, 1, 2 ... (block size = 4 KiB × 2^order). Many blocks in the low columns and zeros on the right means memory is fragmented: plenty of free pages, but not in big contiguous runs.

cat /proc/pagetypeinfo                 # per-migration-type breakdown
cat /sys/kernel/debug/extfrag/extfrag_index   # debugfs: how fragmented for each order

8.2 Symptoms

  • Failed or slow huge page allocation (nr_hugepages cannot reach the target; THP rarely used).
  • page allocation failure: order:N messages in dmesg.
  • High system CPU in kcompactd or in direct compaction.
  • Latency spikes when a process needs a large allocation.

8.3 Mitigation

echo 1 > /proc/sys/vm/compact_memory    # trigger manual compaction (can stall briefly)
  • Reserve huge pages at boot, before memory fragments.
  • Raise vm.min_free_kbytes modestly so reclaim starts earlier and keeps some headroom (careful: too high wastes RAM or triggers OOM).
  • Set THP to madvise to avoid constant compaction (section 7.2).
  • Fix the root cause: restart long-running services that fragment memory, or avoid workloads allocating and freeing many odd-sized kernel objects.

See the reading list for detailed articles on how kernel memory fragmentation works and what compaction does.


9. NUMA

On multi-socket servers (and some large single-socket ones), memory is attached to particular CPU sockets: Non-Uniform Memory Access. Accessing remote memory is slower.

numactl --hardware               # nodes, sizes, distances
numastat -m                      # per-node memory usage
numastat -p <pid>                # per-node usage for a process
lscpu | grep -i numa

Control placement:

numactl --cpunodebind=0 --membind=0 ./app      # pin CPU and memory to node 0
numactl --interleave=all ./database            # spread allocations across nodes

Symptoms of NUMA trouble: one node exhausted while another has plenty of free memory, causing swapping or reclaim on a machine that appears to have free RAM. vm.zone_reclaim_mode (default 0) controls whether the kernel prefers reclaiming locally over using remote memory; leave it at 0 unless you have measured a benefit. Modern kernels can also migrate pages automatically (kernel.numa_balancing).


10. Tuning with sysctl

Parameters live in /proc/sys/vm/, readable and writable via sysctl.

sysctl -a | grep '^vm\.'                       # list all
sudo sysctl vm.swappiness=10                   # set temporarily

Persist in a drop-in file, then reload:

sudo tee /etc/sysctl.d/90-memory.conf <<'EOF'
vm.swappiness = 10
EOF
sudo sysctl --system

Method: change one parameter, measure under a realistic load, and keep only what demonstrably helps. There are no universally right values.

Parameter Default What it does Typical tuning
vm.swappiness 60 Preference for swapping anon pages vs reclaiming cache 10 for databases/latency-sensitive; 60+ for cache-heavy file servers
vm.vfs_cache_pressure 100 Tendency to reclaim dentry/inode caches; >100 reclaims faster, <100 keeps them 50 for servers with millions of files; avoid 0
vm.dirty_background_ratio 10 % of memory of dirty data at which background writeback starts 5 or use dirty_background_bytes on large-RAM systems
vm.dirty_ratio 20 % of dirty data at which writers are forced to write synchronously 10-15; use dirty_bytes on large-RAM systems (percentages get huge)
vm.dirty_expire_centisecs 3000 Age at which dirty data must be written Lower for quicker durability
vm.dirty_writeback_centisecs 500 How often flusher threads wake Rarely changed
vm.min_free_kbytes auto Free memory floor kept for the kernel/atomic allocations Raise slightly for network-heavy or fragmented systems
vm.watermark_scale_factor 10 Gap between watermarks; higher wakes kswapd earlier Raise (e.g., 100-200) to reduce direct reclaim stalls
vm.overcommit_memory / _ratio 0 / 50 Overcommit policy (section 5) 1 for Redis-like fork workloads; 2 for strict environments
vm.max_map_count 65530 Max memory-mapped regions per process Raise to 262144 for Elasticsearch/OpenSearch and similar
vm.nr_hugepages 0 Reserved 2 MiB huge pages See section 7
vm.panic_on_oom / oom_kill_allocating_task 0 / 0 OOM behavior Cluster-specific
vm.zone_reclaim_mode 0 NUMA local reclaim preference Leave at 0
vm.compact_memory write-only Trigger compaction Ad-hoc
vm.drop_caches write-only Drop caches (section 14) Diagnostics only

Tip: on machines with very large RAM, dirty ratios translate into enormous byte amounts (20% of 256 GiB is over 50 GiB), so writers can stall for a long time when the limit is finally hit. Prefer vm.dirty_background_bytes and vm.dirty_bytes, which override the ratios when set.


11. Limiting memory with cgroups and systemd

Control groups cap and account memory per group of processes. They underpin containers (Docker, Podman, Kubernetes) and systemd services.

stat -fc %T /sys/fs/cgroup       # "cgroup2fs" = unified v2 hierarchy (default on current distros)

11.1 cgroup v2 memory files

File Meaning
memory.current Current usage
memory.min / memory.low Protected memory (guaranteed / best-effort)
memory.high Soft limit: exceeding it throttles the group with reclaim (no kill)
memory.max Hard limit: exceeding it triggers the cgroup OOM killer
memory.swap.max Swap allowance (0 disables swap for this group)
memory.stat Detailed breakdown (anon, file, slab, ...)
memory.events Counters for high, max, oom, oom_kill events
memory.pressure PSI for the group

11.2 systemd

# in a service unit or drop-in (systemctl edit myservice)
[Service]
MemoryHigh=1500M
MemoryMax=2G
MemorySwapMax=0
OOMScoreAdjust=200

Ad-hoc test:

systemd-run --scope -p MemoryMax=512M -p MemorySwapMax=0 ./memory-hungry-program
systemctl status myservice            # shows current memory and limits
systemd-cgtop -m                      # top-like view of memory per cgroup

11.3 Containers and Kubernetes

  • A container's memory.max comes from docker run --memory= or the Kubernetes resources.limits.memory. Going over it triggers an OOMKilled (exit code 137) even if the node has RAM to spare.
  • Kubernetes requests.memory influences scheduling and OOM-kill priority (QoS classes: Guaranteed, Burstable, BestEffort), but is not enforced as a hard cap.
  • Page cache counts toward the cgroup's usage but is reclaimed under pressure, so memory.current can look near the limit without being in danger. Look at memory.stat and the working set (working_set_bytes in metrics), not just raw usage.

12. Inspecting hardware memory

12.1 Maximum supported memory

The motherboard's limit comes from the firmware's DMI/SMBIOS tables:

sudo dmidecode -t 16
Physical Memory Array
        Location: System Board Or Motherboard
        Use: System Memory
        Error Correction Type: None
        Maximum Capacity: 32 GB
        Number Of Devices: 4

Maximum Capacity is the total the board supports and Number Of Devices is the number of DIMM slots. Error Correction Type shows whether ECC is supported.

12.2 Memory installed

sudo lshw -short -C memory
H/W path       Device   Class    Description
=============================================
/0/1                    memory   64KiB BIOS
/0/26                   memory   256KiB L1 cache
/0/28                   memory   1MiB L2 cache
/0/29                   memory   6MiB L3 cache
/0/2a                   memory   20GiB System Memory
/0/2a/0                 memory   2GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/1                 memory   2GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/2                 memory   8GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)
/0/2a/3                 memory   8GiB DIMM DDR3 Synchronous 1333 MHz (0.8 ns)

More detail per slot (size, speed, type, manufacturer, part number, which slots are empty):

sudo dmidecode -t memory
sudo dmidecode -t 17          # memory devices only

Compare the mix of DIMMs against what the board supports; mismatched sizes or slot layouts can cause reduced channel interleaving.

12.3 ECC errors

On ECC systems, the kernel's EDAC subsystem reports corrected and uncorrected errors:

edac-util -v
ras-mc-ctl --summary
ras-mc-ctl --errors
dmesg | grep -i -E 'edac|mce|ecc'

Corrected errors that keep repeating on one DIMM indicate a module heading for failure.

12.4 Quick facts

free -h
lsmem                       # memory blocks and online state
cat /proc/meminfo | head -3
getconf PAGESIZE

13. Testing memory (memtest86+)

Faulty RAM causes random crashes, segfaults, and data corruption. memtest86+ boots outside the OS and hammers every address with test patterns.

13.1 Installing on RHEL-family systems (BIOS boot)

sudo yum install memtest86+ memtest-setup      # dnf on newer releases
sudo memtest-setup                             # adds a GRUB entry
sudo grub2-mkconfig -o /boot/grub2/grub.cfg    # regenerate the GRUB config

On UEFI systems the GRUB config lives elsewhere (/boot/efi/EFI/<distro>/grub.cfg). Many current packages of memtest86+ (version 6+) support UEFI directly.

13.2 Debian/Ubuntu

sudo apt install memtest86+
sudo update-grub

Reboot and choose Memtest86+ from the GRUB menu. If the OS package is awkward, boot the official memtest86+ image from a USB stick instead.

13.3 Practical advice

  • Run at least one full pass; for suspected problems, run several passes or overnight. Some faults show up only after hours.
  • Any error is a failure. Reseat DIMMs, then test each module alone in a known-good slot to identify the bad one.
  • If the machine can't be taken offline, memtester (user-space, memtester 1G 3) can test the portion of RAM it can allocate, though it cannot cover memory used by the kernel or other programs.

14. Clearing caches

The kernel can be asked to discard clean cache. This is a diagnostic or benchmarking tool, not a routine "memory cleanup". Dropping caches forces the system to re-read data from disk afterward, so performance is worse until the cache warms up again. Cache is normally released automatically whenever programs need the memory.

# Flush dirty pages to disk first, so more cache is clean and droppable
sync

# Drop the page cache only
echo 1 | sudo tee /proc/sys/vm/drop_caches

# Drop reclaimable slab objects (dentries and inodes)
echo 2 | sudo tee /proc/sys/vm/drop_caches

# Drop both page cache and slab objects
echo 3 | sudo tee /proc/sys/vm/drop_caches

(sudo echo 3 > /proc/... fails because the redirection is done by your shell, not by sudo; use tee or a root shell.)

Legitimate uses: benchmarking cold-cache performance, reproducing a problem that depends on a cold cache, or measuring how much of a process's I/O really hits the disk. Illegitimate use: putting it in cron to make free look nicer. That hurts performance and hides real problems.

To free swapped-out pages, cycle swap instead: swapoff -a && swapon -a (only when there is enough free RAM to absorb them).


15. Finding and debugging memory leaks

A leak is memory a program allocates but never releases, so usage grows over time. First confirm the leak is real, then find where.

15.1 Confirming growth

# Track a process over time
while sleep 60; do ps -o pid,rss,vsz,cmd -p <pid>; done
pidstat -r -p <pid> 60              # sysstat
watch -d -n1 'free -m; grep -i commit /proc/meminfo'
  • RSS steadily rising: a real shortage is developing (resident leak).
  • VSZ/Committed_AS rising, RSS flat: virtual-only leak (allocated, never touched); less harmful but still worth reporting.
  • Rising then plateauing: probably a cache, not a leak.
  • Also consider fragmentation and allocator behavior: some allocators (glibc malloc) retain freed memory in arenas and never return it to the OS, which looks like a leak. Setting MALLOC_ARENA_MAX=2 or using jemalloc/tcmalloc can change the picture.

15.2 valgrind

valgrind --tool=memcheck ./program [program options]
valgrind --tool=memcheck /usr/bin/httpd                  # any binary (note: very slow, a 10-50x slowdown)

# Detailed leak report with allocation stack traces
valgrind --tool=memcheck --leak-check=full --show-leak-kinds=all --track-origins=yes ./program

Leak categories in the summary: definitely lost (real leaks), indirectly lost (reachable only through a lost block), possibly lost, and still reachable (memory alive at exit, usually harmless). Compile with -g (and without heavy optimization) to get file and line numbers.

Related valgrind tools:

valgrind --tool=massif ./program         # heap profiler: who allocates memory over time
ms_print massif.out.<pid>                # view the profile

15.3 Other tools

Tool Use
AddressSanitizer / LeakSanitizer (-fsanitize=address) Compile-time instrumentation, much faster than valgrind; reports leaks at exit
heaptrack Low-overhead heap profiler with a GUI
memleak-bpfcc (BCC/eBPF) Attach to a running process and list outstanding allocations by stack: sudo memleak-bpfcc -p <pid>
pmap -x, /proc/<pid>/smaps See which mappings (heap, anon, libraries) grew
jemalloc/tcmalloc heap profiling Production-friendly profiling for C/C++ services
Language runtime tools Java (jcmd GC.heap_dump, MAT), Go (pprof heap), Python (tracemalloc, objgraph), Node (--inspect heap snapshots)

15.4 What to do next

A leak found in third-party software goes to its maintainers, with the valgrind or profiler output, the version, and steps to reproduce. Fixing it requires changing the source, rebuilding, and re-testing. Until then, contain it: restart on a schedule, or set a cgroup memory limit (MemoryMax=) and let systemd restart the service, so a leak becomes a controlled restart instead of taking down the host.


16. Field notes: leak lab exercise

This exercise makes the two kinds of leak (section 5.2) visible. The original notes used a training tool called bigmem; here is an equivalent you can build anywhere.

16.1 The test program

/* bigmem.c: allocate N MiB and either touch it (resident) or not (virtual only) */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>

int main(int argc, char **argv) {
    int virtual_only = 0, arg = 1;
    if (argc > 1 && strcmp(argv[1], "-v") == 0) { virtual_only = 1; arg++; }
    if (arg >= argc) { fprintf(stderr, "usage: %s [-v] MiB\n", argv[0]); return 1; }

    size_t bytes = (size_t)atol(argv[arg]) * 1024 * 1024;
    char *p = malloc(bytes);
    if (!p) { perror("malloc"); return 1; }
    if (!virtual_only) memset(p, 1, bytes);      /* touching the pages makes them resident */

    printf("allocated %zu MiB (%s); sleeping, Ctrl-C to exit\n",
           bytes >> 20, virtual_only ? "virtual only" : "resident");
    pause();                                      /* never free(p): this is the "leak" */
    return 0;
}
gcc -g -o bigmem bigmem.c

16.2 Step 1: run under valgrind

valgrind --tool=memcheck ./bigmem 256        # resident: allocates and touches 256 MiB
valgrind --tool=memcheck ./bigmem -v 256     # virtual only: allocates 256 MiB but never touches it

Valgrind flags a leak in both cases. It cannot tell you which kind of leak it is.

16.3 Step 2: watch the system from a second terminal

watch -d -n1 'free -m; grep -i commit /proc/meminfo'

This shows physical memory and the committed (promised) totals side by side.

16.4 Step 3: run without valgrind and compare

./bigmem 256          # resident
./bigmem -v 256       # virtual only

Expected observations:

Run Committed_AS used / RSS
./bigmem 256 +256 MiB +256 MiB (physical memory really consumed)
./bigmem -v 256 +256 MiB Only a few MiB (program image, page tables); the untouched pages never became resident

Committed memory rises by the requested amount in both cases; only the resident case consumes real RAM. That distinction is exactly what separates "a leak worth reporting" from "a leak that will take the server down".


17. Troubleshooting workflow

"The server is slow." Work through this list from top to bottom:

  1. Is it swapping? vmstat 1: sustained si/so values, high wa. If yes, memory is short (or something is bloated).
  2. How much is truly available? free -h: look at available, not free. Check cat /proc/pressure/memory for stalls.
  3. Who is using it? ps aux --sort=-rss | head, smem -tk -s pss, systemd-cgtop -m. Include tmpfs (df -h /dev/shm /tmp), since tmpfs files consume RAM and swap.
  4. Kernel memory? slabtop, and compare Slab/SUnreclaim and PageTables in /proc/meminfo. A growing slab points to a kernel or driver leak (or a huge dentry cache).
  5. Is anything being OOM-killed? dmesg -T | grep -i oom; journalctl -k -b | grep -i 'killed process'.
  6. Is the cache being squeezed? High bi in vmstat with low cache and rising disk latency means the hot working set no longer fits.
  7. Fragmentation? cat /proc/buddyinfo; failed huge-page allocations; kcompactd CPU use.
  8. NUMA imbalance? numastat -m shows one node exhausted.
  9. Limits reached? Container/cgroup memory.events and memory.max; overcommit (Committed_AS vs CommitLimit in strict mode).
  10. Hardware? ECC counters in dmesg/ras-mc-ctl; run memtest86+ if crashes look random.
Symptom Likely cause First action
free shows almost no free memory Normal page-cache usage Check available instead
Swap in use but system fast Cold pages parked in swap Nothing; only worry if si/so are active
Constant swap in/out, high wa Working set exceeds RAM Find the memory hog, add RAM, add limits, tune swappiness
Process killed unexpectedly OOM killer or cgroup limit dmesg, check memory.events and limits
RSS grows without bound Leak valgrind, heaptrack, memleak-bpfcc
Periodic latency stalls THP compaction, direct reclaim, or dirty-page throttling Set THP to madvise, raise watermark_scale_factor, use dirty_bytes
page allocation failure: order:N Fragmentation Compaction; reserve huge pages at boot; raise min_free_kbytes
Container OOMKilled (137) with host RAM to spare cgroup limit too low Raise the limit; check working set vs limit
Random crashes or corruption Bad RAM memtest86+, check ECC logs

18. Cheat sheet

Look

free -h                                  # summary (watch "available")
vmstat 1                                 # swap, IO, CPU activity
cat /proc/meminfo                        # everything
cat /proc/pressure/memory                # PSI
ps aux --sort=-rss | head                # top memory users
smem -tk -s pss                          # shared-memory-aware accounting
slabtop -o                               # kernel caches
swapon --show                            # swap devices
cat /proc/buddyinfo                      # fragmentation
numastat -m                              # NUMA usage
dmesg -T | grep -i oom                   # OOM events

Hardware

sudo dmidecode -t 16                     # max capacity and slot count
sudo dmidecode -t 17                     # each DIMM
sudo lshw -short -C memory               # installed memory summary
getconf PAGESIZE                         # base page size

Tune

sudo sysctl vm.swappiness=10
sudo sysctl vm.vfs_cache_pressure=50
sudo sysctl vm.dirty_background_ratio=5 vm.dirty_ratio=15
sudo sysctl vm.nr_hugepages=1024
echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
sudo tee /etc/sysctl.d/90-memory.conf <<< 'vm.swappiness = 10' && sudo sysctl --system

Swap

fallocate -l 4G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
swapoff /swapfile
swapon --show

Limit

systemd-run --scope -p MemoryMax=512M -p MemorySwapMax=0 ./cmd
echo -500 | sudo tee /proc/<pid>/oom_score_adj

Debug leaks

valgrind --tool=memcheck --leak-check=full --show-leak-kinds=all ./prog
valgrind --tool=massif ./prog && ms_print massif.out.*
sudo memleak-bpfcc -p <pid>
watch -d -n1 'free -m; grep -i commit /proc/meminfo'

Caches (diagnostics only)

sync; echo 3 | sudo tee /proc/sys/vm/drop_caches

19. Further reading