Tools for measuring where a Linux system spends its time. Start broad (load, CPU, memory, disk, network) and narrow to a process or device. Network-side tools live in Linux networking.
Reading load and the CPU picture first¶
Most performance investigations start at the CPU, but a busy-looking CPU is frequently a symptom: the processor may just be waiting on disk or network. Before touching any CPU-specific tool, get the rough shape from top and uptime:
- Run queue: every runnable process waits in the run queue for a CPU. Load average counts runnable tasks (plus, on Linux, tasks in uninterruptible I/O wait), so a high load with an idle CPU points at storage, not compute.
- Context switches happen whenever the kernel swaps one process off a CPU for another, and each one costs cache warmth. Interrupts from hardware also cause switches. A rough rule of thumb is that context switches running around ten times the interrupt rate are unremarkable. Far higher ratios suggest too many processes competing for CPU. Compare
csandinin vmstat. - iowait (
wa) is CPU time spent idle while I/O is outstanding. Sustained high values mean look at storage and network, not the processor.
Method¶
- Load and headline numbers: uptime, top, vmstat.
- Which resource: CPU (sysstat
mpstat, turbostat), memory (vmstat, slabtop, numastat), disk (iostat, iotop), network (nethogs and friends). - Which process: top, atop, pmap, strace, perf.
- History: sysstat/sar, pcp, atop logs.
- Benchmark or reproduce: fio, ioping.