vmstat is most powerful as a 1-second sampling tool. This page explains how to read r, b, si/so, bi/bo, and st columns together to tell a complete story about what the server is doing.

What is vmstat?

vmstat (virtual memory statistics) provides a per-second snapshot of the system’s CPU, memory, swap, and block I/O state in a single line of output. Unlike top, it does not refresh a screen — it appends a new line every interval, making it ideal for recording a timeline of server behaviour during a traffic spike or incident. The first line printed is always averaged since boot and should be ignored — all subsequent lines show the actual per-interval delta. The columns that matter most in production diagnosis are: r (runnable processes competing for CPU), b (processes blocked on I/O), si/so (swap in/out — any non-zero value is a warning), bi/bo (block device reads and writes in 1KB units), and st (steal time from the hypervisor on VMs). Reading these columns together rather than in isolation is what makes vmstat genuinely useful. A spike in r with b at zero means CPU contention. A spike in b with low r means disk I/O. Non-zero si/so with rising r means a swap cascade is starting.

Tools and Commands

# Run vmstat every 1 second (first line is since-boot average, ignore it)
vmstat 1

# Capture 30 seconds to a file for later analysis
vmstat 1 30 | tee /tmp/vmstat_spike.txt

# Show memory in megabytes
vmstat -S M 1 10

# Per-CPU statistics with mpstat
mpstat -P ALL 1 5

# Historical CPU data via sar (sysstat package required)
sar -u 1 10

# Review yesterday's sar data
sar -u -f /var/log/sysstat/sa$(date -d yesterday +%d)

Always run with an interval of 1 second. The first output row is the system average since boot — discard it. From the second row onwards, each line represents 1 second of actual activity. mpstat -P ALL adds per-core breakdown. sar from the sysstat package records metrics to disk every 10 minutes by default — invaluable for diagnosing incidents that already happened.

Key Parameters

Flag / ParameterDescriptionSecurity Note
r (run queue)Number of processes waiting for CPU. Values above core count indicate CPU saturation.If r consistently exceeds core count, the server needs CPU capacity added or the process load reduced. Sustained r > 2x core count causes non-linear latency increases.
b (blocked)Number of processes in uninterruptible sleep, typically waiting for I/O. High b with low r means I/O bottleneck.b above 3-5 sustained means disk I/O is saturated. Correlate with the bi/bo columns to confirm direction (reads vs writes) and identify the device with iostat -x.
si / so (swap in/out)Kilobytes per second swapped in (si) and out (so). Any non-zero value on a production server is concerning.si/so non-zero while free memory is available means vm.swappiness is too high. si/so non-zero with near-zero free memory means genuine RAM pressure — adding swap does not fix it.
bi / bo (block I/O)Blocks read (bi) and written (bo) per second from block devices, in 1KB units.Sustained high bo (writes) during a database-heavy period is normal. Sustained high bi (reads) means the database buffer pool is too small and MySQL/PostgreSQL is reading from disk repeatedly.
st (steal)CPU steal time percentage — CPU cycles this VM requested but did not get due to hypervisor limits.st > 0 on any line means hypervisor-level CPU throttling. st > 5% sustained is significant and requires escalating to the hosting provider or upgrading the VM plan.
wa (iowait)Percentage of CPU time spent waiting for I/O to complete. Found in the cpu columns.wa above 20% sustained means the application is waiting on disk faster than the disk can serve it. The fix is database query optimisation (fewer disk reads) or faster storage.

Diagnosis and Fix Workflows

Annotating a vmstat traffic spike: r=8, b=3, si/so both non-zero

During a traffic spike, vmstat output shows r=8 (8 processes competing for CPU), b=3 (3 blocked on I/O), and si/so both at 4/8 KB/s. The combination tells a specific story: the server is genuinely CPU-saturated (r > core count), has 3 processes blocked on I/O simultaneously, and has started swapping. This is early-stage OOM pressure. The si/so swap activity means the kernel has started evicting cold pages under memory pressure. The immediate question is: what is consuming memory? Run free -m and smem -t -k | head -10 to find the memory-hungry process.

# Capture the spike
vmstat 1 30 | tee /tmp/spike.txt
# While it runs, check memory in parallel
watch -n 2 'free -m && echo && smem -t -k | head -10'

Identifying a swap cascade before OOM

A swap cascade starts when si/so first appears in vmstat output while memory is not yet fully exhausted. Catching it early allows action before the OOM killer fires. When you see si/so non-zero for more than 3 consecutive seconds, run smem -t -k --sort=swap to find which process has the most swap usage. If it is a PHP-FPM worker or Apache child, the pool size is too large for available RAM.

# Monitor for swap activity
vmstat 1 | awk '$7>0 || $8>0 {print $0, "SWAP ACTIVE"}'
# Find swap consumer
smem -t -k --sort=swap | head -10

Using sar to diagnose a past incident

A server had high load at 3 AM but by the time you check it at 9 AM, everything looks normal. sar records CPU, memory, and I/O statistics to disk every 10 minutes by default. Review the 3 AM window to see exactly what was happening. The -u flag shows CPU, -b shows I/O, and -r shows memory.

# Review today's data
sar -u
# Review a specific hour (e.g. 03:00)
sar -u -s 03:00:00 -e 04:00:00
# Review memory at 3 AM
sar -r -s 03:00:00 -e 04:00:00

Performance Impact: Missing the Swap Cascade and OOM Risk

The most dangerous pattern in vmstat is a gradual si/so escalation that precedes an OOM killer event. Once the OOM killer fires, it kills a process (often a PHP-FPM worker or MySQL) and the sudden service disruption is visible to users. But the vmstat data collected in the 5-10 minutes before the OOM event would have shown the swap cascade building: si/so starting small, free memory declining, b (blocked) increasing as processes wait for pages to be swapped in. Without continuous monitoring or sar recording, this pre-OOM window is invisible. The second critical pattern is sustained non-zero st (steal) combined with high r. A VM that is being CPU-throttled by the hypervisor while also having a full run queue is doubly constrained — processes queue for CPU slots that the hypervisor is not providing. This combination causes dramatic latency increases and is completely invisible to any application-level metric. Only vmstat’s st column and a comparison between reported CPU usage and actual work done reveals it.

  • Non-zero si/so in vmstat means swap is actively being used — if this appears during peak traffic, OOM killer events are likely imminent.
  • The first line of vmstat output is averaged since boot — it is always wrong for current diagnosis and must be discarded.
  • High b (blocked) with low r (runnable) is an I/O bottleneck, not a CPU bottleneck — do not add CPU capacity to fix it.
  • st (steal) non-zero means the hypervisor is limiting this VM's CPU access — application-level optimisation cannot recover stolen CPU cycles.

Practical Examples

Run vmstat 1 30 during a traffic spike and annotate the output

vmstat 1 30 | tee /tmp/vmstat_spike.txt
# Discard line 1 (since-boot average)
# Look for:
#   r > nproc  => CPU saturation
#   b > 3      => I/O bottleneck
#   si/so > 0  => swap cascade starting
#   st > 0     => VM CPU steal

Save the output to a file so you can analyse it after the incident. The 30-second window captures enough to see whether the pattern is sustained or transient. Look for the correlation between r, b, and si/so columns — they tell the complete story together.

Diagnose a swap cascade with si/so and smem

# Watch for swap activity in vmstat
vmstat 1 | awk '$7>0 || $8>0 {print NR, $0}'

# When triggered, identify the swap consumer
smem -t -k --sort=swap 2>/dev/null | head -10

# Check which process has the highest PSS (proportional memory)
smem -t -k --sort=pss | head -10

smem uses PSS (proportional set size) which accounts for shared memory correctly — more accurate than RSS. Sort by swap to find which process is in swap, then decide whether to restart it, reduce its pool size, or add RAM.

Review past performance with sar

# Install sysstat if not present
apt install sysstat
# Enable recording
systemctl enable sysstat --now

# Review today's CPU data
sar -u
# Review a specific window
sar -u -s 14:00:00 -e 15:00:00
# Review I/O at the same time
sar -b -s 14:00:00 -e 15:00:00

sar stores 30 days of history by default. Use -f /var/log/sysstat/saDD to specify a date (DD = day number). Comparing CPU iowait vs block I/O at the same timestamp confirms whether an I/O spike caused a CPU wait spike.

Troubleshooting Common Issues

Problem: vmstat shows high r consistently but top shows no individual process above 20% CPU

Solution: Many processes each consuming a small CPU percentage can add up. Run ps aux --sort=-%cpu | head -20 and add up the CPU% column. Also check for kernel threads with ps -eo pid,comm,pcpu --sort=-pcpu | head -20 — kernel work like kworker or ksoftirqd can consume CPU that does not appear clearly in top.

Problem: sar command not found or no historical data available

Solution: Install sysstat: apt install sysstat or yum install sysstat. Enable the collection timer: systemctl enable sysstat --now. On Debian/Ubuntu also edit /etc/default/sysstat and set ENABLED=true. Historical data collection begins after this — there is no retroactive data for before installation.

Problem: si/so columns show non-zero values but free memory appears available

Solution: vm.swappiness is too high. With swappiness=60 (default), the kernel proactively moves idle process pages to swap even when RAM is available. Set sysctl vm.swappiness=10 and monitor vmstat — si/so should drop to zero within a few minutes if RAM is genuinely available.

Summary

Run vmstat 1 during incidents and read columns together: r for CPU run queue, b for processes blocked on I/O, si/so for swap in/out pressure, and st for VM hypervisor steal time. Always discard the first output line — it is a system-since-boot average, not a current measurement. Install sysstat and enable it so historical sar data is available for post-incident diagnosis. Any non-zero si/so on a production web server requires immediate investigation.

  • Always run vmstat 1 with an interval — the first line is since-boot average and must be discarded; all useful data starts from line 2.
  • Non-zero si/so columns in vmstat mean active swap use — if free RAM is available, reduce vm.swappiness to 10 via sysctl.
  • Install sysstat (apt install sysstat) and enable it — sar historical data is the only way to diagnose performance incidents that happened hours before you looked.

Is Your Server Running at Full Performance?

INTRAM manages Linux servers with performance tuning built in from day one — correct MySQL configuration, PHP stack selection, nginx or Apache optimisation, and continuous monitoring so slowdowns are caught before users notice.

Explore Managed Hosting

Let’s assess what your business actually needs.

We will use these details only to understand your request and reply appropriately.