A production diagnosis workflow — from noticing high load to pinpointing the offending process with pidstat and strace, then deciding the correct fix.

What is Linux CPU High Usage?

When a Linux server reports high CPU, the number itself is only the starting point. top shows overall CPU split across user, system, iowait, and steal time — but that aggregate hides which process is responsible and why. A server burning 95% CPU because of a runaway PHP worker requires a completely different fix than one showing 80% iowait from a backup job hammering the disk. The diagnosis workflow matters as much as the tools. ps aux --sort=-%cpu gives a one-shot ranked snapshot, but it samples a single moment and misses bursty processes. pidstat -u 1 5 samples every second for five seconds and averages across the window, making it far more reliable for short-lived CPU spikes. Once you have the PID, strace -p <pid> reveals whether the process is stuck in a tight syscall loop — the most common cause of genuine 100% CPU on a single core. This page walks through the full three-step diagnosis: identify, profile, then fix.

Tools and Commands

# Live CPU overview — press 1 to show per-core breakdown
top

# Snapshot sorted by CPU descending
ps aux --sort=-%cpu | head -20

# Sample CPU per process every 1 second for 5 iterations
pidstat -u 1 5

# Sample a specific PID
pidstat -u -p 14823 1 10

# Trace syscalls on a running process (Ctrl+C to stop)
strace -c -p 14823

# Show per-thread CPU for multi-threaded process
pidstat -t -p 14823 1 5

Use top first for the overall picture — the %wa (iowait) and %st (steal) columns immediately tell you whether the bottleneck is CPU-bound or disk/VM-bound. Switch to pidstat for reliable process-level sampling. The -u flag shows CPU; -t expands to threads. strace -c -p accumulates syscall counts and time over its run — a process spending 90% of its time in futex or epoll_wait is waiting, not computing, and the CPU reading from top is misleading.

Key Parameters

Flag / ParameterDescriptionSecurity Note
pidstat -uReport CPU utilisation per process. Shows %usr, %system, %CPU, and CPU affinity.Use with an interval of 1 and count of 5-10 for a meaningful average rather than a point-in-time sample.
pidstat -tExpand to per-thread reporting. Essential for multi-threaded apps where one thread spins.Combine with -p <pid> to avoid flooding output on busy servers with hundreds of processes.
strace -c -pAttach to a running PID and count syscalls by type. Outputs a summary table on exit.strace adds overhead (10-30% slowdown on the traced process). Use -c (count mode) rather than -e trace=all to minimise impact.
top -d 1Set refresh interval to 1 second. Default is 3s which misses short bursts.Press '1' inside top to expand per-core CPU. Steal time (%st) above 5% on a VM indicates CPU contention at the hypervisor level — a hosting-layer problem, not a process-level one.
ps --sort=-%cpuSort ps output by CPU descending. The % shown is averaged since process start, not current.Do not use ps CPU% alone to make kill/renice decisions — a process that ran 100% CPU for 1 second an hour ago will still show an elevated average. Use pidstat for current reality.
renice -n 10 -pLower scheduling priority of a process (nice value 10 = lower priority, 19 = lowest).Renice does not reduce CPU consumption — it reduces scheduling priority. A process at nice 19 will still consume 100% CPU when no other process competes. Use cpulimit or cgroups to cap absolute usage.

Diagnosis and Fix Workflows

Runaway process consuming a full core

You see load average above the core count and top shows a single process at 99% CPU. The first step is to confirm it is genuinely computing and not just blocked in a syscall. Attach strace: if the output shows a rapid repeat of a single syscall like read or write in a tight loop, the process is stuck. Check if it is a worker that can be killed and respawned, or if it requires a code fix.

strace -c -p $(pgrep -f suspicious_process) &
sleep 10
kill %1
# Review the syscall summary table — look for a single call dominating >90% of time

Bursty CPU spike from a scheduled job

A cron job running every 5 minutes can drive CPU to 100% for 30 seconds and then disappear. ps snapshots will miss it. Use pidstat -u 1 60 to watch for 60 seconds across cron fire times. When the spike appears, the PID will show up in pidstat output. Cross-reference with grep CRON /var/log/syslog to correlate timing.

pidstat -u 1 60 | tee /tmp/cpu_sample.txt
grep CRON /var/log/syslog | tail -20

High iowait masquerading as CPU load

A server showing load average 8.5 on 4 cores with top reporting 75% iowait is not CPU-bound — it is I/O-bound. The processes are blocked waiting for disk, not computing. Adding more CPU does nothing. Use iostat -x 1 5 to confirm disk saturation, then identify the I/O-generating process with iotop -o. The fix is at the disk or database layer.

# Confirm iowait is the culprit
iostat -x 1 5
# Find the I/O-generating process
iotop -o -n 5

Cross-referencing pidstat with strace to find a syscall loop

When pidstat shows a process at 85% CPU continuously and the process is not doing obvious heavy computation, run strace -p <pid> without the -c flag for 2 seconds to see the raw syscall stream. A PHP or Python process stuck in a tight select() or read() loop on a socket will be immediately visible. This is often a network timeout misconfiguration causing a process to spin on a non-responding upstream.

# Watch raw syscalls for 2 seconds then interrupt
timeout 2 strace -p 14823 2>&1 | tail -30

Performance Impact: CPU Saturation and Cascading Failure

Sustained CPU saturation above 85% on all cores is not just a performance problem — it is a stability risk. The Linux scheduler must queue runnable processes in the run queue (the r column in vmstat). When the run queue exceeds 2-3x the core count, request latency spikes non-linearly: a server that handles requests in 80ms at 60% CPU may respond in 800ms at 95% CPU as processes wait for scheduler turns. Web servers hit connection timeouts before processes get scheduled. On PHP-FPM or Apache prefork setups, workers pile up as existing requests take longer — filling the pool and triggering 503s for new connections. The iowait dimension compounds this: a MySQL query that takes 5x longer because the disk is saturated holds a PHP worker for that entire duration, draining the worker pool for non-database requests too. CPU steal time above 5% on a VPS means the hypervisor is throttling the VM — no amount of application tuning fixes this; the server needs to be moved or upgraded. Identifying whether the bottleneck is user-CPU, iowait, or steal before applying any fix is the most important step.

  • Run queue exceeding 2x core count causes non-linear latency increases — diagnose immediately rather than waiting for full saturation.
  • CPU steal time above 5% on a VM cannot be fixed at the application layer — it requires a VPS upgrade or migration.
  • High iowait is not a CPU problem — adding vCPUs or renicing processes has no effect when the bottleneck is disk I/O.
  • strace attaches to a live process and adds 10-30% overhead — never leave it running on a production process for more than 30 seconds.
  • ps CPU% is averaged from process start — a process that spiked once will show falsely elevated CPU% for hours.

Practical Examples

Sample CPU per process over 5 seconds with pidstat

pidstat -u 1 5

Samples every second for 5 iterations and prints a final average. The %CPU column shows the real current utilisation — far more reliable than ps for identifying bursty processes. Look for any process consistently above 20% in the average row.

Cross-reference pidstat with strace to find a syscall loop

# Step 1: identify the PID
pidstat -u 1 5 | sort -k8 -rn | head -5

# Step 2: count syscalls on that PID for 10 seconds
strace -c -p 14823 &
sleep 10
kill %1

The strace summary shows which syscall consumes the most time. A process spending 90% of its strace time in futex is waiting on a lock. One spending 90% in read on a non-blocking socket is in a busy-wait loop — the application logic needs fixing.

Per-core CPU breakdown to find uneven load

# Inside top: press '1' to expand per-core
top -d 1
# Or with mpstat:
mpstat -P ALL 1 5

On multi-core servers, overall CPU% can look acceptable while one core is pegged at 100% — single-threaded processes or processes with CPU affinity can cause this. mpstat -P ALL shows each core separately, revealing the imbalance.

Troubleshooting Common Issues

Problem: top shows high CPU but ps shows no process above 5%

Solution: The offending process is spawning and dying faster than ps samples. Run pidstat -u 1 30 for a 30-second window to catch it. Also check if kernel threads are responsible: in top, press H to show threads separately — kernel threads appear with names in brackets like [kworker/0:1].

Problem: CPU is high but strace shows mostly epoll_wait or select

Solution: epoll_wait and select are blocking calls — a process spending most of its time there is actually idle, waiting for I/O or network events. The high CPU reading is a top display artifact if the refresh caught the process mid-compute. Cross-check with pidstat -u 1 10 over a longer window to get a reliable average.

Problem: High CPU disappears by the time diagnosis tools attach

Solution: Use sar -u 1 60 (from sysstat package) to record CPU metrics continuously. Review historical data with sar -u -f /var/log/sysstat/saXX. Install sysstat with apt install sysstat and ensure it is enabled to record metrics every 10 minutes via its systemd timer.

Summary

Start with top to separate user CPU, iowait, and steal — these three tell you which layer to investigate. Use pidstat -u 1 5 instead of ps for reliable process-level data. Once you have a PID, strace -c -p shows whether the process is computing or stuck in a syscall loop. Target: identify the root cause before applying any fix. Renicing without knowing the cause wastes time; killing without understanding may just respawn the same problem.

  • Use pidstat -u 1 5 rather than ps aux — it samples over time and gives reliable averages rather than a single potentially misleading snapshot.
  • Check %wa (iowait) and %st (steal) in top before diagnosing CPU — both can drive high load average without any process consuming real CPU cycles.
  • Run strace -c -p <pid> for 10 seconds to identify whether a high-CPU process is computing or stuck in a repeated syscall loop.

Is Your Server Running at Full Performance?

INTRAM manages Linux servers with performance tuning built in from day one — correct MySQL configuration, PHP stack selection, nginx or Apache optimisation, and continuous monitoring so slowdowns are caught before users notice.

Explore Managed Hosting

Let’s assess what your business actually needs.

We will use these details only to understand your request and reply appropriately.