A practical reference for diagnosing and fixing performance problems across every layer of the Linux web server stack — from CPU and memory to MySQL query execution, PHP-FPM pool sizing, nginx caching, Redis object cache, and kernel-level TCP and I/O tuning.
A server that responds in 200ms converts visitors at measurably higher rates than one that responds in 2000ms. Google’s Core Web Vitals directly penalise slow time-to-first-byte in search rankings. A 1-second improvement in page load time reduces bounce rate by an average of 11% across e-commerce sites. Server performance is not an infrastructure concern — it is a business outcome. This guide covers the complete Linux web server performance stack: CPU diagnosis and process scheduling, load average interpretation and vmstat analysis, MySQL query optimisation and InnoDB configuration, Apache and PHP handler selection, PHP-FPM pool tuning, nginx worker and caching configuration, Redis object cache setup and performance tuning, disk I/O analysis and I/O scheduler selection, memory and swap management, load testing with ApacheBench and k6, and Linux TCP and kernel parameter tuning. Each topic page provides the exact commands used on real production servers, the parameters that matter and why, and the diagnosis workflow to follow when something goes wrong.
CPU and Load Average
High CPU usage on a web server is rarely caused by a single run-away process — it is almost always caused by too many concurrent PHP workers all executing expensive operations simultaneously. Before adding CPU resources, diagnose what is consuming cycles and whether the bottleneck is compute, I/O wait, or a MySQL query holding a PHP worker open.
top, ps aux, perf, cpufreq
Identify which processes consume CPU, distinguish between user-space compute and iowait, and verify whether CPU frequency scaling is throttling the server under load.
View guide →Load Average and System Resource Monitoring
Load average numbers above the CPU count indicate a resource queue: more processes are ready to run than CPUs available to run them. Interpreting load average correctly — distinguishing CPU queue from I/O wait — determines whether the fix is more CPUs, faster disk, or reduced PHP worker count.
uptime, top, sar, iostat
Read and interpret 1/5/15-minute load averages correctly, identify whether high load is CPU-bound or I/O-bound, and use sar for historical load trend analysis.
View guide →vmstat, sar, iostat
Use vmstat to diagnose CPU utilisation split (user/system/idle/iowait), memory pressure, swap activity, and disk I/O rates in a single command.
View guide →MySQL Performance
MySQL is the most common bottleneck in WordPress and PHP application stacks. The three causes of slow MySQL performance — missing indexes on large tables, slow query log disabled so problems go undetected, and InnoDB buffer pool too small to hold the working set in RAM — each have specific diagnostic commands and fixes that do not require schema changes.
mysqldumpslow, slow_query_log, EXPLAIN
Enable the slow query log, identify the queries consuming the most cumulative execution time with mysqldumpslow, and use EXPLAIN to diagnose the execution plan for each slow query.
View guide →EXPLAIN, SHOW INDEX, ALTER TABLE
Add the missing indexes that slow query analysis identifies: composite indexes for multi-column WHERE clauses, covering indexes that prevent full table reads, and how to verify index usage with EXPLAIN.
View guide →innodb_buffer_pool_size, SHOW ENGINE INNODB STATUS
Size the InnoDB buffer pool correctly for the working data set, interpret the SHOW ENGINE INNODB STATUS buffer pool hit ratio, and configure innodb_io_capacity for the storage device.
View guide →Apache and PHP Handler Selection
How Apache handles PHP execution — mod_php vs PHP-FPM via mod_proxy_fcgi — determines server memory consumption, worker concurrency, and PHP version flexibility. Most servers that struggle with RAM pressure under moderate load are running the default mod_php with the prefork MPM, which is the least efficient combination for WordPress.
apache2ctl -M, mpm_event, mod_proxy_fcgi
Compare mod_php with the prefork MPM against PHP-FPM with mpm_event, understand why the difference in memory per connection is 3-5x, and how to migrate without downtime.
View guide →apache2ctl, MaxRequestWorkers, prefork
Calculate the correct MaxRequestWorkers value for the prefork MPM based on available RAM, diagnose the symptoms of an over-provisioned or under-provisioned worker pool.
View guide →php -i, opcache, OPcache statistics
Configure OPcache to eliminate PHP script compilation overhead, verify cache hit rate with opcache_get_status(), and size the cache memory for the application’s file count.
View guide →a2enmod, mpm_event, proxy_fcgi
Decision guide for selecting the PHP handler based on server RAM, expected concurrency, and whether multiple PHP versions are required — with migration steps for each scenario.
View guide →PHP-FPM Performance
PHP-FPM pool configuration directly controls how many concurrent PHP requests the server can handle. Too few workers causes request queueing under load; too many workers exhausts RAM and causes Linux to swap, which is far worse. The slow log reveals which specific PHP functions and scripts account for the most execution time.
php-fpm, pm.max_children, pm.start_servers
Calculate the correct pm.max_children based on available RAM per worker, choose between static and dynamic process management, and monitor pool utilisation with the status page.
View guide →request_slowlog_timeout, php-fpm slow log
Enable the PHP-FPM slow log to capture stack traces for requests exceeding a threshold, identify the specific PHP functions causing slow requests, and correlate with MySQL slow queries.
View guide →nginx Performance
nginx handles far more concurrent connections than Apache with the same RAM because it does not allocate a thread or process per connection. The configuration parameters that matter most for performance are worker count, FastCGI caching (which eliminates PHP execution for anonymous requests entirely), and gzip compression (which reduces bandwidth and response time for text-based content).
worker_processes, worker_connections, nginx.conf
Set worker_processes to match CPU core count, size worker_connections correctly, and configure the event model for maximum connection efficiency on Linux.
View guide →fastcgi_cache, proxy_cache, nginx caching
Configure nginx FastCGI cache to serve WordPress pages without calling PHP, set correct cache bypass rules for logged-in users and POST requests, and verify cache hit rate.
View guide →gzip, gzip_types, gzip_comp_level
Enable gzip for CSS, JavaScript, and HTML responses, select the correct compression level (trade-off between CPU cost and bandwidth saving), and verify compression is applied.
View guide →Disk I/O Performance
Disk I/O is the slowest component in the web server stack after network latency. High iowait in top or vmstat indicates processes are blocked waiting for disk reads or writes — typically MySQL reading uncached data, PHP writing session files, or log rotation causing fsync bursts.
iostat, iotop, fio, hdparm
Use iostat to identify which block devices are saturated, iotop to find which processes are causing the I/O, and fio to benchmark raw disk throughput to distinguish hardware limits from configuration problems.
View guide →I/O scheduler, elevator, udev
Select the correct I/O scheduler for SSD vs spinning disk storage, apply it persistently via udev rules, and verify the improvement with before/after iostat measurements.
View guide →Memory and Swap Management
A server that is swapping is performing at a fraction of its capacity — a MySQL buffer pool page read from NVMe takes 0.1ms, from swap it takes 10ms. Memory management on a web server is about keeping the MySQL buffer pool, PHP-FPM workers, and Redis cache in RAM and out of swap by sizing each component correctly relative to total available memory.
valgrind, smem, /proc/<pid>/maps, pmap
Identify PHP-FPM workers, MySQL, or other processes consuming RAM beyond their expected footprint — the first step before deciding whether to add RAM or tune the configuration.
View guide →dmesg, /var/log/syslog, OOM killer
Read and interpret OOM killer log entries to understand which processes are being killed, why memory is exhausted, and how to prevent recurrence without simply adding RAM.
View guide →swapon, free, vmstat, swappiness
Monitor swap usage in real time, identify what is using swap and why, reduce vm.swappiness to keep production workloads in RAM, and evaluate whether swap size is appropriate.
View guide →Redis Caching
Redis object cache eliminates redundant MySQL queries by storing the results of database lookups in RAM. For a typical WordPress site, Redis reduces MySQL queries per page from 60-100 to 8-15 — a 80-90% reduction that directly improves PHP execution time, server capacity, and TTFB.
redis-server, redis-cli, wp redis enable
Install Redis, configure maxmemory and the allkeys-lru eviction policy, set up a Unix socket for lower latency, and connect WordPress with the Redis Object Cache plugin.
View guide →redis-cli SLOWLOG, CONFIG SET, activedefrag
Diagnose Redis latency with the slow log, fix memory fragmentation with active defrag, tune connection handling for high concurrency, and identify blocking commands with commandstats.
View guide →Load Testing
Load testing before a traffic event is the only way to know whether your server will handle the load — not after the spike causes an outage. ApacheBench provides a fast single-URL throughput baseline; k6 provides scripted multi-step scenarios with pass/fail thresholds that can be integrated into CI/CD pipelines.
ab, ApacheBench
Run a concurrency sweep with ApacheBench to find where requests-per-second peaks and Failed requests starts climbing — this is the server’s actual concurrent capacity limit.
View guide →k6 run, k6 thresholds
Write realistic multi-step load test scripts with think time, set p95 latency thresholds for CI/CD pass/fail gates, and run staged ramp tests to find capacity degradation points.
View guide →Kernel and TCP Tuning
The Linux kernel defaults are conservative and designed for general workloads. A web server under high concurrent load hits three specific limits: the connection accept backlog (silently drops connections when full), TIME_WAIT socket accumulation (causes port exhaustion under sustained connection rate), and file descriptor limits (causes connection refused errors at moderate concurrency). These are configuration problems, not hardware limits.
sysctl, ss, net.core.somaxconn
Raise the connection accept backlog, enable tcp_tw_reuse to prevent port exhaustion from TIME_WAIT accumulation, and tune socket buffers for high-bandwidth connections.
View guide →sysctl, ulimit, I/O scheduler
Set vm.swappiness to prevent production process memory from being swapped to disk, raise LimitNOFILE for all services via systemd overrides, and select the correct I/O scheduler for SSD storage.
View guide →Why Server Performance Directly Affects Revenue and Search Rankings
Server performance translates directly into business outcomes through three measurable channels. The first is search ranking: Google’s Core Web Vitals include Largest Contentful Paint (LCP), which measures how quickly the main content of a page appears. LCP is partially determined by Time to First Byte (TTFB) — the time from the browser making an HTTP request to receiving the first byte of the response. TTFB is directly set by server performance: a server executing PHP in 50ms produces a TTFB of 50-100ms; a server executing the same PHP in 800ms produces a TTFB of 800-1000ms. Google uses Core Web Vitals as a ranking signal, with pages below the ‘Good’ LCP threshold (2.5s) receiving a measurable ranking penalty. The second channel is conversion rate and bounce rate. Multiple large-scale studies from Google, Deloitte, and independent researchers consistently show a 1-second improvement in page load time reduces mobile bounce rate by 8-11% and increases conversions by 1-3%. For a site generating €100,000 per month, a 2% conversion improvement is €2,000 per month — which pays for a significant server upgrade. The third channel is server capacity and cost. A server handling 30 uncached WordPress requests per second requires 10x the hardware to handle 300 simultaneous users that a server handling 300 requests per second (with proper caching) handles on a single machine. The optimisations in this guide — Redis object cache, nginx FastCGI cache, PHP-FPM pool sizing, MySQL InnoDB tuning — are not minor refinements. Combined, they typically improve WordPress request throughput by 10-50x from a default LAMP stack configuration to an optimised production configuration. The pages in this guide document the specific steps to achieve those improvements on real Linux servers.
Frequently Asked Questions
What is the single most impactful performance change for a WordPress server?
For most WordPress sites on a default LAMP stack, enabling nginx FastCGI cache or a full-page caching plugin delivers the largest improvement — it eliminates PHP and MySQL execution entirely for anonymous page views. A server that handles 30 uncached WordPress requests per second can handle 2,000-5,000 requests per second serving cached pages. If FastCGI cache is already in place, Redis object cache is the next highest-impact change, reducing MySQL load by 80-90% for the PHP requests that do reach the server.
How do I find out why my server is slow without downtime?
Start with uptime for load average, then top or htop to identify the process consuming the resource. Check vmstat 2 5 to distinguish CPU saturation from I/O wait. If iowait is high, use iostat -x 2 5 to find the saturated disk device and iotop to identify which process is causing the I/O. If CPU is the constraint, check whether PHP-FPM workers are at pm.max_children (all busy, new requests queuing) or MySQL slow queries are holding PHP workers open. None of these commands cause downtime — they are read-only diagnostics.
How much RAM do I need for a WordPress server?
The minimum viable allocation for a single-site WordPress server with MySQL, PHP-FPM, nginx, and Redis is 2GB RAM. The breakdown: MySQL innodb_buffer_pool_size 512MB, PHP-FPM 20 workers at 30MB each = 600MB, nginx 20MB, Redis 256MB, OS and overhead 200MB — totalling approximately 1.6GB. Below 2GB, either the MySQL buffer pool is too small (data not cached in RAM, causing disk reads) or the PHP-FPM pool is too small (requests queue under load). For sites with multiple databases or high concurrent traffic, 4GB is a more practical minimum.
What is the difference between load average and CPU usage?
CPU usage (percentage shown in top) measures what fraction of available CPU time is being consumed. Load average measures the number of processes in the run queue — either actively executing or waiting for CPU or I/O. A server with 100% CPU usage and load average of 2 on a 2-core machine is CPU-saturated but not overloaded. A server with 20% CPU usage and load average of 8 on a 4-core machine has something blocking processes from running — usually disk I/O (indicated by high iowait in top). The distinction determines whether you need more CPU capacity or faster I/O.
Does improving server performance affect Google Core Web Vitals scores?
Yes, directly. Time to First Byte (TTFB) contributes to Largest Contentful Paint (LCP), one of the three Core Web Vitals. TTFB is entirely determined by server response time. Improving PHP execution time from 800ms to 80ms (through Redis object cache, OPcache, and PHP-FPM tuning) reduces TTFB by approximately 720ms — which can move an LCP score from ‘Needs Improvement’ to ‘Good’. Google Search Console shows TTFB data under the Core Web Vitals report. Total Blocking Time (TBT) and Cumulative Layout Shift (CLS) are primarily frontend concerns, but a slow TTFB causes browsers to start rendering later, which compresses the time available for frontend optimisations.