After optimising the application stack (PHP-FPM, MySQL, nginx, Redis), the kernel itself can still be a bottleneck in three areas: virtual memory management (swap usage, dirty page flushing), file descriptor limits (which cap how many connections and files the server can have open simultaneously), and I/O scheduling (how disk reads and writes are prioritised when multiple processes compete for disk access). This page covers the specific parameters that matter for web server workloads and how to verify they are applied.

The Three Kernel-Level Performance Bottlenecks for Web Servers

Virtual memory management is the first area. Linux uses swap space when physical RAM is exhausted, but it also uses swap proactively based on the vm.swappiness setting. The default value of 60 means the kernel may begin swapping to disk when physical RAM is only 40% consumed — on a web server, swapping is catastrophic because a MySQL query or PHP-FPM process that has been swapped out pauses for 10-100ms while its pages are loaded back from disk. Setting vm.swappiness to 10 keeps processes in RAM unless the server is genuinely out of memory. The dirty page cache is the second sub-area: Linux buffers disk writes in memory (dirty pages) and flushes them periodically. If the dirty page ratio climbs too high, the kernel forces a synchronous flush that stalls all I/O while pages are written out. This causes brief but severe latency spikes for any process touching files during the flush. vm.dirty_ratio and vm.dirty_background_ratio control when flushing begins. File descriptor limits are the second major area. Every TCP connection, open file, and Unix socket consumes a file descriptor. The default system limit (fs.file-max) is usually high enough, but the per-process limit (set via ulimit or systemd’s LimitNOFILE) defaults to 1024 on many distributions. A single nginx worker handling 10,000 simultaneous connections requires 10,000 file descriptors. PHP-FPM, MySQL, and Redis each need their own file descriptors on top. File descriptor exhaustion causes connection refused errors that are identical in symptom to port exhaustion — but the error in logs reads ‘too many open files’. I/O scheduling is the third area. The Linux I/O scheduler determines the order and merging of disk read/write requests. On SSDs (including NVMe), the best-performance scheduler is none (pass-through) or mq-deadline. On spinning disks, deadline or bfq prevents individual processes from starving others. The default scheduler (cfq or bfq depending on distro) is designed for fairness, not throughput.

Tools and Commands

# Check current sysctl values
sysctl vm.swappiness vm.dirty_ratio vm.dirty_background_ratio
sysctl fs.file-max fs.file-nr

# Check per-process file descriptor limits
ulimit -n  # current session
cat /proc/$(pidof nginx | awk '{print $1}')/limits | grep 'open files'
cat /proc/$(pidof -s php-fpm8.1)/limits | grep 'open files'

# Check current open file descriptor count per process
ls /proc/$(pidof -s php-fpm8.1)/fd | wc -l

# Check system-wide open file count
cat /proc/sys/fs/file-nr
# Output: [open] [unused] [max]
# If open approaches max, file descriptor exhaustion is imminent

# Check I/O scheduler for each block device
cat /sys/block/sda/queue/scheduler
cat /sys/block/nvme0n1/queue/scheduler
# Active scheduler shown in [brackets]

# Apply sysctl changes at runtime
sysctl -w vm.swappiness=10
sysctl -w vm.dirty_ratio=15
sysctl -w vm.dirty_background_ratio=5

# Change I/O scheduler at runtime
echo 'mq-deadline' > /sys/block/sda/queue/scheduler

The /proc filesystem exposes live kernel state — /proc/sys/fs/file-nr shows open file descriptors across the entire system in real time. /proc//limits shows the actual limits applied to a running process, accounting for systemd overrides. /sys/block//queue/scheduler shows the active I/O scheduler with the current selection in brackets.

Key Parameters

Flag / ParameterDescriptionSecurity Note
vm.swappinessKernel tendency to swap process memory pages to disk. Range: 0-100. Default: 60.Set to 10 for web servers. Value 10 means the kernel only swaps when free RAM drops to approximately 10% of total — not proactively at 40% as the default does. Value 0 on older kernels disabled swap entirely; on newer kernels (3.5+) it means swap only when memory is truly exhausted.
vm.dirty_ratioPercentage of total RAM that can be dirty pages before a process writing data is forced to wait for flush.Default: 20%. Set to 15%. At 20%, on a 32GB server, 6.4GB of dirty data accumulates before blocking writes. For a web server with < 1GB of active write traffic, this never triggers. Reduce to 15% to start flushing earlier and avoid the synchronous write stall.
vm.dirty_background_ratioPercentage of RAM that triggers the background writeback daemon to begin flushing dirty pages.Default: 10%. Set to 5%. Background flush starts earlier, keeping dirty pages from accumulating to the dirty_ratio threshold that causes blocking. For high I/O servers, reduce further to 3%.
fs.file-maxSystem-wide maximum number of open file handles.Default is usually several hundred thousand on modern kernels. Set to 2097152 (2 million) proactively. The bottleneck is usually the per-process limit (ulimit), not fs.file-max.
LimitNOFILE (systemd)Per-process open file handle limit for systemd-managed services.Add LimitNOFILE=65536 to the [Service] section of /etc/systemd/system/nginx.service.d/override.conf, php-fpm.service.d/override.conf, mysql.service.d/override.conf, and redis-server.service.d/override.conf. Without this, all services use the session default (typically 1024), which nginx exhausts at ~500 concurrent connections.
I/O schedulerBlock device request queue scheduler. Choices: none, mq-deadline, bfq, kyber.Use none or mq-deadline for SSDs and NVMe drives (internal queuing handles optimisation better than the kernel scheduler). Use bfq for spinning disks on shared servers where multiple processes compete for disk. Set permanently in /etc/udev/rules.d/60-ioscheduler.rules.

Diagnosis and Fix Workflows

Diagnose and fix file descriptor exhaustion across the web stack

File descriptor exhaustion produces ‘too many open files’ errors in nginx, PHP-FPM, and MySQL logs. The fix requires raising limits for each process independently via systemd.

# Step 1: Find which service is hitting the limit
grep -r 'too many open files' /var/log/nginx/ /var/log/php*.log /var/log/mysql/

# Step 2: Check the actual limit per process
for svc in nginx php-fpm8.1 mysqld redis-server; do
  pid=$(pidof -s $svc 2>/dev/null)
  if [ -n "$pid" ]; then
    limit=$(awk '/Max open files/{print $4}' /proc/$pid/limits)
    open=$(ls /proc/$pid/fd 2>/dev/null | wc -l)
    echo "$svc: open=$open limit=$limit"
  fi
done

# Step 3: Raise limits for each service via systemd override
for svc in nginx php8.1-fpm mysql redis-server; do
  mkdir -p /etc/systemd/system/${svc}.service.d/
  cat > /etc/systemd/system/${svc}.service.d/limits.conf << 'EOF'
[Service]
LimitNOFILE=65536
EOF
done

systemctl daemon-reload

# Step 4: Restart services to apply new limits
for svc in nginx php8.1-fpm mysql redis-server; do
  systemctl restart $svc
done

# Step 5: Verify new limits
for svc in nginx php-fpm8.1 mysqld redis-server; do
  pid=$(pidof -s $svc 2>/dev/null)
  [ -n "$pid" ] && awk '/Max open files/{print "'$svc': " $4}' /proc/$pid/limits
done
# All should now show: 65536

Set vm.swappiness and verify the server stops swapping

A web server that is swapping degrades dramatically. The symptoms are high iowait in top, slow response times that improve after reducing load, and MySQL query times that spike randomly.

# Check if the server is currently swapping
swapon --show
free -h
# Look at 'used' in the Swap row

# Check iowait (high iowait during swap activity)
top -bn1 | grep -i 'cpu'
# %wa (iowait) above 5% indicates the server is waiting for disk I/O
# during normal operation — often caused by swap read-in

# Check vmstat for swap activity (columns si=swap in, so=swap out)
vmstat 2 5
# si and so columns non-zero = active swapping

# Reduce swappiness immediately
sysctl -w vm.swappiness=10

# Persist the setting
cat >> /etc/sysctl.d/99-webserver.conf << 'EOF'
vm.swappiness = 10
vm.dirty_ratio = 15
vm.dirty_background_ratio = 5
EOF
sysctl -p /etc/sysctl.d/99-webserver.conf

# If the server is currently heavily swapped, consider swapping off then on
# to force swap contents back into RAM (only if RAM permits):
# swapoff -a && swapon -a
# WARNING: This fails if swap contents exceed available free RAM

# Verify swappiness applied:
sysctl vm.swappiness
# Expected: vm.swappiness = 10

Set the optimal I/O scheduler for SSD and NVMe storage

The wrong I/O scheduler on SSD storage adds unnecessary overhead. SSDs do not benefit from request reordering (they have no rotational latency) and their internal queuing handles parallelism better than the kernel’s CFQ scheduler.

# Check current schedulers for all block devices
for dev in /sys/block/*/queue/scheduler; do
  echo "$(echo $dev | cut -d/ -f4): $(cat $dev)"
done
# Example output:
# sda: [mq-deadline] kyber none
# nvme0n1: [none] mq-deadline

# For SSD (sda) — switch to none or mq-deadline
echo 'none' > /sys/block/sda/queue/scheduler
# Verify:
cat /sys/block/sda/queue/scheduler  # should show [none]

# Make persistent via udev rule (survives reboot)
cat > /etc/udev/rules.d/60-ioscheduler.rules << 'EOF'
# SSD drives: use none scheduler
ACTION=="add|change", KERNEL=="sd[a-z]", ATTR{queue/rotational}=="0", ATTR{queue/scheduler}="none"
# NVMe drives: already use none by default, enforce it
ACTION=="add|change", KERNEL=="nvme[0-9]*", ATTR{queue/scheduler}="none"
# HDD (rotational): use mq-deadline for better fairness
ACTION=="add|change", KERNEL=="sd[a-z]", ATTR{queue/rotational}=="1", ATTR{queue/scheduler}="mq-deadline"
EOF

# Reload udev rules
udevadm control --reload-rules && udevadm trigger

Performance Impact: Kernel Parameters as a Force Multiplier for Application Tuning

Kernel tuning does not replace application-level optimisation — it removes the ceiling that kernel defaults impose on an otherwise well-tuned application stack. A PHP-FPM pool set to 200 workers but running with LimitNOFILE=1024 will start rejecting connections once it has 1024 file handles open (100 workers × ~10 FDs each). Raising LimitNOFILE to 65536 removes that artificial ceiling and lets the 200-worker pool operate at full capacity. Similarly, a MySQL server tuned with innodb_buffer_pool_size=8G on a 16GB server will start performing erratically if vm.swappiness=60 causes the kernel to swap out MySQL’s buffer pool pages while RAM is still 60% free — MySQL then re-reads its working set from disk repeatedly, making its own buffer pool optimisation irrelevant. The I/O scheduler change from bfq or cfq to none on NVMe storage is usually invisible in synthetic benchmarks (which run sequential I/O that any scheduler handles well) but becomes significant under the mixed, concurrent I/O pattern of a web server: MySQL doing random reads, PHP-FPM writing session files, nginx writing cache files, and the OS logging — all simultaneously. On NVMe with none scheduler, each process’s I/O requests go directly to the device queue without reordering delay. vm.dirty_ratio tuning matters on servers that write files frequently (upload handlers, cache writers, log-heavy applications). A CMS server writing hundreds of cached pages per minute can accumulate dirty pages faster than the background writeback daemon clears them. When dirty pages hit vm.dirty_ratio, every PHP process that calls write() is stalled until the kernel catches up — producing uniform latency spikes that repeat every few minutes and correlate with disk write traffic.

  • Do not set vm.swappiness=0 on servers that must remain online during memory pressure events — with swappiness=0 on kernels before 3.5, the OOM killer activates as soon as memory is exhausted, potentially killing MySQL or nginx instead of evicting cache pages first.
  • Changing LimitNOFILE requires a service restart — the new limit only applies to processes started after the limit change; running processes keep their original limits.
  • The I/O scheduler change to 'none' removes all kernel-level I/O request merging and reordering — on RAID arrays or virtualized storage where the underlying hardware is a spinning disk, this can reduce throughput; verify with iostat before making permanent.
  • fs.file-max changes take effect immediately via sysctl -w but per-process limits (LimitNOFILE) require a service restart — a high system fs.file-max does not help a service still running with LimitNOFILE=1024.

Practical Examples

Complete kernel tuning profile for a web server

cat > /etc/sysctl.d/99-webserver.conf << 'EOF'
# Virtual memory
vm.swappiness = 10
vm.dirty_ratio = 15
vm.dirty_background_ratio = 5
vm.vfs_cache_pressure = 50

# File handles
fs.file-max = 2097152

# TCP (see also: linux-tcp-tuning-performance)
net.core.somaxconn = 65536
net.ipv4.tcp_max_syn_backlog = 65536
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
net.ipv4.ip_local_port_range = 1024 65535
net.core.netdev_max_backlog = 65536

# Socket buffers
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216
EOF

sysctl -p /etc/sysctl.d/99-webserver.conf

# File descriptor limits for all services
for svc in nginx php8.1-fpm mysql redis-server; do
  mkdir -p /etc/systemd/system/${svc}.service.d/
  printf '[Service]\nLimitNOFILE=65536\n' > \
    /etc/systemd/system/${svc}.service.d/limits.conf
done
systemctl daemon-reload

# Reload all services to apply new limits
systemctl reload nginx
systemctl restart php8.1-fpm mysql redis-server

This single script applies virtual memory, file handle, TCP, and socket buffer tuning in one operation. Run it on a fresh server before load testing to establish baseline performance with kernel defaults already removed as a bottleneck.

Monitor vm.dirty parameters to detect write-stall risk

# Watch dirty page state in real time
watch -n2 'awk '\''{
  if (/Dirty:/) printf "Dirty: %s KB\n", $2
  if (/Writeback:/) printf "Writeback: %s KB\n", $2
}'\'  /proc/meminfo; free -h | grep Mem'

# Check what percentage of RAM is currently dirty:
total_kb=$(awk '/MemTotal/{print $2}' /proc/meminfo)
dirty_kb=$(awk '/^Dirty:/{print $2}' /proc/meminfo)
echo "Dirty: $((dirty_kb * 100 / total_kb))% of RAM"
# If this approaches vm.dirty_ratio (15%), background flush is
# not keeping up — reduce vm.dirty_background_ratio to 3%

# Check for write stalls in dmesg:
dmesg | grep -i 'BDI\|dirty\|writeback' | tail -20

Write stalls produce latency spikes that appear uniform and periodic rather than correlated with traffic spikes. Monitoring dirty page percentage in /proc/meminfo over time reveals whether the background writeback is keeping pace with write traffic.

Troubleshooting Common Issues

Problem: sysctl -p outputs 'setting key net.ipv4.tcp_tw_recycle: No such file or directory'

Solution: tcp_tw_recycle was removed in Linux 4.12. Remove it from /etc/sysctl.d/99-webserver.conf. Use tcp_tw_reuse=1 instead. If the file is from an older guide, check for other removed parameters: tcp_timestamps_use_tstamp, tcp_window_scaling (these are now always-on defaults and the sysctl entries are removed).

Problem: LimitNOFILE change has no effect on the running nginx process

Solution: systemd service overrides only apply to new process launches. nginx reload (SIGHUP) does not restart the master process — use systemctl restart nginx to apply new limits. Verify with: awk ‘/Max open files/{print $4}’ /proc/$(pidof -s nginx)/limits — should now show 65536.

Problem: I/O scheduler change to 'none' does not persist after reboot

Solution: Changes to /sys/block/*/queue/scheduler are volatile and reset at boot. The persistent method is the udev rule in /etc/udev/rules.d/60-ioscheduler.rules (see the workflow above). After writing the rule, run udevadm control –reload-rules && udevadm trigger to apply without rebooting, then verify with cat /sys/block/sda/queue/scheduler.

Summary

Set vm.swappiness=10 to prevent the kernel from swapping web server processes to disk while RAM is still available. Raise LimitNOFILE=65536 for nginx, PHP-FPM, MySQL, and Redis via systemd overrides and restart services to apply. Set vm.dirty_background_ratio=5 to start flushing dirty pages earlier and prevent synchronous write stalls. Change the I/O scheduler to ‘none’ for SSD/NVMe storage via a persistent udev rule.

  • Set LimitNOFILE=65536 via systemd service override for nginx, PHP-FPM, MySQL, and Redis — without this, all services default to 1024 file handles and begin failing with ‘too many open files’ at moderate concurrency.
  • vm.swappiness=10 prevents the kernel from proactively swapping web server process memory when 40% of RAM is still free — MySQL buffer pool and PHP-FPM workers remaining in RAM is more valuable than the swap cache the kernel was trying to build.
  • Change the I/O scheduler to none for SSDs via a udev rule — the kernel’s CFQ/BFQ scheduler adds unnecessary reordering overhead on SSD storage that has no rotational latency to optimise.

Is Your Server Running at Full Performance?

INTRAM manages Linux servers with performance tuning built in from day one — correct MySQL configuration, PHP stack selection, nginx or Apache optimisation, and continuous monitoring so slowdowns are caught before users notice.

Explore Managed Hosting

Let’s assess what your business actually needs.

We will use these details only to understand your request and reply appropriately.