ApacheBench (ab) sends a controlled burst of HTTP requests and reports requests per second, connection time percentiles, and failed requests. It takes 30 seconds to run a meaningful test and requires no configuration files. This page covers how to use it to find the concurrency level where your server starts dropping performance, and how to interpret every number in the output.

What ApacheBench Measures and What It Does Not

ApacheBench works by opening N concurrent TCP connections to your server and sending a total of R HTTP requests as fast as possible. It measures the time from sending the request to receiving the last byte of the response and reports statistics across all completed requests: mean, median, and percentile latencies (50th, 66th, 75th, 80th, 90th, 95th, 98th, 99th, 100th); throughput in requests per second and bytes per second; number of failed requests (non-2xx/3xx responses); and connection time broken into DNS lookup, TCP connect, processing, and total. The output answers: what is the maximum concurrency at which my server still responds under the latency target? This is the number you need before any significant traffic event — a product launch, a marketing campaign, a feature release. What ab does not measure: realistic user behaviour (real users think between requests; ab does not), browser page weight (ab tests one URL at a time), or sustained load over time. For browser-realistic testing and sustained multi-minute load, k6 (the next page) is more appropriate. For WordPress sites, the URL to test matters: test a page that forces PHP execution, not a static asset. Test with the cache warmed (run ab once to warm, then once to measure). Test the logged-in user path separately — it bypasses FastCGI cache and is significantly slower.

Tools and Commands

# Install ApacheBench
apt install apache2-utils

# Basic test: 1000 requests, 10 concurrent
ab -n 1000 -c 10 https://example.com/

# 500 requests, 50 concurrent with Keep-Alive
ab -n 500 -c 50 -k https://example.com/

# POST request with JSON body
ab -n 200 -c 20 -T 'application/json' -p /tmp/payload.json https://example.com/api/endpoint

# Test with a specific Host header (for virtual host testing)
ab -n 500 -c 25 -H 'Host: example.com' https://1.2.3.4/

# Increase timeout to 60 seconds (useful for slow pages)
ab -n 200 -c 20 -s 60 https://example.com/

# Save detailed timing CSV for later analysis
ab -n 1000 -c 50 -e /tmp/ab_results.csv https://example.com/

# Show version
ab -V

The -c flag sets concurrency (parallel connections). The -n flag sets total requests. Start with c=10 and work up — doubling c until requests per second stops increasing or Failed requests starts climbing identifies the concurrency saturation point.

Key Parameters

Flag / ParameterDescriptionSecurity Note
-nTotal number of requests to send.Use at least 1000 requests to get statistically meaningful percentile numbers. Too few requests (< 100) give noisy latency measurements. For percentile accuracy at the 99th percentile, send at least 500 requests.
-cConcurrency: number of requests to send simultaneously.Start at 10 and double: 10, 20, 40, 80. Plot requests/sec and 99th percentile latency at each level. The concurrency where requests/sec flattens or latency spikes is your server's saturation point.
-kEnable HTTP Keep-Alive connections (reuse TCP connections).Use -k for realistic measurements — browsers use Keep-Alive by default. Without -k, ab opens a new TCP connection for every request, which adds TCP handshake overhead and inflates latency numbers beyond what real users experience.
-tMaximum duration of the test in seconds.Use -t 30 (30 seconds) combined with a large -n (e.g. -n 100000) to run a fixed-duration test. This measures sustained throughput rather than total completion time.
-sSocket timeout in seconds for each request. Default: 30.Increase to 60 or 120 when testing slow pages like WordPress admin pages or API endpoints that hit the database. If the timeout is shorter than the server response time, ab reports failed requests that are actually slow responses.
-eWrite a CSV file with percentile data for all completed requests.Useful for comparing multiple test runs objectively. The CSV has one row per percentile (1-100) and the latency (in ms) at that percentile. Plot this data to visualise the latency distribution shape.

Diagnosis and Fix Workflows

Find the concurrency saturation point of your web server

Run a concurrency sweep to find where requests per second peaks and latency begins to degrade. This number is the server’s maximum sustainable concurrency.

# Warm the cache first (do not count this run)
ab -n 100 -c 5 https://example.com/ > /dev/null

# Concurrency sweep: run at increasing levels
for c in 5 10 20 40 80 160; do
  echo "=== Concurrency: $c ==="
  ab -n 500 -c $c -k https://example.com/ 2>&1 | \
    grep -E 'Requests per second|Time per request|Failed requests'
  sleep 5  # let the server recover between tests
done

# Look for the pattern:
# c=5:   Requests per second: 120  Time per request: 42ms  Failed: 0
# c=10:  Requests per second: 185  Time per request: 54ms  Failed: 0
# c=20:  Requests per second: 220  Time per request: 91ms  Failed: 0
# c=40:  Requests per second: 218  Time per request: 184ms Failed: 0
# c=80:  Requests per second: 195  Time per request: 410ms Failed: 23
# c=160: Requests per second: 140  Time per request: 1142ms Failed: 89

# Saturation point: c=20 (throughput peaks; doubling c does not help)
# This tells you: more than 20 concurrent PHP-FPM workers will not improve throughput
# — investigate what is blocking PHP-FPM workers, not add more

Test before and after a configuration change

Use ab to quantify the performance impact of configuration changes: enabling OPcache, switching from mod_php to PHP-FPM, enabling nginx FastCGI cache, or tuning PHP-FPM pool size.

# Establish baseline BEFORE the change
ab -n 1000 -c 20 -k https://example.com/ 2>&1 | \
  tee /tmp/baseline.txt | \
  grep -E 'Requests per second|Time per request|Failed|99%'

# Apply your configuration change
# Example: enable OPcache
php -i | grep opcache.enable  # verify it is off
# edit /etc/php/8.1/fpm/conf.d/10-opcache.ini:
# opcache.enable=1
systemctl reload php8.1-fpm

# Test AFTER the change (warm cache first)
ab -n 100 -c 5 https://example.com/ > /dev/null
ab -n 1000 -c 20 -k https://example.com/ 2>&1 | \
  tee /tmp/after_opcache.txt | \
  grep -E 'Requests per second|Time per request|Failed|99%'

# Compare:
diff /tmp/baseline.txt /tmp/after_opcache.txt
# Typical OPcache improvement: 3-10x requests/sec increase
# PHP page from 200ms -> 20ms mean time per request

Detect connection exhaustion under load

If ab reports a growing number of Failed requests at higher concurrency levels, the server is exhausting a resource. The cause can be PHP-FPM worker pool, MySQL connection limit, or OS file descriptor limits.

# Run ab with 80 concurrent, watch server resources simultaneously
# Terminal 1:
ab -n 2000 -c 80 -k https://example.com/

# Terminal 2 (while ab runs):
# Check PHP-FPM queue
php-fpm8.1 -t 2>&1 | grep 'listen.backlog'
# If listen.backlog is 511 (default) and pm.max_children is 20,
# you can only queue 511 requests before connections are refused

# Check PHP-FPM worker state:
ps aux | grep php-fpm | grep -c 'Waiting for connection'
ps aux | grep php-fpm | grep -c 'Reading headers'
# If 'Reading headers' == pm.max_children, all workers are busy
# Every new request waits in queue or is refused

# Diagnostic: correlate ab's Failed requests with FPM pool status
# Increase pm.max_children and retest
# If failed requests drop proportionally, FPM pool size was the limit

# Check error log for connection refusals:
tail -f /var/log/nginx/error.log | grep -E 'connect|upstream|timeout'
# nginx 502/504 errors during ab test confirm PHP-FPM is exhausted

Performance Impact: What ApacheBench Results Reveal About Your Stack

The requests per second number from ab is not a goal to maximise in isolation — it is a signal about where the bottleneck lies. A server serving cached static HTML might handle 10,000 requests/sec. The same server with WordPress and no caching might handle 30 requests/sec. The gap reveals the cost of dynamic PHP execution. After enabling nginx FastCGI cache, the same WordPress server might handle 5,000 requests/sec for cached pages — proving the stack can handle traffic at nginx speed, not PHP speed. The 99th percentile latency is more operationally important than mean latency. If the mean is 150ms but the 99th percentile is 8000ms, one in a hundred users gets an 8-second page load. This pattern typically indicates a resource queue: most requests are served quickly, but occasionally a request waits for a PHP-FPM worker or MySQL connection that is held by another request. The failed request count from ab needs interpretation: a non-zero value does not always mean requests returned errors. ab counts as failed any response that differs in length from the first response (length-based failure detection). WordPress sessions that rotate nonces, or pages with dynamic timestamps, will trigger length failures even though the page is technically correct. Use -l flag (if supported) or check the response code column to distinguish genuine errors from length variation.

  • Never run ab against a production server during business hours — a test with 100 concurrent connections at 1000 requests is real load that will affect real users sharing the same server.
  • ab does not follow redirects — if your URL redirects (HTTP to HTTPS, www to non-www), ab stops at the redirect response and measures 301 responses, not the actual page. Use the final URL after all redirects.
  • Testing a CDN-fronted domain measures CDN capacity, not your server — temporarily test via the server's IP with a Host header, or use a test URL that bypasses the CDN.
  • The Failed requests counter in ab output counts length mismatches as failures, not only HTTP errors. Verify with the detailed output whether failures are genuine 4xx/5xx responses or length variation.

Practical Examples

Compare WordPress performance with and without FastCGI cache

# Test 1: WordPress without FastCGI cache (pure PHP execution)
# Temporarily disable the cache block in nginx config:
# Comment out: add_header X-Cache-Status $upstream_cache_status;
# and the cache zone directives
nginx -t && systemctl reload nginx

ab -n 300 -c 20 -k https://example.com/ 2>&1 | grep -E 'Requests per second|Time per request'
# Example: Requests per second: 28.45  Time per request: 703ms

# Test 2: WordPress with FastCGI cache enabled (warm cache)
nginx -t && systemctl reload nginx  # re-enable cache
curl -s https://example.com/ > /dev/null  # warm the cache

ab -n 1000 -c 100 -k https://example.com/ 2>&1 | grep -E 'Requests per second|Time per request'
# Example: Requests per second: 4823.10  Time per request: 20ms

# The 170x throughput difference quantifies exactly what FastCGI cache is worth
# to your server capacity

This before/after test gives a concrete number to justify caching infrastructure investment. A server that handles 28 requests/sec without cache cannot survive a traffic spike; 4823 requests/sec with cache can absorb it.

Automate periodic load tests and alert on regression

#!/bin/bash
# /usr/local/bin/perf-check.sh — run from cron daily
URL="https://example.com/"
THRESHOLD_RPS=50  # alert if below this
LOGFILE="/var/log/perf-check.log"

# Warm cache
ab -n 50 -c 5 "$URL" > /dev/null 2>&1

# Measure
RPS=$(ab -n 500 -c 20 -k "$URL" 2>&1 | grep 'Requests per second' | awk '{print int($4)}')
TIMESTAMP=$(date '+%Y-%m-%d %H:%M')

echo "$TIMESTAMP RPS=$RPS" >> "$LOGFILE"

if [ "$RPS" -lt "$THRESHOLD_RPS" ]; then
  echo "ALERT: Server performance degraded. RPS=$RPS (threshold=$THRESHOLD_RPS)" | \
    mail -s "[ALERT] Server performance degraded" admin@example.com
fi

Running a daily ab test and logging the result gives a performance baseline trend. A sudden drop in requests/sec that preceded no deployment indicates a configuration drift, memory pressure, or disk issue.

Troubleshooting Common Issues

Problem: ab reports 'apr_socket_recv: Connection reset by peer' errors

Solution: The server is actively closing connections during the test, usually because the concurrency level exceeds the server’s PHP-FPM pm.max_children combined with the listen.backlog queue. Reduce -c to a level where errors disappear, then incrementally increase it. Check nginx error log for upstream connection refused errors that coincide with ab failures.

Problem: ab shows 0 Failed requests but the server felt slow during the test

Solution: Check the 99th percentile latency in the output — ‘Percentage of requests served within time (ms)’ section. If the 99% row shows 5000ms when the 50% row shows 100ms, the server is queueing some requests. The mean looks fine but tail latency is high. This is typical when PHP-FPM pool is barely large enough — most requests are served quickly, but occasional requests wait for a free worker.

Problem: Requests per second drops to near zero when concurrency exceeds 50

Solution: This is a hard limit being hit — not graceful degradation. Check ulimit -n for the current user (ab uses one file descriptor per connection). If the limit is 1024 and you are testing at c=100 with -n=10000, ab may exhaust its own file descriptors. Run ulimit -n 65536 before running ab. Also check that the server-side ulimit for www-data or nginx is not limiting socket connections.

Summary

Install ab with apt install apache2-utils. Run a concurrency sweep starting at -c 10 and doubling to find where requests per second peaks and Failed requests begins to climb. Use -k for Keep-Alive and always warm the cache with a preliminary run before measuring. The 99th percentile latency matters more than the mean — tail latency reveals queueing problems that averages conceal.

  • Run a concurrency sweep (for c in 10 20 40 80) to find the exact concurrency level where throughput peaks — this is your server’s actual capacity limit, not the requests/sec at a single concurrency level.
  • Always use -k (Keep-Alive) for realistic measurements; without it ab opens a new TCP connection per request and measures TCP overhead, not PHP execution time.
  • Non-zero Failed requests in ab output often reflects response length variation (WordPress dynamic content), not genuine errors — check the HTTP response codes to distinguish before acting on it.

Is Your Server Running at Full Performance?

INTRAM manages Linux servers with performance tuning built in from day one — correct MySQL configuration, PHP stack selection, nginx or Apache optimisation, and continuous monitoring so slowdowns are caught before users notice.

Explore Managed Hosting

Let’s assess what your business actually needs.

We will use these details only to understand your request and reply appropriately.