# Apache HTTP Server - Performance Metrics

This document explains the performance metrics of the Apache HTTP Server
(`httpd`) on RHEL 9.6: what each one means, why it matters, how
`test-apache.sh` measures it, and what to change when a result is poor.

---

## Contents

1. [The metrics at a glance](#1-the-metrics-at-a-glance)
2. [How the tests work](#2-how-the-tests-work)
3. [The metrics in detail](#3-the-metrics-in-detail)
   - [3.1 Startup time](#31-startup-time)
   - [3.2 Throughput](#32-throughput-requests-per-second)
   - [3.3 Latency](#33-latency-response-time)
   - [3.4 Concurrency scaling](#34-concurrency-scaling)
   - [3.5 Keep-alive effect](#35-keep-alive-effect)
   - [3.6 Transfer rate](#36-transfer-rate-bandwidth)
   - [3.7 Error rate](#37-error-rate)
   - [3.8 CPU usage](#38-cpu-usage)
   - [3.9 Memory usage](#39-memory-usage)
   - [3.10 Worker usage](#310-worker-usage)
4. [How the metrics relate to each other](#4-how-the-metrics-relate-to-each-other)
5. [Apache settings that affect performance](#5-apache-settings-that-affect-performance)
6. [Limits of this test](#6-limits-of-this-test)
7. [Glossary](#7-glossary)

---

## 1. The metrics at a glance

| # | Metric | Unit | Question it answers | Better is | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast does Apache come back after a (re)start? | lower | `startup` |
| 2 | Throughput | requests/s | How many requests can Apache serve per second? | higher | `throughput` |
| 3 | Latency | ms | How long does one user wait for an answer? | lower | `latency` |
| 4 | Concurrency scaling | % of peak | Does Apache hold up as more users arrive at once? | higher | `concurrency` |
| 5 | Keep-alive effect | x (ratio) | How much does reusing connections help? | information | `keepalive` |
| 6 | Transfer rate | MB/s | How fast can Apache send large files? | higher | `transfer` |
| 7 | Error rate | % | How many requests fail when Apache is overloaded? | lower | `errors` |
| 8 | CPU usage | % of all CPUs | How much processor time does the load cost? | lower | `cpu` |
| 9 | Memory usage | MB | How much RAM does Apache need, idle and busy? | lower | `memory` |
| 10 | Worker usage | % of MaxRequestWorkers | How close is Apache to running out of workers? | lower | `workers` |

Run one metric with `./test-apache.sh <test name>`, or all of them with
`./test-apache.sh`.

---

## 2. How the tests work

```
   +---------------------+       HTTP over 127.0.0.1        +----------------------+
   |  ApacheBench ("ab") |  ------------------------------>  |  Apache httpd daemon |
   |  simulates N users  |  <------------------------------  |  (systemd service)   |
   +---------------------+          test pages              +----------------------+
             |                                                        |
             |  requests/s, latency,                                  |  CPU, memory: /proc and cgroup
             |  errors, MB/s                                          |  workers: /server-status
             v                                                        v
                      results/<date-time>/report.md  (PASS / FAIL per metric)
```

- **Load generator:** ApacheBench (`ab`), from the RHEL `httpd-tools` package.
  It is started with a number of *clients* (simultaneous users). Each
  client sends a request, waits for the answer, and immediately sends the next
  one, for `DURATION` seconds.
- **Test content:** created by `start-apache.sh` in `/var/www/html/perf-test/`:
  - `small.html`: a 1 KB page, typical of small web requests.
  - `large.bin`: a 1 MB file, typical of a download.
  - `health.txt`: used only to check that Apache is serving.
- **Server-side measurements:** taken once per second while load runs:
  - CPU time from the `httpd.service` cgroup (`cpu.stat`).
  - Memory (PSS) from `/proc/<pid>/smaps_rollup` of every `httpd` process.
  - Busy and idle workers from Apache's status page (`/server-status?auto`).
- **Offline:** all traffic uses the loopback address `127.0.0.1`. No internet
  access and no second machine are needed.
- **Warm-up:** 3 seconds of light, unmeasured load run first, so the first test
  does not start "cold".

---

## 3. The metrics in detail

Each metric below has the same five parts: **what** it is, **why** it matters,
**how** the test measures it, **how to read** the result, and **how to improve** it.

### 3.1 Startup time

**What:** the time from the command that starts Apache until the first
web page is actually served.

**Why:** it decides how long a website is down during a restart, a
configuration change, a package update or a server reboot.

**How it is measured:** Apache is stopped and started `STARTUP_ROUNDS` times
(3 by default) with `systemctl stop httpd` / `systemctl start httpd`. The timer runs
from the start command until `health.txt` answers. The report shows each
round and the average.

**How to read it:**

| Result | Meaning |
|---|---|
| under 500 ms | normal for a plain RHEL Apache |
| 0.5 - 5 s | acceptable; many modules, virtual hosts or TLS certificates to load |
| over 5 s | investigate (slow DNS lookups at start, huge configuration, slow disk) |

**How to improve:** load fewer modules (`/etc/httpd/conf.modules.d/`), avoid
host names that need a DNS lookup in `Listen` / `<VirtualHost>`, and use
`systemctl reload httpd` (graceful) instead of a full restart for config changes.

---

### 3.2 Throughput (requests per second)

**What:** how many complete requests Apache serves per second.

**Why:** this is the capacity of the server. If a site receives 3,000
requests/s at peak, the server needs a throughput well above 3,000.

**How it is measured:** `CONCURRENCY` clients (50 by default) request the
1 KB page with keep-alive for `DURATION` seconds:

```
ab -k -c 50 -t 10 http://127.0.0.1/perf-test/small.html
```

The result is ApacheBench's *Requests per second*.

**How to read it:** it depends strongly on the hardware and on the content.
For static 1 KB files over localhost:

| Server | Typical throughput |
|---|---|
| 2 vCPU virtual machine | 5,000 - 20,000 req/s |
| 8-16 CPU server | 30,000 - 100,000 req/s |
| 32+ CPU server | 100,000+ req/s |

Dynamic pages (PHP, CGI, proxied applications) are typically 10 to 1000 times slower.
The default target in `settings.conf` is 2,000 req/s. Set your own target as
peak expected traffic x 2 (safety margin).

**How to improve:** more CPUs; keep-alive on; enough workers
(`MaxRequestWorkers`); remove unused modules; turn off `HostnameLookups`;
reduce logging on very busy servers; serve static files from a fast disk or
page cache.

---

### 3.3 Latency (response time)

**What:** the time one request takes, from sending it until the full
answer has arrived. Reported as:

| Statistic | Meaning |
|---|---|
| Average | the mean of all requests; can hide slow outliers |
| p50 (median) | half of the requests were faster than this |
| p90 / p95 / p99 | 90 / 95 / 99 of every 100 requests were faster than this |
| Slowest | the single slowest request |

**Why:** latency is what a user *feels*. Percentiles matter more than the
average: with 1,000 visitors, a bad p99 means 10 of them are waiting a long time.

**How it is measured:** the same load as throughput (50 clients, keep-alive).
ApacheBench writes the time for every percentile (0-99) to
`raw/latency-percentiles.csv`, and the report reads p50, p90, p95 and p99 from it.

**How to read it:** for static files over localhost, p95 is usually
below 5 ms. The default targets are **p95 at most 50 ms** and **p99 at most 100 ms**.
A large gap between p50 and p99 means some requests wait. Possible reasons
are all workers busy, CPU saturation, disk I/O or swapping.

**How to improve:** the same as throughput, and also: never let the server swap,
keep `MaxRequestWorkers` high enough that requests do not queue, and check the
worker-usage metric (3.10).

---

### 3.4 Concurrency scaling

**What:** how throughput and latency change as the number of simultaneous
clients grows. By default the steps are 1, 10, 50, 100, 200 and 400 clients.

**Why:** real traffic comes in bursts. A good server keeps its throughput
when the load rises, and latency grows only gradually. A bad one collapses:
throughput falls, latency jumps, errors appear.

**How it is measured:** one throughput run per step (`CONCURRENCY_LEVELS`
in `settings.conf`). The result is the throughput at the highest step as a
percentage of the best (peak) throughput.

**How to read it:**

```
 requests/s
     ^
     |          ____________________   <- healthy: levels off (plateau)
     |        /
     |      /  \_____                  <- unhealthy: collapses under load
     |    /          \____
     |  /
     +----------------------------->  clients
       1   10   50  100  200  400
```

- Throughput rises until the CPUs are busy, then stays flat. **This is normal.**
- Latency grows roughly in proportion to clients once throughput is flat.
  That is queueing, and it is also normal.
- The test **fails** if the highest step keeps less than `TARGET_SCALING_MIN_PCT`
  (50%) of the peak. That means Apache collapses under load.
- The *Keep-alive closes* column counts connections Apache closed after a
  complete answer (see 3.7). A high number means Apache was short of workers.

**How to improve:** raise `MaxRequestWorkers` (and `ServerLimit`) if worker
usage reaches 100%; add CPUs if CPU usage is at 100%; check the
`AH00484: server reached MaxRequestWorkers` message in `/var/log/httpd/error_log`.

---

### 3.5 Keep-alive effect

**What:** the throughput gain from HTTP keep-alive (persistent
connections), shown as a ratio: *requests/s with keep-alive divided by
requests/s without*.

**Why:** without keep-alive, every request needs a new TCP connection
(3-way handshake, then close). With keep-alive, one connection carries many requests.
The saving is larger over a real network (each handshake costs a network
round trip) and much larger with HTTPS (each new connection also needs a TLS
handshake).

**How it is measured:** the same load twice: once with `ab -k`
(keep-alive) and once without.

**How to read it:** a ratio of 2x-10x over localhost is typical. This metric
is **information only** (no pass/fail). It shows why `KeepAlive On` should stay enabled.

**How to improve:** keep `KeepAlive On`. Use a short `KeepAliveTimeout`
(2-5 s) so idle connections do not hold resources. The event MPM (RHEL 9 default)
handles idle keep-alive connections without tying up a worker thread.

---

### 3.6 Transfer rate (bandwidth)

**What:** how many megabytes per second Apache sends while clients download
a 1 MB file.

**Why:** it shows the limit for downloads, images, videos and other large
content. Here requests/s matters less than bytes/s.

**How it is measured:** `CONCURRENCY` clients download `large.bin`
repeatedly with keep-alive. The result is ApacheBench's *Transfer rate*,
converted from KB/s to MB/s.

**How to read it:** over localhost there is no network card in the way, so
the result is usually several GB/s and shows what Apache and the kernel can
do. On a real network the network card is the limit: a 1 Gbit/s link carries
at most about **119 MB/s**, and 10 Gbit/s about **1,190 MB/s**. The default target is
100 MB/s, which a healthy server always passes over localhost.

**How to improve:** keep `EnableSendfile` on (the RHEL default; the kernel
sends the file directly), use `mod_deflate` for text (not for already
compressed files), and use a faster network card on a real network.

---

### 3.7 Error rate

**What:** the percentage of requests that failed while Apache was
**overloaded**:

| Error type | Meaning |
|---|---|
| Connect | the client could not open a connection |
| Receive | the connection broke while reading the answer |
| Wrong length | the answer was cut short or had the wrong size |
| Exceptions | other socket errors |
| Non-2xx | Apache answered with an HTTP error (for example 503 Service Unavailable) |

**Why:** under overload, a good server gets slower but does not fail
requests. Errors mean users see broken pages.

**How it is measured:** `ERROR_TEST_CONCURRENCY` clients (400 by default),
**without** keep-alive: every request opens a new connection, which is the
hardest case. The test also counts new `error`, `crit`, `alert` and `emerg` lines,
and `AH00484` (out of workers) lines, in `/var/log/httpd/error_log`.

Error rate = (failed requests + non-2xx answers) / completed requests x 100

**How to read it:** the target is **at most 0.1%**, and a healthy server shows
0%. Any serious line in the error log should be read and understood.

> **About "keep-alive closes":** In keep-alive mode, Apache sometimes sends a
> complete answer and then closes the connection, for example when a process
> is short of free workers. The client just reconnects. ApacheBench wrongly
> counts each of these as a "Length" failure. `test-apache.sh` counts them
> separately as *keep-alive closes* and does **not** treat them as failures.

**How to improve:** raise `MaxRequestWorkers` / `ServerLimit`; raise the
listen queue (`ListenBacklog` in Apache and `net.core.somaxconn` in the
kernel) if you see Connect errors; raise the open-files limit
(`LimitNOFILE` in a systemd drop-in) for very high connection counts.

---

### 3.8 CPU usage

**What:** the processor time used by all Apache processes, as a percentage
of **all** CPUs (100% = every CPU fully busy with Apache). The report also
shows the same value in CPU cores, and the *CPU time per 1000 requests*
(efficiency).

**Why:** CPU is usually the first limit for a web server. If Apache needs
all CPUs at the expected traffic, there is no room for peaks or for other services.

**How it is measured:** during one sampled load run (50 clients, keep-alive),
the total CPU time of the `httpd.service` cgroup is read before and after
the run (exact, including processes that exit). Without systemd, the CPU
counters of all `httpd` processes are added up instead. The busiest single
second is taken from the per-second samples (`raw/resource-samples.csv`).

**How to read it:** `ab` runs on the same machine and also uses CPU, so
Apache alone rarely reaches 100%. The default target is **at most 90%**.
*CPU time per 1000 requests* is the best number to compare between two
configurations or two servers: lower means more efficient.

**How to improve:** remove unused modules, turn off `HostnameLookups`,
avoid `.htaccess` files (`AllowOverride None`), and use the event MPM (default).
For dynamic content, most CPU is usually in the application (PHP, Java),
not in Apache.

---

### 3.9 Memory usage

**What:** RAM used by all Apache processes together, measured while idle
and at the peak of the load.

**Why:** if Apache and the other services need more RAM than the server has,
Linux starts swapping, and latency becomes very poor. Memory also sets how
many workers you can afford.

**How it is measured:** once per second, the **PSS** (proportional set size)
of every `httpd` process is read from `/proc/<pid>/smaps_rollup` and added up.
PSS divides shared memory (program code, libraries) fairly between the
processes that share it, so the sum is the real total. (Adding up `RSS`, as
`top` or `ps` show it, would count shared memory many times over.)

**How to read it:** a plain RHEL Apache with the event MPM typically uses
30-150 MB for a few hundred workers serving static files. Modules such as
`mod_php` (prefork MPM) can use 20-100 MB **per process**. The default target
is **at most 1024 MB**. If *peak* is much larger than *idle*, Apache started new
processes under load. Check that `ServerLimit` x memory per process fits in RAM.

**How to improve:** use the event MPM (not prefork); run PHP as `php-fpm`
instead of `mod_php`; remove unused modules; set `MaxConnectionsPerChild`
(for example 10000) if memory grows slowly over days (a leak in a module).

---

### 3.10 Worker usage

**What:** a *worker* is one thread that handles one request at a time.
The metric is the largest number of busy workers seen during the load, as a
percentage of `MaxRequestWorkers` (the maximum Apache may run).

**Why:** when all workers are busy, new requests wait in a queue, so latency
jumps and errors start. This metric shows how close the server is to that limit
*before* users notice.

**How it is measured:** during the sampled load run, Apache's status page
(`/server-status?auto`, provided by `mod_status`) is read once per second.
It reports `BusyWorkers` and `IdleWorkers`.

**How to read it:**

| Peak usage | Meaning |
|---|---|
| under 70% | comfortable headroom |
| 70 - 90% | busy; plan more workers or more servers |
| over 90% | **FAIL**: at the limit, requests will queue |

With the event MPM and keep-alive, idle connections do **not** occupy a
worker. So 50 clients requesting small files often need far fewer than 50 busy
workers, because each request takes well under a millisecond. To find the real limit,
raise the load: `CONCURRENCY=400 ./test-apache.sh workers`.

**How to improve:** raise `MaxRequestWorkers`. Raise `ServerLimit` with it
(`MaxRequestWorkers` must not exceed `ServerLimit` x `ThreadsPerChild`),
and check that the extra memory fits (3.9).

---

## 4. How the metrics relate to each other

The metrics are not independent. One simple rule, **Little's law**, links
three of them:

```
clients in the system  =  throughput (requests/s)  x  latency (s)
```

Example: 50 clients and 25,000 requests/s means an average latency of
50 / 25,000 = 0.002 s = **2 ms**. So when throughput stops growing (CPU full),
every extra client adds waiting time. That is why latency rises in the
concurrency test after the plateau.

Typical chain of cause and effect under growing load:

```
more clients --> CPU reaches 100% --> throughput stops growing
             --> workers stay busy longer --> worker usage near 100%
             --> requests queue --> latency (p99) jumps
             --> queue full --> connection errors / 503 answers --> error rate > 0
```

When a result is poor, follow the chain back: high latency with low CPU and
low worker usage points to something outside Apache (disk, network,
a slow application behind Apache).

---

## 5. Apache settings that affect performance

`start-apache.sh` writes these settings to `/etc/httpd/conf.d/zz-perf-test.conf`
from the values in `settings.conf`. Everything else is the RHEL default.

| Apache directive | settings.conf name | Kit value | RHEL default | Affects |
|---|---|---|---|---|
| `ServerLimit` | `SERVER_LIMIT` | 16 | 16 | max child processes, memory |
| `ThreadsPerChild` | `THREADS_PER_CHILD` | 25 | 25 | workers per process |
| `MaxRequestWorkers` | `MAX_REQUEST_WORKERS` | 400 | 400 | max simultaneous requests |
| `StartServers` | `START_SERVERS` | 16 | 3 | workers ready at startup |
| `MinSpareThreads` | `MIN_SPARE_THREADS` | 75 | 75 | idle workers kept ready |
| `MaxSpareThreads` | `MAX_SPARE_THREADS` | 400 | 250 | idle workers before processes are stopped |
| `KeepAlive` | (always On) | On | On | connection reuse |
| `KeepAliveTimeout` | `KEEPALIVE_TIMEOUT` | 5 s | 5 s | how long idle connections stay open |
| `MaxKeepAliveRequests` | `MAX_KEEPALIVE_REQUESTS` | 0 (unlimited) | 100 | requests per connection |
| `ExtendedStatus` | (always On) | On | Off | detail on `/server-status` |

Why the kit changes a few defaults:

- **`StartServers` 16 / `MaxSpareThreads` 400:** all 400 workers are started
  at once and kept. Otherwise the first seconds of each test would measure
  Apache *creating* processes, not *serving* requests, and results would
  differ from run to run.
- **`MaxKeepAliveRequests` 0:** with the default of 100, Apache closes every
  connection after 100 requests. ApacheBench would report each closure as a
  failure (see 3.7).

Other settings worth knowing (not changed by the kit):

| Directive | Effect |
|---|---|
| `HostnameLookups Off` | (default) no DNS lookup per request |
| `EnableSendfile On` | (RHEL default) kernel sends files directly |
| `AllowOverride None` | no `.htaccess` file search on every request |
| `ListenBacklog` | queue length for new connections (default 511) |
| `MaxConnectionsPerChild` | recycle processes after N connections (limits leaks) |
| `Timeout` | how long to wait for a slow client (default 60 s) |

---

## 6. Limits of this test

Be aware of what a localhost static-file test can and cannot tell you:

- **Client and server share one machine.** `ab` uses CPU too, so the
  results are lower than a dedicated server would reach. They are a
  **baseline** for comparing configurations and servers, not an exact capacity.
- **No real network.** Network latency, packet loss and the network card's
  bandwidth are not included. Test from a second machine for those.
- **Static files only.** PHP, CGI, proxied applications and databases are
  much slower and need their own tests.
- **HTTP/1.0 without TLS.** ApacheBench speaks HTTP/1.0 over plain HTTP. HTTPS
  adds CPU cost per connection, and HTTP/2 behaves differently.
- **Short runs.** The default is 10 seconds per test. For a capacity
  baseline, use `DURATION=60` or more and repeat the run 3 times.
- **SELinux.** RHEL 9.6 normally runs SELinux in enforcing mode. The kit uses
  standard RHEL paths, so the standard policy applies. If something is
  blocked, check `ausearch -m AVC -ts recent`.

---

## 7. Glossary

| Term | Meaning |
|---|---|
| **ab** (ApacheBench) | load-testing tool shipped in the `httpd-tools` RPM |
| **client / concurrency** | one simulated user sending requests one after another; concurrency = how many at the same time |
| **event MPM** | Apache's default (RHEL 9) processing model: a few processes, many threads, idle connections handled without a thread |
| **keep-alive** | reusing one TCP connection for several HTTP requests |
| **latency** | time from sending a request until the full answer arrives |
| **MPM** | Multi-Processing Module: how Apache uses processes and threads (event, worker, prefork) |
| **percentile (p95)** | the value that 95% of measurements are below |
| **PSS** | proportional set size: memory with shared pages divided fairly between processes |
| **throughput** | completed requests per second |
| **worker** | one thread that handles one request at a time |
| **cgroup** | Linux control group; systemd puts all `httpd` processes in one, which lets the kit read their total CPU time exactly |
