# Apache Tomcat - Performance Metrics

This document explains the performance metrics of the Apache Tomcat
application server on RHEL 9.6: what each one means, why it matters, how
`test-tomcat.sh` measures it, and what to change when a result is poor.

---

## Contents

1. [The metrics at a glance](#1-the-metrics-at-a-glance)
2. [How Tomcat works, in short](#2-how-tomcat-works-in-short)
3. [How the tests work](#3-how-the-tests-work)
4. [The metrics in detail](#4-the-metrics-in-detail)
   - [4.1 Startup time](#41-startup-time)
   - [4.2 Static page throughput](#42-static-page-throughput)
   - [4.3 Dynamic page throughput](#43-dynamic-page-throughput)
   - [4.4 Latency](#44-latency-response-time)
   - [4.5 Concurrency scaling](#45-concurrency-scaling)
   - [4.6 Keep-alive effect](#46-keep-alive-effect)
   - [4.7 Transfer rate](#47-transfer-rate-bandwidth)
   - [4.8 Compression (gzip)](#48-compression-gzip)
   - [4.9 Worker thread pool](#49-worker-thread-pool)
   - [4.10 Sessions](#410-sessions)
   - [4.11 Error rate](#411-error-rate-under-overload)
   - [4.12 Connections](#412-connections)
   - [4.13 Garbage collection](#413-garbage-collection-gc)
   - [4.14 CPU usage](#414-cpu-usage)
   - [4.15 Memory usage](#415-memory-usage)
5. [How the metrics relate to each other](#5-how-the-metrics-relate-to-each-other)
6. [Tomcat and Java settings that affect performance](#6-tomcat-and-java-settings-that-affect-performance)
7. [Limits of this test](#7-limits-of-this-test)
8. [Glossary](#8-glossary)

---

## 1. The metrics at a glance

| # | Metric | Unit | Question it answers | Better is | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast is Tomcat ready after a (re)start? | lower | `startup` |
| 2 | Static page throughput | requests/s | How many static files can Tomcat serve per second? | higher | `static` |
| 3 | Dynamic page throughput | requests/s | How many application pages (JSP) per second? | higher | `dynamic` |
| 4 | Latency | ms | How long does one user wait for an answer? | lower | `latency` |
| 5 | Concurrency scaling | % of peak | Does Tomcat hold up as more users arrive at once? | higher | `concurrency` |
| 6 | Keep-alive effect | x (ratio) | How much does reusing connections help? | information | `keepalive` |
| 7 | Transfer rate | MB/s | How fast can Tomcat send large files? | higher | `transfer` |
| 8 | Compression | x smaller, % speed | How much does gzip save, and what does it cost? | information | `compression` |
| 9 | Worker thread pool | % of pool limit | Does the thread pool deliver its full capacity with slow pages? | higher | `threads` |
| 10 | Sessions | sessions/s, KB each | How fast are user sessions created, and how much memory do they need? | higher / lower | `sessions` |
| 11 | Error rate | % | How many requests fail when Tomcat is overloaded? | lower | `errors` |
| 12 | Connections | % of maximum | How close is Tomcat to its connection limit? | lower | `connections` |
| 13 | Garbage collection | % of time, ms | How long and how often does Java pause to free memory? | lower | `gc` |
| 14 | CPU usage | % of all CPUs | How much processor time does the load cost? | lower | `cpu` |
| 15 | Memory usage | MB | How much RAM does the Tomcat process need? | lower | `memory` |

Run one metric with `./test-tomcat.sh <test name>`, or all of them with
`./test-tomcat.sh` (about 4 minutes).

---

## 2. How Tomcat works, in short

Tomcat is a Java program that runs Java web applications (servlets and JSP
pages). Knowing its structure makes the metrics easier to understand.

```
                 ONE Java process  (java ... org.apache.catalina.startup.Bootstrap)
   +-----------------------------------------------------------------------------+
   |                                                                             |
   |   HTTP connector (NIO)                        web applications              |
   |   +------------------+   request   +------------------------------------+   |
   |   | acceptor thread  |  --------->  |  worker thread pool                |   |
   |   | poller threads   |             |  (at most maxThreads, default 200)  |   |
   |   | watch ALL open   |  <---------  |  runs servlet / JSP code:          |   |
   |   | connections      |   answer    |  static files, hello.jsp, ...      |   |
   |   +------------------+             +------------------------------------+   |
   |                                                                             |
   |   Java heap: every object lives here (sessions, pages being built ...)      |
   |   Garbage collector (G1): frees unused objects, sometimes pausing threads   |
   +-----------------------------------------------------------------------------+
```

- **One process, many threads.** Unlike Nginx or Apache httpd, Tomcat is a
  single Java process. All work is done by threads inside it.
- **NIO connector.** A few *poller* threads watch every open connection. Only
  when a request has arrived is it handed to a *worker thread* from the pool.
  An idle keep-alive connection therefore costs no thread. Up to
  `maxConnections` (default 8192) connections can be open; up to `maxThreads`
  (default 200) requests are processed at the same time. When all worker
  threads are busy, new requests wait in a queue.
- **acceptCount** (default 100) is the operating system's queue of brand-new
  connections that Tomcat has not yet accepted. When it overflows, the kernel
  drops the connection attempt and the client retries after about one second.
- **Static and dynamic content.** Static files (HTML, images) are sent by
  Tomcat's *DefaultServlet*, which keeps small files in memory. A *JSP* page is
  translated into Java code and compiled the first time it is requested; after
  that every request runs the compiled code.
- **Java heap and garbage collection.** Java objects live in the *heap*
  (`-Xms` start size, `-Xmx` maximum). The *garbage collector* (GC) regularly
  frees objects that are no longer used. During a GC *pause*, application
  threads stop for a moment.
- **JIT warm-up.** Java first runs code slowly, then compiles frequently used
  code into fast machine code (the "just-in-time" compiler). A freshly started
  Tomcat is therefore slower during its first seconds of load.
- **Sessions.** A session is memory Tomcat keeps for one user between requests,
  identified by the `JSESSIONID` cookie. It lives in the heap until it times out.
- **Counters (JMX).** Tomcat and Java publish live counters (busy threads, open
  connections, sessions, heap, GC) through JMX. The kit's page `status.jsp`
  prints them as plain text.

### The kit's own instance

RHEL can run several Tomcat *instances* from one installation. The kit uses
its own, so the normal Tomcat (`/etc/tomcat`, `tomcat.service`) is not changed:

| Item | Value |
|---|---|
| systemd service | `tomcat@perftest.service` (RHEL's `tomcat@.service` template) |
| Instance folder (CATALINA_BASE) | `/var/lib/tomcats/perftest/` |
| Configuration | `/var/lib/tomcats/perftest/conf/server.xml` (written by `start-tomcat.sh`) |
| Java settings | `/etc/sysconfig/tomcat@perftest` (heap, GC, Java version) |
| Test application | `/var/lib/tomcats/perftest/webapps/perf-test/` |
| Logs | `/var/lib/tomcats/perftest/logs/` (catalina, localhost, gc.log) |
| Address | `http://127.0.0.1:8080/perf-test/` (this machine only) |

---

## 3. How the tests work

```
   +------------------------+      HTTP over 127.0.0.1     +------------------------------+
   |  ApacheBench ("ab")    |  --------------------------> |  Tomcat  (tomcat@perftest)   |
   |  several processes,    |  <-------------------------- |  /perf-test/ test application|
   |  simulating N users    |        test pages            +------------------------------+
   +------------------------+                                  |              |
             |                                    status.jsp (JMX)    logs/gc.log
             |  requests/s, latency,             threads, connections,  GC pauses
             |  errors, MB/s                     sessions, heap, GC
             v                                                 |
      results/<date-time>/report.md  (PASS / FAIL / INFO per metric)
             ^
             |  CPU: systemd cgroup (or /proc);  memory: /proc/<pid>/smaps_rollup
```

- **Load generator:** ApacheBench (`ab`), from the RHEL `httpd-tools` package.
  It is started with a number of *clients* (simultaneous users). Each client
  sends a request, waits for the answer, and immediately sends the next one,
  for `DURATION` seconds (10 by default).
- **Several `ab` processes:** one `ab` process uses one CPU core and cannot
  send much more than about 50,000 requests/s. The clients are split over
  `AB_PROCESSES` processes running side by side (by default half of the CPUs,
  at most 8). The report adds up their counts. For the latency percentiles, the
  tables of all processes are merged, each weighted by the number of requests
  it completed.
- **Test application** `/perf-test/`, installed by `start-tomcat.sh`:

  | Page | What it is | Used by |
  |---|---|---|
  | `small.html` | 1 KB static file | static |
  | `hello.jsp` | small dynamic page (0.7 KB), Java code runs on every request | dynamic, latency, concurrency, keepalive, errors, connections, gc, cpu, memory |
  | `text.html` | 40 KB static HTML page | compression |
  | `large.bin` | 1 MB static file | transfer |
  | `slow.jsp?ms=100` | waits 100 ms before answering (like a database call) | threads |
  | `session.jsp` | creates a user session with 1 KB of data | sessions |
  | `status.jsp` | Tomcat and Java counters, as `name=value` lines | all measurements |
  | `health.txt` | "is Tomcat serving?" check | start and startup |

  Only the machine itself may open these pages (`RemoteAddrValve` in the
  application's `META-INF/context.xml`).
- **Server-side measurements**, taken once per second while load runs:
  - CPU time from the `tomcat@perftest.service` cgroup (`cpu.stat`), or from
    `/proc/<pid>/stat` when there is no systemd.
  - Process memory (PSS) from `/proc/<pid>/smaps_rollup`.
  - Busy threads, open connections, sessions and heap from `status.jsp`.
  - GC pauses from Java's GC log (`logs/gc.log`, switched on by the kit with
    `-Xlog:gc`).
- **Warm-up:** `WARMUP_SECONDS` (10) of unmeasured load run before the first
  measured test, so Java's JIT compiler has already optimised the hot code.
  The `startup` test restarts Tomcat, so it runs first and the warm-up after it.
- **Offline:** all traffic uses the loopback address `127.0.0.1`. No internet
  access and no second machine are needed.

---

## 4. The metrics in detail

Each metric below has the same five parts: **what** it is, **why** it matters,
**how** the test measures it, **how to read** the result, and **how to improve** it.

### 4.1 Startup time

**What.** The time from the start command (`systemctl start tomcat@perftest`)
until Tomcat serves its first page.

**Why.** It decides how long a service is unavailable after a restart, an
update or a crash, and how fast a new server can join a cluster. Java
applications are known for slow starts, so this is worth knowing.

**How it is measured.** Tomcat is stopped and started `STARTUP_ROUNDS` (3)
times. For each round the script measures the time until `health.txt`
answers, and then the extra time of the first request to `hello.jsp`.

**How to read it.** An empty Tomcat starts in about one second. Each web
application adds its own start time (a large Spring application can take
10-60 seconds). The first request to a JSP is slower: Java loads the page's
code on first use, and a JSP that was never compiled before is first
translated and compiled (seconds). `start-tomcat.sh` opens every test page
once, so the test measures the "already compiled" case.

**How to improve.**
- Remove web applications that are not needed (`webapps/`).
- Avoid scanning all JAR files for annotations: set
  `tomcat.util.scan.StandardJarScanFilter.jarsToSkip` in
  `conf/catalina.properties`, or `metadata-complete="true"` in `web.xml`.
- Pre-compile JSP pages at build time.
- A slow "Creation of SecureRandom instance" line in `catalina.*.log` means
  the system is short of randomness; RHEL 9 normally is not.

### 4.2 Static page throughput

**What.** The number of requests for a 1 KB static file that Tomcat answers
per second.

**Why.** It shows the raw speed of Tomcat's network layer (connector,
HTTP parsing, DefaultServlet) with almost no application work.

**How it is measured.** `CONCURRENCY` (50) clients request `small.html` with
keep-alive for `DURATION` seconds.

**How to read it.** Small static files come from Tomcat's memory cache, so this
is usually the highest rate in the report. If a server has to deliver many
static files, a web server such as Nginx or Apache httpd in front of Tomcat is
usually still faster and uses less memory.

**How to improve.** Keep-alive on (see 4.6); a larger resource cache
(`<Resources cacheMaxSize="...">` in `context.xml`) for many small files;
more CPU cores.

### 4.3 Dynamic page throughput

**What.** The number of requests for a small JSP page (`hello.jsp`) that
Tomcat answers per second. Every request runs the page's Java code and builds
a new answer.

**Why.** This is the main job of an application server. It is the nearest
thing to "how many users can the application serve".

**How it is measured.** Like 4.2, with `hello.jsp` instead of the static file.
The report also shows the dynamic rate as a percentage of the static rate.

**How to read it.** `hello.jsp` does very little work, so the number shows
Tomcat's own overhead per request (request objects, servlet call, response
writing). A real application page that queries a database is often 10 to
1000 times slower; its speed is set by the application, not by Tomcat.

**How to improve.** Turn JSP development mode off (`JSP_DEVELOPMENT_MODE=false`,
the kit's default; see section 6); keep the access log buffered or off;
give Tomcat enough heap so the GC does not run too often; profile the
application itself.

### 4.4 Latency (response time)

**What.** The time from sending a request until the full answer has arrived,
for the dynamic page, reported as average and percentiles.

**Why.** Users feel latency, not throughput. The slow requests matter most:
p99 = the time within which 99 of every 100 requests finish.

**How it is measured.** `CONCURRENCY` clients request `hello.jsp` with
keep-alive for `DURATION` seconds. `ab` records every request's time; the
tables of all `ab` processes are merged into `raw/latency-percentiles.csv`.

**How to read it.**

| Percentile | Meaning |
|---|---|
| p50 (median) | the typical request |
| p95 | 1 in 20 requests is slower than this |
| p99 | 1 in 100 requests is slower than this |
| slowest | the single worst request (often a one-off) |

In a Java server, a p99 far above p50 is often caused by garbage collection
pauses (compare with the `gc` test) or by JIT compilation early in a run.

**How to improve.** Fix GC first (heap size, see 4.13); avoid overload (see
4.5 and 4.9); keep-alive on; make sure the machine is not swapping.

### 4.5 Concurrency scaling

**What.** How throughput and latency change as the number of simultaneous
clients grows (`CONCURRENCY_LEVELS`: 1 10 50 100 200 500 1000).

**Why.** Real traffic comes in waves. A healthy server keeps its throughput
when more users arrive; only the waiting time grows.

**How it is measured.** The dynamic-page test is repeated at each level. The
verdict compares the throughput at the highest level with the best (peak)
throughput: at least `TARGET_SCALING_MIN_PCT` (50%) must remain.

**How to read it.** Requests/second should rise with clients until the CPUs
are busy, then stay flat. Above `maxThreads` (200) clients some requests wait
for a free worker thread; with fast pages the waits are short, so throughput
stays flat and only latency grows. A strong drop at high levels points to
lock contention in the application, GC pressure, or too few CPU cores.

**How to improve.** More CPU cores; fewer GC pauses; a `maxThreads` value that
suits the application (section 6).

### 4.6 Keep-alive effect

**What.** How much faster Tomcat is when clients reuse one TCP connection for
many requests (keep-alive), compared with a new connection for every request.

**Why.** Opening a TCP connection costs a network round trip and CPU time on
both sides. Browsers use keep-alive; some API clients and load balancers do not.

**How it is measured.** The dynamic-page test runs twice: with keep-alive
(`ab -k`) and without.

**How to read it.** A gain of 1.5-3x on localhost is normal. Over a real
network the gain is larger because every new connection costs a round trip.

**How to improve.** Keep keep-alive on. The Tomcat default
`maxKeepAliveRequests="100"` closes a connection after 100 requests; the kit
uses `-1` (no limit, `MAX_KEEPALIVE_REQUESTS`) because a benchmark client
reaches 100 requests in milliseconds. In production, 100-1000 is a good value.
`keepAliveTimeout` (20 s) decides how long an idle connection stays open.

### 4.7 Transfer rate (bandwidth)

**What.** How many megabytes per second Tomcat sends when clients download a
1 MB file.

**Why.** It matters for downloads, reports and file exports.

**How it is measured.** `CONCURRENCY` clients download `large.bin` repeatedly
for `DURATION` seconds.

**How to read it.** Tomcat sends files larger than 48 KB with **sendfile**:
the kernel copies the file directly from the page cache to the socket, without
passing it through Java. On localhost this reaches many GB/s. Over a real
network, the network card (for example 1 Gbit/s = about 118 MB/s) is the limit.

**How to improve.** Keep `useSendfile` on (the default). The DefaultServlet
setting `sendfileSize` (in KB, default 48) sets the size above which sendfile
is used.

### 4.8 Compression (gzip)

**What.** How much smaller text answers get with gzip, and how many requests
per second Tomcat still serves when it compresses.

**Why.** Compression saves network traffic and makes pages load faster over
slow links, but costs CPU time for every answer.

**How it is measured.** `CONCURRENCY` clients request the 40 KB `text.html`,
first without and then with the header `Accept-Encoding: gzip`. The report
compares answer size and requests/second.

**How to read it.** HTML text usually becomes 5-10x smaller. The drop in
requests/second shows the CPU cost; on a real network the smaller answers
often make the overall result faster.

**Tomcat detail.** Tomcat does **not** compress a file it sends with
sendfile. With the defaults this means static files **larger than 48 KB are
never compressed** (the kit's check: a 120 KB page came back uncompressed).
That is why the test page is 40 KB. To compress large static files, lower
their use of sendfile (DefaultServlet `sendfileSize`), or let a front-end web
server compress them. Dynamic pages (JSP, servlets) are never sent with
sendfile, so they are always compressed when large enough.

**How to improve.** `compression="on"` in the Connector (the Tomcat default is
`off`); `compressionMinSize` (2048 bytes) avoids wasting CPU on tiny answers;
`compressibleMimeType` lists the text types. Do not compress images, video or
ZIP files: they are compressed already.

### 4.9 Worker thread pool

**What.** How well Tomcat's pool of worker threads is used when every request
is slow, and how long requests wait for a free thread.

**Why.** Real application pages often wait for a database or another service.
While they wait, they hold a worker thread. When all `maxThreads` threads are
busy, new requests queue up. This is the most common capacity limit of a Java
application server.

**How it is measured.** `slow.jsp` waits `SLOW_REQUEST_MS` (100 ms) before
answering. The pool can then finish at most
`maxThreads x 1000 / SLOW_REQUEST_MS` = 200 x 10 = **2,000 requests/s**. The
test runs twice:

1. with half as many clients as threads (100): no request should wait;
2. with twice as many clients as threads (400): the pool is full and half of
   the requests wait in the queue.

The verdict: in step 2 Tomcat must reach at least `TARGET_THREAD_POOL_MIN_PCT`
(80%) of the pool limit, with no failed requests. The report also shows the
busy threads (from `status.jsp`) and the waiting time.

**How to read it.**

| Observation | Meaning |
|---|---|
| step 2 reaches ~95% of the limit, busy threads = maxThreads | the pool works; its size is the limit |
| p50 in step 2 about twice SLOW_REQUEST_MS | each request waited about one request-time in the queue |
| p99 in step 2 about 1 second or more | connection attempts were dropped at the start (acceptCount queue full) and retried after 1 s |
| well below 80% of the limit | CPU, locks or GC stop the threads from working |

**How to improve.** If requests wait in the queue *and* the CPU and database
still have spare capacity, raise `maxThreads`. Each thread costs memory (about
1 MB of stack reserved) and CPU switching, and the database must accept the
extra connections, so raise it step by step. Better still: make the slow calls
faster (database indexes, caching).

### 4.10 Sessions

**What.** How many new user sessions Tomcat creates per second, and how much
heap memory each session needs.

**Why.** Every logged-in user (or shopping cart) is a session in the heap
until it times out (30 minutes by default). Too many sessions exhaust the
heap. Clients that ignore the session cookie, such as monitoring scripts,
create a new session with every request and fill the memory.

**How it is measured.** All sessions of `/perf-test` are ended and Java
frees unused memory (full GC). Then `SESSIONS_TEST_COUNT` (50,000) requests go
to `session.jsp`, which stores `SESSION_DATA_BYTES` (1024) bytes in a new
session. `ab` never sends the session cookie back, so every request creates a
new session. Afterwards another full GC runs, and the heap growth is divided
by the number of live sessions. Finally all test sessions are ended again.

**How to read it.** A session with 1 KB of data needs about 1.5 KB of heap:
the rest is Tomcat's own bookkeeping (session object, ID, attribute map).
Memory estimate for production:
`users active within the session timeout x (Tomcat overhead + application data)`.
Example: 20,000 users x 50 KB of application data = about 1 GB of heap.
"Sessions rejected" must be 0 (it grows only if `maxActiveSessions` is set).

**How to improve.** Store less in the session; shorten `session-timeout`;
use `<%@ page session="false" %>` in pages that need no session (the JSP
default is to create one!); do not let monitoring checks create sessions.

### 4.11 Error rate (under overload)

**What.** The share of requests that fail (network errors, wrong answers, HTTP
errors such as 503) when far more clients arrive than Tomcat has threads.

**Why.** Under overload a good server gets slower but keeps answering
correctly. Failed requests mean lost users or broken transactions.

**How it is measured.** `ERROR_TEST_CONCURRENCY` (1000) clients, each opening
a new connection for every request, request `hello.jsp` for `DURATION`
seconds. The test also counts error answers reported by Tomcat itself and new
`SEVERE` lines in `catalina.*.log` and `localhost.*.log`.

**How to read it.** 0% is expected on a healthy host. The slowest requests may
take about 1 or 3 seconds: those are connection attempts that found the
`acceptCount` queue full and were retried by the client's TCP stack. They are
slow but not failed.

**How to improve.** Raise `acceptCount` (and the kernel limit
`net.core.somaxconn`, 4096 on RHEL 9) for sudden bursts of new connections;
raise `maxConnections` or the open-files limit if the logs show "Too many
open files"; add CPU or more Tomcat instances behind a load balancer.

### 4.12 Connections

**What.** The highest number of connections open at the same time, compared
with `maxConnections`.

**Why.** Each keep-alive client holds a connection. When `maxConnections` is
reached, Tomcat stops accepting new connections until one closes, and new
clients wait.

**How it is measured.** `CONNECTIONS_TEST_CLIENTS` (1000) clients request
`hello.jsp` with keep-alive for `DURATION` seconds; `status.jsp` reports the
open connections and busy threads every second. The verdict: peak at most
`TARGET_CONNECTIONS_MAX_PCT` (80%) of `maxConnections`, no failed requests.

**How to read it.** The report shows that 1000 open connections need far fewer
busy threads: with the NIO connector a waiting connection costs no thread.
The count includes the kit's own status connection.

**How to improve.** Raise `maxConnections` (and the open-files limit, set to
65536 by the kit's systemd drop-in); a shorter `keepAliveTimeout` closes idle
connections sooner.

### 4.13 Garbage collection (GC)

**What.** How often Java's garbage collector pauses the application, how long
the pauses are, and what share of time is spent in them.

**Why.** During a GC pause no request makes progress. Frequent or long pauses
show up as high p99 latency and lower throughput. A heap that is too small
leads to constant collections; a heap that is nearly full with live objects
leads to long "full GC" pauses and finally to `OutOfMemoryError`.

**How it is measured.** While `CONCURRENCY` clients request `hello.jsp` for
`DURATION` seconds, the script notes Java's uptime before and after. It then
reads every `Pause` line with a time stamp inside that window from
`logs/gc.log`, for example:

```
[53.080s][info][gc] GC(678) Pause Young (Normal) (G1 Evacuation Pause) 238M->49M(317M) 0.863ms
```

(`238M->49M(317M)` = heap in use before -> after the collection (heap size)).
Verdict: time in pauses at most `TARGET_GC_OVERHEAD_MAX_PCT` (5%), longest
pause at most `TARGET_GC_PAUSE_MAX_MS` (200 ms).

**How to read it.**

| Result | Meaning |
|---|---|
| many short "Pause Young" (< 10 ms), overhead of a few % | normal: short-lived objects are cleaned up cheaply |
| "Pause Full" lines | the heap was full; serious, look at heap size and live data |
| overhead > 10% | heap too small for the allocation rate |
| heap after GC grows run after run | the application keeps objects (possible memory leak) |

**How to improve.** Give Tomcat more heap (`HEAP_MAX_MB`); set
`HEAP_INITIAL_MB` equal to `HEAP_MAX_MB` on dedicated servers (no heap resizing);
reduce what the application keeps in memory (sessions, caches);
`-XX:MaxGCPauseMillis=<ms>` asks G1 for shorter pauses (default goal 200 ms).

### 4.14 CPU usage

**What.** The share of all CPUs used by the Tomcat process under load, and the
CPU time per 1000 requests.

**Why.** It shows how much headroom is left and what one request costs.

**How it is measured.** During the same run as 4.13, the CPU time of the
`tomcat@perftest.service` cgroup (or of the Java process) is read before and
after, and once per second. 100% means every CPU was fully busy with Tomcat.

**How to read it.** The CPU time includes Java's own background work: the JIT
compiler and the garbage collector. `ab` runs on the same machine and uses
CPU as well, so Tomcat seldom reaches 100% in this test. "CPU time per 1000
requests" is the most portable number: it hardly depends on the load level.

**How to improve.** Profile the application; lower GC work (4.13); turn off
JSP development mode; switch off or buffer access logging.

### 4.15 Memory usage

**What.** The memory (RAM) used by the Tomcat Java process, before and under
load, and the part used by the Java heap.

**Why.** It decides how many Tomcat instances or other services fit on a host,
and whether the host risks swapping or the kernel's out-of-memory killer.

**How it is measured.** PSS of the Java process from
`/proc/<pid>/smaps_rollup`, before the load and every second during it;
heap figures from `status.jsp`.

**How to read it.** A Java process's memory is roughly:

```
   process memory  =  heap (up to -Xmx)
                    + metaspace (loaded classes)
                    + code cache (JIT-compiled code)
                    + thread stacks (about 1 MB reserved per thread; used part counted)
                    + Java's own internal memory and libraries
```

The heap grows up to `-Xmx` and Java rarely gives memory back to the system,
so the process size follows the heap settings and the heap's history more
than the current load. Plan hosts with **-Xmx plus 200-500 MB** per Tomcat
instance.

**How to improve.** Size `-Xmx` from measured need (live heap after GC plus
room to work, typically 2-3x the live data); remove unused applications; do not
set `maxThreads` far higher than needed.

---

## 5. How the metrics relate to each other

```
   more clients --> threads busy --> requests wait in queue --> latency up
        |                                  |
        v                                  v
   more connections                 throughput reaches the pool limit (4.9)
                                    or the CPU limit (4.14)

   more requests/s --> more garbage --> more GC pauses (4.13) --> p99 latency up (4.4)
   more sessions   --> larger live heap --> longer GC, more memory (4.10, 4.15)
```

- **Throughput x latency = concurrency** (Little's law): 50 clients at 0.25 ms
  per request give 200,000 requests/s. If latency doubles at the same number of
  clients, throughput halves.
- **Pool limit = maxThreads / request time.** With 200 threads and 100 ms pages,
  Tomcat can never exceed 2,000 requests/s, however fast the CPU is.
- **GC and latency.** Every GC pause adds its length to every request running
  at that moment. Compare the longest GC pause with the p99 latency.
- **Heap, sessions and memory.** Sessions and caches are live data in the heap;
  more live data means more GC work and a larger process.

---

## 6. Tomcat and Java settings that affect performance

All values are in `settings.conf`; `start-tomcat.sh` writes them into the
instance. Tomcat defaults are shown for comparison.

| Setting (settings.conf) | Where it goes | Kit value | Tomcat default | Effect |
|---|---|---|---|---|
| `MAX_THREADS` | Connector `maxThreads` | 200 | 200 | requests processed at the same time |
| `MIN_SPARE_THREADS` | Connector `minSpareThreads` | 10 | 10 | threads kept ready when idle |
| `MAX_CONNECTIONS` | Connector `maxConnections` | 8192 | 8192 | open connections (busy + idle) |
| `ACCEPT_COUNT` | Connector `acceptCount` | 100 | 100 | queue for new connections |
| `CONNECTION_TIMEOUT_MS` | Connector `connectionTimeout` | 20000 | 60000 (20000 in RHEL's server.xml) | wait for a request line |
| `KEEPALIVE_TIMEOUT_MS` | Connector `keepAliveTimeout` | 20000 | = connectionTimeout | idle keep-alive connection lifetime |
| `MAX_KEEPALIVE_REQUESTS` | Connector `maxKeepAliveRequests` | -1 (no limit) | 100 | requests per keep-alive connection |
| `COMPRESSION` | Connector `compression` | on | off | gzip of text answers |
| `COMPRESSION_MIN_BYTES` | Connector `compressionMinSize` | 2048 | 2048 | smallest answer compressed |
| `TEST_ACCESS_LOG` | `AccessLogValve` in Host | off | on in RHEL's server.xml | one log line per request |
| `JSP_DEVELOPMENT_MODE` | JspServlet `development` (conf/web.xml) | false | true | checks JSP files for changes on every request |
| `HEAP_INITIAL_MB` / `HEAP_MAX_MB` | `-Xms` / `-Xmx` | 256 / 1024 | 1/64 and 1/4 of RAM | Java heap size |
| `GC_OPTIONS` | Java options | `-XX:+UseG1GC` | G1 (Java 17) | garbage collector |
| `SESSION_TIMEOUT_MINUTES` | application `web.xml` | 30 | 30 | how long idle sessions are kept |
| `JAVA_VERSION` | `JAVA_HOME` | 17 | RHEL `jre` link | Java version (RHEL 9.6: 1.8, 11, 17, 21) |

Other things that matter:

- **autoDeploy.** The kit sets `autoDeploy="false"`: Tomcat does not keep
  checking `webapps/` for changes. Turn it off in production too.
- **Open files.** Every connection is an open file. The kit's systemd drop-in
  sets `LimitNOFILE=65536` for the instance (Java itself also raises its soft
  limit to the hard limit).
- **Shutdown port.** The kit's `server.xml` has no shutdown port
  (`port="-1"`); systemd stops Tomcat with a TERM signal. This also avoids a
  port conflict with the normal Tomcat (port 8005).
- **SELinux.** Tomcat runs in the SELinux domain `tomcat_t`. In the RHEL 9
  policy that domain is unconfined, so SELinux does not block its ports or
  files. The instance folder is labelled `tomcat_var_lib_t` automatically
  (`restorecon`). Other ports than 8080 are best taken from `http_port_t` or
  `http_cache_port_t` (`semanage port -l | grep http`).

---

## 7. Limits of this test

- **Localhost only.** Client and server share one machine, CPUs and memory.
  There is no real network: transfer rates are far above any network card,
  and `ab` takes CPU time away from Tomcat.
- **A tiny test application.** The pages do almost no work. Real applications
  (databases, frameworks, templates) are much slower per request; this kit
  measures Tomcat's own capacity and behaviour, not an application's.
- **HTTP/1.0 client.** `ab` speaks HTTP/1.0 with keep-alive; no HTTPS, no
  HTTP/2, no browser behaviour (parallel downloads, caching).
- **Short runs.** 10-second runs show steady-state speed, not long-term
  effects such as memory leaks or heap fragmentation. For those, run single
  tests with a larger `DURATION` (for example `DURATION=600 ./test-tomcat.sh gc`).
- **Sessions from one kind of client.** Real session sizes depend on the
  application; the test shows Tomcat's overhead per session.
- **Targets are starting points.** Adjust section 6 of `settings.conf` to the
  hardware and the service level you need.

---

## 8. Glossary

| Term | Meaning |
|---|---|
| acceptCount | queue of new connections waiting to be accepted by Tomcat |
| ab (ApacheBench) | command-line HTTP load generator from `httpd-tools` |
| CATALINA_BASE | folder of one Tomcat instance (conf, webapps, logs, work, temp) |
| CATALINA_HOME | folder of the Tomcat installation (`/usr/share/tomcat`) |
| Connector | the part of Tomcat that receives HTTP requests on a port |
| DefaultServlet | Tomcat's built-in servlet that sends static files |
| GC (garbage collection) | Java's automatic freeing of memory that is no longer used |
| GC pause | moment when the garbage collector stops application threads |
| G1 | "Garbage First", the default garbage collector of Java 17 |
| Heap | memory area where Java objects live; limited by `-Xmx` |
| JIT | "just-in-time" compiler: turns frequently used Java code into machine code |
| JMX / MBean | Java's standard interface for live counters and management |
| JSP | JavaServer Page: HTML with Java code, compiled by Tomcat |
| Keep-alive | reusing one TCP connection for several HTTP requests |
| Latency | time from sending a request until the answer has arrived |
| maxConnections | largest number of connections Tomcat keeps open |
| maxThreads | largest number of worker threads (requests processed at once) |
| NIO | "non-blocking I/O" connector: few threads watch many connections |
| p50 / p95 / p99 | percentiles: 50 / 95 / 99 of every 100 requests were faster |
| PSS | proportional set size: memory of a process, shared pages divided fairly |
| sendfile | kernel feature that copies a file to the network without the application |
| Servlet | Java class that answers HTTP requests inside Tomcat |
| Session | per-user memory kept by Tomcat between requests (`JSESSIONID` cookie) |
| Throughput | requests answered per second |
