# Chrony NTP Server - Performance Metrics

This document explains the performance metrics of the Chrony NTP server on
RHEL 9.6 (`chronyd` from the `chrony` package): what each metric means, why
it matters, how `test-chrony.sh` measures it, and what to change when a
result is poor.

---

## Contents

1. [NTP and Chrony in one minute](#1-ntp-and-chrony-in-one-minute)
2. [The metrics at a glance](#2-the-metrics-at-a-glance)
3. [How the tests work](#3-how-the-tests-work)
4. [The metrics in detail](#4-the-metrics-in-detail)
   - [4.1 Startup time](#41-startup-time)
   - [4.2 Served time quality](#42-served-time-quality)
   - [4.3 Throughput](#43-throughput)
   - [4.4 Response time](#44-response-time)
   - [4.5 Loss](#45-loss)
   - [4.6 Accuracy (time offset seen by clients)](#46-accuracy-time-offset-seen-by-clients)
   - [4.7 Many clients](#47-many-clients)
   - [4.8 Rate limiting](#48-rate-limiting)
   - [4.9 CPU per request](#49-cpu-per-request)
   - [4.10 Memory per client](#410-memory-per-client)
5. [How the metrics relate to each other](#5-how-the-metrics-relate-to-each-other)
6. [How much performance do you need?](#6-how-much-performance-do-you-need)
7. [Chrony settings that affect performance](#7-chrony-settings-that-affect-performance)
8. [Serving time on an offline network](#8-serving-time-on-an-offline-network)
9. [Limits of this test](#9-limits-of-this-test)
10. [Glossary](#10-glossary)

---

## 1. NTP and Chrony in one minute

**NTP** (Network Time Protocol) keeps the clocks of many machines the same.
A client sends a small **request** (one UDP packet, 48 bytes, port 123) to a
server. The server writes its time into the **answer** and sends it back.

```
   NTP client                                             NTP server (chronyd)
  +-----------------------------+                        +------------------------------+
  |                             |   request  (UDP 123)   |                              |
  |  T1  send request  ---------------------------------->  T2  request received        |
  |                             |                        |      ... look up the time ...|
  |  T4  answer received <-----------------------------------  T3  answer sent          |
  |                             |   answer: T1, T2, T3   |                              |
  +-----------------------------+                        +------------------------------+

     round-trip delay  = (T4 - T1) - (T3 - T2)
     offset (how far the client's clock is off) = ((T2 - T1) + (T3 - T4)) / 2
```

A few facts explain almost all NTP server performance:

- **One request, one answer, no connection.** NTP uses UDP. There is no
  login and no session: each request is answered on its own, in a few
  microseconds. This is why one server can serve a very large number of
  clients.
- **Clients ask rarely.** A client asks every **64 to 1024 seconds**
  (chrony's and ntpd's default poll range). 10,000 clients polling every
  64 seconds send only about 160 requests per second.
- **The answer must be fast *and* even.** The client assumes the request
  and the answer took the **same time** on the way. Any difference between
  the two directions (the *asymmetry*) turns directly into a time error of
  half that difference. A server that answers slowly now and then makes
  clients' clocks wobble.
- **chronyd is one process with one thread.** It uses one CPU core. Its
  capacity is how many packets that one core can receive, time-stamp and
  answer per second.
- **The server remembers its clients.** For rate limiting and for
  `chronyc clients`, chronyd keeps a small record for each client address,
  in a table of limited size (`clientloglimit`).
- **Clients check the answer.** Every answer says how good the server's
  time is (leap status, stratum, root delay, root dispersion). Clients ignore
  a server that says it is not synchronized (section 4.2).

---

## 2. The metrics at a glance

| # | Metric | Unit | Question it answers | Better | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast does the server give usable time after a (re)start? | lower | `startup` |
| 2 | Served time quality | stratum, leap | Will clients accept the server's time? | synchronized, low stratum | `quality` |
| 3 | Throughput | answers/s | How many requests per second can the server answer? | higher | `throughput` |
| 4 | Response time | ms | How long does one answer take in normal use? | lower, even | `latency` |
| 5 | Loss | % | How many requests get no usable answer at normal load? | lower | `loss` |
| 6 | Accuracy | µs | How much time error does the server add for its clients? | lower | `accuracy` |
| 7 | Many clients | clients | How many different clients can be served, still fast and complete? | higher | `clients` |
| 8 | Rate limiting | % answered | Is a flooding client slowed down, while normal clients are not? | flooder low, normal 100% | `ratelimit` |
| 9 | CPU per request | µs | How much processor time does one answer cost? | lower | `cpu` |
| 10 | Memory per client | bytes | How much RAM does the server use per client it remembers? | lower | `memory` |

Run one metric with `./test-chrony.sh <test name>`, or all of them with
`./test-chrony.sh`.

---

## 3. How the tests work

```
  +------------------------------------------+               +--------------------------------------+
  |  lib/ntpload.py  (the "clients")         |               |  test chronyd (chronyd-perf-test)    |
  |                                          |               |  /etc/chrony-perf-test.conf          |
  |  sends each request FROM a different     |   loopback    |  local stratum 10, never changes the |
  |  address: 127.10.0.1, 127.10.0.2, ...    | ==== UDP ===> |  clock (-x)                          |
  |  and measures T1..T4 of every answer     |  127.0.0.1    |  answers on 127.0.0.1:11123 only     |
  +------------------------------------------+    :11123     +--------------------------------------+
         |                                                             |
         |  answers/s, response time, offset,                          |  CPU (/proc/<pid>/stat), memory (VmRSS),
         |  loss, stratum / leap status                                |  chronyc serverstats / clients,
         v                                                             v  socket drops (/proc/net/udp)
                     results/<date-time>/report.md   (PASS / FAIL per metric)
```

- **A separate test server.** `start-chrony.sh` runs a **second** chronyd
  with its own configuration (`/etc/chrony-perf-test.conf`) and its own
  service (`chronyd-perf-test.service`). The normal `chronyd.service` and
  `/etc/chrony.conf` are not touched, so the host keeps its own time as before.
- **It never changes the clock.** The test chronyd is started with `-x`:
  it answers NTP requests using the system clock, but never adjusts it.
- **Only this machine can reach it.** It answers on `127.0.0.1`, UDP port
  `11123` (not 123, to avoid any clash with a real NTP server on the host).
  With SELinux, the kit gives that port the label `ntp_port_t`, which is
  what chronyd may use.
- **Offline time source.** The test server has no upstream servers. It
  serves its own clock as `local stratum 10`, exactly as an NTP server on an
  offline network does (section 8).
- **Many clients from one machine.** Every address `127.x.y.z` belongs to
  this machine. `lib/ntpload.py` (Python 3, standard library only) sends each
  request **from** a different such address, so chronyd sees thousands of
  different clients, as on a real network, and keeps a record for each one.
- **Real NTP packets.** Every request is a normal NTP version 4 client
  request. Every answer is checked (mode, origin timestamp, leap status,
  stratum, "Kiss-o'-Death"), and the four timestamps T1..T4 are used exactly
  as an NTP client uses them. T4 is the kernel's receive time of the answer
  (`SO_TIMESTAMPNS`), so Python's own delays do not count.
- **Server-side measurements**, before and after each run:
  - chronyd's CPU time (`/proc/<pid>/stat`) and memory (`VmRSS` in `/proc/<pid>/status`);
  - `chronyc serverstats`: requests received, requests ignored on purpose
    (rate limiting), client records replaced because the table was full;
  - `chronyc clients`: the clients chronyd remembers;
  - the `drops` column of chronyd's socket in `/proc/net/udp`: requests the
    kernel threw away because chronyd was too busy to read them.
- **Loads.** The tests use four kinds of load:

| Load | What it does | Used by |
|---|---|---|
| steady | `LOAD_RATE` (1000) requests/s at random moments from `LOAD_CLIENTS` (500) addresses, `DURATION` (10) s | latency, loss, accuracy, cpu |
| full speed | 2 senders, each keeping 64 requests in flight: every answer lets it send the next request | throughput, cpu |
| steps | 5000 requests/s spread over 10, 1000, 10,000, 50,000 addresses | clients |
| flood + normal | 1 address at 200 requests/s, and 200 addresses at 20 requests/s together | ratelimit |

- The steady load runs once and is shared by the `latency`, `loss`,
  `accuracy` and `cpu` tests; the `cpu` test also re-uses the `throughput` run.
- The `startup`, `ratelimit` and `memory` tests **restart the test server**.
  (Only the test server; the normal chronyd keeps running.)

---

## 4. The metrics in detail

Each metric has the same five parts: **what** it is, **why** it matters,
**how** the test measures it, **how to read** the result, and **how to improve** it.

### 4.1 Startup time

**What:** the time from the command that starts chronyd until the first
answer that clients can **use**: an answer whose leap status is not "not
synchronized" and whose stratum is not 0.

**Why:** a restart happens at every update of `chrony`, every reboot and
every change of the configuration. Until the server gives usable time,
clients cannot correct their clocks. Most clients just keep their last
correction and try again later, so a short outage does no harm. A long one,
or a server that restarts often, lets clocks drift apart.

**How it is measured:** the test server is stopped and started
`STARTUP_ROUNDS` times (5). Before each start, a small program starts asking
for the time every 2 ms. It notes the first answer of any kind, and the
first answer with synchronized time. The report also shows when the start
command itself returned.

**How to read it:**

| Result | Meaning |
|---|---|
| Under 100 ms | normal for a server with `local stratum` (no upstream servers): its time is usable at once |
| 1 - 2 minutes | normal for a server that follows **upstream** NTP servers: it must first measure them several times before it trusts them |
| Over 2 s with `local stratum` | investigate: `journalctl -u chronyd-perf-test`; slow name lookups in the configuration, or a slow disk (`driftfile`, `dumpdir`) |
| Never synchronized | the `local` directive is missing, or `local` waits for sources that cannot be reached (`local ... waitsynced`) |

**How to improve:** on an offline network, use `local stratum N` on the
time server (section 8). With upstream servers, use `iburst` on the `server`
lines (4 quick measurements at start instead of one per poll interval), and
`dumpdir` with the `-r` option so chronyd reloads what it knew before the
restart.

---

### 4.2 Served time quality

**What:** the fields every NTP answer carries about the server's own time:

| Field | Meaning | Good value |
|---|---|---|
| Leap status | 0-2: normal, or a leap second is coming; **3: "not synchronized"** | 0 |
| Stratum | steps from a reference clock: 1 = has a GPS/radio clock, 2 = follows a stratum-1 server, ... **16 = unusable** | as low as the setup allows |
| Reference ID | what the server follows (an IP address, or `127.127.1.1` for its own clock) | as expected |
| Root delay | round-trip time from the server to the reference clock | small |
| Root dispersion | the server's own estimate of its maximum error | small |
| Precision | how finely the server can read its clock | under 1 µs |

**Why:** clients **choose** their servers by these fields. They ignore a
server with leap status 3 or stratum 16, and they prefer servers with a
lower *root distance* (root delay / 2 + root dispersion). A server that
answers quickly but says "not synchronized" is useless.

**How it is measured:** one request is sent and every field of the answer
is shown. `chronyc tracking` of the test server is saved as well.
**PASS** when the leap status is not 3 and the stratum is between 1 and
`TARGET_MAX_STRATUM` (15).

**How to read it:** with `local stratum 10` the answer says stratum 10,
reference `127.127.1.1`, root delay and dispersion near 0. This means only
that the server trusts its **own** clock: it is as right as that clock. On
an offline network nobody can check it against true time, so it is the
administrator's job to set the clock correctly (section 8).

**How to improve:** give the server a better source (a GPS or PTP reference
clock, or a stratum-1 server inside the offline network). In an offline
network without such a source, keep the stratum the same on all servers so
clients do not prefer one wrong clock over another.

---

### 4.3 Throughput

**What:** the highest number of NTP requests the server answers per second.

**Why:** it decides how many clients one server can serve, and how much a
flood (a misconfigured client, a script, an attack) can take before real
clients stop getting answers. Because each client asks only every 64 - 1024
seconds, one chronyd can serve a very large network:

    clients one server can serve  ≈  throughput x poll interval

For example, 200,000 answers/s x 64 s = 12.8 million clients - far more
than any company network. In practice the limit is almost never throughput,
but **bursts**: many clients that start at the same moment (after a power
failure, or clients that all poll on the minute).

**How it is measured:** 2 sender processes, each keeping 64 requests in
flight, send for `DURATION` seconds as fast as chronyd answers. Every answer
lets a sender send the next request, so chronyd always has work, but is not
drowned. The result is the number of answers per second. The report also
shows chronyd's CPU use during the run (up to 100% of one core) and the
senders' own CPU, so you can see which side was the limit.

**How to read it:**

| Result | Meaning |
|---|---|
| chronyd near 100% of one core | this is the server's real limit on this CPU |
| senders at 90%+, chronyd lower | the load generator was the limit (the report says so); try `THROUGHPUT_WORKERS=4` |
| far below 20,000/s | investigate: a very slow CPU, a virtual machine with CPU limits, `-F` / security modules, or many `allow`/`deny` rules |

**How to improve:** a faster CPU core (chronyd uses only one). Avoid huge
`allow`/`deny` lists. For more than one server's worth of load, add more
servers and give clients several `server` lines (or a `pool`): clients then
share the load.

---

### 4.4 Response time

**What:** how long a client waits for an answer: from sending the request
(T1) until the answer arrives (T4), at a normal, steady load. Reported as
percentiles: p50 (median), p95, p99, p99.9 and the slowest.

The report also shows the time spent **inside chronyd** (T3 - T2, from the
answer's own timestamps): from the moment the kernel received the request
until chronyd wrote the answer's transmit time.

**Why:** for NTP the response time matters in two ways:

1. **Evenness matters most.** A client assumes both directions took equally
   long. If some answers are delayed inside the server (after T2 but before
   the packet leaves), that delay becomes a time error. NTP clients reduce
   this by preferring the measurements with the **shortest** delay, but a
   server that is often slow still makes them less accurate.
2. **Timeouts.** A client gives up after a few seconds, so only very large
   delays cause lost answers.

**How it is measured:** the steady load (1000 requests/s at random moments
from 500 addresses, 10 s). The kernel receive time of each answer is used as
T4, so the load generator's own delays do not count.

**How to read it:**

| p99 on this machine (loopback) | Meaning |
|---|---|
| under 0.1 ms | normal for a server that is not overloaded |
| 0.1 - 1 ms | the machine is busy (other programs, virtual machine) |
| over 1 ms | investigate: CPU starved, chronyd swapped out, power saving (deep C-states) |

On a real network, add the network's round-trip time (0.1 - 0.5 ms in a
LAN).

**How to improve:** keep the server machine from being overloaded; give
chronyd a higher priority (`-P 1` or more in `/etc/sysconfig/chronyd`, a
real-time priority) and lock its memory (`-m`) so it is never swapped out.
Use hardware timestamping (`hwtimestamp`) on network cards that support
it: the timestamps are then taken by the card and the server's own delay no
longer matters (section 7).

---

### 4.5 Loss

**What:** the share of requests that got **no usable answer** during the
steady load:

- no answer within 1 second (lost), or
- a "Kiss-o'-Death" answer (stratum 0: the server refuses, e.g. `RATE`), or
- an answer saying "not synchronized".

**Why:** NTP clients survive a lost answer (they ask again at the next
poll), but a server that often drops requests makes clients switch servers
or poll less successfully, and it is a sign that the server is overloaded.
At normal load there should be **no** loss at all.

**How it is measured:** from the steady load. The report also shows
**where** requests were lost:

| Place | Meaning |
|---|---|
| chronyd's receive buffer overflowed | chronyd could not read requests as fast as they came (overload) |
| chronyd ignored them on purpose | rate limiting, or an `allow`/`deny` rule |
| the load generator's own buffer overflowed | the test machine itself was the limit (not the server) |

**How to read it:** **0%** is the only good result at normal load. Any loss
at 1000 requests/s means the server is starved of CPU or misconfigured.

**How to improve:** check `allow` rules (a client that is not allowed gets
no answer), `ratelimit` settings (section 4.8), and the machine's load.

---

### 4.6 Accuracy (time offset seen by clients)

**What:** the **offset** each client computes from an answer:
offset = ((T2 - T1) + (T3 - T4)) / 2 - how far the server's time seems to
be from the client's own.

**Why:** this is what NTP is for. Whatever error the server adds to its
answers ends up in the clients' clocks. The main sources of error inside the
server are:

- **when the timestamps are taken:** T2 should be the moment the request
  arrived, T3 the moment the answer leaves. chronyd uses the kernel's receive
  timestamp for T2. T3 is taken just before sending, so the time needed to
  send adds a small, even error;
- **uneven delays:** a request that waits before it is read makes T2 late
  (chronyd corrects this with the kernel timestamp), an answer that waits
  after T3 makes the client think the answer came back more slowly.

**How it is measured:** the load generator and the test server run on the
**same machine and read the same clock**, so the true offset is exactly 0.
Every microsecond measured is error added by the server and the way it is
reached (the kernel, the loopback interface). The report shows the smallest,
median and largest offset, and the 95th / 99th percentile of the absolute
value. **PASS** when p99 <= `TARGET_OFFSET_P99_MAX_US` (100 µs).

**How to read it:**

| p99 of the absolute offset | Meaning |
|---|---|
| under 10 µs | excellent; the server adds almost no error |
| 10 - 100 µs | good; normal on a busy machine or a virtual machine |
| over 100 µs | investigate: overload (see response time), power saving, a slow virtual CPU |

A real network adds its own asymmetry: typically tens of µs in a switched
LAN, up to around 1 ms across routers and firewalls.

**How to improve:** as for response time (4.4). Hardware timestamping
(`hwtimestamp *`) and the **interleaved mode** (clients use `xleave` on their
`server` line) let chronyd give clients the *exact* time the previous answer
left the network card; together they bring errors down to around 1 µs.

---

### 4.7 Many clients

**What:** the highest number of **different** clients (addresses) the
server served with no failed request and p99 response time within target,
at a fixed total load (`CLIENT_STEPS_RATE`, 5000 requests/s).

**Why:** chronyd keeps a record for each client address. More clients means
a bigger table to search, and once the table is full (`clientloglimit`),
every new client replaces an old record. A server must stay fast and answer
**everyone**, however many clients there are.

**How it is measured:** the same load, spread over 10, 1000, 10,000 and
50,000 client addresses in turn (`CLIENT_STEPS`). For each step: p50 and p99
response time, failed requests, and how many client records chronyd had to
**replace** because its table was full.

**How to read it:** chronyd is designed so that the table lookup takes the
same time for 10 or 50,000 clients: the response time should not grow from
step to step. A non-zero "records replaced" value is **not** an error: those
clients are answered normally. It only means chronyd cannot remember all of
them, which matters for rate limiting (4.8) and for `chronyc clients`.

**How to improve:** raise `clientloglimit` when you want chronyd to
remember (and rate-limit) every client (section 4.10). Use `noclientlog`
when you need neither rate limiting nor `chronyc clients`: then chronyd keeps
no records at all.

---

### 4.8 Rate limiting

**What:** whether chronyd slows down a client that sends far too many
requests, **without** hurting normal clients. Two results:

- the share of the flooding client's requests that were still answered
  (lower is better);
- the share of the normal clients' requests that were answered (must be 100%).

**Why:** an NTP server answers anyone it allows, and its answer is as big as
the request. A broken client (a script in a loop, a misconfigured device
polling every second) or an attacker can waste its capacity. Worse, an
attacker can send requests with a **faked source address**, so the server
floods the victim with answers. Rate limiting protects against both.

**How it is measured:** the test server is restarted with
`ratelimit interval 3 burst 8 leak 2` (`RATELIMIT_LINE`). For `DURATION`
seconds, one client address sends 200 requests/s while 200 normal clients
send 20 requests/s together (each about one request every 10 s, which is
still a busy client: real clients poll every 64 s or more). Then the server
is restarted without rate limiting.

**How to read it:**

| Setting | Meaning |
|---|---|
| `interval 3` | each client may send 1 request per 2^3 = 8 s on average |
| `burst 8` | ... but up to 8 requests at once (clients send 4 quick requests at start with `iburst`) |
| `leak 2` | of the requests over the limit, about 1 in 2^2 = 4 is still answered |

With `leak 2`, about **25%** of the flood is still answered **on purpose**:
if an attacker fakes a real client's address, that real client is slowed
down, not cut off completely. **PASS** when the flooder got at most
`TARGET_RATELIMIT_FLOOD_MAX_PCT` (35%) and normal clients at least
`TARGET_RATELIMIT_NORMAL_MIN_PCT` (100%).

If normal clients lose answers, the limit is too strict for how often they
poll (or many clients share one address behind NAT). If the flooder gets
most answers, the client table may be too small to remember it (see 4.10).

**How to improve:** add `ratelimit` to any NTP server that serves a large or
untrusted network (it is **off** by default in RHEL's `/etc/chrony.conf`).
Choose `interval` below the shortest poll interval of your clients
(`minpoll`): for clients that poll every 64 s (2^6), `interval 3` to `5` is
safe. Many clients behind one NAT address need a higher `burst`. Make
`clientloglimit` large enough to remember every client.

---

### 4.9 CPU per request

**What:** chronyd's CPU time (user + system) per answered request, at full
speed; and chronyd's CPU use at the steady load.

**Why:** it tells how much of the machine the NTP service needs, and what
throughput to expect on a slower CPU: throughput ≈ 1,000,000 µs / CPU per
request. On a machine that also runs other services, it shows that NTP costs
almost nothing.

**How it is measured:** chronyd's CPU time is read from `/proc/<pid>/stat`
before and after the `throughput` and the steady run, and divided by the
number of answers.

**How to read it:** a few µs per answer is normal on a current CPU. At a
low rate the cost per answer is higher (chronyd wakes up for each request),
but the total is tiny: well under 1% of one core at 1000 requests/s. To size
a server:

    CPU needed (share of one core)  ≈  clients / poll interval  x  CPU per request

**How to improve:** the same as for throughput (4.3). Do not run a busy
server with debug logging (`chronyd -d -d`).

---

### 4.10 Memory per client

**What:** how much memory chronyd uses for each client it remembers, and
how much it uses right after start.

**Why:** chronyd's memory is small (a few MB) and **cannot grow without
limit**: the client table never gets bigger than `clientloglimit` (524,288
bytes = 512 KB by default). What this metric really tells you is **how many
clients chronyd can remember** with a given limit - which is what rate
limiting needs.

**How it is measured:** the test server is restarted (empty table), then
`MEMORY_CLIENTS` (20,000) new client addresses each send one request.
chronyd's resident memory (`VmRSS`) is read before and after, and divided by
the number of clients `chronyc clients` lists.

**How to read it:** about 100 bytes of real memory per client is normal.
The table itself reserves **128 bytes per client**, so the default limit of
512 KB holds exactly **4,096** clients; any further client replaces an older
record (the report shows how many replacements happened). The memory right
after start (about 3 - 4 MB) is chronyd's normal size.

**How to improve:** to remember N clients, set

    clientloglimit  ≈  N x 128 bytes

For example, `clientloglimit 4194304` (4 MB) remembers 32,768 clients. The
memory is used only as clients appear, never more than the limit. Or use
`noclientlog` if you need neither rate limiting nor `chronyc clients`.

---

## 5. How the metrics relate to each other

```
      startup  ---->  quality  ---->  clients accept the server at all
                                             |
      throughput  <---- CPU per request     |  (how many requests fit)
            |                                v
            +-----> response time ----> accuracy (the time error clients get)
            |             |
            v             v
          loss  <---- overload, rate limiting
                              ^
      many clients ---> client table (memory per client, clientloglimit)
```

- **Throughput and CPU per request** are two views of the same thing:
  throughput ≈ 1 s / CPU per request.
- **Response time and accuracy** go together: uneven delays inside the
  server become time errors in the clients.
- **Loss** appears when the load comes near the throughput limit, or when
  rate limiting refuses requests.
- **Many clients, rate limiting and memory** are tied by `clientloglimit`:
  rate limiting only works for clients chronyd remembers.

---

## 6. How much performance do you need?

Work it out from the number of clients and how often they ask:

| Your network | Requests per second | Needed |
|---|---|---|
| 500 clients, poll 64 s | ~8 /s | any server |
| 10,000 clients, poll 64 s | ~160 /s | any server |
| 10,000 clients start at the same moment (after a power failure), 4 quick requests each (`iburst`) | 40,000 in a few seconds | throughput above ~20,000/s, or a few seconds of patience |
| 100,000 devices, poll 16 s (aggressive IoT) | ~6,250 /s | one server; set `clientloglimit` for rate limiting |

So in a company network, **throughput is rarely the problem**. What matters
more is:

- **quality**: the server says it is synchronized and has a sensible stratum;
- **accuracy and even response times** under the machine's normal load;
- **no loss** at normal load;
- **rate limiting** configured, with a client table big enough for all clients.

The default targets in `settings.conf` follow this: a modest throughput
(20,000/s), strict loss (0%) and accuracy (100 µs) targets.

---

## 7. Chrony settings that affect performance

All in `/etc/chrony.conf` (or `/etc/sysconfig/chronyd` for command-line
options). See `man chrony.conf` and `man chronyd`.

| Setting | Default on RHEL 9 | Effect |
|---|---|---|
| `allow <subnet>` | none (no NTP server) | chronyd answers NTP clients only when at least one `allow` is set |
| `local stratum N` | off | serve time without any upstream server (offline networks, section 8) |
| `ratelimit interval I burst B leak L` | off | slow down clients that ask too often (4.8) |
| `clientloglimit <bytes>` | 524288 | memory for client records; 128 bytes per client, so 4,096 clients by default (4.10) |
| `noclientlog` | off | keep no client records at all (no rate limiting, no `chronyc clients`) |
| `hwtimestamp <interface>` | off | timestamps taken by the network card: best accuracy (4.6) |
| `-F 2` (in OPTIONS) | on | system-call filter; small safety gain, very small CPU cost |
| `-P <priority>` | off | real-time scheduling priority: even response times on a busy machine |
| `-m` | off | lock chronyd's memory in RAM: never swapped out |
| `port N` | 123 | the NTP port; the kit uses 11123 for its test server |
| `bindaddress` | all addresses | answer only on one address/interface |
| `ntsserverkey`, `ntsservercert` | off | NTS (authenticated NTP): extra CPU per request and a TLS key exchange (TCP 4460) |

Remember to open the firewall for a real server:
`firewall-cmd --permanent --add-service=ntp && firewall-cmd --reload`.

---

## 8. Serving time on an offline network

An air-gapped network has no internet time servers. Its clocks can still be
kept **the same**, which is what most services (Kerberos, TLS, logs, file
servers, clusters) need; whether they are also **right** depends on how the
time server's clock was set.

A minimal time server configuration for an offline network
(`/etc/chrony.conf` on the time server):

```
# No upstream servers: serve this machine's own clock.
local stratum 10

# Who may ask (your offline network).
allow 10.0.0.0/8

# Protect the server, and remember enough clients for that.
ratelimit interval 3 burst 8 leak 2
clientloglimit 4194304         # 4 MB: remembers 32,768 clients

# Keep the clock's frequency correction across restarts.
driftfile /var/lib/chrony/drift
```

Advice:

- **Set the clock right once**, by hand from a trusted watch or a GPS
  receiver (`date -s "2026-09-18 12:00:00"`, then `hwclock --systohc`), then
  let chronyd keep it steady. (`chronyc settime` also works, but only with
  the `manual` directive in chrony.conf.) Without an outside reference it will slowly drift
  (typically a few seconds per week); clients follow it, so they stay equal.
- **Two or three time servers** make the service survive one failure. Use
  `local stratum 10 orphan` on each of them and let them peer with each other
  (`peer` lines): they then agree on one time instead of each serving its own.
- **Clients** use `server <time server> iburst` lines (two or three servers).
- **A GPS or PTP reference clock** (`refclock` directive) turns an offline
  time server into a stratum-1 server with correct time.

This kit's test server uses exactly `local stratum 10`, so its results apply
to such a server.

---

## 9. Limits of this test

- **No real network.** Clients and server run on one machine over the
  loopback interface. A real network adds delay (0.1 - 0.5 ms in a LAN),
  loss and, most important for NTP, **asymmetry**. The results show what
  chronyd itself can do.
- **The load generator shares the machine.** It takes CPU from chronyd,
  especially in the `throughput` test. The report warns when the generator
  itself was the limit.
- **Accuracy is measured against the same clock.** This shows the error
  chronyd and the kernel add; it cannot show whether the server's clock is
  **right** - on an offline network nothing can.
- **No upstream servers.** The test server serves its own clock. The time a
  real server needs to synchronize with upstream servers (minutes) and its
  tracking quality are not measured.
- **NTS (authenticated NTP) is not tested.** It adds a TLS key exchange and
  more CPU per request.
- **Short runs.** Each run lasts `DURATION` (10) seconds. Rare events (a
  slow answer once an hour) need longer runs: `DURATION=300 ./test-chrony.sh latency`.
- **The normal chronyd runs next to it.** The test server shares the machine
  with the host's own chronyd (and anything else running); results are best
  on a quiet machine.

---

## 10. Glossary

| Term | Meaning |
|---|---|
| **NTP** | Network Time Protocol (RFC 5905); keeps clocks the same over a network, UDP port 123 |
| **chronyd / chronyc** | the NTP daemon of RHEL, and its command-line control program |
| **Poll interval** | how often a client asks a server, 2^N seconds (64 - 1024 s by default) |
| **Offset** | how far a clock is from the server's clock |
| **Round-trip delay** | the time a request and its answer spent travelling (not in the server) |
| **Asymmetry** | the difference between the request's and the answer's travel time; half of it becomes a time error |
| **Stratum** | steps from a reference clock: 1 = has one, 2 = follows a stratum-1 server, ... 16 = unusable |
| **Leap status** | 0-2 normal / leap second coming; 3 = "not synchronized" |
| **Root delay / dispersion** | the server's delay to, and estimated error against, its reference clock |
| **Local stratum** | chronyd serves its own clock when it has no better source (`local` directive) |
| **Orphan mode** | several local-stratum servers agree on one of them as the source (`local orphan`) |
| **Kiss-o'-Death (KoD)** | an answer with stratum 0 that tells the client to stop or slow down (e.g. `RATE`) |
| **Rate limiting** | answering only some requests of a client that asks too often (`ratelimit`) |
| **Client log** | chronyd's table of client records (`clientloglimit`, `chronyc clients`) |
| **Interleaved mode** | a mode (`xleave`) where the server sends the exact transmit time of its previous answer |
| **Hardware timestamping** | the network card notes when a packet arrives or leaves (`hwtimestamp`) |
| **NTS** | Network Time Security: authenticated NTP (RFC 8915) |
| **p50 / p95 / p99** | percentiles: 50 / 95 / 99 of 100 values were at or below this |
| **-x** | chronyd option: never change the system clock (used for the test server) |
