# Kerberos Server - Performance Metrics

This document explains the performance metrics of the MIT Kerberos server
(`krb5kdc` and `kadmind`, from the RHEL 9.6 `krb5-server` package): what each
one means, why it matters, how `test-kerberos.sh` measures it, and what to
change when a result is poor.

---

## Contents

1. [Kerberos in one minute](#1-kerberos-in-one-minute)
2. [The metrics at a glance](#2-the-metrics-at-a-glance)
3. [How the tests work](#3-how-the-tests-work)
4. [The metrics in detail](#4-the-metrics-in-detail)
   - [4.1 Startup time](#41-startup-time)
   - [4.2 Logins per second (AS requests)](#42-logins-per-second-as-requests)
   - [4.3 Service tickets per second (TGS requests)](#43-service-tickets-per-second-tgs-requests)
   - [4.4 Latency and errors](#44-latency-and-errors)
   - [4.5 Concurrency (scaling)](#45-concurrency-scaling)
   - [4.6 Logins over TCP](#46-logins-over-tcp)
   - [4.7 Administration operations (kadmind)](#47-administration-operations-kadmind)
   - [4.8 Database dump](#48-database-dump)
   - [4.9 CPU usage](#49-cpu-usage)
   - [4.10 Memory usage](#410-memory-usage)
5. [How the metrics relate to each other](#5-how-the-metrics-relate-to-each-other)
6. [How much performance do you need?](#6-how-much-performance-do-you-need)
7. [KDC settings that affect performance](#7-kdc-settings-that-affect-performance)
8. [Limits of this test](#8-limits-of-this-test)
9. [Glossary](#9-glossary)

---

## 1. Kerberos in one minute

Kerberos is a login system for a whole network. A user (or a machine, or a
service) proves its identity **once** to a central server, the **KDC**
(Key Distribution Center). After that it receives **tickets**, which it shows
to web servers, file servers, SSH servers and so on, without typing a
password again. Every user, host and service is a **principal** in the KDC's
database, for example `alice@EXAMPLE.COM` or `HTTP/web1.example.com@EXAMPLE.COM`.

```
   client                                                         KDC
     |                                                             |
     |  LOGIN = "AS request"  (what kinit, sssd, GDM do)           |
     |  1. AS-REQ  "I am alice, give me a TGT"             ----->  |
     |  <-----  "first prove it" (pre-authentication needed)       |
     |  2. AS-REQ  + proof made with alice's key           ----->  |
     |                              (KDC reads alice from the DB,  |
     |                               writes "last login" to the DB)|
     |  <-----  AS-REP: ticket-granting ticket (TGT)               |
     |                                                             |
     |  SERVICE TICKET = "TGS request" (every new service used)    |
     |  3. TGS-REQ  "with my TGT: a ticket for HTTP/web1"  ----->  |
     |  <-----  TGS-REP: service ticket for HTTP/web1              |
     |                                                             |
     |  4. shows the service ticket to web1 (the KDC is not involved)
```

- A **login** (AS exchange) normally takes **two round trips**. The first
  answer only says *how* the client must prove who it is
  (**pre-authentication**). RHEL 9 uses the **SPAKE** method by default,
  which uses elliptic-curve cryptography on both sides.
- A **service ticket** (TGS exchange) takes **one round trip**. A user asks
  for one for every service they use; tickets are cached until they expire
  (usually 10 or 24 hours).
- Requests travel over **UDP** port 88 (small requests) or **TCP** port 88
  (large requests, or when UDP is blocked).
- The **admin server** `kadmind` (port 749) is used to create, change and
  delete principals, and (port 464) to change passwords. It is not involved
  in logins.

So KDC performance is mostly: **how fast it can look up principals in its
database, do the cryptography, and (for logins) write back to the database.**

---

## 2. The metrics at a glance

| # | Metric | Unit | Question it answers | Better is | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast does the KDC hand out tickets again after a (re)start? | lower | `startup` |
| 2 | Logins per second | logins/s | How many logins (TGTs) per second can the KDC handle at full speed? | higher | `as` |
| 3 | Service tickets per second | tickets/s | How many service tickets per second at full speed? | higher | `tgs` |
| 4 | Latency | ms (p95, p99) | How long does a client wait for a login / a service ticket at a normal load? | lower | `latency` |
| 4b | Errors | % | How many requests fail at a normal load? | lower | `latency` |
| 5 | Concurrency (scaling) | % of peak | Does the KDC keep its speed when many clients push at once? | higher | `concurrency` |
| 6 | Logins over TCP | logins/s | How fast are logins when clients use TCP? | higher | `tcp` |
| 7 | Admin operations | operations/s | How fast can accounts be created, read, changed, deleted? | higher | `kadmin` |
| 8 | Database dump | seconds | How long do a backup and a replication cycle take? | lower | `dump` |
| 9 | CPU usage | % of one core | How much processor time does a normal load cost? | lower | `cpu` |
| 10 | Memory usage | MB | How much RAM does the KDC need? | lower | `memory` |

Run one metric with `./test-kerberos.sh <test name>`, or all of them with
`./test-kerberos.sh`.

---

## 3. How the tests work

```
  +----------------------------------+            +-------------------------------------+
  |  lib/krb5-load.py                |  UDP / TCP |  krb5kdc   (krb5kdc-perf-test)      |
  |  N client processes, each uses   |  port 88   |  realm PERF.TEST, 127.0.0.1 only    |
  |  the system's libkrb5.so.3       | ---------> |       |                             |
  |  (the same code as kinit, sssd)  |            |       v                             |
  +----------------------------------+            |  principal database                 |
                                                  |  /var/kerberos/krb5kdc/             |
  +----------------------------------+  TCP 749   |     principal-perf-test             |
  |  kadmin (admin client)           | ---------> |       ^                             |
  +----------------------------------+            |  kadmind   (kadmind-perf-test)      |
                                                  +-------------------------------------+
         |                                                   |
         |  logins/s, tickets/s, latency, errors             |  CPU, memory:
         v                                                   v  /proc/<krb5kdc pid>/...
                results/<date-time>/report.md   (PASS / FAIL per metric)
```

- **Separate test realm.** `start-kerberos.sh` creates the realm
  `PERF.TEST` with its own configuration file, database and two systemd
  services (`krb5kdc-perf-test`, `kadmind-perf-test`). They listen only on
  `127.0.0.1`. The normal `/etc/krb5.conf`, `kdc.conf`, and the services
  `krb5kdc` and `kadmin` are not touched.
- **Realistic database.** The database holds 1,000 test users, 100 test
  services and 10,000 "filler" principals that are never used. All
  principals **require pre-authentication**, and have keys of the four RHEL 9
  default encryption types, exactly as on a production KDC.
- **Load generator.** `lib/krb5-load.py` is a small Python 3 program. It does
  not build Kerberos messages itself. It calls the **system Kerberos library**
  (`libkrb5.so.3`), so every request is exactly what `kinit` or `sssd` would
  send: real pre-authentication (SPAKE), real encryption, real UDP/TCP
  handling and retries. The test users log in with keys from a keytab (like
  `kinit -k`), so no password typing or password hashing slows the client.
- **Clients.** The generator starts N client processes. Each one sends a
  request, waits for the answer, then sends the next one. It works in two ways:
  - **Full speed** (tests `as`, `tgs`, `concurrency`, `tcp`): every client
    sends its next request immediately. This finds the maximum rate.
  - **Fixed rate** (tests `latency`, `cpu`, `memory`): a steady "normal day"
    load. 100 users log in per second, and each then asks for 4 service
    tickets, so 500 requests per second in total. This is how the latency of
    a *normal* day is measured, not the latency of an overloaded server.
- **Server-side measurements**, taken once per second during the steady
  load: CPU time and memory of all `krb5kdc` processes
  (`/proc/<pid>/stat`, `/proc/<pid>/status`), and the growth of the KDC log.

A complete run takes about 2 to 4 minutes.

---

## 4. The metrics in detail

Each section explains: **what** is measured, **why** it matters, **how** the
test measures it, the **default target**, and **what to do** when the result
is poor.

### 4.1 Startup time

**What:** the time from "start the KDC" until the first client has a ticket.

**Why it matters:** while the KDC is down or starting, nobody can log in and
nobody can open a new service. This matters after a crash, after a patch
reboot, and after a configuration change (the KDC must be restarted to read
`kdc.conf` again). With replica KDCs, clients switch to another KDC, but only
after a timeout of about 1 second per try.

**How it is measured:** the KDC is stopped. A probe client tries to log in
every 10 ms. Then the KDC is started, and the time until the probe's first
successful login is noted. Repeated `STARTUP_ROUNDS` (3) times.

**Target:** at most 5,000 ms. A healthy KDC starts in well under a second,
because it does not load the database into memory; it opens the file and
reads principals on demand.

**If it is poor:** check the log (`/var/log/krb5kdc.log.perf-test` or
`journalctl -u krb5kdc-perf-test`) for slow steps: DNS lookups of the host
name (keep `rdns = false`, give the host an `/etc/hosts` entry), a slow or
remote disk, or a database that must be recovered after a crash.

### 4.2 Logins per second (AS requests)

**What:** the number of successful logins the KDC completes per second when
`CLIENTS` (8) clients log in as fast as they can. One login = one TGT.

**Why it matters:** this is the capacity of the KDC for **login storms**:
Monday 08:00, when everyone logs in; the moment after a network outage, when
thousands of machines and services renew their tickets at once; or a cluster
job start where hundreds of nodes run `kinit` at the same time.

**How it is measured:** `krb5-load.py as --clients 8`, for `DURATION` (10) s.
Every login is a full pre-authenticated AS exchange (two round trips). The
client processes use only a small part of their CPU time (column
"Client processes busy"); if they get close to 100%, the report warns that
the load generator, not the KDC, may have been the limit.

**Target:** at least 500 logins/s.

**Typical results:** on the build machine, a default RHEL 9 KDC (single
process, db2, SPAKE, last-login writes on) reached about **1,500 to 1,800
logins/s**. See [section 7](#7-kdc-settings-that-affect-performance) for how
much each setting changes this.

**If it is poor:**
- Look at CPU usage. If the KDC process is at 100% of one core, it is
  CPU-bound. Use worker processes (`KDC_WORKERS`), see section 7.
- If the CPU is *not* busy, the KDC waits for the disk: every login writes
  the "last successful login" time into the database. Use a faster disk,
  the LMDB back end, or switch the write off (`DISABLE_LAST_SUCCESS=true`).
- A policy with `pw_max_fail` (account lockout) causes a database write for
  every *failed* login as well.

### 4.3 Service tickets per second (TGS requests)

**What:** the number of service tickets the KDC hands out per second at full
speed.

**Why it matters:** in a real network there are many more service ticket
requests than logins. Every user asks for one ticket per service used (web
sites, file shares, mail, SSH hosts, databases), and NFS with Kerberos
(`sec=krb5`) asks for tickets for each user on each client.

**How it is measured:** each client logs in once (not measured), then asks
again and again for tickets for `HTTP/web0001.perf.test` ...
`HTTP/web0100.perf.test`. The ticket is *not* cached, so every request really
reaches the KDC.

**Target:** at least 1,000 tickets/s.

**Typical results:** about 2 to 2.5 times the login rate, because a TGS
exchange is one round trip, needs no pre-authentication, and does not write
to the database.

**If it is poor:** the same remedies as for logins (worker processes, LMDB).
Very large tickets (many authorization data, PACs) cost more CPU per request.

### 4.4 Latency and errors

**What:** the time a client waits from sending its first request until it has
its ticket, at a steady, normal load, and the share of requests that fail.
Reported separately for **logins** and for **service tickets**, as average,
median (p50), p90, **p95**, **p99**, p99.9 and slowest.

**Why it matters:** this is what a user notices. A slow KDC makes the login
screen hang, makes the first opening of every web site or share slow, and
makes scripts and batch jobs slow. Percentiles matter more than the average:
"p99 = 50 ms" means 1 in 100 requests takes 50 ms or longer, and a user
who opens 20 services a day meets that case often. **Errors** mean that
someone could not log in or could not open a service.

**How it is measured:** the steady load of section 3 (`LOGIN_RATE` = 100
logins/s, each followed by `TGS_PER_LOGIN` = 4 service tickets) for
`DURATION` seconds. Every request is timed by the client. The client library
itself would re-send a lost UDP request after 1 second; such a request shows
up as a slow request, not as an error. Errors are requests that failed in the
end (for example "Cannot contact any KDC", "Clock skew too great").

**Targets:** logins p95 ≤ 20 ms and p99 ≤ 50 ms; service tickets p95 ≤ 10 ms
and p99 ≤ 25 ms; errors ≤ 0.1%.

**Typical results:** on localhost, a login takes about 2 to 4 ms and a
service ticket about 1 ms. Across a real network, add one network round trip
per request for a service ticket, and two per request for a login.

**If it is poor:**
- p99 much higher than p95: the KDC is sometimes blocked, usually by disk
  writes (last-login writes, a busy disk, `fsync` of the log), or the machine
  is busy with other work.
- Everything slow: the KDC is near its capacity (compare the steady load with
  the result of 4.2), or the CPU is slow.
- Around 1,000 ms delays: UDP packets are lost and the library re-sends them.
  Check `netstat -su` for "receive buffer errors", and the firewall.
- Errors: read the error texts in the report. "Clock skew too great" = time
  not synchronised (chronyd); "Cannot contact any KDC" = KDC down or overloaded.

### 4.5 Concurrency (scaling)

**What:** how the login rate and the waiting time change when more and more
clients push at the same time: `CLIENT_LEVELS` = 1, 2, 4, 8, 16, 32, 64
clients, each at full speed.

**Why it matters:** in a login storm, thousands of clients arrive at once.
A good server then works at its peak rate and every client waits a little
longer. A bad server **slows down under pressure** (lock contention, dropped
UDP packets and re-sends) or starts to fail requests.

**How it is measured:** the full-speed login test at every level. The verdict
compares the rate at the highest level with the peak rate of all levels.

**Target:** at the highest level, at least 70% of the peak rate.

**How to read the table:** with a single-process KDC the rate stops rising
at about 2 to 4 clients (that is when the KDC is busy all the time). After
that, the average waiting time grows in proportion to the number of clients:
twice as many clients, twice the wait. That is expected and fine. A falling
rate, or a "Failed %" above 0, is not.

**If it is poor:** use worker processes (`KDC_WORKERS`); with worker
processes use LMDB, because db2 serialises every write with one lock; raise
the UDP receive buffer (`net.core.rmem_default`) if the kernel drops packets.

### 4.6 Logins over TCP

**What:** the login rate when every request is sent over TCP instead of UDP.

**Why it matters:** clients switch to TCP when a request is bigger than
`udp_preference_limit` (1,465 bytes), for example with FAST armoring, PKINIT
(smart cards) or large tickets. Some sites force TCP because firewalls or
load balancers handle it better. Each round trip then needs its own TCP
connection: set-up, request, answer, close.

**How it is measured:** the same as 4.2, with the client setting
`udp_preference_limit = 1` (file `perf-test-krb5-tcp.conf` placed in front
of the normal client configuration). The report also shows the rate as a
percentage of the UDP rate from test `as`, if that test ran.

**Target:** at least 300 logins/s.

**Typical results:** on localhost TCP is about as fast as UDP (80 to 140% of
it; the differences are mostly measurement noise). Over a real network, the
extra round trip of the TCP handshake makes each request slower.

**If it is poor:** check for a limit on connections: firewall connection
tracking (`nf_conntrack` table full), `net.core.somaxconn`, or many sockets
in TIME_WAIT on the client side.

### 4.7 Administration operations (kadmind)

**What:** how many administration operations per second the admin server
`kadmind` handles, for four kinds of operation:

| Operation | Command | Database |
|---|---|---|
| add | `addprinc -randkey` | creates a principal with new random keys (write) |
| read | `getprinc` | reads a principal (read only) |
| change | `cpw -randkey` | gives a principal new random keys (write) |
| delete | `delprinc -force` | deletes a principal (write) |

**Why it matters:** identity management systems (FreeIPA, scripts that
create accounts for new staff or students, provisioning of hosts and
service keytabs) talk to `kadmind`. Bulk jobs (thousands of new accounts at
the start of a semester, key rotation of all hosts) depend on this speed.

**How it is measured:** `KADMIN_OPERATIONS` (1,000) principals per kind, each
kind as one batch in one `kadmin` session (one admin login), over the network
protocol to `kadmind`. The time includes the session start. The verdict uses
the slowest kind, and a kind that did not complete every principal counts as
0 operations/s.

**Target:** at least 100 operations/s for the slowest kind.

**If it is poor:** write operations depend on disk speed (every change is
written to the database file); `kadmind` handles one request at a time. A
password quality policy with a large `dict_file` makes password changes
slower.

### 4.8 Database dump

**What:** the time `kdb5_util dump` needs to write the whole principal
database to a file, for the database size of the test (11,100 principals by
default).

**Why it matters:** the dump is the first step of every **backup** and of
classic **replication** to replica KDCs (`kdb5_util dump` → `kprop` →
`kdb5_util load` on the replica). It runs while the KDC keeps serving. If it
takes long, the replicas lag behind, and changes (new users, new passwords)
reach them later.

**How it is measured:** `DUMP_ROUNDS` (3) dumps while the KDC is running;
the average is reported. The dump file is deleted afterwards because it
contains all keys (encrypted with the master key).

**Target:** at most 30 s.

**Typical results:** much less than a second for 11,000 principals
(about 60,000+ principals per second). The dump time grows linearly with the
number of principals, so the result tells you the time for your real
database size too.

**If it is poor:** slow disk or a busy disk; a very large database with many
old principals (delete unused ones); with frequent changes, use incremental
propagation (`iprop_enable`) instead of full dumps.

### 4.9 CPU usage

**What:** the CPU time used by the KDC during the steady load, as a
percentage of one CPU core.

**Why it matters:** by default the KDC is **one single process with one
thread**, so it can never use more than one core. **100% means the KDC is at
its limit**, no matter how many cores the machine has. The CPU usage at a
normal load shows how much headroom is left for a login storm: 34% at the
normal load means a storm of about 3 times the normal load can still be
handled.

**How it is measured:** the CPU ticks of all `krb5kdc` processes
(`/proc/<pid>/stat`) before and after the steady load; also sampled every
second. The report also shows the CPU time per 1,000 requests and the bytes
written to the KDC log per request (the KDC logs every request).

**Target:** at most 50% of one core at the steady load.

**If it is poor:** use worker processes (`KDC_WORKERS`). Check the encryption
types: the `aes-sha2` types (RHEL 9 default) and SPAKE cost more CPU than the
older `aes-sha1` types and the encrypted-timestamp method, but they are more
secure; do not weaken them just for speed. Logging to a slow target also
costs CPU.

### 4.10 Memory usage

**What:** the memory in RAM (resident set size) of all KDC processes, idle
and at the peak of the steady load.

**Why it matters:** a KDC needs little memory, so this metric is mainly a
**leak check**: memory that grows during a test and does not return is a
problem for a server that runs for months.

**How it is measured:** `/proc/<pid>/status` (VmRSS) before the steady load
and every second during it.

**Target:** at most 256 MB.

**Typical results:** about 8 to 20 MB for a single-process KDC; it grows a
few MB when the load starts and then stays flat. The KDC keeps no state per
user between requests; the database is read through the operating system's
page cache, which does not count here. With worker processes, each worker
adds about 6 MB (4 workers: about 35 MB in total).

**If it is poor:** memory that keeps growing across runs: compare
`raw/resource-samples.csv` of several runs; check that you have the latest
RHEL `krb5` errata.

---

## 5. How the metrics relate to each other

```
   capacity (4.2 logins/s, 4.3 tickets/s)
        |
        |  normal load / capacity = how busy the KDC is  --->  4.9 CPU usage
        v
   latency (4.4) stays low while the KDC is far from its capacity,
   and grows fast when it gets close (requests start to queue)
        |
        v
   concurrency (4.5): above capacity, extra clients only wait longer;
   a good KDC keeps its peak rate and fails no requests
```

- **Capacity vs. CPU:** for a single-process KDC, capacity ≈ (steady-load
  rate) ÷ (CPU usage at that rate). Example: 500 requests/s at 34% CPU means
  about 1,500 requests/s at 100%. That matches the full-speed tests.
- **Logins vs. service tickets:** a login costs about twice as much as a
  service ticket (two round trips, pre-authentication, a database write).
- **Disk vs. CPU:** logins (last-login write) and all kadmind write operations
  also depend on disk speed; service tickets and reads do not.
- **Dump vs. database size:** both the dump time and (a little) the login
  time grow with the number of principals.

---

## 6. How much performance do you need?

Estimate your peak load and compare it with the measured capacity.
Plan to use **at most about 30 to 50%** of the capacity at the peak, so
that a single KDC can carry the load when a replica is down.

| Situation | Rough load |
|---|---|
| 10,000 users log in within 15 minutes (morning) | 10,000 ÷ 900 s ≈ **11 logins/s**, plus ~5 to 20 service tickets per user ≈ 60 to 220 tickets/s |
| 2,000 Linux hosts (sssd), each renewing its host ticket and user tickets | low steady load, but **all at once** after a network or KDC outage: 2,000+ logins within seconds |
| 500-node HPC cluster, job start with `kinit` on every node | 500 logins in 1 to 2 seconds |
| NFS with `sec=krb5`: 200 clients × 50 users | 10,000 service tickets, spread over the ticket lifetime; again a peak after outages |

Even a single-process KDC with default settings (about 1,500 logins/s and
3,000 tickets/s here) handles most sites. Outage recovery and batch jobs are
what usually reach the limit, and short peaks are absorbed by client
re-sends. Always run at least **two KDCs** (a primary and a replica) for
availability, not for speed.

---

## 7. KDC settings that affect performance

All these can be changed in `settings.conf` (then run `./start-kerberos.sh`
again). The table shows logins per second measured on the build machine
(32 cores, localhost, 16 clients, 4 s per run; the machine was shared, so
expect ±15%):

| Configuration | Logins/s | vs. default |
|---|---:|---:|
| **RHEL 9 default**: single process, db2, SPAKE, last-login writes on | ~1,770 | 1.0× |
| `SPAKE_PREAUTH=no` (encrypted-timestamp pre-authentication) | ~2,580 | 1.5× |
| `DISABLE_LAST_SUCCESS=true` (no database write per login) | ~2,180 | 1.2× |
| both of the above | ~3,290 | 1.9× |
| `DB_BACKEND=lmdb` | ~2,650 | 1.5× |
| `KDC_WORKERS=4` (db2) | ~5,580 | 3.2× |
| `KDC_WORKERS=4` + `DISABLE_LAST_SUCCESS=true` | ~7,820 | 4.4× |
| `KDC_WORKERS=4` + `DB_BACKEND=lmdb` | ~8,720 | 4.9× |
| `KDC_WORKERS=4` + lmdb + no SPAKE + no last-login writes | ~15,950 | 9.0× |

What the settings mean, and their trade-offs:

- **`KDC_WORKERS`** (`krb5kdc -w N`): starts N worker processes that share
  the requests. The single biggest improvement when the KDC is CPU-bound.
  Use about the number of cores you can spare.
- **`DB_BACKEND=lmdb`**: the LMDB database (`db_library = klmdb`) allows many
  readers at the same time and writes faster than the classic db2 file,
  especially with worker processes. It is part of RHEL 9's `krb5-server`.
  Moving an existing realm needs a dump and load.
- **`DISABLE_LAST_SUCCESS=true`** (`disable_last_success` in `[dbmodules]`):
  the KDC no longer records the time of the last successful login. Faster,
  but you lose that information (audits, finding unused accounts).
  The similar `disable_lockout = true` also removes the write on *failed*
  logins, but switches off account lockout.
- **`SPAKE_PREAUTH=no`**: uses the older encrypted-timestamp
  pre-authentication. SPAKE (RHEL 9 default) protects better against offline
  password guessing, so **keep SPAKE** unless you have measured that you need
  the speed and understand the risk.
- **Encryption types** (`SUPPORTED_ENCTYPES`): the RHEL 9 default prefers
  `aes256-cts-hmac-sha384-192`. Do not add old types (RC4, DES) for speed;
  the system crypto policy (`update-crypto-policies`) may forbid them anyway.
- **Logging:** the KDC writes one log line per request (about 470 bytes
  with the default encryption types). On a slow or network file system, log to a local disk; rotate the
  log with logrotate.
- **Hardware and OS:** a fast local disk (SSD) for `/var/kerberos/krb5kdc`,
  enough UDP receive buffer (`sysctl net.core.rmem_default`), time
  synchronisation (chronyd). A KDC must not share its machine with other
  busy services; it is also a security-critical host.

---

## 8. Limits of this test

- **Loopback only.** All traffic goes over `127.0.0.1`: no network delay, no
  packet loss, no firewall. Real latencies are higher by the network round
  trip time (see 4.4).
- **One machine.** The load generator and the KDC share the CPUs. On a small
  machine, the generator takes CPU away from the KDC. The report warns when
  the client processes were close to fully busy.
- **Shared or busy machines give noisy numbers.** Run the test twice and
  compare, and do not run other heavy work at the same time.
- **Keytab logins only.** Users log in with keys from a keytab, as services
  and `kinit -k` do. With a password, the *client* also spends time turning
  the password into a key; the KDC's work is the same.
- **Not covered:** PKINIT (smart cards), OTP/RADIUS, FAST armoring,
  cross-realm trusts, Active Directory, FreeIPA's LDAP back end
  (`kdb` plugin `ipadb`), replication with `kpropd` or `iprop`, and the
  password-change service (port 464).
- **Closed-loop full-speed tests.** Each client waits for its answer before
  it sends the next request, so the KDC is never flooded beyond what the
  clients can wait for. Real login storms from thousands of machines can be
  harder, because clients also re-send after timeouts.
- **Test realm.** The database has 11,100 principals by default. Set
  `FILLER_PRINCIPALS` to match your real database size for more realistic
  dump and startup times.

---

## 9. Glossary

| Term | Meaning |
|---|---|
| **KDC** | Key Distribution Center, the Kerberos server (`krb5kdc`) that hands out tickets |
| **Realm** | a Kerberos "domain", written in capitals (`PERF.TEST`, `EXAMPLE.COM`) |
| **Principal** | an account in the KDC database: a user, host or service (`alice@REALM`, `HTTP/web1.example.com@REALM`) |
| **Ticket** | proof of identity for one service, valid for a limited time, encrypted so only that service can read it |
| **TGT** | ticket-granting ticket: the ticket received at login, used to ask for all other tickets |
| **AS request / AS exchange** | "Authentication Service": the login, which returns a TGT (what `kinit` does) |
| **TGS request / TGS exchange** | "Ticket Granting Service": asking for a service ticket with the TGT |
| **Pre-authentication** | the client proves it knows its key *before* the KDC sends a TGT; stops offline password guessing |
| **SPAKE** | a modern pre-authentication method (RHEL 9 default), based on elliptic-curve cryptography |
| **Keytab** | a file that holds the keys of one or more principals, so they can log in without a password |
| **kadmind** | the administration server: create, change, delete principals; change passwords |
| **kprop / iprop** | replication of the database to replica KDCs: full copies (kprop) or incremental changes (iprop) |
| **db2 / LMDB** | the two database back ends included in RHEL's `krb5-server` |
| **Worker processes** | `krb5kdc -w N`: N KDC processes that share the requests, to use more than one CPU core |
| **p95 / p99** | 95th / 99th percentile: 95% / 99% of the requests were faster than this |
| **Round trip** | one request sent to the KDC plus its answer |
| **RSS** | resident set size: the part of a process's memory that is in RAM |
