# OpenSSH Server - Performance Metrics

This document explains the performance metrics of the OpenSSH server on
RHEL 9.6 (`sshd` from the `openssh-server` package): what each metric means,
why it matters, how `test-ssh.sh` measures it, and what to change when a
result is poor.

---

## Contents

1. [SSH in one minute](#1-ssh-in-one-minute)
2. [The metrics at a glance](#2-the-metrics-at-a-glance)
3. [How the tests work](#3-how-the-tests-work)
4. [The metrics in detail](#4-the-metrics-in-detail)
   - [4.1 Startup time](#41-startup-time)
   - [4.2 Login time](#42-login-time)
   - [4.3 Authentication methods](#43-authentication-methods)
   - [4.4 Login rate](#44-login-rate)
   - [4.5 Burst of logins (MaxStartups)](#45-burst-of-logins-maxstartups)
   - [4.6 Concurrent sessions](#46-concurrent-sessions)
   - [4.7 Command over an open connection](#47-command-over-an-open-connection)
   - [4.8 Keystroke echo](#48-keystroke-echo)
   - [4.9 Transfer speed](#49-transfer-speed)
   - [4.10 SFTP](#410-sftp)
   - [4.11 CPU](#411-cpu)
   - [4.12 Memory](#412-memory)
   - [4.13 Reload](#413-reload)
5. [How the metrics relate to each other](#5-how-the-metrics-relate-to-each-other)
6. [How much performance do you need?](#6-how-much-performance-do-you-need)
7. [sshd settings that affect performance](#7-sshd-settings-that-affect-performance)
8. [SSH on an offline network](#8-ssh-on-an-offline-network)
9. [Limits of this test](#9-limits-of-this-test)
10. [Glossary](#10-glossary)

---

## 1. SSH in one minute

**SSH** (Secure Shell) gives a user or a program an encrypted connection to
a remote machine: a shell, a single command, or a file transfer (SFTP, scp).
The server is **sshd**. Every login goes through the same five phases:

```
   ssh client                                                  sshd (server)
  +-------------------------+                          +-------------------------------+
  | 1. TCP connect          | ----- SYN / SYN-ACK ---> | listener accepts, starts a    |
  |                         |                          | process for this connection   |
  | 2. Greeting             | <--- "SSH-2.0-OpenSSH" - | both sides send their version |
  | 3. Key exchange         | <======================> | agree on keys (curve25519 ...),|
  |    (check host key)     |                          | sign with the HOST key        |
  |                         |      - - - encrypted from here on - - -                  |
  | 4. Authentication       | ---- user key / pwd ---> | authorized_keys or PAM        |
  | 5. Session              | <======================> | PAM session, start the shell  |
  |    command / shell      |                          | or command, then log out      |
  +-------------------------+                          +-------------------------------+
```

A few facts explain almost all SSH server performance:

- **A login is expensive, a logged-in connection is cheap.** A login needs
  public-key cryptography, several round trips, PAM and new processes. Once
  logged in, sending a keystroke or a block of data is fast.
- **Every connection has its own processes.** sshd's "listener" only
  accepts connections. For each one it starts separate processes (one as
  root, one as the user), so one broken session never hurts another, and
  the server's capacity grows with CPU cores.
- **sshd protects itself against floods.** `MaxStartups` limits how many
  connections may be *in the middle of logging in*. Above that, sshd
  refuses new connections at random - also legitimate ones from automation
  tools (section 4.5).
- **One connection uses one CPU core per side** for encryption. The speed of
  one file transfer is the speed of one core (section 4.9).
- **Much of a login is not sshd itself.** PAM (passwords, SSSD, LDAP,
  Kerberos, systemd-logind), DNS lookups and the user's shell start-up all
  run inside the login. On an offline network, a timeout there adds
  seconds to every login (section 8).

---

## 2. The metrics at a glance

| # | Metric | Unit | Question it answers | Better | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast can users log in again after a (re)start? | lower | `startup` |
| 2 | Login time | ms | How long does one complete login take, and which phase is slow? | lower | `login` |
| 3 | Authentication methods | ms | Do keys, RSA keys or passwords make a login slower? | lower | `auth` |
| 4 | Login rate | logins/s | How many complete logins per second does the server handle? | higher | `rate` |
| 5 | Burst of logins | clients | How many clients can log in *at the same moment*? | higher | `burst` |
| 6 | Concurrent sessions | sessions | Can many sessions stay open at once, all working? | all | `concurrent` |
| 7 | Command over an open connection | ms | How fast is a command when the connection is re-used? | lower | `session` |
| 8 | Keystroke echo | ms | How quickly does a typed letter appear? | lower | `interactive` |
| 9 | Transfer speed | MB/s | How fast is data through one connection, per cipher? | higher | `transfer` |
| 10 | SFTP | MB/s, files/s | How fast are large and small file copies? | higher | `sftp` |
| 11 | CPU | ms per login, s per GB | How much processor time does a login or a GB cost? | lower | `cpu` |
| 12 | Memory | MB, KB per session | How much RAM does sshd use at rest and per session? | lower | `memory` |
| 13 | Reload | ms, sessions | Does a configuration reload interrupt anyone? | fast, all survive | `reload` |

Run one metric with `./test-ssh.sh <test name>`, or all of them with
`./test-ssh.sh`.

---

## 3. How the tests work

```
  +--------------------------------------------+              +------------------------------------+
  |  lib/sshload.py  (the "clients")           |              |  test sshd  (sshd-perf-test)       |
  |                                            |   loopback   |  /etc/ssh/sshd_config_perf_test    |
  |  starts the real OpenSSH "ssh" / "sftp"    | ==== TCP ==> |  own host keys, RHEL crypto policy |
  |  programs, many at once, and times them    |  127.0.0.1   |  only user "sshperf" may log in    |
  |  (ssh -F .client/ssh_config sshperf-test)  |    :11022    |  listens on 127.0.0.1:11022 only   |
  +--------------------------------------------+              +------------------------------------+
         |                                                                  |
         |  login phases (from "ssh -v"), logins/s,                         |  CPU of all sshd processes
         |  refused logins, echo time, MB/s                                 |  (/proc/<pid>/stat), memory
         v                                                                  v  (PSS, /proc/<pid>/smaps_rollup)
                         results/<date-time>/report.md   (PASS / FAIL per metric)
```

- **A separate test server.** `start-ssh.sh` runs a **second** sshd with
  its own configuration (`/etc/ssh/sshd_config_perf_test`), its own host
  keys and its own service (`sshd-perf-test.service`). The normal
  `sshd.service` and `/etc/ssh/sshd_config` are not touched, so your own
  SSH connections to the host are never interrupted.
- **Same settings as a normal RHEL server.** The test configuration uses
  RHEL 9's defaults and the system-wide **crypto policy**
  (`update-crypto-policies --show`), so the same ciphers and key-exchange
  methods are used as by the real sshd.
- **Only this machine can reach it.** It listens on `127.0.0.1`, TCP port
  `11022` (not 22). With SELinux, the kit gives that port the label
  `ssh_port_t`, which is what sshd may use.
- **A test user.** The tests log in as the local user `sshperf`, created
  by `start-ssh.sh` with an ed25519 key, an RSA 3072 key and a random
  password. Only this user may log in to the test server (`AllowUsers`).
- **Real clients.** Every login is a real `ssh` (or `sftp`) process with
  the normal RHEL client and crypto policy, exactly what users and tools
  run. `lib/sshload.py` (Python 3, standard library only) starts them, many
  at once when needed, and times them. The login phases are taken from the
  output of `ssh -v`.
- **Server-side measurements**, before and after each run:
  - CPU time of the listener and every per-connection process, including
    processes that already ended (a parent collects the CPU time of its
    finished children);
  - memory (PSS) of all sshd processes, so shared memory is counted once.
- **Loads.**

| Load | What it does | Used by |
|---|---|---|
| one after another | `LOGIN_SAMPLES` (30) logins, each timed by phase | login, auth |
| parallel clients | `RATE_WORKERS` (8) clients log in again and again for `DURATION` (10) s | rate, cpu |
| all at once | 5, 10, 20, 50, 100 clients released at the same moment | burst |
| hold open | `CONCURRENT_SESSIONS` (200) sessions opened, kept open, checked | concurrent, memory |
| one connection | 200 commands over one open connection; 500 keystrokes | session, interactive |
| bulk data | `TRANSFER_MB` (1024) MB per cipher and direction; sftp files | transfer, sftp, cpu |

- The `startup` and `reload` tests **restart / reload the test server**
  (only the test server; the normal sshd keeps running).
- In the report, examples are given from a validation run on a 32-core
  machine ("Example"). Your numbers will differ; compare them with your
  targets, and with earlier runs of the same machine.

---

## 4. The metrics in detail

Each metric has the same five parts: **what** it is, **why** it matters,
**how** the test measures it, **how to read** the result, and **how to improve** it.

### 4.1 Startup time

**What:** the time from the command that starts sshd until the first
complete login works. The report also shows when the server sent its first
greeting (it accepts connections).

**Why:** sshd is restarted at every update of `openssh-server`, every reboot
and after some configuration changes. While it is down, nobody can log in -
including the administrator who just restarted it.

**How it is measured:** the test server is stopped and started
`STARTUP_ROUNDS` times (5). A small program tries to connect every 5 ms,
from before the start until the greeting arrives; then one real login is
made.

**How to read it:**

| Result | Meaning |
|---|---|
| Greeting under 50 ms, first login about 1 login time later | normal: sshd has no warm-up (Example: greeting 15 ms, first login 160 ms) |
| First start after installation: seconds | normal once: `sshd-keygen` creates the host keys (RSA takes longest) |
| Over 1 s every time | investigate: `journalctl -u sshd`; slow `Include` files, `AuthorizedKeysCommand`, a slow disk |
| Never answers | configuration error: run `sshd -t` |

**How to improve:** there is rarely anything to improve. Important instead:
**a stop or restart does not end open sessions** (`KillMode=process` in
`sshd.service`), so restarting sshd over SSH is safe - as long as the new
configuration is valid (`sshd -t`, section 4.13).

---

### 4.2 Login time

**What:** the time of one complete non-interactive login
(`ssh server true`): connect, key exchange, authentication, run the command,
log out. The test splits it into the five phases of section 1.

**Why:** every `ssh` command, every `scp`, every Ansible task without
connection re-use and every monitoring check pays this time. It is also
what a user feels when opening a terminal.

**How it is measured:** `LOGIN_SAMPLES` (30) logins one after another, with
`ssh -v`. The moment each phase ends is taken from the debug output:

| Phase | Ends when `ssh -v` prints | What happens | Example |
|---|---|---|---:|
| 1. TCP connect | `Connection established` | client starts, reads its configuration, TCP handshake | 4 ms |
| 2. Greeting | `Remote protocol version` | sshd starts the connection's process, both send their version | 5 ms |
| 3. Key exchange | `SSH2_MSG_NEWKEYS received` | key agreement (curve25519), host key signature and check | 10 ms |
| 4. Authentication | `Authenticated to` | the user's key is offered, checked in `authorized_keys`, signature verified; PAM `auth`/`account` | 55 ms |
| 5. Session | (ssh exits) | channel opened, PAM session (logind, limits, SELinux context), `true` runs, logout | 62 ms |
| **Total** | | | **140 ms** |

**How to read it:**

| Result | Meaning |
|---|---|
| 100 - 300 ms on a LAN | normal for RHEL 9 |
| A phase of about 40 ms that "does nothing" | TCP waits for a delayed acknowledgement (see below) - not a server problem |
| Authentication or session over 1 s | a timeout: DNS, GSSAPI/Kerberos, SSSD/LDAP, or `pam_*` modules (section 8) |
| Key exchange over 100 ms | slow CPU, or a heavy key exchange (`diffie-hellman-group18-sha512`, RSA 4096+ host key) |
| Exactly 5, 10, 25, 30 s | a DNS or network timeout; almost always `UseDNS yes` or GSSAPI (section 8) |

*About the 40 ms waits:* in the validation run, the authentication and the
session phase each spent about 40 ms waiting, not working. `strace` of
sshd showed it waiting for the client's next small packet, which the
client's TCP stack held back (Nagle's algorithm) until the server's
"delayed ACK" timer (40 ms on Linux) had expired. On a real network the
same effect can appear, mixed with the network's round-trip time. It is
the reason an SSH login rarely takes less than about 100 ms.

**How to improve:** re-use connections (section 4.7) - this removes phases
1-4 entirely. Keep `UseDNS no` (RHEL default). On hosts without Kerberos,
set `GSSAPIAuthentication no` on clients. Keep the PAM stack local and fast.

---

### 4.3 Authentication methods

**What:** the login time with an **ed25519 key**, an **RSA 3072 key** and a
**password**.

**Why:** the method changes who does the work. With keys, sshd verifies one
signature. With a password, sshd hands over to **PAM**, which checks
`/etc/shadow` - or asks SSSD, LDAP or Kerberos. The password is also where
network-dependent delays appear.

**How it is measured:** the same login `LOGIN_SAMPLES` times with each
method. The password comes from `lib/askpass.sh` (`SSH_ASKPASS_REQUIRE=force`),
so no keyboard is needed.

**How to read it:**

| Result | Meaning |
|---|---|
| All three within a few ms | normal with local users (Example: 140 ms / 140 ms / 143 ms p50) |
| Password much slower | the PAM stack is slow: SSSD, LDAP or Kerberos in `password-auth`, or `pam_faillock` writing to a slow disk |
| RSA much slower than ed25519 | the *client* is slow (signing with RSA costs more than verifying) |

**How to improve:** use keys (ed25519) for people and automation. For
central accounts, make SSSD cache credentials (`cache_credentials = True`)
so logins work, and stay fast, when the directory server is down.

---

### 4.4 Login rate

**What:** complete logins per second from `RATE_WORKERS` (8) clients that
log in again and again.

**Why:** automation, monitoring checks, backup scripts and CI jobs log in
far more often than people. A server that runs out of login capacity makes
all of them slow at once.

**How it is measured:** 8 clients, each running `ssh sshperf-test true` in
a loop for `DURATION` seconds. The report also shows the CPU used per login,
and how many logins per second this machine's cores could do for sshd alone.

**How to read it:** with a fixed number of clients, each waiting for its
own login, the rate is about **clients / login time**
(Example: 8 / 0.14 s = 55 logins/s). It shows that the server keeps its
login time under parallel load. The *capacity* of the machine is higher;
it is limited by CPU (Example: 48 ms CPU per login -> about 670 logins/s on
32 cores, section 4.11).

8 clients is below the `MaxStartups` start value (10) on purpose: more
parallel logins than that are refused at random (section 4.5).

**How to improve:** re-use connections (ControlPersist) so fewer logins are
needed; raise `MaxStartups` if many logins arrive in parallel; faster PAM.
To measure a higher rate, raise both:
`MAX_STARTUPS=100:30:200 ./start-ssh.sh && RATE_WORKERS=64 ./test-ssh.sh rate`.

---

### 4.5 Burst of logins (MaxStartups)

**What:** how many clients can log in **at exactly the same moment**
without one being refused.

**Why:** real load comes in bursts: `ansible -f 50`, a parallel ssh tool, a
cron job that runs on 200 machines at 02:00, all users after a network
outage. sshd's `MaxStartups` refuses connections above a limit, and the
refused client just fails with:

```
kex_exchange_identification: read: Connection reset by peer
kex_exchange_identification: Connection closed by remote host
```

**How it is measured:** for each step (5, 10, 20, 50, 100 clients), all
`ssh` processes are started first and wait; then all are released at once.
The result is the highest step where every client logged in.

**How to read it:** `MaxStartups start:rate:full` (RHEL default `10:30:100`)
counts connections that are **not yet logged in**. Up to *start* (10) all
are accepted. Above it, each new connection is refused with *rate* (30%)
probability, rising to 100% at *full* (100). Example from the validation run:

| Clients at once | `MaxStartups 10:30:100` (default) | `MaxStartups 100:30:200` |
|---:|---:|---:|
| 10 | 10 of 10 | 10 of 10 |
| 20 | 15 of 20 | 20 of 20 |
| 50 | 32 of 50 | 50 of 50 |
| 100 | 54 of 100 | 100 of 100 |

So with the default, **more than 10 simultaneous logins will randomly fail.**
This is intended protection, not a fault - but it surprises many automation
setups.

**How to improve:**

- Automation hosts / jump hosts: `MaxStartups 100:30:200` (or at least the
  number of parallel clients, e.g. Ansible `forks`).
- Keep `LoginGraceTime` short (e.g. `30`) so half-open or hostile
  connections do not hold the slots for the default 120 s.
- Clients: retry (`ConnectionAttempts 3` in `ssh_config`), and re-use
  connections (section 4.7).
- Note: OpenSSH 9.8 and newer (RHEL 10) also have `PerSourcePenalties` and
  `PerSourceMaxStartups`, which limit per client address. RHEL 9.6 has
  OpenSSH 8.7, which does not.

---

### 4.6 Concurrent sessions

**What:** whether `CONCURRENT_SESSIONS` (200) sessions can be open at the
same time, and all still work.

**Why:** jump hosts, terminal servers, file servers with SFTP users and
build machines keep many sessions open for hours. Each session costs
processes, memory and file descriptors.

**How it is measured:** 200 sessions are opened (at most 8 logging in at a
time, so MaxStartups never interferes). Each runs `cat` on the server. At
the end, "ping" is sent through every session; a session that sends it back
works end to end.

**How to read it:** PASS means every session opened and still worked. Where
sessions fail, the ssh errors are in `raw/concurrent-ssh-errors.txt`. Typical
limits, in the order they are usually reached:

| Limit | Where | Symptom |
|---|---|---|
| processes per user | `ulimit -u`, `/etc/security/limits.d/20-nproc.conf` (4096) | `fork: Resource temporarily unavailable` |
| tasks per user (systemd) | `UserTasksMax` in `logind.conf`, `TasksMax=` of `user-UID.slice` | same |
| memory | section 4.12 | OOM killer ends sessions |
| `MaxSessions` (10) | per **connection**, only for multiplexed sessions | `channel N: open failed` |

**How to improve:** raise the limit that was reached. `MaxSessions` does
**not** limit the number of users or connections - only sessions sharing one
connection.

---

### 4.7 Command over an open connection

**What:** the time of one command sent through a connection that is
already logged in (`ControlMaster` / `ControlPersist`).

**Why:** this is how Ansible (default `ControlPersist=60s`), many scripts and
IDE remote plugins work. Only a new channel, the PAM session and the command
remain - no TCP connect, key exchange or authentication.

**How it is measured:** one master connection is opened, then 200
`ssh sshperf-test true` run through it.

**How to read it:** compare with the login time. Example: **21 ms** instead
of 140 ms for a full login, about 7x faster. What remains is mostly the PAM
session and starting the user's shell.

**How to improve:** use it. In `~/.ssh/config`:

```
Host *
    ControlMaster auto
    ControlPath ~/.ssh/cm-%C
    ControlPersist 10m
```

If many sessions share one connection, raise `MaxSessions` on the server
(default 10).

---

### 4.8 Keystroke echo

**What:** the time from typing one character until it comes back from the
server - what an interactive user feels.

**Why:** the character is not shown locally: it travels to the server, the
shell echoes it, and it travels back. Slow echo makes a remote shell feel
"sticky".

**How it is measured:** one session runs `cat`; single bytes are sent 100
times per second (a very fast typist), 500 in total, and each round trip is
timed.

**How to read it:** on the same machine it measures only sshd's own
delay: Example p50 0.09 ms, p99 0.23 ms. On a real network add the round-trip
time (`ping`). Users notice delays above about 100 ms. Occasional spikes
point to CPU starvation (look at `cpu` and the server's load) or to the
server's disk (shell history, logging).

**How to improve:** sshd turns on `TCP_NODELAY` for interactive sessions by
itself. Keep the server from being CPU-saturated; on high-latency links,
tools that echo locally (mosh) help, but are not part of RHEL.

---

### 4.9 Transfer speed

**What:** MB/s through **one** SSH connection, upload and download, for
every cipher allowed by the crypto policy. No disk is involved.

**Why:** backups, `scp`/`rsync -e ssh`, database dumps and log shipping all
go through SSH. The cipher decides how much CPU each byte costs.

**How it is measured:** `dd if=/dev/zero bs=1M count=1024 | ssh -c <cipher>
sshperf-test 'cat > /dev/null'` and the reverse, for each cipher.

**How to read it:** one connection encrypts on one CPU core per side, so the
result is **one core's speed**. Example (Ryzen 9 5950X, `DEFAULT` policy):

| Cipher | MB/s | Note |
|---|---:|---|
| `aes256-gcm@openssh.com` | ~1,050 | default, uses AES-NI, integrity built in |
| `aes128-gcm@openssh.com` | ~1,140 | slightly faster than AES-256 |
| `aes256-ctr` / `aes128-ctr` | ~780 | needs a separate MAC (`hmac-sha2-256-etm`) |
| `chacha20-poly1305@openssh.com` | ~640 | the best choice on CPUs *without* AES-NI |

A result far below a 1 Gbit/s network (≈ 115 MB/s) is a problem; on
10 Gbit/s networks one SSH stream often cannot fill the link.

**How to improve:** use GCM ciphers (the default). Run several transfers in
parallel for more throughput (several cores). Do not use `-C`
(compression) on fast networks - it costs more CPU than it saves; it helps
only on slow links with compressible data. In FIPS mode
(`fips-mode-setup`) `chacha20-poly1305` is not allowed; this does not
matter on CPUs with AES-NI.

---

### 4.10 SFTP

**What:** the speed of `sftp`: one large file (`SFTP_FILE_MB`, 256 MB) up
and down, and many small files (`SFTP_SMALL_FILES`, 1000 x 4 KB) up.

**Why:** SFTP is the usual way to exchange files with a Linux server
(WinSCP, FileZilla, `sftp` in scripts). Small files behave very differently
from large ones.

**How it is measured:** `sftp -b <batch file>` against the test user's home
directory (on disk). Each sftp run includes one login; the time of an sftp
run that only logs in is subtracted.

**How to read it:** Example: large file 727 MB/s up, 890 MB/s down; small
files about 4,400 files/s.

- **Large files** are limited by the cipher (section 4.9) and the disk.
- **Small files** are limited by **round trips**: each file needs several
  requests (open, write, close) that each wait for an answer. On a network
  with 1 ms round-trip time, 3 round trips per file allow at most ~330 files/s,
  whatever the server's speed.

**How to improve:** pack many small files into one archive (`tar`) before
sending; or run several sftp sessions in parallel. For large files, larger
request windows (`sftp -R 128 -B 262144`) help on long-distance links.

---

### 4.11 CPU

**What:** sshd CPU time **per login** (from the `rate` run) and **per GB**
transferred (from the `transfer` run, default cipher).

**Why:** it tells how many logins and how much traffic a server can carry,
and how much an SSH gateway or SFTP server needs to be sized for.

**How it is measured:** CPU time (user + system) of the listener and all
per-connection processes, from `/proc/<pid>/stat`, before and after the run.
It includes the processes that already ended (their parent collected their
CPU time) and the command each login ran.

**How to read it:** Example: **48 ms per login**, **1.0 s per GB**. A login
costs far more than its cryptography alone: sshd starts several processes
(the per-connection sshd, the user-side sshd, the user's shell), runs PAM,
and writes the login records. Rules of thumb:

    CPU cores needed  ≈  logins per second x CPU per login
                       +  MB per second x (CPU per GB / 1024)

    Example: 20 logins/s x 0.048 s + 500 MB/s x (1.0 s / 1024) ≈ 0.96 + 0.49 ≈ 1.5 cores

The ssh *client* also uses CPU (Example: 12 ms per login) - on the same
machine in this test, on the users' machines in real life.

**How to improve:** fewer logins (connection re-use); lighter PAM
(`pam_motd` scripts, `pam_lastlog`, slow `pam_exec` hooks); AES-GCM ciphers.

---

### 4.12 Memory

**What:** sshd's memory at rest (only the listener) and per open session.

**Why:** tells how many sessions fit into the server's RAM.

**How it is measured:** during the `concurrent` run: the **PSS** (proportional
set size) of all sshd processes and the commands they run, from
`/proc/<pid>/smaps_rollup`. PSS splits shared memory (program code, libraries)
between the processes that share it, so adding them up gives the real total.

**How to read it:** Example: **4 MB at rest**, **about 3 MB per session**
(the session's sshd processes plus a small command). A real interactive
session adds the user's shell (bash: 3-5 MB) and whatever the user runs.

    memory for N sessions  ≈  at rest + N x per session
    Example: 500 sessions  ≈  4 MB + 500 x 3 MB  ≈  1.5 GB (plus the users' own programs)

**How to improve:** memory per session is small and hardly tunable; plan
RAM for the number of sessions plus the users' programs.

---

### 4.13 Reload

**What:** after `systemctl reload sshd`: how long until new logins work,
and whether open sessions survive.

**Why:** configuration changes (new `AllowGroups`, ciphers, `Match` blocks)
are applied with a reload - often over an SSH session that must not break.

**How it is measured:** `RELOAD_SESSIONS` (20) sessions are opened, the test
server is reloaded (`SIGHUP`: sshd starts its own program again and reads the
configuration and host keys), then the time until a new listening socket
exists and until a new login works is measured. Finally every open session
is checked.

**How to read it:** Example: listening again after 19 ms, a new login
works after 159 ms (one login time), **20 of 20 sessions survived**. Open
sessions belong to their own processes, so a reload or even a restart
never ends them.

**How to improve / be careful:** a reload with an **invalid** configuration
stops the listener: no new logins until it is fixed, and systemd retries
only after 42 s (`RestartSec=42s`). Always run `sshd -t` first:

```bash
sshd -t && systemctl reload sshd
```

---

## 5. How the metrics relate to each other

```
      login time  ---- x parallel clients ---->  login rate  ---- limited by ---->  CPU per login
          |                                          |
          |  re-use the connection                   |  more than MaxStartups at once
          v                                          v
   command over an open connection            burst: logins refused at random

      open sessions  ---- x memory per session ---->  RAM needed
      cipher speed   ---- x parallel transfers ---->  throughput (until disk/network)
      sftp small files ---- limited by ---->  round trips (network latency)
```

- **Login time and login rate:** with N clients each waiting for its login,
  rate ≈ N / login time. More clients raise the rate until the **CPU** is
  full or **MaxStartups** refuses them.
- **Connection re-use changes everything:** a command over an open
  connection costs a fraction of a login, in time and CPU.
- **Transfer speed is per connection:** it is the speed of one core; SFTP
  large files follow it, SFTP small files follow the round-trip time instead.

---

## 6. How much performance do you need?

Examples for an offline company network. Compare them with your report.

| Use of the server | Logins | Sessions open | Data | What to check |
|---|---|---|---|---|
| Normal server, admins log in | a few per hour | < 20 | small | login time; `reload` (safe changes) |
| Ansible target, 50 forks | bursts of 50 | 50 short | small | **burst** (`MaxStartups >= 50`), login time |
| Monitoring via SSH (every 60 s, 1000 checks) | ~17 per second | few | small | login rate, CPU per login; better: ControlPersist |
| Jump host for 300 users | bursts at 9:00 | 300 long | small | burst, concurrent, memory, keystroke echo |
| SFTP server for backups, 1 Gbit/s | a few | 10 | 100 MB/s+ | transfer, sftp, CPU per GB, disk |
| Build/CI host, git over SSH | 5-50 per second | many short | medium | login rate, CPU per login, burst |

Rough capacity of one server from the report:

- **logins per second** ≈ CPU cores x 1000 / (CPU per login in ms)
- **simultaneous logins** without refusals = first `MaxStartups` value
- **sessions** ≈ free RAM / (memory per session + the users' programs)
- **MB/s per connection** = transfer speed; more with parallel connections

---

## 7. sshd settings that affect performance

All in `/etc/ssh/sshd_config` (or a file in `/etc/ssh/sshd_config.d/`).
Check the effective values with `sshd -T`, and a change with `sshd -t`.

| Setting | RHEL 9 default | Effect | When to change |
|---|---|---|---|
| `MaxStartups` | `10:30:100` | logins in progress before sshd refuses at random | automation, jump hosts: `100:30:200` |
| `LoginGraceTime` | `120` | seconds an unauthenticated connection may hold a slot | `30` on exposed servers |
| `MaxSessions` | `10` | sessions over one multiplexed connection | clients with heavy ControlMaster use |
| `UseDNS` | `no` | reverse DNS lookup of every client | keep `no`, especially offline |
| `GSSAPIAuthentication` | `yes` | Kerberos login; can cause DNS/KDC timeouts | `no` where Kerberos is not used |
| `UsePAM` | `yes` | account checks, limits, SELinux, logind | keep `yes` on RHEL; keep PAM modules fast |
| `PrintMotd` / `pam_motd` | `no` / on | message of the day, scripts in `/etc/motd.d` | avoid slow motd scripts |
| `Ciphers`, `KexAlgorithms` | from the crypto policy | CPU per login and per byte | change the policy, not sshd_config: `update-crypto-policies` |
| `Compression` | `yes` (only if the client asks) | CPU vs bytes | keep; clients should not use `-C` on LANs |
| `ClientAliveInterval` | `0` | detect dead clients | `300` on servers with many long sessions |
| `AuthorizedKeysCommand` | not set | external key lookup (e.g. SSSD) at every login | make sure it is fast and cached |

On RHEL 9, cipher and key-exchange settings in `sshd_config` are
**overridden** by the crypto policy include; change them with
`update-crypto-policies --set <POLICY>` (and a reboot) instead.

---

## 8. SSH on an offline network

Most "SSH is slow" problems on an offline network are **timeouts** of
services that try to reach something that is not there:

| Symptom | Cause | Fix |
|---|---|---|
| Every login waits exactly 5, 10 or 30 s before the password prompt | the **client** looks up Kerberos principals / DNS names (`GSSAPIAuthentication yes` in `/etc/ssh/ssh_config.d/50-redhat.conf`) | client: `GSSAPIAuthentication no` (or a working DNS and KDC) |
| Wait after the password / key | the **server** does a reverse DNS lookup (`UseDNS yes`) | `UseDNS no` (RHEL 9 default) |
| Wait after authentication | PAM: SSSD cannot reach LDAP/AD, `pam_faillock`, a slow `pam_exec` | SSSD `cache_credentials`, short timeouts; fix the directory server |
| Wait before the shell prompt | `/etc/profile.d/*` or `.bashrc` scripts that reach the network (subscription-manager, proxies) | remove or guard those scripts |
| `kex_exchange_identification` errors under load | MaxStartups refusals (section 4.5) | raise `MaxStartups` |

Also useful offline:

- **Host keys**: distribute `ssh_known_hosts` (or SSH certificates) with
  your configuration management, so users are never asked to trust a key.
- **Time**: SSH does not need exact time, but **SSH certificates** and
  **Kerberos** do. Run an NTP server inside the offline network (see the
  Chrony kit).
- **Crypto policy**: `FIPS` mode removes chacha20 and curve25519; the
  `transfer` and `login` tests show the effect.

---

## 9. Limits of this test

- **Clients and server share one machine.** Both sides' encryption uses the
  same CPUs, and there is no network. Real logins add the network's
  round-trip time several times (about 6-8 round trips per login).
- **Loopback TCP.** The ~40 ms delayed-ACK waits seen in the login phases
  (section 4.2) depend on the TCP stack and the network; they can look
  different across a real network.
- **Local user.** The test user is in `/etc/passwd`. With SSSD, LDAP or
  Kerberos users, the authentication and session phases take longer - test
  with such a user by changing `TEST_USER` (the user must accept the kit's key).
- **Non-interactive sessions.** The tests run `true` or `cat`, not a login
  shell, so `/etc/profile`, `.bash_profile` and motd scripts are not
  measured. `ssh -t server 'exit'` shows their effect by hand.
- **OpenSSH versions.** RHEL 9.6 has OpenSSH 8.7. The kit was validated with
  OpenSSH 9.9 (AlmaLinux 9.8), which starts a separate `sshd-session`
  program per connection and has `PerSourcePenalties` (switched off by the
  kit, since all test clients share 127.0.0.1). CPU and memory per session
  can differ a little between versions.
- **The test server's systemd service.** The service
  (`sshd-perf-test.service`) is a copy of RHEL's `sshd.service`; on a host
  without systemd (a container) sshd is started directly.

---

## 10. Glossary

| Term | Meaning |
|---|---|
| **sshd** | the OpenSSH server daemon |
| **Listener** | the main sshd process; it only accepts connections and starts a process for each |
| **Greeting / banner** | the first line the server sends: `SSH-2.0-OpenSSH_8.7` |
| **Key exchange (KEX)** | both sides agree on encryption keys (e.g. `curve25519-sha256`) and the server proves its identity with its host key |
| **Host key** | the server's own key pair (`/etc/ssh/ssh_host_*_key`); clients remember it in `known_hosts` |
| **Authentication** | the user proves who they are: a key in `~/.ssh/authorized_keys`, a password (PAM), or Kerberos (GSSAPI) |
| **PAM** | Pluggable Authentication Modules (`/etc/pam.d/sshd`): passwords, account checks, limits, SELinux, logind |
| **Session / channel** | a shell, command or SFTP subsystem inside one connection; one connection can carry several |
| **ControlMaster / ControlPersist** | client feature: keep one logged-in connection open and run later commands through it |
| **MaxStartups** | limit of connections that have not logged in yet; above it sshd refuses new ones at random |
| **Cipher** | the encryption algorithm for data (`aes256-gcm@openssh.com` ...) |
| **MAC** | message authentication code, the integrity check for ciphers without one built in (`-ctr`) |
| **Crypto policy** | RHEL's system-wide choice of allowed algorithms (`DEFAULT`, `FIPS`, `FUTURE`, `LEGACY`) |
| **SFTP** | the SSH file transfer protocol (`sftp`, WinSCP); runs as a subsystem inside an SSH session |
| **PSS** | proportional set size: memory with shared pages split between the processes sharing them |
| **p50 / p95 / p99** | percentiles: 50 / 95 / 99 of 100 measurements were at or below this value |
| **Round trip (RTT)** | the time for a message to go to the other side and an answer to come back |
| **Nagle / delayed ACK** | two TCP features that save packets; together they can delay small messages by ~40 ms |
