# DNS Server - Performance Metrics

This document explains the performance metrics of a DNS server, using BIND
(`named`, the DNS server shipped with RHEL 9.6) as the example: what each metric
means, why it matters, how `test-dns.sh` measures it, and what to change when
a result is poor.

---

## Contents

1. [The metrics at a glance](#1-the-metrics-at-a-glance)
2. [How the tests work](#2-how-the-tests-work)
3. [The metrics in detail](#3-the-metrics-in-detail)
   - [3.1 Startup time](#31-startup-time)
   - [3.2 Throughput](#32-throughput-queries-per-second)
   - [3.3 Latency](#33-latency-response-time)
   - [3.4 Concurrency scaling](#34-concurrency-scaling)
   - [3.5 Query loss](#35-query-loss)
   - [3.6 TCP performance](#36-tcp-performance)
   - [3.7 Cache performance](#37-cache-performance)
   - [3.8 Zone transfer time](#38-zone-transfer-time-axfr)
   - [3.9 CPU usage](#39-cpu-usage)
   - [3.10 Memory usage](#310-memory-usage)
4. [How the metrics relate to each other](#4-how-the-metrics-relate-to-each-other)
5. [BIND settings that affect performance](#5-bind-settings-that-affect-performance)
6. [Limits of this test](#6-limits-of-this-test)
7. [Glossary](#7-glossary)

---

## 1. The metrics at a glance

| # | Metric | Unit | Question it answers | Better is | Test name |
|---|---|---|---|---|---|
| 1 | Startup time | ms | How fast does DNS come back after a (re)start? | lower | `startup` |
| 2 | Throughput | queries/s (QPS) | How many queries can the server answer per second? | higher | `throughput` |
| 3 | Latency | ms | How long does one client wait for an answer? | lower | `latency` |
| 4 | Concurrency scaling | % of peak | Does the server hold up as more queries arrive at once? | higher | `concurrency` |
| 5 | Query loss | % | How many queries get no answer, or a failure, at a high rate? | lower | `loss` |
| 6 | TCP performance | queries/s | How well does the server handle DNS over TCP? | higher | `tcp` |
| 7 | Cache performance | ms, queries/s | How much faster are answers the server already remembers? | lower hit latency | `cache` |
| 8 | Zone transfer time | s | How fast can a secondary server copy the whole zone? | lower | `axfr` |
| 9 | CPU usage | % of all CPUs | How much processor time does the load cost? | lower | `cpu` |
| 10 | Memory usage | MB | How much RAM does named need, idle and busy? | lower | `memory` |

Run one metric with `./test-dns.sh <test name>`, or all of them with
`./test-dns.sh`.

**The two jobs of a DNS server.** Most metrics apply to both, but it helps to
know which one a test looks at:

- **Authoritative server:** owns a zone (here `perf.test`) and answers
  questions about it from its own data. Tests: throughput, latency,
  concurrency, loss, tcp, axfr.
- **Caching (recursive) server:** answers questions about *other* zones by
  asking other servers, then remembers the answer. Test: cache.

---

## 2. How the tests work

```
  +----------------------+   queries over 127.0.0.1   +--------------------------------+
  |  lib/dnsload.py      |  ------------------------> |  named (BIND)                  |
  |  load generator,     |  <------------------------ |                                |
  |  N queries in flight |          answers           |  view "perf-main"              |
  +----------------------+                            |    perf.test      own zone     |
             |                                        |    upstream.test  forwarded -+ |
             |  queries/s, latency,                   |                              | |
             |  lost queries, answer codes            |  view "perf-upstream"        | |
             v                                        |    upstream.test  <----------+ |
  results/<date-time>/report.md                       |    (plays "the internet")      |
  (PASS / FAIL per metric)                            +--------------------------------+
                                                        CPU, memory: cgroup and /proc
                                                        counters: statistics channel
```

- **Load generator:** `lib/dnsload.py`, part of this kit. The well-known tools
  `dnsperf` and `resperf` are not in RHEL (only in EPEL), so the kit brings its
  own. It needs only Python 3, which every RHEL 9 system has. It reads the same
  query-file format as `dnsperf` (`name type` per line).
  - It keeps a fixed number of queries **in flight** (sent, not yet answered):
    as soon as an answer arrives, it sends the next query. This is the DNS
    version of "number of simultaneous users".
  - Or it sends at a fixed **rate** (queries per second), like steady real traffic.
  - It runs several worker processes (`LOAD_PROCESSES`), each on its own CPU core.
    On a modern CPU one process sends roughly 100,000 - 150,000 queries per second.
- **Test data:** created by `start-dns.sh` in `/var/named/perf-test/`:
  - zone `perf.test`: `ZONE_RECORDS` hosts (100,000 by default), each with an
    IPv4 (A) and an IPv6 (AAAA) record, plus MX, TXT and CNAME records.
  - zone `upstream.test`: a wildcard zone in which every name exists. It is
    served by a second *view* of the same named process, which plays the part
    of the internet for the cache test (see the picture above).
- **Query mix:** most tests use a fixed list of 100,000 queries, mixed like
  everyday traffic (the list is saved in `raw/queries-mixed.txt`):

  | Share | Query | Answer |
  |---:|---|---|
  | 70% | `host<N>.perf.test A` | IPv4 address |
  | 10% | `host<N>.perf.test AAAA` | IPv6 address |
  | 5% | `perf.test MX` | mail server |
  | 5% | `perf.test TXT` | text record |
  | 10% | `missing<N>.perf.test A` | NXDOMAIN ("this name does not exist") |

- **Server-side measurements:**
  - CPU time from the `named.service` cgroup (`cpu.stat`), once per second.
  - Memory (PSS) from `/proc/<pid>/smaps_rollup`, once per second.
  - named's own counters (queries received, cache hits, cache memory) from
    its *statistics channel*, a small built-in web page with JSON data on
    `127.0.0.1:8953`.
- **Offline:** all traffic uses the loopback address `127.0.0.1`. No internet
  access and no second machine are needed.
- **Warm-up:** 3 seconds of light, unmeasured load run first, so the first test
  does not start "cold".

---

## 3. The metrics in detail

Each metric below has the same five parts: **what** it is, **why** it matters,
**how** the test measures it, **how to read** the result, and **how to improve** it.

### 3.1 Startup time

**What:** the time from the command that starts named until the first DNS
answer is sent.

**Why:** while named starts, clients get no answers, and every service that
needs name lookups waits. Startup time grows with the size of the zones,
because named reads and checks every zone file before it answers.

**How it is measured:** named is stopped and started `STARTUP_ROUNDS` times
(3 by default) with `systemctl stop named` / `systemctl start named`. The timer
runs from the start command until a query for `perf.test SOA` is answered.
The test also times one `rndc reload` (read the configuration and changed
zones again without a restart).

**How to read it:**

| Result | Meaning |
|---|---|
| under 1 s | normal for small and medium zones (up to ~100,000 hosts) |
| 1 - 10 s | acceptable for large zones or many zones |
| over 10 s | investigate: huge zones, slow disk, many DNSSEC keys, very large configuration |

A **reload** should be much faster than a restart, and named keeps answering
with the old data while it reloads.

**How to improve:** use `rndc reload <zone>` for changes instead of a restart;
split very large zones; for big zones store them in the faster binary format
(`masterfile-format raw;`, converted with `named-compilezone -F raw`); keep the
zone files on a fast disk.

---

### 3.2 Throughput (queries per second)

**What:** how many DNS queries named answers per second, at full speed.

**Why:** this is the capacity of the server. If your clients send 20,000
queries per second at peak, the server needs a throughput well above that,
because traffic comes in bursts.

**How it is measured:** the load generator keeps `OUTSTANDING` queries (100 by
default) in flight for `DURATION` seconds, using the query mix above. The
result is the number of answers divided by the time.

**How to read it:** it depends strongly on the CPU. As a rough guide for an
authoritative server with the query mix above:

| Result | Meaning |
|---|---|
| over 100,000 QPS | a modern multi-core server, far more than most sites need |
| 20,000 - 100,000 QPS | a small VM or a busy server; fine for most companies |
| under 20,000 QPS | investigate (CPU, threads, logging of every query) |

For reference, a busy company DNS server typically sees a few hundred to a
few thousand queries per second.

The report also shows the **load generator's CPU**. When it is close to 100%,
the load generator (not named) was the limit; raise `LOAD_PROCESSES`.

**How to improve:** make sure named uses all CPU cores (it does by default,
one thread per core); turn off query logging (`querylog no;`); avoid
per-query work such as response policy zones (RPZ) or complex ACLs if not
needed; see section 5.

---

### 3.3 Latency (response time)

**What:** the time from sending one query until its answer arrives.

**Why:** almost every network action starts with a DNS lookup: opening a
web page, sending a mail, connecting to a database. A slow DNS answer delays
all of them. Users feel latency, not throughput.

**How it is measured:** a steady `LATENCY_TEST_QPS` queries per second (10,000
by default) are sent for `DURATION` seconds, a normal load rather than the
maximum. Every answer's delay is recorded, and the report shows the average
and the **percentiles**.

**Why percentiles and not only the average:** "p95 = 0.5 ms" means 95 of every
100 answers came within 0.5 ms. The average hides the slow answers; p95 and
p99 show how bad it is for the unlucky clients.

**How to read it:** over localhost, the network adds almost nothing, so this
is named's own processing time.

| Result (p95) | Meaning |
|---|---|
| under 1 ms | normal for an authoritative server answering from memory |
| 1 - 5 ms | acceptable; named is busy or the machine is shared |
| over 5 ms | investigate: CPU saturated, swapping, heavy logging |

On a real network, add the network round-trip time (typically 0.2 - 1 ms in a
data centre).

**How to improve:** keep CPU usage below ~70% at peak; avoid swapping (memory);
turn off query logging; place the server close (in network terms) to its clients.

---

### 3.4 Concurrency scaling

**What:** how throughput and latency change as more queries are in flight at
the same time (1, 10, 50, 100, 200, 500, 1000 by default).

**Why:** real traffic comes in bursts. A good server delivers *more* answers
per second as load rises, until the CPUs are full, and then stays at that
level. A poor one *collapses*: it answers fewer queries as load grows.

**How it is measured:** the throughput test is repeated at each level in
`OUTSTANDING_LEVELS`. The verdict compares the throughput at the highest level
with the best (peak) throughput of all levels.

**How to read it:**

```
 queries/s
    ^            ____________________    <- healthy: rises, then levels off
    |          /
    |        /       \
    |      /           \_____            <- collapse: falls under overload
    |    /
    +------------------------------->  queries in flight
```

- Throughput rising and then flat is healthy. The level where it flattens is
  the point where the server (or load generator) is fully busy.
- Latency grows in proportion once throughput is flat (see section 4).
- A drop below `TARGET_SCALING_MIN_PCT` (50%) of peak means the server
  collapses under overload.
- A few "failed" (lost) queries at the highest levels usually mean the UDP
  receive buffers were full; see 3.5.

**How to improve:** more CPU cores; larger UDP receive buffers (see 3.5);
check that no firewall connection tracking is involved (see section 6).

---

### 3.5 Query loss

**What:** the share of queries that got **no answer** (lost) or a **failure
answer** (SERVFAIL, REFUSED), while a fixed high rate is sent.

**Why:** DNS normally uses UDP, which has no delivery guarantee. When a server
is overloaded, queries are dropped silently. The client only notices after a
timeout (typically 1 - 5 seconds) and then asks again, so one lost query costs
the user seconds, not milliseconds. Loss is the most user-visible sign of an
overloaded DNS server.

**How it is measured:** `LOSS_TEST_QPS` queries per second (50,000 by default)
are sent for `DURATION` seconds. A query without an answer after
`QUERY_TIMEOUT` seconds (2) counts as lost. The report also shows where the
loss happened:

| Report line | What it tells you |
|---|---|
| Queries named received | named's own counter. Lower than "sent" = dropped **before** named (kernel buffer full) |
| Dropped by the kernel | UDP receive-buffer overflows (`RcvbufErrors` in `/proc/net/snmp`) |
| SERVFAIL / REFUSED | named answered, but with an error |

NXDOMAIN ("name does not exist") is a correct answer and is **not** counted
as a failure.

**How to read it:** the target is at most 0.1%. Set `LOSS_TEST_QPS` to your
expected peak (or twice it, for a safety margin). Any loss at your normal
peak rate is a problem.

**How to improve:**
- Kernel drops: raise the UDP buffers, for example
  `sysctl -w net.core.rmem_max=8388608 net.core.rmem_default=8388608`.
- named drops: more CPU, or more servers behind a load balancer / anycast.
- SERVFAIL on a caching server: the upstream servers are slow or unreachable.
- REFUSED: the client is not allowed by `allow-query` / `allow-recursion`.

---

### 3.6 TCP performance

**What:** throughput over TCP, in two ways: connections reused for many
queries, and a new connection for every query (the worst case, which has a
target).

**Why:** DNS uses TCP when an answer does not fit in a UDP packet (common with
DNSSEC and large TXT records), for zone transfers, and for clients that
require it. DNS-over-TLS also runs on TCP. A TCP connection needs a handshake
and uses server memory, so TCP costs much more than UDP.

**How it is measured:** three runs with `OUTSTANDING` queries in flight: UDP
(for comparison), TCP with reused connections, and TCP with a new connection
per query. Each in-flight query has its own connection.

**How to read it:** TCP with a new connection per query usually reaches 10 -
30% of the UDP rate; with reused connections 30 - 60%. This is normal. A very
low value, or many failed queries, points to the `tcp-clients` limit.

**How to improve:** raise `tcp-clients` (BIND default 150; the kit uses
`TCP_CLIENTS` = 1000); keep answers small so they fit in UDP (minimal
responses, `minimal-responses yes;`); make sure `net.core.somaxconn` is large
enough for many new connections.

---

### 3.7 Cache performance

**What:** how fast a caching server answers names it already has in its cache
(a **hit**), compared with names it must first fetch from another server (a
**miss**).

**Why:** a caching server answers most queries from memory. The hit ratio and
the speed of hits decide how fast the network feels for everyone; misses are
slow because they wait for other servers on the internet.

**How it is measured:**

1. The cache is emptied (`rndc flush`).
2. **Miss phase:** `CACHE_TEST_NAMES` (50,000) names that were never asked
   before are each asked once. Every one is forwarded to the "upstream" view
   and then stored in the cache.
3. **Hit phase:** the same names are asked again for `DURATION` seconds. Every
   one is now answered from the cache.

The report shows speed and latency of both phases, and named's own counters:
cache hit ratio, number of cached names and cache memory.

**How to read it:**

- Hits should be many times faster than misses, and the hit-phase hit ratio
  should be ~100%.
- The "upstream" here is on the same machine, so a miss costs only named's own
  work. **On a real network each miss also waits 10 - 100 ms for the
  internet**, so in production the difference is far larger.
- In production, a good caching server has a hit ratio of 80 - 95%.

**How to improve:** give the cache enough memory (`max-cache-size`); keep
named running (a restart empties the cache); consider `prefetch` (on by
default in BIND 9.16), which refreshes popular names before they expire.

---

### 3.8 Zone transfer time (AXFR)

**What:** the time to transfer the complete zone `perf.test` to a client, as a
secondary server would do.

**Why:** secondary DNS servers keep a copy of the zone and get it through a
zone transfer. Its speed decides how fast a new secondary is ready and how
long a full refresh takes. Very large zones can take minutes.

**How it is measured:** `dig perf.test AXFR` is run `AXFR_ROUNDS` times (3).
The report shows seconds, records, megabytes and records per second.

**How to read it:** hundreds of thousands to millions of records per second
over localhost are normal. Over a real network, bandwidth and distance also count.

**How to improve:** use incremental transfers (IXFR, only the changes) for day
to day updates, which BIND does automatically for dynamic zones; use
`transfer-format many-answers;` (the default); give the transfer enough
bandwidth.

---

### 3.9 CPU usage

**What:** the share of all CPU cores that named used during a full-speed run,
and the CPU time per query.

**Why:** CPU is normally the resource that limits a DNS server. Knowing the
CPU cost per query tells you how many queries per second the machine can
handle before it is full.

**How it is measured:** during one full-speed run, named's CPU time is read
from its systemd cgroup (`cpu.stat`) once per second. `100%` means all cores
were fully busy with named.

**How to read it:**

- "CPU time per query" x "queries per second" = CPU needed. Example: 20 µs per
  query x 50,000 QPS = 1 second of CPU per second = **1 full core**.
- The load generator runs on the same machine and uses CPU too; named can
  use at most the cores that are left.
- Plan to stay below ~70% at your normal peak.

**How to improve:** turn off query logging; remove unneeded features that run
for every query (RPZ, DNS64, complex views); more or faster cores.

---

### 3.10 Memory usage

**What:** the memory named uses, idle and under load, plus how much of it is
the answer cache.

**Why:** named keeps all zones and the cache in RAM. If it runs out, the
system swaps (latency explodes) or the kernel stops the process.

**How it is measured:** PSS memory of the named process, read from
`/proc/<pid>/smaps_rollup` once per second during a full-speed run. PSS counts
memory shared with other programs (like system libraries) only by named's
fair share. Cache memory comes from named's own counters.

**How to read it:**

- Memory = zones + cache + a fixed part per thread. The RHEL build of BIND
  reserves large internal tables at start (built with `--with-tuning=large`),
  so even an idle named with a small zone uses several hundred MB on a
  machine with many cores.
- The cache grows with the number of *different* names looked up, up to
  `max-cache-size`. Run the `cache` test before `memory` to see a filled cache.
- Memory should stay flat under load; steady growth during an authoritative
  load test would point to a leak.

**How to improve:** set `max-cache-size` to a fixed value on caching servers
(for example `max-cache-size 2g;`); reduce `NAMED_THREADS` on machines with
many cores and little memory; remove unused zones.

---

## 4. How the metrics relate to each other

The metrics are not independent. One simple rule, **Little's law**, links
three of them:

```
queries in flight  =  throughput (queries/s)  x  latency (s)
```

Example: 100 queries in flight and 200,000 queries/s means an average latency
of 100 / 200,000 = 0.0005 s = **0.5 ms**. So when throughput stops growing
(CPU full), every extra query in flight only adds waiting time. That is why
latency rises in the concurrency test after the plateau.

Typical chain of cause and effect under growing load:

```
more queries --> CPU reaches 100% --> throughput stops growing
             --> queries wait in the socket buffer --> latency (p99) jumps
             --> buffer full --> queries dropped --> query loss > 0
             --> clients time out and repeat --> even more queries
```

The last step makes DNS overload dangerous: lost queries are repeated by the
clients, which adds load exactly when the server is already full. Keep a
safety margin between your peak rate and the measured throughput.

For a caching server, the **cache hit ratio** sits in front of all of this:
at a 90% hit ratio, only 1 in 10 queries costs a slow upstream lookup.

---

## 5. BIND settings that affect performance

`start-dns.sh` writes `/etc/named/perf-test.conf` using the values in
`settings.conf`. Everything else is the BIND default.

| named.conf option | settings.conf name | Kit value | BIND default | Affects |
|---|---|---|---|---|
| `named -n` (threads) | `NAMED_THREADS` | one per CPU core | one per CPU core | throughput, CPU, memory |
| `max-cache-size` | `MAX_CACHE_SIZE` | 90% | 90% of RAM | cache memory, hit ratio |
| `recursive-clients` | `RECURSIVE_CLIENTS` | 1000 | 1000 | parallel cache misses |
| `tcp-clients` | `TCP_CLIENTS` | 1000 | 150 | parallel TCP connections |
| `dnssec-validation` | (always no) | no | auto | CPU per cache miss |
| `notify` | (always no) | no | yes | messages to secondaries |

Why the kit changes a few defaults:

- **`tcp-clients` 1000:** with the default of 150, the TCP test would measure
  the connection limit instead of named's speed.
- **`dnssec-validation no`:** an offline machine cannot refresh the DNSSEC
  trust anchors, and the test zones are not signed. On a production caching
  server, keep DNSSEC validation on.

Other settings worth knowing (not changed by the kit):

| Option | Effect |
|---|---|
| `querylog` | logs every query; very useful for debugging, but costs noticeable throughput |
| `minimal-responses yes;` | smaller answers: fewer bytes, fewer TCP fallbacks |
| `rate-limit { ... };` | Response Rate Limiting against abuse; limits answers per client |
| `prefetch` | refreshes popular cache entries before they expire (on by default) |
| `max-udp-size` / `edns-udp-size` | largest UDP answer before switching to TCP |
| `masterfile-format raw;` | faster loading of very large zones |
| `clients-per-query`, `max-clients-per-query` | how many identical cache misses wait for one upstream lookup |

Operating-system settings:

| Setting | Effect |
|---|---|
| `net.core.rmem_max`, `net.core.rmem_default` | UDP receive buffer size; too small = lost queries under bursts |
| `net.core.somaxconn` | queue for new TCP connections |
| firewall (nftables) connection tracking | every UDP query creates a tracking entry; on busy DNS servers, exclude port 53 from tracking or raise `nf_conntrack_max` |

---

## 6. Limits of this test

Be aware of what a localhost test can and cannot tell you:

- **Client and server share one machine.** The load generator uses CPU too,
  so the results are lower than a dedicated server would reach. They are a
  **baseline** for comparing configurations and servers, not an exact capacity.
- **The load generator is written in Python.** It is fast enough for most
  servers (about 100,000+ queries/s per process), and the report warns when
  it was the limit. For extreme rates, `dnsperf` from EPEL can be used with
  the same query files (`raw/queries-mixed.txt`).
- **No real network.** Network latency, packet loss and firewalls are not
  included. In particular, the firewall's connection tracking is not in the
  path for loopback traffic, but it often is in production (see section 5).
  Test from a second machine for those.
- **The "internet" is simulated.** Cache misses are answered by the same
  machine, so they are much faster than real ones.
- **No DNSSEC.** Signed zones make answers larger (more TCP) and validation
  costs CPU on caching servers.
- **Other programs on the machine.** Anything else using the CPU (builds,
  virtual machines, backups) makes the results worse and uneven, latency
  most of all. Test on an otherwise idle machine.
- **Short runs.** The default is 10 seconds per test. For a capacity baseline,
  use `DURATION=60` or more and repeat the run 3 times.
- **SELinux.** RHEL 9.6 normally runs SELinux in enforcing mode. The kit uses
  standard RHEL paths and ports (the statistics port 8953 is already allowed
  for named), so the standard policy applies. If something is blocked, check
  `ausearch -m AVC -ts recent`.

---

## 7. Glossary

| Term | Meaning |
|---|---|
| **A / AAAA record** | the IPv4 / IPv6 address of a name |
| **authoritative server** | a server that owns a zone and answers for it from its own data |
| **AXFR** | full zone transfer: a copy of the whole zone, sent over TCP |
| **BIND / named** | the DNS server software in RHEL; `named` is the program (daemon) |
| **cache hit / miss** | the answer was / was not already in the server's memory |
| **caching (recursive) server** | a server that looks up names on other servers for its clients and remembers the answers |
| **cgroup** | Linux control group; systemd puts named in one, which lets the kit read its CPU time exactly |
| **forwarding** | sending a query on to one fixed other server instead of looking it up on the internet directly |
| **in flight (outstanding)** | queries sent but not yet answered |
| **latency** | time from sending a query until its answer arrives |
| **NXDOMAIN** | the answer "this name does not exist" (a correct answer, not an error) |
| **percentile (p95)** | the value that 95% of measurements are below |
| **PSS** | proportional set size: memory with shared pages divided fairly between processes |
| **QPS** | queries per second: throughput of a DNS server |
| **rndc** | the command-line tool that controls a running named (reload, flush, status, stop) |
| **SERVFAIL / REFUSED** | error answers: "the server failed" / "you are not allowed to ask" |
| **statistics channel** | named's built-in web page with counters (JSON), here on `127.0.0.1:8953` |
| **TSIG** | a shared secret key used to sign DNS messages; here it routes forwarded queries to the "upstream" view |
| **TTL** | time to live: how long an answer may be kept in a cache |
| **UDP / TCP** | the two transports of DNS; UDP is fast but has no delivery guarantee |
| **view** | a separate "virtual DNS server" inside one named process, chosen by who is asking |
| **zone** | a part of the DNS name space managed together, stored in one zone file (here `perf.test`) |
