# Guide

How to test the kernel on an offline RHEL 9.6 machine, what the three suites
actually test, and how to read what comes back.

- [The short version](#the-short-version)
- [Getting the bundle onto the machine](#getting-the-bundle-onto-the-machine)
- [The three suites](#the-three-suites)
- [Running the tests](#running-the-tests)
- [Reading the report](#reading-the-report)
- [Settings](#settings)
- [Choosing what runs](#choosing-what-runs)
- [Things that will bite you](#things-that-will-bite-you)
- [Troubleshooting](#troubleshooting)

## The short version

On the connected machine:

```bash
./download-rpms.sh --kver 5.14.0-570.62.1.el9_6.x86_64
./build-ltp.sh
./make-bundle.sh
```

On the offline machine, as root:

```bash
./install-offline.sh
./test-kernel.sh --profile smoke     # prove the plumbing works, ~5 min
./test-kernel.sh                     # the real run, up to 2 hours
cat results/latest/report.md
```

## Getting the bundle onto the machine

### 1. Find out what kernel the target runs

On the offline machine:

```bash
uname -r      # e.g. 5.14.0-570.62.1.el9_6.x86_64
```

This matters more than it looks. `kernel-modules-internal` carries
`Requires: kernel-uname-r = <that exact string>`, because the KUnit modules
are compiled against that one kernel build and load into nothing else. Get
the wrong one and every KUnit module fails with a vermagic error.

### 2. Fetch the RPMs

```bash
./download-rpms.sh --kver 5.14.0-570.62.1.el9_6.x86_64
```

Two packages:

| Package | Gives you | Tied to the kernel? |
|---|---|---|
| `kernel-modules-internal` | the KUnit test modules | **yes**, exactly |
| `kernel-selftests-internal` | `/usr/libexec/kselftests` | no, it is userspace |

**On real RHEL, get these from your own subscription, Satellite or the
installation media**, on a connected RHEL machine of the same z-stream:

```bash
dnf download kernel-modules-internal kernel-selftests-internal
```

Then drop the `.rpm` files into `bundle/rpms/` yourself.

`download-rpms.sh` falls back to the AlmaLinux vault, which is fine for a
lab: AlmaLinux builds from the same sources. But the modules are signed by
AlmaLinux's key, not Red Hat's, so **on a machine with Secure Boot enabled
they will not load**. `install-offline.sh` checks the vendor and says so.

Neither package is required. LTP runs without either, and an engine whose
files are missing is skipped rather than failing the run.

### 3. Build LTP

```bash
./build-ltp.sh
```

Ten minutes or so. It installs the build dependencies, compiles LTP, strips
the debug symbols (1.6 GB down to about 470 MB) and writes
`bundle/ltp-<version>-<arch>.tar.gz`, which unpacks to `/opt/ltp`.

Watch the configure summary it prints. A library reported as `no` means a
whole group of tests silently will not exist in the bundle, and nobody
notices until the target reports far fewer tests than expected.

Build it on the same major release as the target. EL9 is EL9 — a binary
built on 9.8 runs on 9.6, because both are glibc 2.34 — but do not build on
EL8 or EL10 and carry it over.

### 4. Pack and carry

```bash
./make-bundle.sh
```

Then on the offline machine:

```bash
sha256sum -c kernel-test-*.tar.gz.sha256
tar xzf kernel-test-*.tar.gz
cd kernel-test
./install-offline.sh
```

`install-offline.sh` touches no network. It installs only the packages that
are genuinely missing (handing `dnf` a directory of RPMs is an *upgrade*
request and fails on anything already at that version), unpacks LTP, creates
the users LTP's unprivileged tests need, and then **checks that what it
installed actually works** — it loads the KUnit core, lists the selftests and
starts LTP's runner. Run it with `--check` to see the state of the machine
without changing anything.

## The three suites

### KUnit — tests inside the kernel

KUnit tests are kernel modules. Loading the module runs its test suites
during `module_init` and the results come back as KTAP, both in the kernel
log and under `/sys/kernel/debug/kunit/<suite>/results`.

They test kernel internals directly: `overflow_kunit` checks the overflow
helper macros, `cpumask_kunit` the CPU mask code, `siphash_kunit` the hash,
`resource_kunit` resource allocation. They are fast — dozens of suites in
seconds — and they are the only one of the three that tests code paths you
cannot reach from userspace.

RHEL 9 ships about 70 of them, already built, in `kernel-modules-internal`.

**RHEL builds the kernel with `CONFIG_KUNIT_DEFAULT_ENABLED` unset**, so
KUnit refuses to run until it is switched on. The engine loads the core with
`modprobe kunit enable=1` and drops an `options kunit enable=1` file into
`/etc/modprobe.d` for the duration of the run, because a test module can pull
the core in as a dependency and an auto-loaded core gets no options. Without
this every suite silently reports nothing and the kernel log just says
`kunit: disabled`.

### kselftest — the kernel's own functional tests

`/usr/libexec/kselftests`, from `kernel-selftests-internal`: about 285 tests
in 14 collections (`mm`, `cgroup`, `net`, `netfilter`, `bpf`, `livepatch`,
`iommu`, `tc-testing`, ...). These are ordinary programs and shell scripts
that exercise the kernel through real syscalls, and they emit TAP.

Red Hat also ships `prepare_system.sh`, which sets up whatever a collection
needs before it runs, in three escalating tiers: `-s` safe, `-m` also loads
modules, `-d` also makes destructive changes. The engine picks the tier from
the profile, and only ever reaches `-d` under `--profile full --unsafe`.
Red Hat's own warning about `-d` is worth repeating: afterwards the machine
is not useful for anything but selftests.

The engine runs **one test per invocation** rather than a whole collection at
a time. A collection run is faster, but one hung test then costs the whole
collection and there is no way to give a single test its own timeout.

### LTP — breadth

The Linux Test Project: about 2700 programs grouped into 64 scenario files.
`syscalls` alone is over a thousand tests covering the syscall surface and
its error paths. `cve` reproduces specific published vulnerabilities, which
is the one to run when a security erratum lands.

LTP is driven with **kirk**, LTP's own runner, because it writes a JSON
report that can be parsed exactly. `runltp` is used automatically if kirk
cannot start.

## Running the tests

```bash
./test-kernel.sh --check                 # say what would run, run nothing
./test-kernel.sh --profile smoke         # minutes
./test-kernel.sh                         # standard, the default
./test-kernel.sh --profile full --unsafe # everything
./test-kernel.sh --engines "kunit ltp"   # pick engines
./test-kernel.sh -t before-upgrade       # name the run folder
```

Exit status: `0` nothing failed, `1` something did, `2` it could not run at
all.

Start with `--check`. It prints the kernel, SELinux mode, Secure Boot state,
and whether each engine's files are present, without running anything.

### Profiles

| | `smoke` | `standard` | `full` |
|---|---|---|---|
| Time | minutes | under 2 hours | hours |
| KUnit | all modules except drm/vkms/ntb | same | same |
| kselftest | `cachestat`, `memfd` | `+ mm cgroup netfilter iommu net/mptcp` | every collection |
| LTP | `smoketest` | 18 scenarios | 45 scenarios |
| Reconfigures the network | no | no | **yes** |
| Offlines CPUs | no | no | **yes** |
| Taints the kernel | KUnit modules only | KUnit modules only | **yes**, livepatch |
| Needs `--unsafe` | no | no | **yes** |

Exactly what each profile runs is in `conf/`, as plain text lists you can
edit. `conf/ltp-skip.list` and `conf/kselftest-deny.list` hold the tests that
are never run while `KT_SAFE=1` — the ones that drive the machine out of
memory on purpose, fill the filesystem, or offline CPUs and do not always put
them back.

## Reading the report

`results/latest/report.md`:

- **Verdict** — pass or fail, and the tally.
- **By engine** and **By collection** — where the failures are concentrated.
- **Failures** — every failure with its detail line.
- **Slowest tests** — what to cut if a run has to fit a window.
- **Kernel log** — this is the section to read first.

Five results are possible:

| | meaning |
|---|---|
| `PASS` | the test passed |
| `FAIL` | the test ran and did not pass |
| `SKIP` | the test decided it does not apply here — missing hardware, a kernel option that is off, not enough memory |
| `ERROR` | the test could not run: it would not start, or LTP called it `brok` |
| `TIMEOUT` | it was still running when its timeout expired and was killed |

**A lot of SKIPs is normal.** A RHEL kernel has many options off, and a test
whose feature is absent reports itself as skipped. That is the test working
correctly.

**Some FAILs are normal too**, and this is the most important thing to
understand about kernel testing: upstream selftests and LTP are written
against mainline, RHEL's kernel is not mainline, and some tests need
hardware or a network setup this machine does not have. A first run on a
healthy machine typically has a handful of failures.

This is exactly why `test-patch.sh` exists. The number that means something
is not "how many failed" but "what changed".

### The kernel log section

`*.badness` files hold anything matching `BUG:`, `Oops`, `WARNING: CPU`,
`soft lockup`, `KASAN:`, `refcount_t:` or `Call Trace:` that appeared while
a particular test ran.

**A kernel BUG or Oops matters more than any number of failed tests.** A test
that fails tells you a behaviour is wrong; an Oops tells you the kernel just
corrupted its own state, and every result after it on that boot is suspect.

## Settings

Everything lives in `settings.conf`, and every value can be overridden by an
environment variable or a command line flag. The order is: command line beats
environment, environment beats the file.

```bash
KSELFTEST_TIMEOUT=600 ./test-kernel.sh          # per-test timeout
LTP_SUITES="syscalls cve" ./test-kernel.sh      # just these scenarios
KUNIT_MODULES="overflow_kunit" ./test-kernel.sh # just this module
KT_RESULTS_DIR=/srv/results ./test-kernel.sh
```

The ones worth knowing:

| | |
|---|---|
| `KT_PROFILE` | `smoke`, `standard` or `full` |
| `KT_ENGINES` | which of `kunit kselftest ltp` to run |
| `KT_SAFE` | `1` refuses tests marked disruptive; `full` needs `0` |
| `KT_FAIL_FAST` | stop at the first failure |
| `KUNIT_SKIP` | regex of modules never loaded (drm and ntb by default) |
| `KSELFTEST_TIMEOUT` | seconds per selftest |
| `LTP_TIMEOUT` | seconds per LTP test |
| `LTP_SUITE_TIMEOUT` | seconds for a whole LTP scenario |
| `LTP_TMPDIR` | LTP's scratch space — needs several GB |
| `PATCH_FLAKY_LIST` | tests whose failures are not counted as regressions |

## Choosing what runs

Run one thing while you are investigating:

```bash
KUNIT_MODULES="overflow_kunit"          ./test-kernel.sh --engines kunit
KSELFTEST_COLLECTIONS="mm"              ./test-kernel.sh --engines kselftest
LTP_SUITES="cve"                        ./test-kernel.sh --engines ltp
```

`LTP_SUITES="cve"` after a security erratum is the highest-value single run
in this kit.

To change a profile permanently, edit the list in `conf/`. `selftest.sh`
checks every name in those files against what is actually installed, so a
typo is caught before a two-hour run wastes its time on it.

## Things that will bite you

**Secure Boot.** With Secure Boot on, the kernel loads only modules signed by
a key it trusts. KUnit modules from AlmaLinux will not load on a Red Hat
machine. kselftest and LTP are unaffected. `install-offline.sh` reports the
Secure Boot state; `./test-kernel.sh --check` does too.

**The wrong `kernel-modules-internal`.** It must match `uname -r` exactly.
`install-offline.sh` refuses a mismatch rather than letting you find out from
a vermagic error later. `kernel-selftests-internal` is userspace and is only
warned about.

**SELinux.** The target is usually enforcing and the build machine usually is
not. A clean run under permissive does not mean enforcing will be clean.
`env.txt` records the mode for every run. If a test fails only under
enforcing, check `ausearch -m AVC -ts recent`.

**Disk space.** LTP's filesystem tests want several GB in `LTP_TMPDIR`
(`/var/tmp/ltp-scratch` by default). Under 1 GB and they report failures
that are about the disk, not the kernel. Both the installer and the engine
warn.

**The kernel you booted is not the one you installed.** After a kernel RPM
lands, the machine still runs the old kernel until it reboots — and it may
not boot the newest one. `grubby --default-kernel` says which it will take.
`test-patch.sh` checks this for you and refuses to compare results from the
wrong kernel.

**LTP version changes.** Upgrading LTP renames and adds tests, which show up
in a patch comparison as "new" and "gone" and bury the real answer. The
version is pinned in `build-ltp.sh`. Use the same LTP on both sides of a
comparison.

**A venv earlier in `PATH`.** LTP's `kirk` starts with
`#!/usr/bin/env python3`. On a machine where some tool's virtualenv comes
first in `PATH`, that is the wrong interpreter. Everything here calls
`/usr/bin/python3` explicitly; override with `KT_PYTHON` if your system
python is somewhere else.

## Troubleshooting

**Every KUnit suite reports nothing, and `dmesg` says `kunit: disabled`.**
Something else holds the kunit core loaded with `enable=0`. `rmmod` it, or
add `kunit.enable=1` to the kernel command line and reboot. The engine
detects this case and says so rather than reporting a silent zero.

**`kunit.ko` is installed but will not load.** Either it was built for a
different kernel (`modinfo -F vermagic kunit` against `uname -r`) or Secure
Boot is refusing the signature. `dmesg | tail` says which.

**kselftest reports "No such collection".** The shipped
`kselftest-list.txt` starts with a few lines of build log, including a
colourised `llvm: [ OFF ]` line. The engine filters those out; if you call
`run_kselftest.sh` yourself, filter with
`grep -E '^[A-Za-z0-9_./-]+:[^[:space:]]+$'`.

**A whole LTP scenario is `ERROR`.** Look at `results/latest/ltp/<name>.out`.
Usually kirk could not start, or the scenario needs something the profile did
not prepare.

**Everything in one collection fails.** Check `prepare_system.sh` ran — the
run log says which tier was used. `net` and `tc-testing` need the `-d` tier,
which only `--profile full --unsafe` reaches.

**A run takes far longer than the profile says.** Check the "Slowest tests"
table, then lower `LTP_TIMEOUT` and `KSELFTEST_TIMEOUT`, or cut scenarios out
of the profile list.
