# Testing a patch

The question a patch test answers is **not** "does the kernel pass its
tests". It never entirely does: upstream selftests and LTP are written
against mainline, RHEL's kernel is not mainline, and some tests need
hardware or a network this machine does not have. A healthy RHEL 9.6 machine
has a handful of failures on a first run.

The question is **"what changed"**. So `test-patch.sh` runs the same tests
twice — once before the patch, once after — and reports only the difference.

- [How it works](#how-it-works)
- [The modes](#the-modes)
- [Walkthroughs](#walkthroughs)
- [Reading the comparison](#reading-the-comparison)
- [Flaky tests](#flaky-tests)
- [Unattended runs](#unattended-runs)
- [In a pipeline](#in-a-pipeline)
- [Troubleshooting](#troubleshooting)

## How it works

```
  baseline  ->  apply  ->  [reboot]  ->  verify  ->  after  ->  compare
     |            |                         |          |          |
  run the      install                  is the      run the    diff, and
  tests       the patch               patched      SAME tests   report
                                       kernel                   only what
                                       running?                 changed
```

A kernel patch usually needs a reboot, so the work is split into phases and
the progress is kept in a state file (`/var/lib/kernel-test/state`). The run
picks up where it left off when the machine comes back.

```bash
./test-patch.sh start --mode rpm --patch-dir /srv/erratum
reboot
./test-patch.sh continue
```

Other commands: `status` shows where a run has got to, `abort` forgets it
(it does **not** undo the patch), and `baseline`, `apply`, `after`,
`compare` each run one phase on their own.

Exit status: `0` no regression, `1` regression found, `2` it could not run,
`3` paused and waiting for you (a reboot, or a manual patch).

### Two rules the script enforces

**Both runs must test the same things.** `--profile` and `--engines` are
recorded in the baseline and reused. Comparing a `smoke` run against a
`standard` one produces a page of "new tests" and no answer.

**The machine must really be running the patched kernel.** After a kernel
RPM lands, the machine still runs the old kernel until it reboots — and it
may not boot the newest one. The `verify` phase compares `uname -r` against
what was installed and refuses to go on if they differ. This one check saves
more wasted runs than everything else in the script.

## The modes

| `--mode` | For | Reboot |
|---|---|---|
| `rpm` | a kernel erratum, z-stream update, or locally rebuilt kernel RPM | yes |
| `kpatch` | a kpatch / livepatch RPM | no |
| `module` | one out-of-tree module you are building yourself | no |
| `source` | `.patch` files applied to a kernel tree, then built | yes |
| `manual` | anything else — the script pauses while you apply it | up to you |

### `rpm`

The common case on an air-gapped machine: someone hands you a directory of
RPMs from an erratum.

```bash
./test-patch.sh start --mode rpm --patch-dir /srv/kernel-erratum
```

It runs `dnf install` with `--disablerepo='*'` (nothing to reach for on an
offline machine) and `--nogpgcheck`, falling back to `upgrade` for packages
already present. It records the package changes in `pkgs.diff`, notices
whether a new kernel appeared, and if one did, tells you to reboot.

If `dnf` fails, it is almost always a dependency that is not in the patch
directory. Add it and run `./test-patch.sh continue`.

### `kpatch`

A livepatch takes effect immediately, so there is no reboot and `uname -r`
does not change.

```bash
./test-patch.sh start --mode kpatch --patch-dir /srv/kpatch-rpms
```

The whole thing runs in one go. The loaded patch list is saved to
`kpatch-list.txt` in the comparison folder. If no livepatch ends up loaded
the script says so loudly, because otherwise you would be comparing a kernel
against itself and calling the result "no regression".

### `module`

For a driver or module you are building yourself.

```bash
./test-patch.sh start --mode module --module-src /srv/mydriver
```

Builds against `/lib/modules/$(uname -r)/build` (so `kernel-devel` must be
installed) and `insmod`s the result. It checks for module signature
enforcement first and stops with a clear message rather than letting
`insmod` fail with `-EKEYREJECTED`.

### `source`

Applies `.patch` files to a kernel tree, builds it, installs it.

```bash
./test-patch.sh start --mode source \
    --source-dir /usr/src/linux-5.14.0 \
    --patch-dir  /srv/patches
```

This needs `gcc`, `make`, `flex`, `bison` and `bc`, so it belongs on a build
machine, not a locked-down production host. Every patch is applied
`--dry-run` first and the whole series is refused if any one of them does not
apply, because a half-applied series is much harder to undo than a refused
one. The build is seeded from the running kernel's `/boot/config-$(uname -r)`
so the result is comparable with the baseline.

### `manual`

The script records the baseline, then stops and waits.

```bash
./test-patch.sh start --mode manual --profile standard
# ... apply the patch however you like, reboot if you need to ...
./test-patch.sh continue
```

Use this when nothing else fits: a vendor installer, a firmware change, a
sysctl or boot-parameter change you want to measure, a hand-built kernel.

## Walkthroughs

### A kernel erratum on an offline machine

```bash
# 1. Baseline and install.  ~2 hours for the standard profile.
./test-patch.sh start --mode rpm --patch-dir /srv/RHSA-2026-1234

# 2. It stops and tells you to reboot.  Check what you will boot into:
grubby --default-kernel
reboot

# 3. Same tests again, then the comparison.
./test-patch.sh continue

# 4. Read it.
cat results/latest/comparison.md
```

Use `--profile smoke` first if you have not run this on the machine before —
it proves the plumbing in five minutes instead of finding a problem two
hours in.

### A security erratum, quickly

LTP's `cve` scenario reproduces specific published vulnerabilities. For a
security fix, that plus `syscalls` is the highest-value narrow run:

```bash
LTP_SUITES="cve syscalls" \
  ./test-patch.sh start --mode rpm --patch-dir /srv/RHSA-2026-1234 \
                        --engines "kunit ltp"
```

### Just this subsystem

If the patch only touches memory management:

```bash
KSELFTEST_COLLECTIONS="mm cgroup" LTP_SUITES="mm hugetlb" \
  ./test-patch.sh start --mode manual
```

Narrowing is a trade: a faster run that cannot see a regression somewhere
else. For anything going to production, run `standard` at least once.

## Reading the comparison

`results/latest/comparison.md`:

| Section | What it means |
|---|---|
| **Regressions** | passed before, does not pass now. **This is the answer.** |
| **Regressions in known-flaky tests** | the same, but the test is on the flaky list. Check by hand. |
| **Fixed by the patch** | failed before, passes now. Often exactly what the patch was for. |
| **Failing before and after** | already broken. Not this patch's doing. |
| **Appeared / disappeared** | usually a package or LTP version that changed with the patch, not a deleted test |

The header table shows the kernel on each side. If they are the same and you
expected them to differ, the comparison says so — that is the "you did not
actually reboot into it" case.

An empty Regressions section is the result you want. It does not mean the
kernel is perfect; it means the patch did not break anything these suites
can see.

`comparison.json` has the same content for a script, plus `verdict`, which is
`OK` or `REGRESSION`.

The comparison folder also keeps `report-before.md`, `report-after.md`,
`patch-state.txt`, `pkgs.diff` and `apply.log`, so it stands on its own as
the record of what happened.

## Flaky tests

Some tests fail now and then for reasons that have nothing to do with a
patch: timing on a loaded machine, how much memory happens to be free,
whether the network is quiet. Left alone they show up as regressions and
hide the real ones.

`conf/flaky.list` holds patterns for those, one `engine:collection:test` per
line, globs allowed:

```
ltp:*:leapsec01
ltp:*:clock_nanosleep*
kselftest:net/mptcp:*
```

They are reported in their own section rather than counted as regressions,
and they do not affect the exit status.

Add to this list only after you have seen a test fail on an *unchanged*
kernel — twice. Everything in the list is a regression you have agreed in
advance not to be told about.

Confirm a suspicious regression by re-running just that collection on the
patched kernel:

```bash
KSELFTEST_COLLECTIONS="mm" ./test-kernel.sh --engines kselftest -t confirm
```

## Unattended runs

`--auto-reboot` reboots the machine by itself and installs a one-shot
systemd unit so the run continues after it comes back:

```bash
./test-patch.sh start --mode rpm --patch-dir /srv/erratum --auto-reboot
```

The unit (`kernel-test-resume.service`) removes itself once the comparison is
written. Progress goes to `/var/lib/kernel-test/resume.log`.

Only do this on a machine you are willing to have reboot without asking.

## In a pipeline

```bash
./test-patch.sh start --mode rpm --patch-dir "$PATCH_DIR" --auto-reboot
# after the reboot the unit runs 'continue'; poll for the verdict:
./test-patch.sh status
```

Once it is done:

```bash
jq -r .verdict results/latest/comparison.json      # OK | REGRESSION
jq -r '.regressions[] | "\(.test): \(.before) -> \(.after)"' \
      results/latest/comparison.json
```

`test-patch.sh compare` exits `1` on a regression, so it gates a stage
directly. Keep the whole `results/latest/` folder as the build artifact — the
verdict without `dmesg-full.txt` and the `*.badness` files is not enough to
investigate anything later.

## Troubleshooting

**"expected to be running X but this is Y".** The machine booted the wrong
kernel. `grubby --set-default /boot/vmlinuz-X`, reboot, and
`./test-patch.sh continue` again.

**"a baseline already exists".** A run is already in progress. `status` to
see it, `abort` to forget it. `abort` does not undo a patch that was already
installed.

**"no new kernel package appeared".** The RPMs installed but none of them
was a kernel. Fine for a userspace patch; if you expected a new kernel,
check `pkgs.diff` in the state directory.

**Dozens of "new" and "gone" tests.** The two runs did not test the same
things. Either the profile differed, or `kernel-selftests-internal` or LTP
changed version with the patch. The header table shows the profile and the
engines for both runs.

**The comparison says OK but you do not believe it.** Check that the after
run actually ran: `results/latest/report-after.md` has the test count. A run
where an engine was skipped for a missing package still produces a
comparison, and a comparison of nothing against nothing is `OK`.
