# Guide: determining an ELF's file and process access on RHEL 9.6

This is the working method behind the scripts in `tools/`. Read it once; after
that `doc/CHEATSHEET.md` is enough.

---

## 1. The order of operations

Work from the least invasive to the most invasive, because each step tells you
how dangerous the next one is.

| # | step | executes the binary? | command |
|---|---|---|---|
| 1 | Identify and hash it | no | `file`, `sha256sum`, `rpm -qf` |
| 2 | Static capability review | no | `tools/elf-static.sh ./bin` |
| 3 | Contained dynamic run | **yes** | `tools/elf-trace.sh -- ./bin` |
| 4 | Low-overhead confirmation | yes | `tools/elf-preload.sh -- ./bin` |
| 5 | Observe it in place | it is already running | `tools/elf-snapshot.sh -p PID` |
| 6 | Continuous production watch | no extra run | `tools/elf-audit.sh start ./bin` |

Step 2 is not optional. If static analysis shows `socket`, `connect` and
`execve` imports, you isolate the network *before* step 3.

---

## 2. Static: what the binary *can* do

`tools/elf-static.sh ./bin` collects, from the file only:

**Identity** — size, SHA-256, `file(1)` type, and `rpm -qf` + `rpm -V`. If the
binary claims to belong to a package but `rpm -V` says its hash changed, stop
and treat it as modified.

**Linkage** — the ELF type (`EXEC` vs `DYN`), the program interpreter, the
`DT_NEEDED` libraries, and `RPATH`/`RUNPATH`. A `RUNPATH` pointing somewhere
writable is a library-hijack path; check with
`ls -ld $(readelf -d ./bin | grep RUNPATH | grep -oP '\[\K[^]]+')`.

**Imported symbols, bucketed** — this is the fastest read on intent:

| bucket | what its presence means |
|---|---|
| `open`/`fopen`/`stat`/`unlink`/`rename` | it does file I/O; the write-capable names matter most |
| `fork`/`execve`/`posix_spawn`/`system`/`popen` | it starts other programs |
| `dlopen`/`dlsym` | **ldd will not show you everything** — you must trace it |
| `socket`/`connect`/`getaddrinfo` | it talks on the network |
| `getpwnam`/`crypt`/`pam_*` | it looks at accounts or authenticates |
| `setuid`/`prctl`/`ptrace` | it changes privilege or inspects other processes |

**Embedded strings** — path-shaped strings are candidate files. They are hints,
not evidence: a path may be dead code, and a computed path (`"%s/%s.conf"`)
never shows up whole.

**Blind spots to keep in mind**

- A statically linked binary has no dynamic symbol table; buckets come back
  empty. Fall back to `strings` and to the syscall-instruction count that the
  script reports (`objdump -d | grep syscall`).
- Go and Rust binaries issue raw syscalls. Imports tell you nothing. Trace them.
- A packed binary shows very few sections and unreadable strings. Its unpacked
  behaviour only appears at runtime.

---

## 3. Dynamic: what it actually did

### 3.1 The strace invocation that matters

`tools/elf-trace.sh` runs:

```bash
strace -f -ttt -T -y -yy -s 256 \
       -e trace=%file,%desc,%process,%network -e signal=none \
       -o strace.log -- ./binary
```

| flag | why it is there |
|---|---|
| `-f` | follow `fork`/`clone` — **without it you miss every child process** |
| `-y` | annotate each fd with the path it points at, so `read(3)` becomes `read(3</etc/passwd>)` |
| `-yy` | also resolve sockets to `TCP:[10.0.0.1:443]` and devices |
| `-s 256` | show enough of each string argument that long paths are not truncated |
| `%file` | every syscall that takes a path |
| `%desc` | every fd operation — this is how read/write intent is recovered |
| `%process` | `fork`, `clone`, `execve`, `wait`, `kill` |
| `%network` | `socket`, `connect`, `bind`, `sendto` |
| `-ttt -T` | timestamps and per-call duration; needed to line a trace up with a log |

Variants worth knowing:

```bash
# just the summary, to see where the time goes
strace -f -c -- ./binary

# attach to something already running, for 30 s
tools/elf-trace.sh -T 30 -p "$(pgrep -f mydaemon)"

# a daemon that forks away: trace the service manager's child instead
systemctl stop mysvc
strace -f -y -o /tmp/svc.log -- /usr/sbin/mysvc -D
```

### 3.2 Reading the report

`REPORT.md` sections, and what each is for:

- **Files written, created or deleted** — the blast radius. Anything here
  outside the program's own data directory is a finding.
- **Files read** — ignore the loader noise (`/etc/ld.so.cache`, `locale-archive`,
  `gconv`); look for config, credentials and data.
- **Paths probed but absent (ENOENT)** — the most under-used section. A program
  that stats `/etc/foo/secrets.conf` will *use* it if you create it. This is how
  you discover undocumented configuration and plugin directories.
- **Shared objects loaded at runtime** — compared against `ldd`. A `.so` here
  that `ldd` did not list came from `dlopen` (or from NSS/PAM). If it lives in a
  writable directory, that is a hijack.
- **Process tree** — who forked whom, with the exact `argv` of each `execve`.
  Watch for `/bin/sh -c` with a string that contains anything user-controlled.
- **Socket activity** — `connect` to a hardcoded address is the interesting case.
- **Failed syscalls by errno** — `EACCES` tells you what it wanted but was
  denied; that is the SELinux/permission story in one table.

### 3.3 Attribution: which *process* touched what

Every table has a `pids` column, and `processes.csv` gives the parent of each
pid plus every image it `execve`d. To answer "which process wrote this file":

```bash
grep '/var/lib/thing/data' out/*/files.csv     # -> path,access,syscalls,pids
awk -F, '$1==<pid>' out/*/processes.csv        # -> that pid's image and parent
```

A pid can exec several images in sequence (`sh` → `uname`); `all_images` keeps
the whole chain, which is why the report can flag a shell that no longer exists
by the time the process exits.

---

## 4. When ptrace is not an option

### 4.1 LD_PRELOAD interposer (`tools/elf-preload.sh`)

Replaces libc's `open`, `fopen`, `execve`, `popen`, `dlopen`, `connect`… and
logs each call. Overhead is a function call, not a context switch, so you can
leave it on during a load test.

It **cannot** see: statically linked binaries, setuid/setgid binaries (the
loader drops `LD_PRELOAD`), raw-syscall programs (Go), or a program that
deliberately scrubs its environment. It is a confirmation tool, not a
containment tool — do not rely on it for anything adversarial.

### 4.2 fanotify (`tools/elf-fanotify.sh`)

Watches whole mounts, so it sees file access by *every* process, including a
daemon your target merely sends a request to. That is the one thing strace
structurally cannot do. The cost is that you must filter out unrelated system
noise, and fanotify reports only `comm(pid)` — cross-reference the pids with the
process tree from the strace run.

### 4.3 Kernel audit (`tools/elf-audit.sh`)

The production answer. Rules are filtered by `-F exe=/path/to/binary`, so the
noise stays bounded, and records survive reboots.

```bash
sudo tools/elf-audit.sh start /usr/local/bin/suspect mykey
# ... let it run for a day ...
sudo tools/elf-audit.sh report mykey
sudo tools/elf-audit.sh stop mykey
```

Caveats: rules are lost on reboot unless written to `/etc/audit/rules.d/`;
`open`/`openat` rules on a busy binary can fill the audit log, so keep
`-F success=1` and watch `auditctl -s` for `lost`.

### 4.4 bpftrace (optional, `rpms/optional-bpftrace/`)

When you want system-wide tracing with near-zero overhead and full control:

```bash
bpftrace -e 'tracepoint:syscalls:sys_enter_openat /comm == "suspect"/ {
    printf("%s %s\n", comm, str(args->filename)); }'
bpftrace -e 'tracepoint:sched:sched_process_exec { printf("%s -> %s\n", comm, str(args->filename)); }'
```

RHEL 9 ships kernel BTF, so these run without kernel headers. Not available
inside an unprivileged container.

---

## 5. RHEL 9.6 specifics

**SELinux.** In enforcing mode a denial looks exactly like a missing file: the
trace shows `EACCES`, not "SELinux". Confirm before you blame the binary:

```bash
ausearch -m AVC -ts recent | audit2why
```

Do not disable SELinux to make a trace work — run the trace in permissive mode
(`setenforce 0`) only in a lab, and say so in the write-up, because permissive
mode changes which code paths the program reaches.

**ptrace_scope.** `sysctl kernel.yama.ptrace_scope` — `0` lets you attach to any
of your processes, `1` (the common hardened setting) only to descendants, `2`
root only, `3` disables attach entirely. Launching *under* strace works at
every level except 3; attaching with `-p` does not. Check it before you conclude
a binary is "blocking" the tracer.

**Containers.** `podman run --cap-add=SYS_PTRACE` is required for strace inside
a container. fanotify and audit rules need the host, not a container.

**Systemd services.** A service's file access includes what systemd does on its
behalf — `ProtectHome`, `PrivateTmp`, `ReadOnlyPaths` change which paths exist
from the process's point of view. `systemctl show mysvc -p PrivateTmp,ProtectSystem`
before you interpret a `/tmp` path: under `PrivateTmp=yes` the real path is
`/tmp/systemd-private-*/tmp/…`.

**Immutable/offline hosts.** If you cannot install anything, use
`build/10-build-sources.sh`; the binaries land in `build/bin/` and need no root
to build, only to run fanotify.

---

## 6. Installing on RHEL 9.6 without breaking the host

`build/00-install-rpms.sh` installs **only packages that are absent**, and never
upgrades one the system already has. That rule exists because of a real failure:

```
Problem 1: nothing provides elfutils-debuginfod-client(x86-64) = 0.192-6.el9_6.alma.1
           needed by elfutils-0.192-6.el9_6.alma.1
Problem 2: cannot install both elfutils-libelf-0.192-6.el9_6.alma.1 and
           elfutils-libelf-0.192-5.el9 from @System
```

Two distinct mistakes produced that, and both are worth knowing:

1. **A rebuilt package is not a drop-in for the vendor's.** AlmaLinux's
   `elfutils` carries a `.alma.1` release tag and its subpackages require each
   other at *exactly* that NEVRA. Install one on RHEL and it demands the whole
   rebuilt set, while RHEL's installed `elfutils-libs` demands the originals.
   Neither side can win. The fix is not `--force` or `--allowerasing` — it is to
   not ship that package at all.
2. **Handing `dnf` a directory of RPMs is an upgrade request.** Anything in the
   directory that matches an installed package gets pulled into the transaction.
   Filtering the list down to genuinely missing packages avoids the whole class.

So the bundle contains no `elfutils`, and nothing else that is part of a base
RHEL 9.6 install. `strace` and `ltrace` need `libdw.so.1` and `libelf.so.1`,
which RHEL 9.6 already provides through `elfutils-libs` and `elfutils-libelf`.

**If the host has no compiler.** `make -C target` failing with
`make: cc: No such file or directory` is expected on a locked-down server, and
the kit handles it: the Makefiles detect the missing compiler and install the
binaries from `prebuilt/` instead, and `build/20-use-prebuilt.sh` does the same
without make. Those binaries were compiled against glibc 2.34 - RHEL 9's
baseline - so they run on any RHEL 9.

gcc is deliberately **not** bundled, for the same reason elfutils is not:
`glibc-devel` requires an exact `glibc` NEVRA, so pulling in a toolchain would
force an upgrade of the host's C library. If you do want to build from source,
install gcc from your own RHEL media, where the versions match.

**If the host has no binutils.** `elf-static.sh` needs `readelf`, `objdump` and
`strings`. When they are absent it falls back to `tools/elf_static_py.py`, which
reproduces the header, the dynamic section, the undefined-symbol buckets, the
hardening flags and the embedded strings using pyelftools - one noarch RPM, no
awkward dependencies. Only section 4 (direct syscall sites) is lost, because
that genuinely needs a disassembler. The report says so at the top when it is
running in that mode.

**The one gap you may hit.** `binutils` (which gives you `readelf`, `objdump`,
`strings`) requires `libdebuginfod.so.1` from `elfutils-debuginfod-client`. That
package is present on a normal RHEL 9.6 server but absent from minimal images.
If the installer reports it:

```bash
dnf install elfutils-debuginfod-client      # from your RHEL media or Satellite
#   - or skip binutils; these give you the same commands:
dnf install gcc                             # binutils comes with it
build/10-build-sources.sh                   # strace + lsof, no RPMs at all
```

Verified on `registry.access.redhat.com/ubi9/ubi:9.6`: with
`elfutils-debuginfod-client` present all 10 packages install offline and the
self-test reports 44 pass / 0 fail (the 2 skips are fanotify and auditd, which a
container cannot provide). On a minimal image without it, 9 of 10 install and
only `binutils` is reported missing.

## 7. Common traps

| symptom | cause | what to do |
|---|---|---|
| Report shows almost nothing | the binary forked and exited immediately | you forgot `-f`, or the real work happens in a child that outlives the trace — use `-f` and `-T` |
| `ldd` output disagrees with the trace | `dlopen` | trust the trace; that is the point of section 6 of the report |
| Every path is relative and unresolved | traced from a different cwd | pass `--cwd` to `parse_strace.py`, or run the trace from the program's working directory |
| `strace: attach: ptrace(PTRACE_SEIZE, …): Operation not permitted` | yama ptrace_scope, or not the owner | check `sysctl kernel.yama.ptrace_scope`, run as root |
| Trace stops after `execve` | the binary is setuid | root can still trace it; `LD_PRELOAD` cannot |
| A file appears with no opening syscall | it was inherited from the parent | look at `/proc/PID/fd` at start, or trace the parent |
| Timing changes the behaviour | strace slows the program 10–100× | use the interposer or bpftrace instead |
| `fatrace` produces nothing | fanotify cannot mark overlayfs | run it on the host, not in a container |
| `make: cc: No such file or directory` | no compiler on the host | nothing to fix - the Makefile installs `prebuilt/` instead; `build/20-use-prebuilt.sh` if make is missing too |
| `readelf: command not found` | no binutils | `elf-static.sh` falls back to pyelftools automatically |
| `/proc/PID/exe` says `coreutils` | RHEL ships coreutils as one multi-call binary | read `cmdline`; `elf-snapshot.sh` prints a note when this happens |
| `/tmp` paths do not exist afterwards | `PrivateTmp=yes` | inspect via `nsenter -t PID -m ls /tmp` |

---

## 8. A worked example: httpd

Everything above is exercised end to end by `examples/httpd-example.sh`, which
analyses Apache - a daemon that loads its code with `dlopen`, forks workers that
each run a pool of threads, and execs CGI programs that are not httpd at all.
`doc/EXAMPLE-HTTPD.md` is a captured run.

It starts its **own** instance on 127.0.0.1:8008 under its own ServerRoot, so a
system httpd keeps serving untouched, and it stops that instance on exit.

```bash
examples/httpd-example.sh                  # ~40 s
examples/httpd-example.sh -p 8009 -o /var/tmp/httpd-analysis
```

Four things it demonstrates that a synthetic target cannot:

1. **`ldd` is nearly useless on a plugin host.** `ldd /usr/sbin/httpd` lists 10
   libraries; the running process maps a module set that `ldd` never mentions,
   because every one arrives through `dlopen`.
2. **Threads are not processes.** The event MPM forks a few workers, each
   running ~26 threads. strace reports a thread as just another pid, so a naive
   reading says httpd forked 80 processes. `parse_strace.py` reads the
   `CLONE_THREAD` flag and reports 3 processes + 81 threads.
3. **The exec'd image is the CGI script, not the shell.** httpd calls
   `execve("whoami.cgi")` and the kernel resolves the `#!` line, so `/bin/sh`
   never appears as an exec - only the script, and then its own children.
4. **Name-service lookups are part of the request path.** Each request that
   resolves a user tries `nscd` and then SSSD over unix sockets. Your web
   server's latency depends on daemons that never appear in its config.

If you are analysing the *system* httpd rather than this sandbox instance, see
the "Applying this to a system httpd" section the example writes, and mind
`PrivateTmp`, SELinux labels, and the fact that workers `setuid` away from root.

## 9. Turning a trace into a statement you can defend

A finding is only as good as its evidence chain:

1. Record the binary's SHA-256 and the exact command line you traced.
2. Keep `strace.log` alongside the report; every table is reproducible from it.
3. State the environment: host, kernel, SELinux mode, user, container or not.
4. Say what you exercised. "It does not touch the network" is wrong; "over the
   three runs below, covering startup, a reload and a shutdown, it opened no
   sockets" is right.
5. Re-run with `--cwd` and arguments varied before you generalise.

The checklist at the end of every `ANALYSIS.md` mirrors these points.
