ref:main

fix(runner): a recycled PID is not a running runner #63

open colechristensen cole.christensen@gmail.com wants to merge fix/runner-pid-reuse into main

A runner on a dev host refused to start, every five seconds, forever:

Error: another anvil runner appears to be running (PID 1158).
If it isn't, remove /home/chaos/.anvil-runner/runner-default.pid

PID 1158 was minio server /data --console-address :9001, running as root.

The runner that wrote that PID file died without cleaning up; the kernel later handed 1158 to MinIO. is_alive then answered the only question it can — kill(1158, 0) returns EPERM, “exists but not yours” — which alive_from_kill counts as alive, deliberately and correctly for the case it was written for:

// EPERM means the process exists but belongs to somebody else. Reading it
// as "gone" would have a non-root `runner status` declare a live daemon
// crashed and start a second one on the same work dir.
errno == libc::EPERM

Nothing downstream asked the question that actually mattered: not “does a process hold this PID” but “is it ours”. And because the answer never changes while MinIO holds the PID, systemd’s Restart=always could not clear it. The restart counter reached 60.

The fix

inspect now demotes a live PID to stale only on a positive identification of something else, read from /proc/<pid>/comm.

World-readable is the point. The case this exists for is a PID reused by a root-owned process that this account can neither signal nor ptrace — so /proc/<pid>/exe, which would be more precise, is exactly what that permission boundary denies us. comm is readable regardless of owner, which is how the diagnosis identified MinIO in the first place.

Unknown — non-Linux Unix, and all of Windows — stays Live, preserving today’s behaviour there. The two errors are not symmetric: refusing to start is visible and recoverable, whereas overwriting a live runner’s PID file runs two daemons against one work directory. Matching on the executable name rather than argv is the same asymmetry: a stray short-lived anvil CLI inheriting the PID is an acceptable and self-correcting false positive.

The decision is taken as parameters rather than probed inline, so the matrix is testable — the same reason alive_from_kill was split out. Unknown is unreachable on Linux, and a live unsignallable process cannot be conjured without root.

Also: status read a different file than the daemon wrote

anvil runner status reported this host’s healthy runner as not running, naming a stale PID from a file nothing had written since August:

Process running no (stale PID 242076 in /home/chaos/.anvil-runner/runner.pid)

The service-unit templates emitted runner-{instance}.pid, while instance_path("default") returned the legacy runner.pid. The templates write a home-relative %h/... path and so cannot call instance_path; the name they formatted inline had drifted from it.

instance_file_name is now the single rule both use, with a test that they cannot drift again. Default still maps to runner.pid — hosts installed before instances existed have that file, and renaming it would strand their running daemon.

For existing installs: a unit generated before this change still says runner-default.pid, so runner status stays wrong there until anvil runner service install is re-run.

Verified against the process that caused it

With 1158 still held by root’s MinIO:

$ cat /proc/1158/comm → minio (Uid: 0)
$ ~/.local/bin/anvil runner status ... → Process running yes (PID 1158)
$ ./target/debug/anvil runner status ... → Process running no (stale PID 1158 ...)

The host is unblocked and the runner is polling again.

Tests

The regression test uses the test process itself as the stranger — alive, real, and not an anvil binary — so it reproduces without privileges. Plus the full classify matrix (including the Unknown fallback that is otherwise unreachable on Linux), identity_from_comm, identity of a live non-anvil child and of a reaped PID, and the two naming invariants.

One existing test changed meaning rather than breaking: inspect_live_pid_returns_live wrote its own PID and asserted Live, which was only ever true because liveness was the whole answer. It is now inspect_live_pid_belonging_to_another_program_is_stale, asserting the opposite for the same input, because the contract changed and that input is precisely the bug.

cargo fmt --check, cargo clippy --all-targets -D warnings, and cargo test (733 tests) all clean.

Not fixed here

For system-scope installs the daemon writes /run/anvil-runner-<instance>/runner.pid while status still resolves a home path, so the same class of mismatch remains. Fixing it needs status to resolve through the install record, on two platforms I can’t exercise from here — better done deliberately than blind.

Separately, ~/.anvil-runner/ had ~70 orphaned id-*.pid files accumulated since July that nothing prunes. Harmless after this change (they now classify as stale correctly), but nothing cleans them up.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JB4ESzf5areVYwCFdueNtQ

Created Sep 04, 2026 at 23:07 UTC