SIGN IN SIGN UP

Retire agents whose process died, safe against PID reuse (#44)

Retire agents whose process died, safe against PID reuse (#44)

If an agent's process dies without its PTY wrapper reporting it (reboot,
crash, kill -9), its row stayed live until the heartbeat timed out, and
forever if a reused PID made it look alive. The original per-command
startup scan checked plain PID liveness, which has the same reuse problem,
so this replaces it with process-identity tracking inside existing cleanup.

- Record a process identity with every tracked PID (start time, plus boot
  ID on Linux) on the PTY, ConPTY, headless and orphan-recovery paths.
  Failing to record the PID fails the launch and reaps the spawned process.
  If no identity can be read, the PID is stored without one and treated as
  before this change.
- Stale cleanup stops rows whose PID is gone or now belongs to a different
  process, without waiting for the heartbeat timeout or wake grace. It uses
  a guarded stop that sends no signal, removes the row only if it still
  holds the checked process, and waits while child rows exist. Rows are
  stopped with a snapshot, so `hcom r` works.
- Moving a PID from a placeholder to the real agent row (on session bind,
  session switch and orphan recovery) is atomic with the launch context and
  bindings, and works inside the transaction Kimi's session bind already
  holds. A failed move rolls back instead of leaving two rows claiming one
  process.
- Every CLI command and an open TUI run this check, at most once per 30s.
  kill and stop never signal a PID that now belongs to another process.

---------

Co-authored-by: aannoo <aannoo@users.noreply.github.com>
S
Solar Lorian committed
c4e172709f912390814b37364e3049cfdbde10a8
Parent: 962cbef
Committed by GitHub <noreply@github.com> on 9/30/2026, 5:46:54 PM