Retire agents whose process died, safe against PID reuse (#44)
Retire agents whose process died, safe against PID reuse (#44) If an agent's process dies without its PTY wrapper reporting it (reboot, crash, kill -9), its row stayed live until the heartbeat timed out, and forever if a reused PID made it look alive. The original per-command startup scan checked plain PID liveness, which has the same reuse problem, so this replaces it with process-identity tracking inside existing cleanup. - Record a process identity with every tracked PID (start time, plus boot ID on Linux) on the PTY, ConPTY, headless and orphan-recovery paths. Failing to record the PID fails the launch and reaps the spawned process. If no identity can be read, the PID is stored without one and treated as before this change. - Stale cleanup stops rows whose PID is gone or now belongs to a different process, without waiting for the heartbeat timeout or wake grace. It uses a guarded stop that sends no signal, removes the row only if it still holds the checked process, and waits while child rows exist. Rows are stopped with a snapshot, so `hcom r` works. - Moving a PID from a placeholder to the real agent row (on session bind, session switch and orphan recovery) is atomic with the launch context and bindings, and works inside the transaction Kimi's session bind already holds. A failed move rolls back instead of leaving two rows claiming one process. - Every CLI command and an open TUI run this check, at most once per 30s. kill and stop never signal a PID that now belongs to another process. --------- Co-authored-by: aannoo <aannoo@users.noreply.github.com>
S
Solar Lorian committed
c4e172709f912390814b37364e3049cfdbde10a8
Parent: 962cbef
Committed by GitHub <noreply@github.com>
on 9/30/2026, 5:46:54 PM