SIGN IN SIGN UP

fix(container-watcher): cancel ctx in Stop; don't close workerChan (send-on-closed panic) (#837)

Stop() closed cw.workerChan while enrichAndProcess could be mid-send (the select
case and the BlockEvents blocking send), so an in-flight send racing Stop()
panicked with "send on closed channel" on teardown — observed on node-agent pods
during DaemonSet rollouts.

Root cause: Stop() did not actually stop the producer (eventProcessingLoop) or
consumer (workerPoolLoop) goroutines. cw.ctx was the caller's context
(context.Background() in cmd/main.go, never cancelled) and Stop() cancelled
nothing, so both loops only died with the process. close(workerChan) was a
stand-in for shutdown that never worked: there is no range over workerChan, and
a select-reader on a closed channel busy-spins instead of exiting — it only
created the send-on-closed hazard.

Fix:
- Derive a cancelable context in Start (cw.ctx, cw.cancel = WithCancel(ctx)) and
  call cw.cancel() at the top of Stop, so both goroutines genuinely terminate.
  This also makes Stop() correct mid-process (previously it would leak both
  goroutines and silently drop events on the released worker pool).
- Drop close(cw.workerChan); the channel is GC'd with the watcher.
- Make the BlockEvents blocking send ctx-aware so a full channel can't block
  forever after the consumer has exited.

Docs-exempt: internal shutdown-race bug fix; no API/config/documented-behavior change

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
M
Matthias Bertschy committed
5bda7013c146be5f85af004cd93449db9e32e7b9
Parent: ca2feee
Committed by GitHub <noreply@github.com> on 6/19/2026, 12:08:32 PM