SIGN IN SIGN UP

kprobes: Fix permanent hang when flushing the kprobe optimizer

Writing 0 to /proc/sys/debug/kprobes-optimization while a kprobe is
jump-optimized never returns. The writer sleeps in D state forever with
kprobe_sysctl_mutex held, so any later read or write of that sysctl
hangs as well. For example, with vfs_read+9 as an optimizable address
in this build:

  # cd /sys/kernel/tracing
  # echo 'p:myprobe vfs_read+9' >> kprobe_events
  # echo 1 > events/kprobes/myprobe/enable
  # # wait until /sys/kernel/debug/kprobes/list shows [OPTIMIZED]
  # echo 0 > /proc/sys/debug/kprobes-optimization

  INFO: task sh:246 blocked for more than 10 seconds.
  Call Trace:
   <TASK>
   __schedule+0x1176/0x4f70
   schedule+0xdc/0x2c0
   schedule_timeout+0x17b/0x260
   wait_for_completion+0x173/0x3c0
   wait_for_kprobe_optimizer_locked+0xbc/0x130
   proc_kprobes_optimization_handler+0x156/0x1b0
   proc_sys_call_handler+0x324/0x490
   vfs_write+0x52d/0xfe0
   ksys_write+0xff/0x200
   do_syscall_64+0x106/0x630
   entry_SYSCALL_64_after_hwframe+0x77/0x7f
   </TASK>
  ...
  INFO: task cat:265 is blocked on a mutex likely owned by task sh:246.

wait_for_kprobe_optimizer_locked() reinitializes optimizer_completion,
asks the optimizer thread to flush and sleeps in wait_for_completion().
The thread drains the (un)optimizing lists, but calls complete() only
if completion_done() is true, i.e. if the completion is already done,
which never happens while someone waits. disarm_all_kprobes() and
kprobe_trace_self_tests_init() wait the same way.

Calling complete() unconditionally would not be enough: the waiter
drops kprobe_mutex while it sleeps, and nothing else serializes the
sysctl handler against the debugfs "enabled" file. A second flusher
that still finds the lists non-empty, e.g. because a disabled probe is
queued for unoptimizing, reinitializes the completion under the first:

  sysctl write                      debugfs "enabled" write
  unoptimize_all_kprobes()
    wait_for_kprobe_optimizer_locked()
      init_completion(c)
      mutex_unlock(&kprobe_mutex)
      wait_for_completion(c)
                                    disarm_all_kprobes()
                                      wait_for_kprobe_optimizer_locked()
                                        init_completion(c)
                                          // c->wait is reset, the first
                                          // waiter is off the queue
                                        mutex_unlock(&kprobe_mutex)
                                        wait_for_completion(c)
  kprobe_optimizer()
    complete(c)
      // wakes the debugfs writer only

where c is &optimizer_completion. Lining up the two writes during an
optimizer pass loses the sysctl writer this way.

Replace the completion with a counter of optimizer passes, bumped at the
end of each pass and signalled with wake_up_var_locked(), both under
kprobe_mutex. A flusher samples the count and waits with
wait_var_event_mutex(), which drops kprobe_mutex only while sleeping, so
a new count means a whole pass ran in the meantime. Nothing is
reinitialized, so several flushers can sleep in the wait at once.

Link: https://lore.kernel.org/all/20260924092142.199198-1-parri.andrea@gmail.com/

Fixes: 73c12f209462 ("kprobes: Use dedicated kthread for kprobe optimizer")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
A
Andrea Parri committed
5bfa9f1a9dcb6ecb607adbc1c0226605c972935b
Parent: 1d653a1
Committed by Masami Hiramatsu (Google) <mhiramat@kernel.org> on 9/25/2026, 2:35:23 PM