SIGN IN SIGN UP

[inductor] Compile combo-kernel benchmark candidates in the async-compile pool (#198881) (#198881)

Summary:

`speedup_by_combo_kernel` benchmarks each candidate subkernel to decide whether
merging them is profitable, and every one of those benchmarks compiled inline on
the scheduler thread. Triton's frontend is pure Python and holds the GIL, so
`CachingAutotuner._precompile_worker`'s per-config loop never overlaps with
anything: the candidate group compiles strictly serially even with idle cores.
Threads don't fix it -- compiling 48 real kernels from an affected model measured
11.2s on 1 thread and 8.7s on 16 -- so the work has to go to processes.

`speedup_by_fusion` already has the right shape: call `Scheduler.compile_kernel`
for every candidate up front, which codegens once and hands the compile to the
async-compile process pool, then benchmark the returned modules. Do the same here.
Only the benchmarks move after codegen; the codegen sequence is unchanged, and
benchmarking doesn't touch the graph state successive codegens observe. With no
pool, `compile_kernel` returns a `None` future and the compile happens lazily in
`benchmark_codegened_module`, exactly as `benchmark_fused_nodes` did before.

Compiling in a worker means a failure now arrives wrapped in `SubprocException`
rather than `CompilationError`, which would silently disable the triton#2151
loop-carried-variable workaround and turn a skipped candidate into a failed
compile. That workaround had three copies in this file, so it becomes one
`_is_loop_carried_compile_error` predicate matching both sides of the pool
boundary. That also closes the same gap in `speedup_by_fusion`, which already
compiles through the pool.

Differential Revision: D121464299

Pull Request resolved: https://github.com/pytorch/pytorch/pull/198881
Approved by: https://github.com/karthickai
B
Bin Bao committed
01949e998f4ecac6f0e3661285bbd12ce1467779
Parent: 36a88a5
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com> on 10/1/2026, 3:10:20 AM