[inductor] Compile combo-kernel benchmark candidates in the async-compile pool (#198881) (#198881)
Summary: `speedup_by_combo_kernel` benchmarks each candidate subkernel to decide whether merging them is profitable, and every one of those benchmarks compiled inline on the scheduler thread. Triton's frontend is pure Python and holds the GIL, so `CachingAutotuner._precompile_worker`'s per-config loop never overlaps with anything: the candidate group compiles strictly serially even with idle cores. Threads don't fix it -- compiling 48 real kernels from an affected model measured 11.2s on 1 thread and 8.7s on 16 -- so the work has to go to processes. `speedup_by_fusion` already has the right shape: call `Scheduler.compile_kernel` for every candidate up front, which codegens once and hands the compile to the async-compile process pool, then benchmark the returned modules. Do the same here. Only the benchmarks move after codegen; the codegen sequence is unchanged, and benchmarking doesn't touch the graph state successive codegens observe. With no pool, `compile_kernel` returns a `None` future and the compile happens lazily in `benchmark_codegened_module`, exactly as `benchmark_fused_nodes` did before. Compiling in a worker means a failure now arrives wrapped in `SubprocException` rather than `CompilationError`, which would silently disable the triton#2151 loop-carried-variable workaround and turn a skipped candidate into a failed compile. That workaround had three copies in this file, so it becomes one `_is_loop_carried_compile_error` predicate matching both sides of the pool boundary. That also closes the same gap in `speedup_by_fusion`, which already compiles through the pool. Differential Revision: D121464299 Pull Request resolved: https://github.com/pytorch/pytorch/pull/198881 Approved by: https://github.com/karthickai
B
Bin Bao committed
01949e998f4ecac6f0e3661285bbd12ce1467779
Parent: 36a88a5
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com>
on 10/1/2026, 3:10:20 AM