qwen4exp : halve the indexer score memory (#29825)
* qwen4exp : halve the indexer score memory The indexer scored all heads in one product and rectified a copy of it, so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the largest buffers of the graph at long context. Each head now gets its own product, rectified and summed in place into one [n_pool, n_tokens] score. * qwen4exp: let the allocator reuse the indexer score buffers Address review from CISC: use plain ggml_add and ggml_relu in the indexer head loop. The graph allocator already runs them in place when their source has no other consumer, so the _inplace variants are not needed. The compute buffer and the speed are unchanged. * cuda: support 4 heads in the lightning indexer Dispatch 4 heads to the vector kernel, too few for a wmma tile, and accept them in supports_op. test-backend-ops covers 4 heads. * metal: take the lightning indexer head count as a function constant The kernel reads the head count from a function constant and zero fills the last head tile, so any head count runs and 64 heads is unchanged. * qwen4exp: compute the indexer score with the lightning indexer Address review from am17an: the unweighted sum of the rectified head scores scaled by 1/sqrt(head_dim) is the lightning indexer with every head weight set to that scale, so the indexer calls ggml_lightning_indexer on the pooled keys with an f16 pool mask. The keys are read once for all heads and no per head score is materialized. * vulkan: tile the lightning indexer over keys and tokens A workgroup scores 64 keys against 8 tokens: the keys are staged once in shared memory, the queries one head at a time, and each invocation owns one key for two tokens, so no dot product needs a cross invocation reduction. The subgroup variant and the flat dispatch are gone, the grid is keys x tokens x streams. * vectorize vulkan loads and use fp16 dot product --------- Co-authored-by: Ruben Ortlam <rortlam@redhat.com>
P
Pascal committed
889edf43ddae0cfe9a4564a882764dc879759870
Parent: 99b9548
Committed by GitHub <noreply@github.com>
on 10/3/2026, 5:19:00 AM