perf: rework the stage-1 shortlist pipeline — 4–12× faster stage 1 (5.6–7.1× e2e combined with #169) (#170)
* perf(search): rework the stage-1 shortlist pipeline Stage 1 (candidate generation) rebuilt phase by phase, each change equivalence-tested against the path it replaces: - centroid GEMM: column-block-parallel (par_cdot), bit-identical to the single dot call — the dim-long reduction stays inside one block. - IVF probe: running-threshold chunk scan (probe_top_k_scan) replaces a K-length buffer fill + select_nth per query token; same top-k set by value, one sequential read per row. - candidate gather: doc ids are dense, so a word bitmap dedups in O(postings) and scanning its set bits emits the same sorted list the sort+dedup produced, without the sort. - approximate flood: centroid scores quantized to u8 in a fused transpose+quantize pass (PLAID's own centroid-interaction trick) and scored chunk-outer with a register-resident 16-lane max accumulator. Quantization is monotone per lane; validated at nDCG@10 parity to four decimals against exact flooding on real-embedding corpora. - prune: partial-select the surviving n_full_scores before sorting only that prefix; identical top list and order. - code reads: a lazily-built u32 side-array of the token codes, so the flood reads 4 B/token instead of the mmap's 8. Stage 1 is self-contained: no new public API beyond the codes_u32 accessor, no configuration, and the returned shortlist is identical to the one the previous pipeline produced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * perf(search): per-doc parallel exact scoring The exact-scoring loops (dense and batched paths) processed the shortlist in fixed 128-doc parallel chunks. The chunking served a decompression-memory rationale that no longer holds — in-flight decompressed docs are bounded by the rayon thread count either way — and at n_decompress = 1024 it created only 8 tasks, underfilling and straggling the pool: 6.1 effective threads measured on a 10-core M4. Plain per-doc par_iter lets rayon split adaptively; 9.4 effective threads on the same measurement. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(search): harden optimized stage one * fix(search): avoid retaining duplicate code cache * fix(search): avoid unaligned code copies * fix(search): clamp n_ivf_probe to the centroid count before sizing probe scratch An oversized n_ivf_probe means "probe every cell" and was accepted by every release through v1.6.5 (colgrep passes COLGREP_N_IVF_PROBE through unclamped). The reworked probe path fed the raw parameter to Vec::with_capacity for its scratch buffers, so usize-scale values died with a capacity-overflow panic and large ones attempted multi-GB allocations. Clamp to the centroid count at the effective-probe computation, where the subset arm already clamps to the eligible pool. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Raphael Sourty <raphael.sourty@gmail.com>
C
Chao-Chun (Joe) Hsu committed
2c6ca2f8afd75c6502a92fa292b500f3127efe7e
Parent: 21903f1
Committed by GitHub <noreply@github.com>
on 8/17/2026, 9:43:49 AM