SIGN IN SIGN UP

perf: rework the stage-1 shortlist pipeline — 4–12× faster stage 1 (5.6–7.1× e2e combined with #169) (#170)

* perf(search): rework the stage-1 shortlist pipeline

Stage 1 (candidate generation) rebuilt phase by phase, each change
equivalence-tested against the path it replaces:

- centroid GEMM: column-block-parallel (par_cdot), bit-identical to the
  single dot call — the dim-long reduction stays inside one block.
- IVF probe: running-threshold chunk scan (probe_top_k_scan) replaces a
  K-length buffer fill + select_nth per query token; same top-k set by
  value, one sequential read per row.
- candidate gather: doc ids are dense, so a word bitmap dedups in
  O(postings) and scanning its set bits emits the same sorted list the
  sort+dedup produced, without the sort.
- approximate flood: centroid scores quantized to u8 in a fused
  transpose+quantize pass (PLAID's own centroid-interaction trick) and
  scored chunk-outer with a register-resident 16-lane max accumulator.
  Quantization is monotone per lane; validated at nDCG@10 parity to four
  decimals against exact flooding on real-embedding corpora.
- prune: partial-select the surviving n_full_scores before sorting only
  that prefix; identical top list and order.
- code reads: a lazily-built u32 side-array of the token codes, so the
  flood reads 4 B/token instead of the mmap's 8.

Stage 1 is self-contained: no new public API beyond the codes_u32
accessor, no configuration, and the returned shortlist is identical to
the one the previous pipeline produced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(search): per-doc parallel exact scoring

The exact-scoring loops (dense and batched paths) processed the shortlist
in fixed 128-doc parallel chunks. The chunking served a
decompression-memory rationale that no longer holds — in-flight
decompressed docs are bounded by the rayon thread count either way — and
at n_decompress = 1024 it created only 8 tasks, underfilling and
straggling the pool: 6.1 effective threads measured on a 10-core M4.
Plain per-doc par_iter lets rayon split adaptively; 9.4 effective threads
on the same measurement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(search): harden optimized stage one

* fix(search): avoid retaining duplicate code cache

* fix(search): avoid unaligned code copies

* fix(search): clamp n_ivf_probe to the centroid count before sizing probe scratch

An oversized n_ivf_probe means "probe every cell" and was accepted by
every release through v1.6.5 (colgrep passes COLGREP_N_IVF_PROBE through
unclamped). The reworked probe path fed the raw parameter to
Vec::with_capacity for its scratch buffers, so usize-scale values died
with a capacity-overflow panic and large ones attempted multi-GB
allocations. Clamp to the centroid count at the effective-probe
computation, where the subset arm already clamps to the eligible pool.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Raphael Sourty <raphael.sourty@gmail.com>
C
Chao-Chun (Joe) Hsu committed
2c6ca2f8afd75c6502a92fa292b500f3127efe7e
Parent: 21903f1
Committed by GitHub <noreply@github.com> on 8/17/2026, 9:43:49 AM