feat: asymmetric residual scoring — int8 query × fused LUT MaxSim, 5–8× faster rescoring, identical NDCG@10 (#169)
* feat(residual): asymmetric int8-query x LUT scoring for residual indexes Score stored residual codes directly — int8 query x fused byte->weights LUT plus the centroid term stage-1 already computed — instead of decompress->f32 GEMM. Fused kernels (NEON tbl+SDOT, AVX2 pshufb+maddubs, AVX-512 VNNI vpdpbusd) expand each doc token's packed bytes once in registers, amortized over all query rows, and fold the float epilogue 4/8/16 query rows at a time against a centroid-major score matrix. Every path computes the identical integer accumulator and the identical float epilogue expression: the parity suite asserts bit-equality with the scalar reference across nq x nbits x dim, per kernel and through the dispatcher, and a semantics test pins the normalized scoring to the decompress reference. The safe dispatcher hard-asserts every SIMD precondition, including the padded query-plane stride the kernels' past-dim chunk loads rely on. Quality measured at |dNDCG@10| <= 0.002 across 3 ColBERT checkpoints x 3 BEIR corpora x nbits 4/2/1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(search): residual_asym — wire fused LUT scoring into both search paths SearchParameters::residual_asym (default off; same index, A/B-able per search) routes exact scoring through the fused kernels. The dense path transposes stage-1's [nq, K] centroid matrix once per query into the kernels' centroid-major layout with a cache-blocked transpose; the batched-centroid path packs the sparse centroid scores it already computed into a compact centroid-major matrix plus a per-doc code remap, so the asym path survives num_centroids > centroid_batch_size unchanged. Per-token 1/||reconstruction|| norms (the float path's renormalize) are computed once per index on first use and cached; the initializer runs on a dedicated rayon pool because a global-pool par_iter inside get_or_init can deadlock when reached from inside a worker. Integration tests: self-retrieval per nbits, float-path ranking agreement, score tracking, batched-vs-dense parity, scalar dispatch for non-multiple-of-8 dims, and the float fallback above the fused kernels' dim ceiling. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(api,colgrep): expose residual_asym and warn on silent kernel fallback The flag shipped unreachable. SearchParameters::residual_asym defaults to false and both consumer surfaces built their parameters with `..Default::default()`, so neither the HTTP API nor colgrep could turn it on — anyone benchmarking through them measured the float path with no way to enable the feature. SearchParamsRequest gains `residual_asym` (Option<bool>, additive and backwards compatible); colgrep gains COLGREP_RESIDUAL_ASYM alongside its three existing search-time env knobs. The arm also falls back silently: to float for binary indexes, dims above MAX_DIM, and codecs whose LUT will not factor, and to the scalar kernel when the CPU lacks AVX2 or NEON `dotprod`. Every fallback returns correct scores, so the only symptom is that a ~5x rescore measures ~1.2x — which reads as "the optimization does not work" rather than "it did not run". Measured on one machine and one index, varying only the build target (SciFact, 5183 docs, nbits 4, 300 real queries): native aarch64 dispatched to neon-sdot and rescored in 3.16 ms vs 15.75 ms float (4.99x); the same tree built x86_64 under Rosetta dispatched to scalar and measured 22.75 ms vs 27.76 ms (1.22x). prepare_score_query now reports the outcome once per process — a warning when the caller asked for the fused kernel and will not get it, plus the success case under NEXT_PLAID_REPORT_KERNEL=1 so a benchmark can record which kernel produced its numbers. residual_lut::simd_dispatch_available mirrors the kernel's own dispatch condition so the report cannot drift from what actually runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(example): asym_rescore_check — prove which rescore kernel dispatched Anyone measuring this PR needs to know which kernel ran before trusting a ratio. The SIMD paths dispatch at runtime and fall back silently, so a build that never reaches them still returns correct scores and simply measures a small speedup — which reads as "the optimization does not work". This example names the dispatched kernel first, loudly flagging the scalar case, then times Stage 2's two arms on identical inputs under rayon, as a real search runs them: decompress + f32 MaxSim against maxsim_residual_lut_i8 straight over the packed codes. On an M4 it reports 3.0x with neon-sdot and 0.72x — asym slower than float — when built x86_64 under Rosetta, so the failure mode is unmistakable. No dataset and no index build: it constructs a codec and packed residuals directly and runs in about a second. Uses only the crate's public API. The synthetic corpus is RAM-resident, where the float arm costs ~44 ns/token against ~67 measured on a real mmap-backed index, so its ratio is a floor rather than a reproduction of the headline numbers. Its job is the kernel line; the quoted speedups come from real corpora. Standalone and safe to delete once you have run it — nothing else references it, so reverting this one commit removes it cleanly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(index): don't block a rayon worker inside the inv-norms initializer `search_many_mmap(parallel = true)` with `residual_asym` on a cold index deadlocked: all workers parked, 0% CPU, forever. That is the first request a batch-serving deployment makes, so it was reachable in normal use — a single-query server never hit it, which is why it survived the demo. Every worker races into `residual_inv_norms`'s `OnceLock::get_or_init`. The winner called `ThreadPool::install` on a dedicated one-shot pool, and while blocked there it *stays available to the outer pool's work-stealing* — so it stole a sibling query, that query re-entered `get_or_init` on the same thread, and the initializer could never finish. The comment there described this exact failure mode and claimed the dedicated pool prevented it; the pool prevents the compute from depending on the global pool, but `install` itself reopens the hole. The invariant is narrower than the old comment assumed: the initializer must not block the calling worker in *any* way that leaves the outer pool free to schedule onto it. That rules out a global-pool `par_iter` and `install` alike. So hand the parallel compute to a plain OS thread and block on joining it — a thread join parks without stealing — while still running the work on a dedicated pool, since by that point every other global worker may already be parked on this OnceLock with nobody left to run an injected job. Kept lazy rather than built eagerly at load: the cache is 4 bytes per token (~400 MB on a 100M-token index) for a feature that is default off, and binary indexes never use it at all. Also preferred over computing outside `get_or_init`, which is correct but has every racing worker redo the whole job — worst precisely on the batch path that triggers this. Regression test drives 128 queries through the parallel path on a cold 400-doc index; it hangs on the previous initializer and passes on this one. It carries a watchdog because the failure mode is a hang rather than a panic, which would otherwise stall CI instead of failing it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(python-sdk): expose residual_asym on SearchParams and the search CLI Adding the field to the HTTP API was not enough to make it reachable: the Python client models search parameters as its own dataclass whose `to_dict` whitelists keys, so `residual_asym` was dropped before the request was ever built. Anyone driving the API through the SDK — which is how the CLI works — still had no way to turn the feature on. `SearchParams.residual_asym` is `Optional[bool]` defaulting to None and is omitted from the payload entirely when unset, rather than sent as false, so callers who never express an opinion keep following the server's default if it later changes. `next-plaid search` gains a matching `--residual-asym/--no-residual-asym`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(search): harden asymmetric residual setup * fix(search): bound asymmetric norm memory * mmap residual inverse norms * fix: satisfy clippy for residual norm update * feat(index): prewarm LUT sidecars for legacy indexes * fix(index): bound legacy LUT prewarm memory * fix(search): skip the subset probe select when n_ivf_probe is 0 The subset arm's partial select computes n_probe - 1, so n_ivf_probe = 0 with a subset underflowed to usize::MAX and panicked inside select_nth_unstable_by (present since v1.6.5; the no-subset path was fixed incidentally by the stage-1 rework's probe scan, which handles n = 0 internally). Zero cells probed now returns an empty result on both paths. * feat(search)!: asymmetric residual scoring becomes the residual scoring path There is now one way to search a residual index. The residual_asym parameter is removed from SearchParameters, the REST API, the Python SDK, and colgrep's COLGREP_RESIDUAL_ASYM env var — none of which ever shipped in a release. Dispatch is automatic: residual indexes score through the fused int8xLUT kernels; binary indexes, dims above MAX_DIM, and codecs without norm tables keep their existing paths as silent internal fallbacks. NEXT_PLAID_FLOAT_RESCORE=1 remains as an undocumented process-wide emergency escape (not public API). To make the default land with its speedup rather than just its scores, load() now auto-builds the inverse-norm sidecar for pre-sidecar indexes (the one-time cost of prewarm_residual_lut_sidecar, ~1s per 3M tokens); failure to write — e.g. a read-only index directory — degrades to the per-document norm fallback with a one-time warning instead of erroring. Kernel dispatch reporting is now entirely opt-in via NEXT_PLAID_REPORT_KERNEL: a fallback is normal operation, not a broken request. Equivalence tests lose the flag they toggled, so their float reference is now computed from the public decompression API (get_document_embeddings + maxsim_score) — the exact computation the internal fallback runs — and the oversize-dim test asserts bit-identical scores against it. The read-only degradation gets a unix-gated test. * chore(example): drop asym_rescore_check The example existed to prove which rescore kernel an explicit residual_asym request resolved to. With the flag removed and dispatch automatic, NEXT_PLAID_REPORT_KERNEL=1 on any search prints the resolved kernel, which is all the example demonstrated. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Raphael Sourty <raphael.sourty@gmail.com>
C
Chao-Chun (Joe) Hsu committed
8f01b07cc65fbc1b006735660cfafdd9e96328cb
Parent: bbf8e5a
Committed by GitHub <noreply@github.com>
on 8/18/2026, 8:57:50 AM