SIGN IN SIGN UP

feat: asymmetric residual scoring — int8 query × fused LUT MaxSim, 5–8× faster rescoring, identical NDCG@10 (#169)

* feat(residual): asymmetric int8-query x LUT scoring for residual indexes

Score stored residual codes directly — int8 query x fused byte->weights
LUT plus the centroid term stage-1 already computed — instead of
decompress->f32 GEMM. Fused kernels (NEON tbl+SDOT, AVX2 pshufb+maddubs,
AVX-512 VNNI vpdpbusd) expand each doc token's packed bytes once in
registers, amortized over all query rows, and fold the float epilogue
4/8/16 query rows at a time against a centroid-major score matrix.

Every path computes the identical integer accumulator and the identical
float epilogue expression: the parity suite asserts bit-equality with the
scalar reference across nq x nbits x dim, per kernel and through the
dispatcher, and a semantics test pins the normalized scoring to the
decompress reference. The safe dispatcher hard-asserts every SIMD
precondition, including the padded query-plane stride the kernels' past-dim
chunk loads rely on. Quality measured at |dNDCG@10| <= 0.002 across
3 ColBERT checkpoints x 3 BEIR corpora x nbits 4/2/1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(search): residual_asym — wire fused LUT scoring into both search paths

SearchParameters::residual_asym (default off; same index, A/B-able per
search) routes exact scoring through the fused kernels. The dense path
transposes stage-1's [nq, K] centroid matrix once per query into the
kernels' centroid-major layout with a cache-blocked transpose; the
batched-centroid path packs the sparse centroid scores it already
computed into a compact centroid-major matrix plus a per-doc code remap,
so the asym path survives num_centroids > centroid_batch_size unchanged.

Per-token 1/||reconstruction|| norms (the float path's renormalize) are
computed once per index on first use and cached; the initializer runs on
a dedicated rayon pool because a global-pool par_iter inside get_or_init
can deadlock when reached from inside a worker.

Integration tests: self-retrieval per nbits, float-path ranking
agreement, score tracking, batched-vs-dense parity, scalar dispatch for
non-multiple-of-8 dims, and the float fallback above the fused kernels'
dim ceiling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(api,colgrep): expose residual_asym and warn on silent kernel fallback

The flag shipped unreachable. SearchParameters::residual_asym defaults to
false and both consumer surfaces built their parameters with
`..Default::default()`, so neither the HTTP API nor colgrep could turn it
on — anyone benchmarking through them measured the float path with no way
to enable the feature. SearchParamsRequest gains `residual_asym`
(Option<bool>, additive and backwards compatible); colgrep gains
COLGREP_RESIDUAL_ASYM alongside its three existing search-time env knobs.

The arm also falls back silently: to float for binary indexes, dims above
MAX_DIM, and codecs whose LUT will not factor, and to the scalar kernel
when the CPU lacks AVX2 or NEON `dotprod`. Every fallback returns correct
scores, so the only symptom is that a ~5x rescore measures ~1.2x — which
reads as "the optimization does not work" rather than "it did not run".
Measured on one machine and one index, varying only the build target
(SciFact, 5183 docs, nbits 4, 300 real queries): native aarch64 dispatched
to neon-sdot and rescored in 3.16 ms vs 15.75 ms float (4.99x); the same
tree built x86_64 under Rosetta dispatched to scalar and measured 22.75 ms
vs 27.76 ms (1.22x).

prepare_score_query now reports the outcome once per process — a warning
when the caller asked for the fused kernel and will not get it, plus the
success case under NEXT_PLAID_REPORT_KERNEL=1 so a benchmark can record
which kernel produced its numbers. residual_lut::simd_dispatch_available
mirrors the kernel's own dispatch condition so the report cannot drift
from what actually runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(example): asym_rescore_check — prove which rescore kernel dispatched

Anyone measuring this PR needs to know which kernel ran before trusting a
ratio. The SIMD paths dispatch at runtime and fall back silently, so a build
that never reaches them still returns correct scores and simply measures a
small speedup — which reads as "the optimization does not work".

This example names the dispatched kernel first, loudly flagging the scalar
case, then times Stage 2's two arms on identical inputs under rayon, as a
real search runs them: decompress + f32 MaxSim against
maxsim_residual_lut_i8 straight over the packed codes. On an M4 it reports
3.0x with neon-sdot and 0.72x — asym slower than float — when built x86_64
under Rosetta, so the failure mode is unmistakable.

No dataset and no index build: it constructs a codec and packed residuals
directly and runs in about a second. Uses only the crate's public API.

The synthetic corpus is RAM-resident, where the float arm costs ~44 ns/token
against ~67 measured on a real mmap-backed index, so its ratio is a floor
rather than a reproduction of the headline numbers. Its job is the kernel
line; the quoted speedups come from real corpora.

Standalone and safe to delete once you have run it — nothing else references
it, so reverting this one commit removes it cleanly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(index): don't block a rayon worker inside the inv-norms initializer

`search_many_mmap(parallel = true)` with `residual_asym` on a cold index
deadlocked: all workers parked, 0% CPU, forever. That is the first request a
batch-serving deployment makes, so it was reachable in normal use — a
single-query server never hit it, which is why it survived the demo.

Every worker races into `residual_inv_norms`'s `OnceLock::get_or_init`. The
winner called `ThreadPool::install` on a dedicated one-shot pool, and while
blocked there it *stays available to the outer pool's work-stealing* — so it
stole a sibling query, that query re-entered `get_or_init` on the same
thread, and the initializer could never finish. The comment there described
this exact failure mode and claimed the dedicated pool prevented it; the pool
prevents the compute from depending on the global pool, but `install` itself
reopens the hole.

The invariant is narrower than the old comment assumed: the initializer must
not block the calling worker in *any* way that leaves the outer pool free to
schedule onto it. That rules out a global-pool `par_iter` and `install`
alike. So hand the parallel compute to a plain OS thread and block on joining
it — a thread join parks without stealing — while still running the work on a
dedicated pool, since by that point every other global worker may already be
parked on this OnceLock with nobody left to run an injected job.

Kept lazy rather than built eagerly at load: the cache is 4 bytes per token
(~400 MB on a 100M-token index) for a feature that is default off, and
binary indexes never use it at all. Also preferred over computing outside
`get_or_init`, which is correct but has every racing worker redo the whole
job — worst precisely on the batch path that triggers this.

Regression test drives 128 queries through the parallel path on a cold
400-doc index; it hangs on the previous initializer and passes on this one. It
carries a watchdog because the failure mode is a hang rather than a panic,
which would otherwise stall CI instead of failing it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(python-sdk): expose residual_asym on SearchParams and the search CLI

Adding the field to the HTTP API was not enough to make it reachable: the
Python client models search parameters as its own dataclass whose `to_dict`
whitelists keys, so `residual_asym` was dropped before the request was ever
built. Anyone driving the API through the SDK — which is how the CLI works —
still had no way to turn the feature on.

`SearchParams.residual_asym` is `Optional[bool]` defaulting to None and is
omitted from the payload entirely when unset, rather than sent as false, so
callers who never express an opinion keep following the server's default
if it later changes. `next-plaid search` gains a matching
`--residual-asym/--no-residual-asym`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(search): harden asymmetric residual setup

* fix(search): bound asymmetric norm memory

* mmap residual inverse norms

* fix: satisfy clippy for residual norm update

* feat(index): prewarm LUT sidecars for legacy indexes

* fix(index): bound legacy LUT prewarm memory

* fix(search): skip the subset probe select when n_ivf_probe is 0

The subset arm's partial select computes n_probe - 1, so n_ivf_probe = 0
with a subset underflowed to usize::MAX and panicked inside
select_nth_unstable_by (present since v1.6.5; the no-subset path was
fixed incidentally by the stage-1 rework's probe scan, which handles
n = 0 internally). Zero cells probed now returns an empty result on both
paths.

* feat(search)!: asymmetric residual scoring becomes the residual scoring path

There is now one way to search a residual index. The residual_asym
parameter is removed from SearchParameters, the REST API, the Python
SDK, and colgrep's COLGREP_RESIDUAL_ASYM env var — none of which ever
shipped in a release. Dispatch is automatic: residual indexes score
through the fused int8xLUT kernels; binary indexes, dims above MAX_DIM,
and codecs without norm tables keep their existing paths as silent
internal fallbacks. NEXT_PLAID_FLOAT_RESCORE=1 remains as an
undocumented process-wide emergency escape (not public API).

To make the default land with its speedup rather than just its scores,
load() now auto-builds the inverse-norm sidecar for pre-sidecar indexes
(the one-time cost of prewarm_residual_lut_sidecar, ~1s per 3M tokens);
failure to write — e.g. a read-only index directory — degrades to the
per-document norm fallback with a one-time warning instead of erroring.
Kernel dispatch reporting is now entirely opt-in via
NEXT_PLAID_REPORT_KERNEL: a fallback is normal operation, not a broken
request.

Equivalence tests lose the flag they toggled, so their float reference
is now computed from the public decompression API
(get_document_embeddings + maxsim_score) — the exact computation the
internal fallback runs — and the oversize-dim test asserts bit-identical
scores against it. The read-only degradation gets a unix-gated test.

* chore(example): drop asym_rescore_check

The example existed to prove which rescore kernel an explicit
residual_asym request resolved to. With the flag removed and dispatch
automatic, NEXT_PLAID_REPORT_KERNEL=1 on any search prints the resolved
kernel, which is all the example demonstrated.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Raphael Sourty <raphael.sourty@gmail.com>
C
Chao-Chun (Joe) Hsu committed
8f01b07cc65fbc1b006735660cfafdd9e96328cb
Parent: bbf8e5a
Committed by GitHub <noreply@github.com> on 8/18/2026, 8:57:50 AM