SIGN IN SIGN UP

perf(binary): skip residual computation in encode_index_chunk for binary indexes (#159)

Binary indexes store the embeddings' sign bits and never read residuals,
but encode_index_chunk still ran compress_and_residuals_cpu — a full f32
clone of the batch plus a centroid subtraction per token — and threw the
residuals away. Mirror create_index_files: in binary mode compute only
the nearest-centroid codes the IVF needs.

~9% faster per chunk on Apple silicon (2k docs / 64k tokens / dim 128:
42.2ms -> 38.7ms) and drops an O(tokens x dim) f32 allocation (~32 MB
per chunk at that shape) from colgrep's binary build path.

Follow-up from the #155 review; regression test asserts the binary chunk
is bit-identical to create_index_files' artifact (sign bits + IVF codes).
R
Raphael Sourty committed
4ff801eef11004e20a6ffb62591b6aaeb6859aec
Parent: dba1539
Committed by GitHub <noreply@github.com> on 7/23/2026, 12:10:57 PM