perf(binary): skip residual computation in encode_index_chunk for binary indexes (#159)
Binary indexes store the embeddings' sign bits and never read residuals, but encode_index_chunk still ran compress_and_residuals_cpu — a full f32 clone of the batch plus a centroid subtraction per token — and threw the residuals away. Mirror create_index_files: in binary mode compute only the nearest-centroid codes the IVF needs. ~9% faster per chunk on Apple silicon (2k docs / 64k tokens / dim 128: 42.2ms -> 38.7ms) and drops an O(tokens x dim) f32 allocation (~32 MB per chunk at that shape) from colgrep's binary build path. Follow-up from the #155 review; regression test asserts the binary chunk is bit-identical to create_index_files' artifact (sign bits + IVF codes).
R
Raphael Sourty committed
4ff801eef11004e20a6ffb62591b6aaeb6859aec
Parent: dba1539
Committed by GitHub <noreply@github.com>
on 7/23/2026, 12:10:57 PM