SIGN IN SIGN UP

Decode whole words in bulk in bitmapContainer's many-iterator

nextMany walked set bits one at a time. When the current word is exhausted
it now decodes whole words with the AVX-512 kernel for as long as the
caller's buffer has room for the 64 values a single word could produce,
falling back to the word-at-a-time loop for the partial word. bitset stays
zero across the bulk step, preserving the invariant that word `base` has
been fully consumed.

The word count per step is room/64, which cannot overflow the buffer
whatever the data looks like.

Xeon Gold 6548N, ns per value, 4096-entry buffer:

    cardinality 4096    1.240 -> 0.418  (2.97x)
    cardinality 16384   1.348 -> 0.130  (10.4x)
    cardinality 32768   1.229 -> 0.097  (12.7x)
    cardinality 65536   1.021 -> 0.074  (13.8x)

BenchmarkRealDataNextMany over the whole corpus:

    census-income     10.19 ms -> 3.41 ms  (2.99x)
    weather_sept_85   20.22 ms -> 7.42 ms  (2.72x)

A 64-entry buffer is the worst case, one word per bulk step; it is still
slightly faster (1.21x), not a regression.

nextMany64 is left alone: []uint64 output needs a byte-to-quadword widen,
a different kernel with a different trade-off.

Claude-Session: https://claude.ai/code/session_0123ePBefqPjrCvhxdwWFakj
D
Daniel Lemire committed
f97fbb15f5ccb062d0f8c40cc83841c0a09c4400
Parent: 11cdbaa