Decode bitmap containers into array containers with VPCOMPRESSB
fillArray and the fillArrayAND/ANDNOT/XOR helpers all carried a
"TODO: rewrite in assembly". Add the uint16 counterpart of the kernel
introduced for ToArray: one VPCOMPRESSB per word, widened 32 at a time with
VPMOVZXBW and written with a masked store, so exactly popcount(word) values
are written and the output needs no slack.
fillArraySkipVector additionally tests eight words at a time with VPTESTMQ
and skips empty groups, which wins only on very sparse containers;
fillArray picks between the two on cardinality, crossover just under 768.
The scalar loops keep their exact previous bodies, with only an early-out
added in front, so the non-AVX-512 path is textually unchanged.
Xeon Gold 6548N, ns per value:
fillArray card=64 10.21 -> 2.567 (3.98x)
fillArray card=1024 1.797 -> 1.118 (1.61x)
fillArray card=4096 0.911 -> 0.372 (2.45x)
fillArrayAND out=1044 1.678 -> 1.076 (1.56x)
fillArrayAND out=4076 0.903 -> 0.392 (2.30x)
The sparsest two-input case (out=249) is about 9% slower: at that density
the cost is scanning both containers, not emitting the values, so the
kernel has nothing to win. A size threshold made it worse rather than
better because inserting the guard shifts the scalar loop's alignment. It
does not show end to end -- andBitmap at that size is 1.01x -- and a
VPTESTMQ group-skip variant should fix it properly as a follow-up.
Claude-Session: https://claude.ai/code/session_0123ePBefqPjrCvhxdwWFakj D
Daniel Lemire committed
f05cd7e0d3e604d554febe31bcf8779901c9772d
Parent: 11cdbaa