Use AVX512_VPOPCNTDQ for the population-count kernels
The AVX2 kernels build a population count out of a VPSHUFB nibble lookup
plus VPSADBW and handle 256 bits per iteration. AVX512_VPOPCNTDQ provides
VPOPCNTQ, which counts all eight 64-bit lanes of a ZMM register in one
instruction, so the new kernels handle 2048 bits per iteration with four
accumulators to keep the adds off a single dependency chain.
One 8 KB container on a Xeon Gold 6548N:
popcntSlice 179.9 ns -> 44.7 ns (4.02x vs AVX2)
popcntAndSlice 225.9 ns -> 68.5 ns (3.30x vs AVX2)
popcntOrSlice 226.9 ns -> 68.5 ns (3.31x vs AVX2)
These back GetCardinality, computeCardinality, AndCardinality,
OrCardinality, XorCardinality, Intersects, and the cardinality half of
every bitmap-container boolean operation.
Detection goes through golang.org/x/sys/cpu, so the OSXSAVE and XCR0
opmask/ZMM checks are handled upstream and GODEBUG=cpu.avx512vpopcntdq=off
disables the path at run time. CPUs without VPOPCNTQ keep the AVX2 path.
Claude-Session: https://claude.ai/code/session_0123ePBefqPjrCvhxdwWFakj D
Daniel Lemire committed
f1aa81dcb5f2a2efe67b08f1d0bd012a9df428ed
Parent: 92bd7e5