Optimize AVX2 popcount inner loop and CPU feature detection
David Sparks reviewed popcnt_avx2_amd64.s and suggested the improvements
implemented here; the credit for all of them is his.
The main win is in COUNTBLOCK. VPSADBW computes |a-b| per byte and sums
each group of 8, so it can absorb the per-byte add that VPADDB was doing:
feeding it the two nibble counts directly yields (B + lo) - (B - hi) =
lo + hi, removing VPADDB and its latency from the hot loop. That needs
two lookup tables, one biased up by B and one subtracted from B. The bias
must satisfy 4 <= B <= 251 so that neither table wraps as unsigned bytes
and a >= b always holds, making the absolute value a no-op; B = 15 is
free because Ymask already holds 15 in every byte.
The rest:
- Generic (VEX-encoded) AVX instructions accept an unaligned memory
source, so the second input of the And/Or/Xor/Mask loops is read
straight out of memory rather than loaded into a register first. Only
VPANDN's non-negated operand may come from memory, so that loop loads
m into the register and reads s from memory.
- Ydata/Yhi/Yc2 and Ylo/Yc1 have disjoint live ranges and now share
registers; six architectural registers suffice instead of ten.
- lutmask shrinks from 64 to 24 bytes: VBROADCASTI128 duplicates the
16-byte table into both 128-bit lanes and VPBROADCASTB splats the
nibble mask from a single byte.
- SHRQ and ANDQ already set ZF, so the TESTQ instructions that followed
them were redundant.
- SETUP moves below the branch that skips the vector loop entirely.
- The 64-bit tail loops use POPCNTQ with a memory source and a 32-bit
loop counter (32-bit avoids the partial-register merge a DECB needs).
- XORL rather than XORQ for zeroing: no REX prefix, and some cores do
not recognize XORQ as a zeroing idiom.
- _hasAVX2 complements each feature word and TESTs it instead of
AND/CMP, saving a large immediate per check, and all three checks now
fall through to a single SETEQ that stores ZF into the bool result.
Measured on an Intel Xeon Gold 6548N with go1.26.3. Microbenchmarks,
count=10:
PopcntSlice1024AVX2 205.0n -> 180.0n -12.20%
PopcntAndSlice1024AVX2 241.6n -> 225.7n -6.58%
Popcount 14.66n -> 13.73n -6.35%
And through the public API on dense bitmaps (32 bitmap containers,
count=8), which is what the popcount slice routines actually serve:
Bitmap.AndCardinality 9.404µ -> 8.655µ -7.96%
Bitmap.OrCardinality 128.6µ -> 120.8µ -6.02%
The pure-Go fallbacks are unchanged and measure flat, as expected. The
assembly text shrinks from 987 to 914 bytes and the rodata blob from 64
to 24. Verified against the Go reference for all five routines over the
existing differential test plus 20000 randomized rounds with unaligned
sub-slices and all-ones/all-zeros inputs. D
Daniel Lemire committed
2eed3b6342f506da0ebd045b4e05f3c3321bdec0
Parent: ff191ee