Use VBIT for the arm64 NEON AndNot popcount kernel
_popcntMaskSliceNEON computes popcount(s &^ m). It was forming s &^ m as
an invert followed by an AND: a vector of all-ones was materialised once
in V15, and each iteration did four VEOR to produce ~m and then four VAND
to apply it. That is eight vector operations per 64-byte block to express
what the hardware can do in four.
arm64 has BIC (and-not) for exactly this, but the Go assembler has no
mnemonic for the vector form. VBIT does the same job in the same single
instruction. BIT ("bitwise insert if true") selects bit by bit between
its destination and one source under the control of the other:
Vd<i> = Vm<i> ? Vn<i> : Vd<i> i.e. Vd = (Vm & Vn) | (^Vm & Vd)
Holding Vn at zero reduces that to "clear every bit of Vd that is set in
Vm", which is Vd &^= Vm. So with Vd = s and Vm = m the mask is applied in
place, and V15 is now simply kept at zero to serve as Vn instead of
holding all-ones for the invert.
The inner loop drops from eight ops to four. The routine was ALU-bound
rather than load-bound, so this shows up directly:
popcntMaskSlice, 1024 words 137.7ns -> 107.3ns 1.28x
62 GB/s -> 80 GB/s
Measured on an Apple M4 Max, Go 1.24.4. The And/Or/Xor/Slice kernels are
untouched and unchanged. Container-level AndNot gains far less (~1%),
because andNotBitmap's cost is dominated by the scalar loop that writes
the result, not by this cardinality pass.
The observation that arm64 can do this in one instruction is due to
David Sparks, who also prompted the UADALP note now recorded in the file
header: folding byte lanes with UADALP would save an instruction per
iteration, but it has no Go assembler mnemonic and issues at roughly 2.4
per cycle on an M4 against 4 per cycle for VADD/VCNT/VUADDW, making it a
~20% loss for the plain popcount kernel. D
Daniel Lemire committed
82d1a83f24b2bd339cf2a666c64448311d2acc89
Parent: b2081ae