SIGN IN SIGN UP

fastaggregation: avoid runContainer16.lazyIOR slow path in FastOr

(*Bitmap).lazyOR's same-key branch calls c1.lazyIOR(c2). When c1 is
a runContainer16, lazyIOR falls back to ior -> inplaceUnion -> Add ->
searchRange, which is O(N · logR) per element merged. This makes FastOr
catastrophically slow over inputs containing run-encoded blocks: a
synthetic benchmark with 15 bitmaps of ~6000-bit runs takes ~335 ms/op
before this change.

Pre-promote runContainer16 slots to bitmapContainer before the inner
lazyIOR call. The bitmapContainer's lazy union is O(1024) regardless
of cardinality, so the K-way fan-in stays linear in K. Mirrors the
explicit toBitmapContainer pre-promotion in parallel.go ParHeapOr and
the Java BitmapContainer.lazyor(RunContainer) path. Issue #81 has been
tracking the deeper fix (proper runContainer16.lazyIOR/lazyOR
implementations) since 2016; this is the surgical workaround at the
single FastOr call site.

BenchmarkFastOrRunContainers (added):
  before:  335 ms/op    12 KB/op   447 allocs/op
  after:   637 µs/op   335 KB/op   257 allocs/op    (~526x)

The output container type for slots that started as runContainer16
inputs is now bitmapContainer (or arrayContainer after repairAfterLazy
if sparse). repairAfterLazy does not re-encode to runs; callers that
want a run-optimised result should call RunOptimize() on the FastOr
return value. ParHeapOr already behaves the same way.

The full v2 test suite passes.

Refs: https://github.com/RoaringBitmap/roaring/issues/81
T
tamirms committed
482ea719feb65f4d7b3722bbbb0e8cc50bdaf5d7
Parent: 6d3d113