fastaggregation: avoid runContainer16.lazyIOR slow path in FastOr
(*Bitmap).lazyOR's same-key branch calls c1.lazyIOR(c2). When c1 is a runContainer16, lazyIOR falls back to ior -> inplaceUnion -> Add -> searchRange, which is O(N · logR) per element merged. This makes FastOr catastrophically slow over inputs containing run-encoded blocks: a synthetic benchmark with 15 bitmaps of ~6000-bit runs takes ~335 ms/op before this change. Pre-promote runContainer16 slots to bitmapContainer before the inner lazyIOR call. The bitmapContainer's lazy union is O(1024) regardless of cardinality, so the K-way fan-in stays linear in K. Mirrors the explicit toBitmapContainer pre-promotion in parallel.go ParHeapOr and the Java BitmapContainer.lazyor(RunContainer) path. Issue #81 has been tracking the deeper fix (proper runContainer16.lazyIOR/lazyOR implementations) since 2016; this is the surgical workaround at the single FastOr call site. BenchmarkFastOrRunContainers (added): before: 335 ms/op 12 KB/op 447 allocs/op after: 637 µs/op 335 KB/op 257 allocs/op (~526x) The output container type for slots that started as runContainer16 inputs is now bitmapContainer (or arrayContainer after repairAfterLazy if sparse). repairAfterLazy does not re-encode to runs; callers that want a run-optimised result should call RunOptimize() on the FastOr return value. ParHeapOr already behaves the same way. The full v2 test suite passes. Refs: https://github.com/RoaringBitmap/roaring/issues/81
T
tamirms committed
482ea719feb65f4d7b3722bbbb0e8cc50bdaf5d7
Parent: 6d3d113