Skip to content

Union run containers by run instead of by value - #575

Open
gitRasheed wants to merge 1 commit into
RoaringBitmap:masterfrom
gitRasheed:perf/run-union
Open

gitRasheed wants to merge 1 commit into
RoaringBitmap:masterfrom
gitRasheed:perf/run-union

Conversation

@gitRasheed

Copy link
Copy Markdown

Description

runContainer16.inplaceUnion added the other container one value at a time. An in-place Or on run containers got slower with every value. It now inserts up to 16 runs in place and merges above that, like inplaceIntersect does.

Type of Change

  • Performance improvement
  • Test improvements

Changes Made

What was changed?

  • inplaceUnion inserts runs in place or merges.
  • union copies runs in blocks.
  • toArrayContainer allocates once.
  • iorRun16 merges above 4 runs.
  • The union and intersect helpers check the next run before searching.
  • New union tests and three real-data benchmarks.

Why was it changed?

An in-place Or over every adjacent pair in weather_sept_85_srt took 85 ms.

How was it changed?

A few runs are cheaper to insert than a full merge. Many runs are cheaper to merge. Without the in-place path, folding small bitmaps into a large accumulator was 2.3x slower than master on c7i.

Testing

go test ./... passes on c7i.xlarge and c8g.xlarge. The root package also passes with -tags=gofuzz.

Formatting

go fmt, make unconvert and git diff --check are clean.

Fuzzing

smat, 300 seconds: 1.5M executions, no failures.

Performance Impact

Real datasets after RunOptimize, less time than master (geometric mean):

Instance Or in place Or fold And in place
c7i.xlarge 72% 85% 6%
c8g.xlarge 73% 85% 6%

Sorted datasets gain the most. The in-place pass over weather_sept_85_srt drops from 85 ms to 3.0 ms.

Benchmark c7i master c7i this PR c8g master c8g this PR
RunContainerInplaceUnion, 4 runs 3.73 µs 1.92 µs 4.02 µs 2.30 µs
RunContainerInplaceUnion, 1,024 runs 665.8 µs 19.6 µs 637.1 µs 19.4 µs
ArrayContainerIorRun16, 4 runs 2.83 µs 2.83 µs 2.76 µs 2.77 µs
ArrayContainerIorRun16, 100 runs 61.5 µs 4.75 µs 57.8 µs 3.75 µs

Some small shapes are slower, mostly operands whose runs the receiver already holds. They cost a few µs more, up to 7.6 µs instead of 3.8 µs on c7i. Skipping runs the receiver already contains before merging would avoid this. I left it out to keep the PR small.

Breaking Changes

None.

runContainer16.inplaceUnion added the other container one value at a
time. It now inserts up to 16 runs in place and merges above that.
toArrayContainer allocates once, iorRun16 merges above 4 runs, and the
union and intersect helpers check the next run before searching.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant