Skip to content

Add GPU voxel grid filter - #6489

Open
kai-waang wants to merge 2 commits into
PointCloudLibrary:masterfrom
kai-waang:gpu_voxel_grid
Open

kai-waang wants to merge 2 commits into
PointCloudLibrary:masterfrom
kai-waang:gpu_voxel_grid

Conversation

@kai-waang

Copy link
Copy Markdown
Contributor

As discussed in #5882, this PR adds pcl::gpu::VoxelGrid in a new pcl_gpu_filters module, allowing voxel-grid downsampling of pcl::PointXYZ clouds within GPU processing pipelines. It supports device input/output as well as host-cloud overloads. The implementation uses Thrust to sort points by voxel index.

Benchmarks are also added for comparison: CPU pcl::VoxelGrid, GPU filtering with device input/output, and GPU filtering with output download. The mug and milk datasets use a leaf size of 0.01 in each dimension. File loading and initial upload are excluded from timing; the download case includes transferring the output back to the host. The benchmarks were run on a system with two Intel Xeon Platinum 8488C CPUs, using a RTX PRO 6000.

Results from a Release build:

Dataset CPU GPU GPU with download
mug 5.92 ms 0.328 ms 0.372 ms
milk 7.73 ms 0.331 ms 0.453 ms
Raw benchmark output
2026-10-05T12:23:18+08:00
Running /home/wk/programs/pcl/build-release/benchmarks/benchmark_gpu_filters_voxel_grid
Run on (192 X 3800 MHz CPU s)
CPU Caches:
  L1 Data 48 KiB (x96)
  L1 Instruction 32 KiB (x96)
  L2 Unified 2048 KiB (x96)
  L3 Unified 107520 KiB (x2)
Load Average: 8.47, 8.54, 9.04
***WARNING*** CPU scaling is enabled, the benchmark real time measurements may be noisy and will incur extra overhead.
---------------------------------------------------------------------------
Benchmark                                 Time             CPU   Iterations
---------------------------------------------------------------------------
BM_VoxelGridCpu_mug                    5.92 ms         5.92 ms           10
BM_VoxelGridGpu_mug                   0.328 ms        0.328 ms          211
BM_VoxelGridGpuWithDownload_mug       0.372 ms        0.372 ms          186
BM_VoxelGridCpu_milk                   7.73 ms         7.69 ms            9
BM_VoxelGridGpu_milk                  0.331 ms        0.331 ms          206
BM_VoxelGridGpuWithDownload_milk      0.453 ms        0.453 ms          129

The GPU VoxelGrid filter is consistently faster than the CPU version across the point-cloud sizes I tested. Larger inputs were built by translating and merging copies of the milk point cloud (milk_cartoon_all_small_clorox.pcd), with non-finite points removed.
voxel_grid_scaling

The GPU version also shows a clear speed advantage across different leaf size:

voxel_grid_leaf_size
Output on different leaf size
Dataset mug: 307200 points, 209280 finite, is_dense=0
Bounding box (m): [-0.45643 -0.51074  0.69001] to [0.71518 0.17923  2.5927]
Leaf (m)   Output points     CPU (ms)     GPU (ms)   GPU+download (ms)
   0.002          120149      10.7900       0.5232              0.9694
   0.005           32921       8.9288       0.3408              0.4779
   0.010            9389       5.8868       0.3342              0.3772
   0.020            2620       5.3566       0.3287              0.3475
   0.050             507       5.0164       0.3316              0.3450
   0.100             154       3.7166       0.3502              0.3536
   0.200              53       3.4257       0.3968              0.3998

Dataset milk: 307200 points, 241407 finite, is_dense=0
Bounding box (m): [-1.0608 -0.8692  0.5010] to [1.1525 0.2197 2.0630]
Leaf (m)   Output points     CPU (ms)     GPU (ms)   GPU+download (ms)
   0.002          172905      12.7743       0.5259              1.1216
   0.005           66280      11.1402       0.5150              0.7855
   0.010           25253       7.6405       0.3377              0.4446
   0.020            7756       6.3814       0.3326              0.3683
   0.050            1509       5.5383       0.3298              0.3426
   0.100             425       4.2330       0.3372              0.3447
   0.200             135       4.2167       0.3593              0.3657

Add PointXYZ voxel-grid downsampling for device and host inputs, with
centroid computation and reusable device buffers. Preserve host metadata
and handle non-finite points and voxel-index overflow.

Register GPU benchmarks in the existing benchmark module and add focused
regression coverage for the filter.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant