PolySplat: Workload-Regime-Aware Rasterization for 3D Gaussian Splatting
Abstract
Rasterization dictates the interactive budget of 3D Gaussian Splatting (3DGS). However, the comparative speed of modern CUDA rasterizers is typically evaluated on a narrow canonical benchmark, masking severe regime-dependent performance reversals. This limited scope hides three critical kernel-level bottlenecks: warp lane underutilization, exposed global-memory latency during pixel shading, and terminal-tail penalties from hardware thread scheduling. In this paper, we present PolySplat, a workload-regime-aware 3DGS rasterizer designed to overcome these inefficiencies. PolySplat introduces a warp-saturating tile-key emitter with adaptive three-way dispatch, asynchronous shared-memory staging in the render kernel, and a persistent kernel architecture with centralized atomic dispatch to strictly bound tail latency. To rigorously validate our system, we introduce an extended 76-target benchmark that achieves comprehensive workload-regime coverage by stratifying targets across extreme Gaussian counts (up to 56.5M), high resolutions (up to 9K), and six diverse scene categories. Evaluated on this comprehensive suite, PolySplat achieves dataset-balanced geometric-mean speedups of 1.26x to 6.48x over state-of-the-art rasterizers at lossless visual quality. Notably, under sustained 60Hz interaction at extreme resolutions, PolySplat strictly meets 1-vsync display deadlines where existing renderers suffer from massive queue divergence and input-to-photon lag.