CC-GS: Low-Memory 3D Gaussian Splatting Training via CPU-GPU Block-Wise Context Compositing
Jian Xu ⋅ Siyi Wu ⋅ Yi Li ⋅ Bingzhe Li ⋅ Sian Jin ⋅ Wei Niu ⋅ Sheng Di ⋅ Yuede Ji ⋅ Miao Yin
Abstract
Training a single 3D Gaussian Splatting (3DGS) scene routinely requires tens of gigabytes of GPU memory, placing standard 3DGS training beyond many low-memory GPU budgets commonly found on laptops and edge devices. Existing low-memory approaches mainly follow two directions: block-wise training and frustum-culling-based host offloading. Block-wise methods train spatial blocks independently and merge them afterward. However, because 3DGS renders each image through depth-ordered alpha compositing over all visible Gaussians, independent block training reduces optimization to block-local objectives, thereby losing full-image loss supervision. Frustum-culling-based host offloading retains full-image loss supervision by jointly rendering all visible Gaussians while uploading only the view-visible subset to the GPU. However, its GPU memory consumption still scales with the size of the visible set, which can exceed the memory budget of low-end GPUs for dense scenes or wide-coverage views. We propose CC-GS, a low-memory block-wise 3DGS training framework based on CPU-GPU block-wise context compositing, featuring three key designs. $\textbf{i) Block-wise context compositing}$: CC-GS partitions Gaussians into capacity-bounded blocks and represents non-optimized visible blocks using cached per-pixel color and residual transmittance, with depth used only for ordering, enabling full-image loss supervision while processing one capacity-bounded block at a time. $\textbf{ii) Two-pass block-wise training}$: each training view is processed through a lightweight context pass for caching compositing context and a differentiable block-optimization pass for updating individual blocks. $\textbf{iii) Asynchronous CPU-GPU execution}$: CC-GS overlaps block gathering, host-device transfers, rendering, backpropagation, gradient offloading, and CPU-side updates through an asynchronous CPU-GPU pipeline. Experiments on standard 3DGS benchmarks show that CC-GS keeps the measured peak GPU memory below 2GB for all evaluated scenes while matching the rendering quality of standard 3DGS. Compared with standard 3DGS, CC-GS, incurs at most a $1.54\times$ dataset-level training slowdown.
Chat is not available.
Successful Page Load