Geometry Matters in Packed 3D Attention
Amil Khan ⋅ Anantajit Subrahmanya ⋅ B.S. Manjunath
Abstract
Packed high-token 3D fields are not just long-sequence modeling: once a volume is packed into tokens, sequence order is no longer the geometry of the sample. This creates a general failure mode for attention over packed geometric data: sequence-index positional mechanisms can learn the memory layout rather than the object. We introduce Pack3D, a geometry-aware local--global attention operator that separates memory layout from attention geometry. Pack3D keeps local neighborhood attention exact, compresses only distant context into contiguous 3D block summaries, and indexes attention with axial RoPE on true 3D token-grid coordinates rather than packed sequence positions. With RoPE applied before block compression, each query interacts with a geometry-aware mixture of block tokens, so compressed global context remains spatially interpretable after packing. A synthetic repacking benchmark makes the failure mode explicit: learned positional embeddings and packed-sequence RoPE fail under layout changes, whereas axial grid-coordinate RoPE remains invariant. On Allen WTC--11 hiPSC microscopy, our main claim-grade domain, Pack3D improves matched learned-position local--global attention by 0.0978 best macro Dice and 0.1248 final macro Dice at 262k tokens and depth 8. In an Allen-15 joint comparison, Pack3D also improves over a parameter-matched TokenGrid 3D U-Net on both best and final macro Dice. Inference-time ablations show that the trained model genuinely uses its compressed summaries, and the fused forward/backward implementation is $6.49\times$ lower latency than dense SDPA at 65k tokens.
Chat is not available.
Successful Page Load