TACOcc: Target-Adaptive Cross-Modal Fusion with Sequential Volume Rendering for 3D Semantic Occupancy Prediction
Luyao Lei, Shuo Xu, Yifan Bai, Zelin Yang, Yuanbo Guo, Xing Wei
Abstract
Multi-modal 3D semantic occupancy prediction remains challenged by two fundamental issues: (i) geometric--semantic misalignment introduced by fixed-neighborhood fusion under heterogeneous sensing distributions, and (ii) feature degradation with prediction inconsistency in dynamic scenes caused by sparse supervision. We propose TACOcc, a framework coupling a target-adaptive, bidirectional symmetric fusion module with sequential volume rendering supervision. The fusion module predicts a query-wise neighborhood size via a differentiable Gumbel-Softmax strategy, expanding the receptive field for large objects to enrich context while contracting it for small objects to suppress noise, thereby achieving precise cross-modal alignment. To stabilize predictions under sparse labels and motion, we introduce temporally enhanced Gaussian rendering that aggregates multi-frame dependencies, initializes dual-source geometric anchors, and transfers multi-view photometric constraints from images to 3D occupancy features. A velocity-adaptive temporal bandwidth further mitigates flicker in fast-motion cases. Experiments on nuScenes and SemanticKITTI demonstrate strong performance, including 28.9% mIoU on nuScenes, particularly improving small-object categories and long-range regions. These results highlight that scale-aware bidirectional fusion and temporally grounded volumetric supervision form an effective recipe for robust multi-modal occupancy perception.