2026
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
CVPR 2026
General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effecti