VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain--i.e., precisely calibrated multi-view camera poses--to fuse multi-view information into a global scene representation, limiting deployment in real-world scenes. We target a more practical setting: Sensor-