← Search

Saurabh Nair

3 accepted papers

2026

LA-Pose: Latent Action Pretraining Meets Pose Estimation

CVPR 2026

This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alternative to the current trend of fully supervised training with 3D annotations. Concretely, we employ inverse- and forward-dynamics models to learn la

Cited by 0SourceScholar
2025

Rig3R: Rig-Aware Conditioning and Discovery for 3D Reconstruction

NeurIPS 2025spotlight

Estimating agent pose and 3D scene structure from multi-camera rigs is a central task in embodied AI applications such as autonomous driving. Recent learned approaches such as DUSt3R have shown impressive performance in multiview settings. However, these models treat images as unstructured collectio…

Cited by 0SourceScholar
2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…