← Search

Pujith Kachana

5 accepted papers

2026

LA-Pose: Latent Action Pretraining Meets Pose Estimation

CVPR 2026

This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alternative to the current trend of fully supervised training with 3D annotations. Concretely, we employ inverse- and forward-dynamics models to learn la

Cited by 0SourceScholar
2025

IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes

ICRA 2025

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language input. One such application is indoor navigation using natural la

Cited by 3SourcecodeScholar
2025

Rig3R: Rig-Aware Conditioning and Discovery for 3D Reconstruction

NeurIPS 2025spotlight

Estimating agent pose and 3D scene structure from multi-camera rigs is a central task in embodied AI applications such as autonomous driving. Recent learned approaches such as DUSt3R have shown impressive performance in multiview settings. However, these models treat images as unstructured collectio…

Cited by 0SourceScholar
2025

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

IROS 2025

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is often challenging due to the diversity of scenes, large number of fine-grained objects, and complex free-form nature of

Cited by 8SourcecodeScholar