← Search

Raktim Gautam Goswami

4 accepted papers

2025

OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation

NeurIPS 2025poster

Visual imitation learning enables robotic agents to acquire skills by observing expert demonstration videos. In the one-shot setting, the agent generates a policy after observing a single expert demonstration without additional fine-tuning. Existing approaches typically train and evaluate on the sam…

Cited by 0SourceScholar
2025

RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training

CVPR 2025highlight

Vision-based pose estimation of articulated robots with unknown joint angles has applications in collaborative robotics and human-robot interaction tasks. Current frameworks use neural network encoders to extract image features and downstream layers to predict joint angles and robot pose. While imag…

2024

Floor Plan Based Active Global Localization and Navigation Aid for Persons With Blindness and Low Vision

RA-L 2024

Navigation of an agent, such as a person with blindness or low vision, in an unfamiliar environment poses substantial difficulties, even in scenarios where prior maps, like floor plans, are available. It becomes essential first to determine the agent's pose in the environment. The task's complexity

Cited by 2SourceScholar
2024

SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition

RA-L 2024

Large-scale LiDAR mappings and localization leverage place recognition techniques to mitigate odometry drifts, ensuring accurate mapping. These techniques utilize scene representations from LiDAR point clouds to identify previously visited sites within a database. Local descriptors, assigned to each

Cited by 10SourcecodeScholar