← Search

Rohit Mohan

14 accepted papers

2026

ForecastOcc: Vision-Based Semantic Occupancy Forecasting

ICRA 2026poster

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and dynamic objects, while semantic information remains largely a…

2026

UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation

RSS 2026poster

LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a critical failure mode. Under adverse conditions, degradation or failure of the camera sensor can significantly compromise the reliability of the perception sys…

Cited by 0SourceScholar
2025

Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters

IROS 2025

LiDAR semantic segmentation models are typically trained from random initialization as universal pre-training is hindered by the lack of large, diverse datasets. Moreover, most point cloud segmentation architectures incorporate custom network layers, limiting the transferability of advances from vis

Cited by 7SourceScholar
2025

Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware Learning

IROS 2025

Autonomous vehicles that navigate in open-world environments may encounter previously unseen object classes. However, most existing LiDAR panoptic segmentation models rely on closed-set assumptions, failing to detect unknown object instances. In this work, we propose ULOPS, an uncertainty-guided ope

Cited by 3SourceScholar
2024

AmodalSynthDrive: A Synthetic Amodal Perception Dataset for Autonomous Driving

RA-L 2024

Unlike humans, who can effortlessly estimate the entirety of objects even when partially occluded, modern computer vision algorithms still find this aspect extremely challenging. Leveraging this amodal perception for autonomous driving remains largely untapped due to the lack of suitable datasets. T

Cited by 16SourceScholar
2024

Progressive Multi-Modal Fusion for Robust 3D Object Detection

CoRL 2024poster

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspect…

Cited by 3SourceScholar
2024

Syn-Mediverse: A Multimodal Synthetic Dataset for Intelligent Scene Understanding of Healthcare Facilities

RA-L 2024

Safety and efficiency are paramount in healthcare facilities where the lives of patients are at stake. Despite the adoption of robots to assist medical staff in challenging tasks such as complex surgeries, human expertise is still indispensable. The next generation of autonomous healthcare robots hi

Cited by 9SourceScholar
2022

Amodal Panoptic Segmentation

CVPR 2022poster

Humans have the remarkable ability to perceive objects as a whole, even when parts of them are occluded. This ability of amodal perception forms the basis of our perceptual and cognitive understanding of our world. To enable robots to reason with this capability, we formulate and propose a novel tas…

Cited by 48PDFScholar
2022

Panoptic Nuscenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking

RA-L 2022

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing these tasks using LiDAR point clouds provides reliable pred

Cited by 243SourceScholar
2019

Robot Localization in Floor Plans Using a Room Layout Edge Extraction Network

IROS 2019poster

Indoor localization is one of the crucial enablers for deployment of service robots. Although several successful techniques for indoor localization have been proposed, the majority of them relies on maps generated from data gathered with the same sensor modality used for localization. Typically, ted…

Cited by 56SourceScholar