← Search

Suman Saha

8 accepted papers

2025

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

CVPR 2025poster

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth o…

2024

Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding

ECCV 2024poster

"3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task has recently risen in popularity. A common formulation to tackle 3D visual grou…

2023

EDAPS: Enhanced Domain-Adaptive Panoptic Segmentation

ICCV 2023poster

With autonomous industries on the rise, domain adaptation of the visual perception stack is an important research direction due to the cost savings promise. Much prior art was dedicated to domain-adaptive semantic segmentation in the synthetic-to-real context. Despite being a crucial output of the p…

Cited by 14PDFcodeScholar
2021

Learning To Relate Depth and Semantics for Unsupervised Domain Adaptation

CVPR 2021poster

We present an approach for encoding visual task relationships to improve model performance in an Unsupervised Domain Adaptation (UDA) setting. Semantic segmentation and monocular depth estimation are shown to be complementary tasks; in a multi-task learning setting, a proper encoding of their relati…

Cited by 70PDFcodeScholar
2021

Three Ways To Improve Semantic Segmentation With Self-Supervised Depth Estimation

CVPR 2021poster

Training deep networks for semantic segmentation requires large amounts of labeled training data, which presents a major challenge in practice, as labeling segmentation masks is a highly labor-intensive process. To address this issue, we present a framework for semi-supervised semantic segmentation,…

Cited by 115PDFcodeScholar
2020

Reparameterizing Convolutions for Incremental Multi-Task Learning without Task Interference

ECCV 2020poster

Multi-task networks are commonly utilized to alleviate the need for a large number of highly specialized single-task networks. However, two common challenges in developing multi-task models are often overlooked in literature. First, enabling the model to be inherently incremental, continuously incor…

2017

Online Real-Time Multiple Spatiotemporal Action Localisation and Prediction

ICCV 2017poster

We present a deep-learning framework for real-time multiple spatio-temporal (S/T) action localisation and classification. Current state-of-the-art approaches work offline, and are too slow to be useful in real-world settings. To overcome their limitations we introduce two major developments. Firstly…

Cited by 384PDFScholar