← Search

Sayan Deb Sarkar

7 accepted papers

2025

CrossOver: 3D Scene Cross-Modal Alignment

CVPR 2025highlight

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment.…

2025

GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer

NeurIPS 2025poster

Transferring appearance to 3D assets using different representations of the appearance object - such as images or text - has garnered interest due to its wide range of applications in industries like gaming, augmented reality, and digital content creation. However, state-of-the-art methods still fai…

Cited by 0SourcecodeScholar
2023

HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language

ACL 2023findings

This paper presents “HaVQA”, the first multimodal dataset for visual question answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, t…

2023

SGAligner: 3D Scene Alignment with Scene Graphs

ICCV 2023poster

Building 3D scene graphs has recently emerged as a topic in scene representation for several embodied AI applications to represent the world in a structured and rich manner. With their increased use in solving downstream tasks (e.g., navigation and room rearrangement), can we leverage and recycle th…

Cited by 14PDFcodeScholar
2022

Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose Estimation

CVPR 2022oral

We propose a robust and accurate method for estimating the 3D poses of two hands in close interaction from a single color image. This is a very challenging problem, as large occlusions and many confusions between the joints may happen. State-of-the-art methods solve this problem by regressing a heat…

Cited by 174PDFcodeScholar
2021

Monte Carlo Scene Search for 3D Scene Understanding

CVPR 2021poster

We explore how a general AI algorithm can be used for 3D scene understanding to reduce the need for training data. More exactly, we propose a modification of the Monte Carlo Tree Search (MCTS) algorithm to retrieve objects and room layouts from noisy RGB-D scans. While MCTS was developed as a game-p…

Cited by 29PDFcodeScholar
2020

General 3D Room Layout from a Single View by Render-and-Compare

ECCV 2020poster

We present a novel method to reconstruct the 3D layout of a room—walls, floors, ceilings—from a single perspective view in challenging conditions, by contrast with previous single-view methods restricted to cuboid-shaped layouts. This input view can consist of a color image only, but considering a de…