← Search

Moitreya Chatterjee

18 accepted papers

2026

AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

CVPR 2026

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industri

Cited by 0SourceScholar
2026

LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction

CVPR 2026

Recent feed-forward reconstruction models like VGGT and \pi^3 achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming methods address this through learned memory mechanisms or caus

Cited by 0SourcecodeScholar
2026

Point4Cast: Streaming Dynamic Scene Reconstruction and Forecasting

CVPR 2026

Understanding how the 3D world evolves over time is a fundamental task in computer vision, essential for embodied settings, autonomous driving, etc. It requires not only the reconstruction of the observed scene but also the anticipation of how the scene dynamics will unfold in the future. While the

Cited by 0SourceScholar
2026

Understanding Dynamic Compute Allocation in Recurrent Transformers

ICML 2026poster

Token-level adaptive computation seeks to reduce inference cost by allocating more computation to harder tokens and less to easier ones. However, prior work is primarily evaluated on natural-language benchmarks using task-level metrics, where token-level difficulty is unobservable and confounded wit…

Cited by 0SourceScholar
2025

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing

CVPR 2025poster

Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost o…

2024

A Probability-guided Sampler for Neural Implicit Surface Rendering

ECCV 2024poster

"Several variants of Neural Radiance Fields (NeRFs) have significantly improved the accuracy of synthesized images and surface reconstruction of 3D scenes/objects. In all of these methods, a key characteristic is that none can train the neural network with every possible input data, specifically, ev…

2024

CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments

AAAI 2024technical

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied Navigation framework in which the agent may interact with a human/o…

Cited by 5SourcePDFScholar
2024

Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling

CVPR 2024highlight

Extensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences two drawbacks limit their ubiquity: (i) a significant reduction in reconstruction…

Cited by 4SourcePDFScholar
2022

Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio

ICASSP 2022accepted

The distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker…

Cited by 0SourceScholar
2022

Learning Audio-Visual Dynamics Using Scene Graphs for Audio Source Separation

NeurIPS 2022accept

There exists an unequivocal distinction between the sound produced by a static source and that produced by a moving one, especially when the source moves towards or away from the microphone. In this paper, we propose to use this connection between audio and visual dynamics for solving two challengin…

Cited by 12SourcePDFScholar
2021

A Hierarchical Variational Neural Uncertainty Model for Stochastic Video Prediction

ICCV 2021poster

Predicting the future frames of a video is a challenging task, in part due to the underlying stochastic real-world phenomena. Prior approaches to solve this task typically estimate a latent prior characterizing this stochasticity, however do not account for the predictive uncertainty of the (deep le…

Cited by 18PDFScholar
2021

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

AAAI 2021technical

Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning…

Cited by 50SourcePDFScholar
2021

Visual Scene Graphs for Audio Source Separation

ICCV 2021poster

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sound sources or avoid modeling object interactions that may be useful to better ch…

Cited by 42PDFcodeScholar
2016

Deep Neural Networks with Inexact Matching for Person Re-Identification

NeurIPS 2016poster

Person Re-Identification is the task of matching images of a person across multiple camera views. Almost all prior approaches address this challenge by attempting to learn the possible transformations that relate the different views of a person from a training corpora. Then, they utilize these trans…

2015

Acoustic and para-verbal indicators of persuasiveness in social multimedia

ICASSP 2015accepted

Persuasive communication and interaction play an important and pervasive role in many aspects of our lives. With the rapid growth of social multimedia websites such as YouTube, it has become more important and useful to understand persuasiveness in the context of online social multimedia content. In…

Cited by 0SourceScholar