← Search

Narendra Ahuja

26 accepted papers

2026

Finding Distributed Object-Centric Properties in Self-Supervised Transformers

CVPR 2026

Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in \texttt [CLS] token attention maps of the final layer. However, these maps often contain spurious activations resulting in poor localization of objects. This is because the \textt

Cited by 0SourceScholar
2026

RigMo: Unifying Rig and Motion Learning for Generative Animation

CVPR 2026

Despite significant progress in 4D generation, rig and motion--the core structural and dynamic components of animation--are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent pr

Cited by 0SourceScholar
2025

PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling

ICCV 2025poster

Skinning and rigging are fundamental components in animation, articulated object reconstruction, motion transfer, and 4D generation. Existing approaches predominantly rely on Linear Blend Skinning (LBS), due to its simplicity and differentiability. However, LBS introduces artifacts such as volume lo…

Cited by 0SourcePDFScholar
2025

Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation

NeurIPS 2025spotlight

We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts --- structural components aligned with…

Cited by 0SourceScholar
2024

CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen

AAAI 2024technical

Addressing Out-Of-Distribution (OOD) Segmentation and Zero-Shot Semantic Segmentation (ZS3) is challenging, necessitating segmenting unseen classes. Existing strategies adapt the class-agnostic Mask2Former (CA-M2F) tailored to specific tasks. However, these methods cater to singular tasks, demand tr…

Cited by 13SourcePDFScholar
2024

Learning Implicit Representation for Reconstructing Articulated Objects

ICLR 2024poster

3D Reconstruction of moving articulated objects without additional information about object structure is a challenging problem. Current methods overcome such challenges by employing category-specific skeletal models. Consequently, they do not generalize well to articulated objects in the wild. We tr…

2024

S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video

ICML 2024poster

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational resources and training time, and require additional human annotat…

2023

Long-Distance Gesture Recognition Using Dynamic Neural Networks

IROS 2023poster

Gestures form an important medium of communication between humans and machines. An overwhelming majority of existing gesture recognition methods are tailored to a scenario where humans and machines are located very close to each other. This short-distance assumption does not hold true for several ty…

Cited by 2SourceScholar
2022

Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio

ICASSP 2022accepted

The distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker…

Cited by 3SourceScholar
2022

Learning Audio-Visual Dynamics Using Scene Graphs for Audio Source Separation

NeurIPS 2022accept

There exists an unequivocal distinction between the sound produced by a static source and that produced by a moving one, especially when the source moves towards or away from the microphone. In this paper, we propose to use this connection between audio and visual dynamics for solving two challengin…

Cited by 12SourcePDFScholar
2021

A Hierarchical Variational Neural Uncertainty Model for Stochastic Video Prediction

ICCV 2021poster

Predicting the future frames of a video is a challenging task, in part due to the underlying stochastic real-world phenomena. Prior approaches to solve this task typically estimate a latent prior characterizing this stochasticity, however do not account for the predictive uncertainty of the (deep le…

Cited by 18PDFScholar
2021

Visual Scene Graphs for Audio Source Separation

ICCV 2021poster

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sound sources or avoid modeling object interactions that may be useful to better ch…

Cited by 42PDFcodeScholar
2018

DeepMVS: Learning Multi-View Stereopsis

CVPR 2018poster

We present DeepMVS, a deep convolutional neural network (ConvNet) for multi-view stereo reconstruction. Taking an arbitrary number of posed images as input, we first produce a set of plane-sweep volumes and use the proposed DeepMVS network to predict high-quality disparity maps. The key contribution…

2017

Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution

CVPR 2017poster

Convolutional neural networks have recently demonstrated high-quality reconstruction for single-image super-resolution. In this paper, we propose the Laplacian Pyramid Super-Resolution Network (LapSRN) to progressively reconstruct the sub-band residuals of high-resolution images. At each pyramid lev…

Cited by 3356PDFScholar
2017

Robust Visual Tracking Using Oblique Random Forests

CVPR 2017poster

Random forest has emerged as a powerful classification technique with promising results in various vision tasks including image classification, pose estimation and object detection. However, current techniques have shown little improvements in visual tracking as they mostly rely on piece wise orthog…

Cited by 99PDFcodeScholar
2016

A Comparative Study for Single Image Blind Deblurring

CVPR 2016spotlight

Numerous single image blind deblurring algorithms have been proposed to restore latent sharp images under camera motion. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real blurred images. It is thus unclear how these algorithms would perform on images…

Cited by 524PDFScholar
2015

Structural Sparse Tracking

CVPR 2015poster

Sparse representation has been applied to visual tracking by finding the best target candidate with minimal reconstruction error by use of target templates. However, most sparse representation based trackers only consider holistic or local representations and do not make full use of the intrinsic st…

Cited by 216SourcePDFScholar
2015

Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations From Surveillance Videos

ICCV 2015poster

We present a novel approach for jointly estimating tar- gets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by…

Cited by 85PDFScholar