← Search

Rahul Sukthankar

21 accepted papers

2026

Efficiently Reconstructing Dynamic Scenes One D4RT at a Time

CVPR 2026

Understanding and reconstructing the complex geometry and motion of dynamic 4D scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward network designed to efficiently solve this task. D4RT utilizes a unified transformer archi

Cited by 0SourceScholar
2022

Discrete Representations Strengthen Vision Transformer Robustness

ICLR 2022poster

Vision Transformer (ViT) is emerging as the state-of-the-art architecture for image recognition. While recent studies suggest that ViTs are more robust than their convolutional counterparts, our experiments find that ViTs are overly reliant on local features (\eg, nuisances and texture) and fail to…

Cited by 54SourcePDFScholar
2022

UFO Depth: Unsupervised learning with flow-based odometry optimization for metric depth estimation

ICRA 2022poster

We propose an efficient method for unsupervised learning of metric depth estimation from a single image in the context of unconstrained videos captured from UAVs. We combine the accuracy of an analytical solution based on odometry with the power of deep learning. First, we show how to correct the no…

Cited by 6SourceScholar
2021

Neural Descent for Visual 3D Human Pose and Shape

CVPR 2021poster

We present deep neural network methodology to reconstruct the 3d pose and shape of people, including hand gestures and facial expression, given an input RGB image. We rely on a recently introduced, expressive full body statistical 3d human model, GHUM, trained end-to-end, and learn to reconstruct it…

Cited by 75PDFScholar
2021

Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus

AAAI 2021technical

We address the challenging problem of semi-supervised learning in the context of multiple visual interpretations of the world by finding consensus in a graph of neural networks. Each graph node is a scene interpretation layer, while each edge is a deep net that transforms one layer at one node into…

Cited by 11SourcePDFScholar
2021

THUNDR: Transformer-Based 3D Human Reconstruction With Markers

ICCV 2021poster

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and t…

Cited by 83PDFScholar
2020

GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models

CVPR 2020oral

We present a statistical, articulated 3D human shape modeling pipeline, within a fully trainable, modular, deep learning framework. Given high-resolution complete 3D body scans of humans, captured in various poses, together with additional closeups of their head and facial expressions, as well as ha…

Cited by 423PDFcodeScholar
2020

SelfieDroneStick: A Natural Interface for Quadcopter Photography

IROS 2020poster

A physical selfie stick extends the user's reach, enabling the acquisition of personal photos that include more of the background scene. Similarly, a quadcopter can capture photos from vantage points unattainable by the user; but teleoperating a quadcopter to good viewpoints is a difficult task. Thi…

Cited by 1SourcecodeScholar
2020

Speech2Action: Cross-Modal Supervision for Action Recognition

CVPR 2020poster

Is it possible to guess human action from dialogue alone? In this work we investigate the link between spoken words and actions in movies. We note that movie screenplays describe actions, as well as contain the speech of characters and hence can be used to learn this correlation with no additional s…

Cited by 78PDFScholar
2020

Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows

ECCV 2020poster

Monocular 3D human pose and shape estimation is challenging due to the many degrees of freedom of the human body and the difficulty to acquire training data for large-scale supervised learning in complex visual scenes where humans with diverse shape and appearance, appear against complex backgrounds…

Cited by 160SourcePDFScholar
2019

Relational Action Forecasting

CVPR 2019oral

This paper focuses on multi-person action forecasting in videos. More precisely, given a history of H previous frames, the goal is to detect actors and to predict their future actions for the next T frames. Our approach jointly models temporal and spatial interactions among different actors by const…

Cited by 100PDFScholar
2018

AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions

CVPR 2018poster

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person o…

Cited by 1319SourcePDFScholar
2018

Actor-centric Relation Network

ECCV 2018poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level and model temporal context with 3D ConvNets. Here, we go one step further and model spatio-temporal relations to capture the interactions between human actors, relevant objects and scene…

Cited by 280SourcePDFScholar
2018

Rethinking the Faster R-CNN Architecture for Temporal Action Localization

CVPR 2018poster

We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive field alignment using a multi-scale architecture that can accom…

Cited by 846SourcePDFScholar
2017

Cognitive Mapping and Planning for Visual Navigation

CVPR 2017poster

We introduce a neural architecture for navigation in novel environments. Our proposed architecture learns to map from first-person views and plans a sequence of actions towards goals in the environment. The Cognitive Mapper and Planner (CMP) is based on two key ideas: a) a unified joint architecture…

Cited by 876PDFScholar
2016

Discovering the Physical Parts of an Articulated Object Class From Multiple Videos

CVPR 2016poster

We propose a motion-based method to discover the physical parts of an articulated object class (e.g. head/torso/leg of a horse) from multiple videos. The key is to find object regions that exhibit consistent motion relative to the rest of the object, across multiple videos. We can then learn a locat…

Cited by 14PDFScholar
2015

Articulated Motion Discovery Using Pairs of Trajectories

CVPR 2015poster

We propose an unsupervised approach for discovering characteristic motion patterns in videos of highly articulated objects performing natural, unscripted behaviors, such as tigers in the wild. We discover consistent patterns in a bottom-up manner by analyzing the relative displacements of large numb…

Cited by 50SourcePDFScholar
2015

MatchNet: Unifying Feature and Metric Learning for Patch-Based Matching

CVPR 2015poster

Motivated by recent successes on learning feature representations and on learning feature comparison functions, we propose a unified approach to combining both for training a patch matching system. Our system, dubbed MatchNet, consists of a deep convolutional network that extracts features from pa…

2015

Robust Video Segment Proposals With Painless Occlusion Handling

CVPR 2015poster

We propose a robust algorithm to generate video segment proposals. The proposals generated by our method can start from any frame in the video and are robust to complete occlusions. Our method does not assume specific motion models and even has a limited capability to generalize across videos. We bu…

Cited by 35SourcePDFScholar