← Search

Jitendra Malik

164 accepted papers

2026

D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping

ICLR 2026poster

Simulation provides a cost-effective and flexible platform for data generation and policy learning to develop robotic systems. However, bridging the gap between simulation and real-world dynamics remains a significant challenge, especially in physical parameter identification. In this work, we intro…

Cited by 0SourcecodeScholar
2026

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations

ICRA 2026poster

Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on lab…

2026

DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos

ICML 2026spotlight

Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels.…

Cited by 81SourceScholar
2026

Learning Dexterous Manipulation Skills from Imperfect Simulations

ICRA 2026poster

Reinforcement learning and sim-to-real transfer have made significant progress in dexterous manipulation. However, progress remains limited by the difficulty of simulating complex contact dynamics and multisensory signals, especially tactile feedback. In this work, we propose DexScrew, a sim-to-real…

2026

Learning to Grasp Anything By Playing with Random Toys

ICLR 2026poster

Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop generalizable dexterous manipulation skills by mastering a small set of simple toys and then applying that knowledge to more c…

Cited by 0SourceScholar
2026

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

ICRA 2026poster

Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available in research labs. Can we leverage them to train bimanual robot policies? We pre…

2026

Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning

ICML 2026poster

Diffusion models are increasingly used as powerful conditional generators, yet real deployments often involve multiple target distributions arising from different tasks, e.g., diverse prompt domains in text-to-image generation, or multiple environments in robotics with diffusion policies. This natur…

Cited by 0SourceScholar
2026

Rodrigues Network for Learning Robot Actions

ICLR 2026oral

Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the underlying kinematic structure of articulated systems. To this end, we propose the **Neural Rodrigues Operator**, a lear…

Cited by 0SourceScholar
2026

SAM 3D Body: Robust Full-Body Human Mesh Recovery

CVPR 2026

We introduce SAM 3D Body (3DB), a promptable model for single-image full-body 3D human mesh recovery (HMR) that demonstrates state-of-the-art performance, with strong generalization and consistent accuracy in diverse in-the-wild conditions. 3DB estimates the human pose of the body, feet, and hands.

Cited by 0SourcecodeScholar
2026

TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment

RSS 2026poster

Human demonstrations collected by wearable devices (e.g., tactile gloves) provide fast and dexterous supervision for policy learning, and are guided by rich, natural tactile feedback. However, a key challenge is how to transfer human-collected tactile signals to robots despite the differences in sen…

Cited by 0SourceScholar
2026

Tracking by Predicting 3-D Gaussians Over Time

CVPR 2026

We propose Video Gaussian Masked Autoencoders (Video-GMAE), a self-supervised approach for representation learning that encodes a sequence of images into a set of Gaussian splats moving over time. Representing a video as a set of Gaussians enforces a reasonable inductive bias: that 2-D videos are of

Cited by 0SourcecodeScholar
2026

ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

RSS 2026poster

Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner. Although vision-based methods have advanced rapidly, tactile sensing remains crucial for fine-grained control—particularly in unstructured or visually occluded se…

Cited by 0SourceScholar
2025

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time

EMNLP 2025

This paper presents AlphaOne ( 𝛼1 ), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. 𝛼1 first introduces 𝛼 moment, which represents the scaled thinking phase with a universal parameter 𝛼 .Within this scaled pre- 𝛼 moment phase, it dynamically sc

2025

An Empirical Study of Autoregressive Pre-training from Videos

ICCV 2025poster

We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a…

Cited by 0SourcePDFScholar
2025

AutoEval Done Right: Using Synthetic Data for Model Evaluation

ICML 2025poster

The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data can be used to decrease the number of human annotations required for this purpose in a process called autoevaluation. We suggest efficient and statistically pr…

2025

DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy

NeurIPS 2025spotlight

Garment manipulation is a critical challenge due to the diversity in garment categories, geometries, and deformations. Despite this, humans can effortlessly handle garments, thanks to the dexterity of our hands. However, existing research in the field has struggled to replicate this level of dexteri…

Cited by 0SourcecodeScholar
2025

DexterityGen: Foundation Controller for Unprecedented Dexterity

RSS 2025poster

Teaching robots dexterous manipulation skills, such as tool use, presents a significant challenge. Current approaches can be broadly categorized into two strategies: human teleoperation (for imitation learning) and sim-to-real reinforcement learning. The first approach is difficult as it is hard fo…

Cited by 9PDFScholar
2025

Estimating Body and Hand Motion in an Ego-sensed World

CVPR 2025highlight

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand parameters that capture a device wearer's actions in the allocentric…

Cited by 5SourcePDFScholar
2025

From Simple to Complex Skills: The Case of In-Hand Object Reorientation

ICRA 2025

Learning policies in simulation and transferring them to the real world has become a promising approach in dexterous manipulation. However, bridging the sim-to-real gap for each new task requires substantial human effort, such as careful reward engineering, hyperparameter tuning, and system identifi

Cited by 16SourceScholar
2025

Hand-Object Interaction Pretraining from Videos

ICRA 2025

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and reta

Cited by 46SourcecodeScholar
2025

Learning In-Hand Translation Using Tactile Skin with Shear and Normal Force Sensing

ICRA 2025

Recent progress in reinforcement learning (RL) and tactile sensing has significantly advanced dexterous manipulation. However, these methods often utilize simplified tactile signals due to the gap between tactile simulation and the real world. We introduce a sensor model for tactile skin that enable

Cited by 23SourceScholar
2025

Learning Visuotactile Skills With Two Multifingered Hands

ICRA 2025

Aiming to replicate human-like dexterity, perceptual experiences, and motion patterns, we explore learning from human demonstrations using a bimanual system with multifingered hands and visuotactile data. Two significant challenges exist: the lack of an affordable and accessible teleoperation system

Cited by 119SourcecodeScholar
2025

OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

ICML 2025poster

Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained vision-language models (VLMs) as visual and language features are independently fed into downstream policies, degrading the p…

2025

Poly-Autoregressive Prediction for Modeling Interactions

CVPR 2025poster

We introduce a simple framework for predicting the behavior of an agent in multi-agent settings. In contrast to autoregressive (AR) tasks, such as language processing, our focus is on scenarios with multiple agents whose interactions are shaped by physical constraints and internal motivations. To th…

Cited by 0SourcePDFScholar
2025

Reconstructing People, Places, and Cameras

CVPR 2025highlight

We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibrated multi-view images featuring people. Our approach combines data-driven scene…

2025

RoboVerse: A Unified Platform, Benchmark and Dataset for Scalable and Generalizable Robot Learning

RSS 2025poster

Data scaling and standardized evaluation benchmarks have driven remarkable advances in natural language processing and computer vision. However, in robotics, scaling up data and establishing evaluation protocols pose significant challenges. Directly collecting real-world data is inefficient and reso…

Cited by 0PDFScholar
2025

Scaling Properties of Diffusion Models For Perceptual Tasks

CVPR 2025poster

In this paper, we argue that iterative computation with diffusion models offers a powerful paradigm for not only generation but also visual perception tasks. We unify tasks such as depth estimation, optical flow, and amodal segmentation under the framework of image-to-image translation, and show how…

Cited by 4SourcePDFScholar
2025

Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids

CoRL 2025poster

Learning generalizable robot manipulation policies, especially for complex multi-fingered humanoids, remains a significant challenge. Existing approaches primarily rely on extensive data collection and imitation learning, which are expensive, labor-intensive, and difficult to scale. Sim-to-real rein…

Cited by 0SourceScholar
2025

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

CoRL 2025oral

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from high-fidelity simulators, other modalities (e.g. sound) can be notorio…

Cited by 0SourceScholar
2025

Visual Imitation Enables Contextual Humanoid Control

CoRL 2025oral

How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably the simplest way is to _just show them_—casually capture a human motion video and feed it to humanoids. We introduce **VideoMimic**, a real-to-sim-to-real pipeline that mines everyday…

Cited by 0SourceScholar
2025

Whole-Body Conditioned Egocentric Video Prediction

NeurIPS 2025poster

We train models to predict ego-centric video from human actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions…

Cited by 0SourceScholar
2024

Adaptive Human Trajectory Prediction via Latent Corridors

ECCV 2024poster

"Human trajectory prediction is typically posed as a zero-shot generalization problem: a predictor is learnt on a dataset of human motion in training scenes, and then deployed on unseen test scenes. While this paradigm has yielded tremendous progress, it fundamentally assumes that trends in human be…

Cited by 4SourcePDFScholar
2024

Conformal Decision Theory: Safe Autonomous Decisions from Imperfect Predictions

ICRA 2024poster

We introduce Conformal Decision Theory, a framework for producing safe autonomous decisions despite imperfect machine learning predictions. Examples of such decisions are ubiquitous, from robot planning algorithms that rely on pedestrian predictions, to calibrating autonomous manufacturing to exhibi…

Cited by 33SourceScholar
2024

Conformal Policy Learning for Sensorimotor Control under Distribution Shifts

ICRA 2024poster

This paper focuses on the problem of detecting and reacting to changes in the distribution of a sensorimotor controller’s observables. The key idea is the design of policies that can take conformal quantiles as input, to detect distribution shifts with formal statistical guarantees, which we define…

Cited by 5SourceScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning

CVPR 2024poster

Large pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning which is highly memory-intensive for tasks with high-resolution data e.g. video understanding small object detection and point cloud analysis…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

GOAT: GO to Any Thing

RSS 2024poster

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system cap…

2024

Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots

ICLR 2024poster

We present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and mot…

Cited by 111SourcePDFScholar
2024

Humanoid Locomotion as Next Token Prediction

NeurIPS 2024spotlight

We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction of sensorimotor sequences. To account for the multi-modal nature of the data, we perform prediction in a modal…

Cited by 55SourcePDFScholar
2024

Learning Vision-based Pursuit-Evasion Robot Policies

ICRA 2024poster

Learning strategic robot behavior—like that required in pursuit-evasion interactions—under real-world constraints is extremely challenging. It requires exploiting the dynamics of the interaction, and planning through both physical state and latent intent uncertainty. In this paper, we transform this…

Cited by 15SourceScholar
2024

Lessons from Learning to Spin “Pens”

CoRL 2024poster

In-hand manipulation of pen-like objects is a most basic and important skill in our daily lives, as many tools such as hammers and screwdrivers are similarly shaped. However, current learning-based methods struggle with this task due to a lack of high-quality demonstrations and the significant gap b…

Cited by 16SourcecodeScholar
2024

Manipulator as a Tail: Promoting Dynamic Stability for Legged Locomotion

ICRA 2024poster

For locomotion, is an arm on a legged robot a liability or an asset for locomotion? Biological systems evolved additional limbs beyond legs that facilitates postural control. This work shows how a manipulator can be an asset for legged locomotion at high speeds or under external perturbations, where…

Cited by 5SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Re-evaluating the Need for Visual Signals in Unsupervised Grammar Induction

NAACL 2024findings

Are multimodal inputs necessary for grammar induction? Recent work has shown that multimodal training inputs can improve grammar induction. However, these improvements are based on comparisons to weak text-only baselines that were trained on relatively little textual data. To determine whether multi…

2024

Reconstructing Hands in 3D with Transformers

CVPR 2024poster

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery HaMeR follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and robustness compared to previous work. The key to HaMeR's success lies…

2024

Sequential Modeling Enables Scalable Learning for Large Vision Models

CVPR 2024poster

We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this we define a common format "visual sentences" in which we can represent raw images and videos as well as annotated data sources such as semantic se…

2024

What Matters to You? Towards Visual Representation Alignment for Robot Learning

ICLR 2024poster

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs, their rewards will inevitably use visual representations. Recently there has been excitement in using representations from pre-trained visual model…

Cited by 8SourcePDFScholar
2024

xT: Nested Tokenization for Larger Context in Large Images

ICML 2024poster

Modern computer vision pipelines handle large images in one of two sub-optimal ways: down-sampling or cropping. These two methods incur significant losses in the amount of information and context present in an image. There are many downstream applications in which global context matters as much as h…

2023

Decoupling Human and Camera Motion From Videos in the Wild

CVPR 2023poster

We propose a method to reconstruct global human trajectories from videos in the wild. Our optimization method decouples the camera and human motion, which allows us to place people in the same world coordinate frame. Most existing methods do not model the camera motion; methods that rely on the back…

2023

General In-hand Object Rotation with Vision and Touch

CoRL 2023poster

We introduce Rotateit, a system that enables fingertip-based object rotation along multiple axes by leveraging multimodal sensory inputs. Our system is trained in simulation, where it has access to ground-truth object shapes and physical properties. Then we distill it to operate on realistic yet noi…

Cited by 106SourceScholar
2023

Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

ICML 2023oral

Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effective accuracies and attractive FLOP counts, the added complexity actually makes these transformers slower than their vani…

2023

Humans in 4D: Reconstructing and Tracking Humans with Transformers

ICCV 2023poster

We present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances the state of the art and shows the capability to analyze unusual poses that have in t…

Cited by 221PDFcodeScholar
2023

Learning a Single Near-hover Position Controller for Vastly Different Quadcopters

ICRA 2023poster

This paper proposes an adaptive near-hover position controller for quadcopters, which can be deployed to quadcopters of very different mass, size and motor constants, and also shows rapid adaptation to unknown disturbances during runtime. The core algorithmic idea is to learn a single policy that ca…

Cited by 23SourceScholar
2023

MAViL: Masked Audio-Video Learners

NeurIPS 2023poster

We present Masked Audio-Video Learners (MAViL) to learn audio-visual representations with three complementary forms of self-supervision: (1) reconstructing masked raw audio and video inputs, (2) intra-modal and inter-modal contrastive learning with masking, and (3) self-training to predict aligned a…

2023

Multi-skill Mobile Manipulation for Object Rearrangement

ICLR 2023top-25%

We study a modular approach to tackle long-horizon mobile manipulation tasks for object rearrangement, which decomposes a full task into a sequence of subtasks. To tackle the entire task, prior work chains multiple stationary manipulation skills with a point-goal navigation skill, which are learned…

2023

Multiview Compressive Coding for 3D Reconstruction

CVPR 2023poster

A central goal of visual recognition is to understand objects and scenes from a single image. 2D recognition has witnessed tremendous progress thanks to large-scale learning and general-purpose representations. But, 3D poses new challenges stemming from occlusions not depicted in the image. Prior wo…

2023

Navigating to Objects Specified by Images

ICCV 2023poster

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modula…

Cited by 41PDFScholar
2023

On the Benefits of 3D Pose and Tracking for Human Action Recognition

CVPR 2023poster

In this work we study the benefits of using tracking and 3D poses for action recognition. To achieve this, we take the Lagrangian view on analysing actions over a trajectory of human motion rather than at a fixed point in space. Taking this stand allows us to use the tracklets of people to predict t…

2023

Robot Learning with Sensorimotor Pre-training

CoRL 2023oral

We present a self-supervised sensorimotor pre-training approach for robotics. Our model, called RPT, is a Transformer that operates on sequences of sensorimotor tokens. Given a sequence of camera images, proprioceptive robot states, and actions, we encode the sequence into tokens, mask out a subset,…

Cited by 54SourceScholar
2023

Speculative Decoding with Big Little Decoder

NeurIPS 2023poster

The recent emergence of Large Language Models based on the Transformer architecture has enabled dramatic advancements in the field of Natural Language Processing. However, these models have long inference latency, which limits their deployment and makes them prohibitively expensive for various real-…

2023

Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?

NeurIPS 2023poster

We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we syste…

Cited by 161SourcePDFScholar
2022

ABO: Dataset and Benchmarks for Real-World 3D Object Understanding

CVPR 2022poster

We introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with complex geometries and physically-based materials that correspond to real, househ…

Cited by 225PDFcodeScholar
2022

Adapting Rapid Motor Adaptation for Bipedal Robots

IROS 2022poster

Recent advances in legged locomotion have en-abled quadrupeds to walk on challenging terrains. However, bipedal robots are inherently more unstable and hence it's harder to design walking controllers for them. In this work, we leverage recent advances in rapid adaptation for locomotion control, and…

Cited by 59SourcecodeScholar
2022

Coupling Vision and Proprioception for Navigation of Legged Robots

CVPR 2022poster

We exploit the complementary strengths of vision and proprioception to develop a point-goal navigation system for legged robots, called VP-Nav. Legged systems are capable of traversing more complex terrain than wheeled robots, but to fully utilize this capability, we need a high-level path planner i…

Cited by 73PDFcodeScholar
2022

Differentiable Stereopsis: Meshes From Multiple Views Using Differentiable Rendering

CVPR 2022poster

We propose Differentiable Stereopsis, a multi-view stereo approach that reconstructs shape and texture from few input views and noisy cameras. We pair traditional stereopsis and modern differentiable rendering to build an end-to-end model which predicts textured 3D meshes of objects with varying top…

Cited by 33PDFcodeScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Image-to-Image Regression with Distribution-Free Uncertainty Quantification and Applications in Imaging

ICML 2022spotlight

Image-to-image regression is an important learning task, used frequently in biological imaging. Current algorithms, however, do not generally offer statistical guarantees that protect against a model’s mistakes and hallucinations. To address this, we develop uncertainty quantification techniques wit…

2022

In-Hand Object Rotation via Rapid Motor Adaptation

CoRL 2022poster

Generalized in-hand manipulation has long been an unsolved challenge of robotics. As a small step towards this grand goal, we demonstrate how to design and learn a simple adaptive controller to achieve in-hand object rotation using only fingertips. The controller is trained entirely in simulation on…

Cited by 115SourcecodeScholar
2022

Legged Locomotion in Challenging Terrains using Egocentric Vision

CoRL 2022oral

Animals are capable of precise and agile locomotion using vision. Replicating this ability has been a long-standing goal in robotics. The traditional approach has been to decompose this problem into elevation mapping and foothold planning phases. The elevation mapping, however, is susceptible to fai…

Cited by 242SourcecodeScholar
2022

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

CVPR 2022poster

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved version of MViT that incorporates decomposed relative positional embeddings and residual pooling connections. We instantia…

Cited by 950PDFcodeScholar
2022

MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition

CVPR 2022oral

While today's video recognition systems parse snapshots or short clips accurately, they cannot connect the dots and reason across a longer range of time yet. Most existing video architectures can only process <5 seconds of a video without hitting the computation or memory bottlenecks. In this paper,…

Cited by 261PDFcodeScholar
2022

Open-World Instance Segmentation: Exploiting Pseudo Ground Truth From Learned Pairwise Affinity

CVPR 2022poster

Open-world instance segmentation is the task of grouping pixels into object instances without any pre-determined taxonomy. This is challenging, as state-of-the-art methods rely on explicit class semantics obtained from large labeled datasets, and out-of-domain evaluation performance drops significan…

Cited by 54PDFcodeScholar
2022

PONI: Potential Functions for ObjectGoal Navigation With Interaction-Free Learning

CVPR 2022oral

State-of-the-art approaches to ObjectGoal navigation (ObjectNav) rely on reinforcement learning and typically require significant computational resources and time for learning. We propose Potential functions for ObjectGoal Navigation with Interaction-free learning (PONI), a modular approach that dis…

Cited by 178PDFcodeScholar
2022

Real-World Robot Learning with Masked Visual Pre-training

CoRL 2022oral

In this work, we explore self-supervised visual pre-training on images from diverse, in-the-wild videos for real-world robotic tasks. Like prior work, our visual representations are pre-trained via a masked autoencoder (MAE), frozen, and then passed into a learnable control module. Unlike prior work…

Cited by 272SourceScholar
2022

Reversible Vision Transformers

CVPR 2022oral

We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory footprint from the depth of the model, Reversible Vision Transformers enable memory efficient scaling of transformer architectures. We adapt two popular models, name…

Cited by 68PDFcodeScholar
2022

Squeezeformer: An Efficient Transformer for Automatic Speech Recognition

NeurIPS 2022accept

The recently proposed Conformer model has become the de facto backbone model for various downstream speech tasks based on its hybrid attention-convolution architecture that captures both local and global features. However, through a series of systematic studies, we find that the Conformer architectu…

2022

Tracking People by Predicting 3D Appearance, Location and Pose

CVPR 2022oral

We present an approach for tracking people in monocular videos by predicting their future 3D representations. To achieve this, we first lift people to 3D from a single frame in a robust manner. This lifting includes information about the 3D pose of the person, their location in the 3D space, and the…

Cited by 74PDFcodeScholar
2021

Active 3D Shape Reconstruction from Vision and Touch

NeurIPS 2021poster

Humans build 3D understandings of the world through active object exploration, using jointly their senses of vision and touch. However, in 3D shape reconstruction, most recent progress has relied on static datasets of limited sensory data such as RGB images, depth maps or haptic readings, leaving th…

2021

From Goals, Waypoints & Paths to Long Term Human Trajectory Forecasting

ICCV 2021poster

Human trajectory forecasting is an inherently multimodal problem. Uncertainty in future trajectories stems from two sources: (a) sources that are known to the agent but unknown to the model, such as long term goals and (b) sources that are unknown to both the agent & the model, such as intent of oth…

Cited by 343PDFcodeScholar
2021

Habitat 2.0: Training Home Assistants to Rearrange their Habitat

NeurIPS 2021spotlight

We introduce Habitat 2.0 (H2.0), a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios. We make comprehensive contributions to all levels of the embodied AI stack – data, simulation, and benchmark tasks. Specifically, we present: (i) R…

2021

Learning Long-term Visual Dynamics with Region Proposal Interaction Networks

ICLR 2021poster

Learning long-term dynamics models is the key to understanding physical common sense. Most existing approaches on learning dynamics from visual input sidestep long-term predictions by resorting to rapid re-planning with short-term models. This not only requires such models to be super accurate but a…

2021

Minimizing Energy Consumption Leads to the Emergence of Gaits in Legged Robots

CoRL 2021poster

Legged locomotion is commonly studied and expressed as a discrete set of gait patterns, like walk, trot, gallop, which are usually treated as given and pre-programmed in legged robots for efficient locomotion at different speeds. However, fixing a set of pre-programmed gaits limits the generality of…

Cited by 138SourcecodeScholar
2021

Multiscale Vision Transformers

ICCV 2021poster

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multiscale Transformers have several channel-resolution scale stages. Starting from the input resolution and a small channel dim…

Cited by 1664PDFcodeScholar
2021

Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets From 3D Scans

ICCV 2021poster

Computer vision now relies on data, but we know surprisingly little about what factors in the data affect performance. We argue that this stems from the way data is collected. Designing and collecting static datasets of images (or videos) locks us in to specific design choices and limits us to post-…

Cited by 299PDFcodeScholar
2021

SEAL: Self-supervised Embodied Active Learning using Exploration and 3D Consistency

NeurIPS 2021poster

In this paper, we explore how we can build upon the data and models of Internet images and use them to adapt to robot vision without requiring any extra labels. We present a framework called Self-supervised Embodied Active Learning (SEAL). It utilizes perception models trained on internet images to…

Cited by 96SourcePDFScholar
2021

State-Only Imitation Learning for Dexterous Manipulation

IROS 2021poster

Modern model-free reinforcement learning methods have recently demonstrated impressive results on a number of problems. However, complex domains like dexterous manipulation remain a challenge due to the high sample complexity. To address this, current approaches employ expert demonstrations in the f…

Cited by 135SourceScholar
2021

Tracking People with 3D Representations

NeurIPS 2021poster

We present a novel approach for tracking multiple people in video. Unlike past approaches which employ 2D representations, we focus on using 3D representations of people, located in three-dimensional space. To this end, we develop a method, Human Mesh and Appearance Recovery (HMAR) which in addition…

2021

Uncertainty Sets for Image Classifiers using Conformal Prediction

ICLR 2021spotlight

Convolutional image classifiers can achieve high predictive accuracy, but quantifying their uncertainty remains an unresolved challenge, hindering their deployment in consequential settings. Existing uncertainty quantification techniques, such as Platt scaling, attempt to calibrate the network’s pr…

Cited by 414SourcePDFScholar
2020

3D Shape Reconstruction from Vision and Touch

NeurIPS 2020poster

When a toddler is presented a new toy, their instinctual behaviour is to pick it up and inspect it with their hand and eyes in tandem, clearly searching over its surface to properly understand what they are playing with. At any instance here, touch provides high fidelity localized information while…

2020

Deep Isometric Learning for Visual Recognition

ICML 2020poster

Initialization, normalization, and skip connections are believed to be three indispensable techniques for training very deep convolutional neural networks and obtaining state-of-the-art performance. This paper shows that deep vanilla ConvNets without normalization nor skip connections can also be tr…

2020

It is not the Journey but the Destination: Endpoint Conditioned Trajectory Prediction

ECCV 2020poster

Human trajectory forecasting with multiple socially interact-ing agents is of critical importance for autonomous navigation in human environments, e.g., for self-driving cars and social robots. In this work, we present Predicted Endpoint Conditioned Network (PECNet) for flexible human trajectory pre…

2020

Long-term Human Motion Prediction with Scene Context

ECCV 2020poster

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off. Existing works on predicting human motion do not pay attention to…

2020

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

ECCV 2020poster

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insigh…

2020

Robust Learning Through Cross-Task Consistency

CVPR 2020oral

Visual perception entails solving a wide set of tasks (e.g., object detection, depth estimation, etc). The predictions made for different tasks out of one image are not independent, and therefore, are expected to be 'consistent'. We propose a flexible and fully computational framework for learning w…

Cited by 187PDFcodeScholar
2020

Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation

CoRL 2020

Vision-based robotics often factors the control loop into separate components for perception and control. Conventional perception components usually extract hand-engineered features from the visual input that are then used by the control component in an explicit manner. In contrast, recent advances

Cited by 0SourcePDFScholar
2020

Side-Tuning: A Baseline for Network Adaptation via Additive Side Networks

ECCV 2020poster

When training a neural network for a desired task, one may prefer to adapt a pre-trained network rather than starting from randomly initialized weights. Adaptation can be useful in cases when training data is scarce, when a single learner needs to perform multiple tasks, or when one wishes to encode…

Cited by 257SourcePDFScholar
2020

Which Tasks Should Be Learned Together in Multi-task Learning?

ICML 2020poster

Many computer vision applications require solving multiple tasks in real-time. A neural network can be trained to solve multiple tasks simultaneously using multi-task learning. This can save computation at inference time as only a single network needs to be evaluated. Unfortunately, this often leads…

2019

3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera

ICCV 2019poster

A comprehensive semantic understanding of a scene is important for many applications - but in what space should diverse semantic information (e.g., objects, scene categories, material types, 3D shapes, etc.) be grounded and what should be its structure? Aspiring to have one unified structure that ho…

Cited by 413PDFcodeScholar
2019

Combining Optimal Control and Learning for Visual Navigation in Novel Environments

CoRL 2019

Model-based control is a popular paradigm for robot navigation because it can leverage a known dynamics model to efficiently plan robust robot trajectories. However, it is challenging to use model-based methods in settings where the environment is a priori unknown and can only be observed partially

Cited by 0SourcePDFScholar
2019

Habitat: A Platform for Embodied AI Research

ICCV 2019oral

We present Habitat, a platform for research in embodied artificial intelligence (AI). Habitat enables training embodied agents (virtual robots) in highly efficient photorealistic 3D simulation. Specifically, Habitat consists of: (i) Habitat-Sim: a flexible, high-performance 3D simulator with configu…

Cited by 2011PDFcodeScholar
2019

Learning Independent Object Motion From Unlabelled Stereoscopic Videos

CVPR 2019poster

We present a system for learning motion maps of independently moving objects from stereo videos. The only annotations used in our system are 2D object bounding boxes which introduce the notion of objects in our system. Unlike prior learning based approaches which have focused on predicting dense opt…

Cited by 37PDFScholar
2019

Learning Individual Styles of Conversational Gesture

CVPR 2019poster

Human speech is often accompanied by hand and arm gestures. We present a method for cross-modal translation from "in-the-wild" monologue speech of a single speaker to their conversational gesture motion. We train on unlabeled videos for which we only have noisy pseudo ground truth from an automatic…

Cited by 403PDFcodeScholar
2019

Learning to Navigate Using Mid-Level Visual Priors

CoRL 2019

How much does having visual priors about the world (e.g. the fact that the world is 3D) assist in learning to perform downstream motor tasks (e.g. navigating a complex environment)? What are the consequences of not utilizing such visual priors in learning? We study these questions by integrating a g

2019

Mesh R-CNN

ICCV 2019poster

Rapid advances in 2D perception have led to systems that accurately detect objects in real-world images. However, these systems make predictions in 2D, ignoring the 3D structure of the world. Concurrently, advances in 3D shape prediction have mostly focused on synthetic benchmarks and isolated objec…

Cited by 569PDFScholar
2019

ShapeMask: Learning to Segment Novel Objects by Refining Shape Priors

ICCV 2019oral

Instance segmentation aims to detect and segment individual objects in a scene. Most existing methods rely on precise mask annotations of every category. However, it is difficult and costly to segment objects in novel categories because a large number of mask annotations is required. We introduce Sh…

Cited by 156PDFcodeScholar
2018

AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions

CVPR 2018poster

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person o…

Cited by 1319SourcePDFScholar
2018

End-to-End Recovery of Human Shape and Pose

CVPR 2018poster

We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape…

2018

Factoring Shape, Pose, and Layout From the 2D Image of a 3D Scene

CVPR 2018poster

The goal of this paper is to take a single 2D image of a scene and recover the 3D structure in terms of a small set of factors: a layout representing the enclosing surfaces as well as a set of objects represented in terms of shape and pose. We propose a convolutional neural network-based approach to…

Cited by 157SourcePDFScholar
2018

Gibson Env: Real-World Perception for Embodied Agents

CVPR 2018poster

Perception and being active (having a certain level of motion freedom) are closely tied. Learning active perception and sensorimotor control in the physical world is cumbersome as existing algorithms are too slow to efficiently learn in real-time and robots are fragile and costly. This has given ris…

2018

Learning Category-Specific Mesh Reconstruction from Image Collections

ECCV 2018poster

We present a learning framework for recovering the 3D shape, camera, and texture of an object from a single image. The shape is represented as a deformable 3D mesh model of an object category where a shape is parameterized by a learned mean shape and per-instance predicted deformation. Our approach…

2018

More Than a Feeling: Learning to Grasp and Regrasp Using Vision and Touch

RA-L 2018

For humans, the process of grasping an object relies heavily on rich tactile feedback. Most recent robotic grasping work, however, has been based only on visual input, and thus cannot easily benefit from feedback after initiating contact. In this letter, we investigate how a robot can learn to use t

Cited by 396SourceScholar
2018

Multi-View Consistency as Supervisory Signal for Learning Shape and Pose Prediction

CVPR 2018poster

We present a framework for learning single-view shape and pose prediction without using direct supervision for either. Our approach allows leveraging multi-view observations from unknown poses as supervisory signal during training. Our proposed training setup enforces geometric consistency between t…

Cited by 232SourcePDFScholar
2018

Taskonomy: Disentangling Task Transfer Learning

CVPR 2018poster

Do visual tasks have a relationship, or are they unrelated? For instance, could having surface normals simplify estimating the depth of an image? Intuition answers these questions positively, implying existence of a structure among visual tasks. Knowing this structure has notable uses; it is the con…

2018

Visual Memory for Robust Path Following

NeurIPS 2018oral

Humans routinely retrace a path in a novel environment both forwards and backwards despite uncertainty in their motion. In this paper, we present an approach for doing so. Given a demonstration of a path, a first network generates an abstraction of the path. Equipped with this abstraction, a second…

Cited by 64SourcePDFScholar
2018

Zero-Shot Visual Imitation

ICLR 2018oral

The current dominant paradigm for imitation learning relies on strong supervision of expert actions to learn both 'what' and 'how' to imitate. We pursue an alternative paradigm wherein an agent first explores the world without any expert supervision and then distills its experience into a goal-condi…

2017

Cognitive Mapping and Planning for Visual Navigation

CVPR 2017poster

We introduce a neural architecture for navigation in novel environments. Our proposed architecture learns to map from first-person views and plans a sequence of actions towards goals in the environment. The Cognitive Mapper and Planner (CMP) is based on two key ideas: a) a unified joint architecture…

Cited by 876PDFScholar
2017

Combining self-supervised learning and imitation for vision-based rope manipulation

ICRA 2017poster

Manipulation of deformable objects, such as ropes and cloth, is an important but challenging problem in robotics. We present a learning-based system where a robot takes as input a sequence of images of a human manipulating a rope from an initial to goal configuration, and outputs a sequence of actio…

Cited by 367SourceScholar
2017

Learning Shape Abstractions by Assembling Volumetric Primitives

CVPR 2017poster

We present a learning framework for abstracting complex shapes by learning to assemble objects using 3D volumetric primitives. In addition to generating simple and geometrically interpretable explanations of 3D objects, our framework also allows us to automatically discover and exploit consistent st…

Cited by 402PDFcodeScholar
2017

Learning to Optimize

ICLR 2017poster

Algorithm design is a laborious process and often requires many iterations of ideation and validation. In this paper, we explore automating algorithm design and present a method to learn an optimization algorithm. We approach this problem from a reinforcement learning perspective and represent any p…

Cited by 69SourceScholar
2017

Multi-View Supervision for Single-View Reconstruction via Differentiable Ray Consistency

CVPR 2017oral

We study the notion of consistency between a 3D shape and a 2D observation and propose a differentiable formulation which allows computing gradients of the 3D shape given an observation from an arbitrary view. We do so by reformulating view consistency using a differentiable ray consistency (DRC) te…

Cited by 646PDFScholar
2016

Human Pose Estimation With Iterative Error Feedback

CVPR 2016spotlight

Hierarchical feature extractors such as Convolutional Networks (ConvNets) have achieved impressive performance on a variety of classification tasks using purely feedforward processing. Feedforward architectures can learn rich representations of the input space but do not explicitly model dependencie…

Cited by 1074PDFcodeScholar
2016

Learning to Poke by Poking: Experiential Learning of Intuitive Physics

NeurIPS 2016oral

We investigate an experiential learning paradigm for acquiring an internal model of intuitive physics. Our model is evaluated on a real-world robotic manipulation task that requires displacing objects to target locations by poking. The robot gathered over 400 hours of experience by executing more th…

Cited by 641SourcePDFScholar
2015

Aligning 3D Models to RGB-D Images of Cluttered Scenes

CVPR 2015poster

The goal of this work is to represent objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instances in the scene and then using a convolutional neural network (CNN) to predict the pose of the object. This CNN is tra…

Cited by 320SourcePDFScholar
2015

Category-Specific Object Reconstruction From a Single Image

CVPR 2015poster

Object reconstruction from a single image -- in the wild -- is a problem where we can make progress and get meaningful results today. This is the main message of this paper, which introduces an automated pipeline with pixels as inputs and 3D surfaces of various rigid categories as outputs in images…

Cited by 421SourcePDFScholar
2015

Deformable Part Models are Convolutional Neural Networks

CVPR 2015poster

Deformable part models (DPMs) and convolutional neural networks (CNNs) are two widely used tools for visual recognition. They are typically viewed as distinct approaches: DPMs are graphical models (Markov random fields), while CNNs are "black-box" non-linear classifiers. In this paper, we show tha…

Cited by 600SourcePDFScholar
2015

Depth From Shading, Defocus, and Correspondence Using Light-Field Angular Coherence

CVPR 2015poster

Light-field cameras are now used in consumer and industrial applications. Recent papers and products have demonstrated practical depth recovery algorithms from a passive single-shot capture. However, current light field capture devices have narrow baselines and constrained spatial resolution; theref…

Cited by 263SourcePDFScholar
2015

Hypercolumns for Object Segmentation and Fine-Grained Localization

CVPR 2015poster

Recognition algorithms based on convolutional networks (CNNs) typically use the output of the last layer as feature representation. However, the information in this layer may be too coarse to allow precise localization. On the contrary, earlier layers may be precise in localization but will not capt…

Cited by 2008SourcePDFScholar