← Search

Deepak Pathak

89 accepted papers

2026

3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning

ICML 2026poster

We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint p…

Cited by 0SourceScholar
2026

Beyond Scalar Rewards: Learning from Text Feedback in LLM Post-Training

ICML 2026poster

The success of RL for LLM post-training stems from an unreasonably uninformative source: a single bit of information per rollout as binary reward or preference label. At the other extreme, distillation offers dense supervision but requires demonstrations, which are costly and difficult to scale. We …

Cited by 0SourceScholar
2026

Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling

ICLR 2026oral

We introduce Latent Particle World Model (LPWM), a self-supervised object-centric world model scaled to real-world multi-object datasets and applicable in decision-making. LPWM autonomously discovers keypoints, bounding boxes, and object masks directly from video data, enabling it to learn rich scen…

Cited by 0SourcecodeScholar
2026

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

ICML 2026poster

We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet question–answer (QA) pairs—a major bottleneck going forward, since such data is limited in scale and concentrated mainly in …

Cited by 0SourceScholar
2026

YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Prediction

CVPR 2026

Crop yield prediction requires substantial data to train scalable models. However, creating yield prediction datasets is constrained by high acquisition costs, heterogeneous data quality, and data privacy regulations. Consequently, existing datasets are scarce, low in quality, or limited to regional

Cited by 0SourcecodeScholar
2025

Deep Reactive Policy: Learning Reactive Manipulator Motion Planning for Dynamic Environments

CoRL 2025poster

Generating collision-free motion in dynamic, partially observable environments is a fundamental challenge for robotic manipulators. Classical motion planners can compute globally optimal trajectories but require full environment knowledge and are typically too slow for dynamic scenes. Neural motion…

Cited by 0SourceScholar
2025

Demonstrating LEAP Hand V3: Low-Cost, Easy-to-Assemble, High-Performance Hand for Robot Learning

RSS 2025poster

Replicating human-like dexterity in robotic hands has been a long-standing challenge in robotics. Recently, with the rise of robot learning and humanoids, the demand for dexterous robot hands to be reliable, affordable, and easy to reproduce has grown significantly. To address these needs, we presen…

Cited by 0PDFScholar
2025

DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies

RSS 2025poster

Many believe that large-scale datasets for robotics could be a key enabler of dexterous robotic policies that can generalize across diverse environments. While teleoperation provides high-fidelity datasets, its high cost limits its scalability. Instead, what if people could use their own hands, just…

Cited by 1PDFScholar
2025

Diffusion Beats Autoregressive in Data-Constrained Settings

NeurIPS 2025poster

Autoregressive (AR) models have long dominated the landscape of large language models, driving progress across a wide range of tasks. Recently, diffusion-based language models have emerged as a promising alternative, though their advantages over AR models remain underexplored. In this paper, we syst…

Cited by 0SourcecodeScholar
2025

FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learning

RSS 2025poster

Many contact-rich tasks humans perform, such as box pickup or hammering, rely on force feedback for reliable execution. However, this force information, which is readily available in most robot arms, is not commonly used in teleoperation and policy learning. Consequently, robot behavior is often lim…

Cited by 1PDFScholar
2025

Local Policies Enable Zero-Shot Long-Horizon Manipulation

ICRA 2025

Sim2real for robotic manipulation is difficult due to the challenges of simulating complex contacts and generating realistic task distributions. To tackle the latter problem, we introduce ManipGen, which leverages a new class of policies for sim2real transfer: local policies. Locality enables a vari

Cited by 31SourcecodeScholar
2025

Neural MP: A Neural Motion Planner

IROS 2025

The current paradigm for motion planning generates solutions from scratch for every new problem, which consumes significant amounts of time and computational resources. For complex, cluttered scenes, motion planning approaches can often take minutes to produce a solution, while humans are able to ac

Cited by 0SourcecodeScholar
2024

Bimanual Dexterity for Complex Tasks

CoRL 2024poster

To train generalist robot policies, machine learning methods often require a substantial amount of expert human teleoperation data. An ideal robot for humans collecting data is one that closely mimics them: bimanual arms and dexterous hands. However, creating such a bimanual teleoperation system wit…

Cited by 21SourcecodeScholar
2024

Continuously Improving Mobile Manipulation with Autonomous Real-World RL

CoRL 2024poster

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards object interactions and prevents stagnation near goal states, 2…

Cited by 3SourcecodeScholar
2024

Demonstrating Learning from Humans on Open-Source Dexterous Robot Hands

RSS 2024poster

Emulating human-like dexterity with robotic hands has been a long-standing challenge in robotics. In recent years, machine learning has demanded robot hands to be reliable, inexpensive and easy-to-reproduce. For the past few years we have been investigating how to address these demands. We will demo…

Cited by 0SourcePDFScholar
2024

Evaluating Text-to-Visual Generation with Image-to-Text Generation

ECCV 2024poster

"Despite significant progress in generative AI, comprehensive evaluation remains challenging because of the lack of effective metrics and standardized benchmarks. For instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce relia…

2024

Language Models as Black-Box Optimizers for Vision-Language Models

CVPR 2024poster

Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However many VLMs rely on proprietary data and are not open-source which restricts the use of white-box approaches for fine-tuning. As such…

2024

Legolas: Deep Leg-Inertial Odometry

CoRL 2024poster

Estimating odometry, where an accumulating position and rotation is tracked, has critical applications in many areas of robotics as a form of state estimation such as in SLAM, navigation, and controls. During deployment of a legged robot, a vision system's tracking can easily get lost. Instead, usin…

Cited by 1SourcecodeScholar
2024

On the Surprising Effectiveness of Attention Transfer for Vision Transformers

NeurIPS 2024poster

Conventional wisdom suggests that pre-training Vision Transformers (ViT) improves downstream performance by learning useful representations. Is this actually true? We investigate this question and find that the features and representations learned during pre-training are not essential. Surprisingly…

2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Revisiting the Role of Language Priors in Vision-Language Models

ICML 2024poster

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are trained for next-word generation given an image. We explore their zero-shot perform…

2024

SPIN: Simultaneous Perception Interaction and Navigation

CVPR 2024poster

While there has been remarkable progress recently in the fields of manipulation and locomotion mobile manipulation remains a long-standing challenge. Compared to locomotion or static manipulation a mobile system must make a diverse range of long-horizon tasks feasible in unstructured and dynamic env…

2023

Affordances From Human Videos as a Versatile Representation for Robotics

CVPR 2023poster

Building a robot that can understand and learn to interact by watching humans has inspired several vision problems. However, despite some successful results on static datasets, it remains unclear how current models can be used on a robot directly. In this paper, we aim to bridge this gap by leveragi…

Cited by 165SourcePDFScholar
2023

DEFT: Dexterous Fine-Tuning for Hand Policies

CoRL 2023poster

Dexterity is often seen as a cornerstone of complex manipulation. Humans are able to perform a host of skills with their hands, from making food to operating tools. In this paper, we investigate these challenges, especially in the case of soft, deformable objects as well as complex, relatively long…

Cited by 0SourcecodeScholar
2023

Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback

NeurIPS 2023poster

The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can be great test-time adapters for discriminative models. Our me…

2023

Efficient RL via Disentangled Environment and Agent Representations

ICML 2023oral

Agents that are aware of the separation between the environments and themselves can leverage this understanding to form effective representations of visual input. We propose an approach for learning such structured representations for RL algorithms, using visual knowledge of the agent, which is ofte…

Cited by 6SourcePDFScholar
2023

Internet Explorer: Targeted Representation Learning on the Open Web

ICML 2023poster

Vision models typically rely on fine-tuning general-purpose models pre-trained on large, static datasets. These general-purpose models only capture the knowledge within their pre-training datasets, which are tiny, out-of-date snapshots of the Internet---where billions of images are uploaded each day…

2023

LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning

RSS 2023poster

Dexterous manipulation has been a long-standing challenge in robotics. While machine learning techniques have shown some promise, results have largely been currently limited to simulation. This can be mostly attributed to the lack of suitable hardware. In this paper, we present LEAP Hand, a low-cost…

2023

Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning With Multimodal Models

CVPR 2023poster

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples may not be sufficient to characterize an entire concept class…

2023

PlayFusion: Skill Acquisition via Diffusion from Language-Annotated Play

CoRL 2023poster

Learning from unstructured and uncurated data has become the dominant paradigm for generative approaches in language or vision. Such unstructured and unguided behavior data, commonly known as play, is also easier to collect in robotics but much more difficult to learn from due to its inherently mult…

Cited by 53SourcecodeScholar
2023

Test-time Adaptation with Slot-Centric Models

ICML 2023poster

Current visual detectors, though impressive within their training distribution, often fail to parse out-of-distribution scenes into their constituent entities. Recent test-time adaptation methods use auxiliary self-supervised losses to adapt the network parameters to each test example independently…

2023

Your Diffusion Model is Secretly a Zero-Shot Classifier

ICCV 2023poster

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus…

Cited by 267PDFcodeScholar
2022

Adapting Rapid Motor Adaptation for Bipedal Robots

IROS 2022poster

Recent advances in legged locomotion have en-abled quadrupeds to walk on challenging terrains. However, bipedal robots are inherently more unstable and hence it's harder to design walking controllers for them. In this work, we leverage recent advances in rapid adaptation for locomotion control, and…

Cited by 59SourcecodeScholar
2022

Continual Learning with Evolving Class Ontologies

NeurIPS 2022accept

Lifelong learners must recognize concept vocabularies that evolve over time. A common yet underexplored scenario is learning with class labels that continually refine/expand old classes. For example, humans learn to recognize ${\tt dog}$ before dog breeds. In practical settings, dataset ${\it versio…

Cited by 12SourcePDFScholar
2022

Coupling Vision and Proprioception for Navigation of Legged Robots

CVPR 2022poster

We exploit the complementary strengths of vision and proprioception to develop a point-goal navigation system for legged robots, called VP-Nav. Legged systems are capable of traversing more complex terrain than wheeled robots, but to fully utilize this capability, we need a high-level path planner i…

Cited by 73PDFcodeScholar
2022

Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion

CoRL 2022oral

An attached arm can significantly increase the applicability of legged robots to several mobile manipulation tasks that are not possible for the wheeled or tracked counterparts. The standard modular control pipeline for such legged manipulators is to decouple the controller into that of manipulation…

Cited by 165SourcecodeScholar
2022

HERD: Continuous Human-to-Robot Evolution for Learning from Human Demonstration

CoRL 2022poster

The ability to learn from human demonstration endows robots with the ability to automate various tasks. However, directly learning from human demonstration is challenging since the structure of the human hand can be very different from the desired robot gripper. In this work, we show that manipulati…

Cited by 9SourceScholar
2022

Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

ICML 2022spotlight

Can world knowledge learned by large language models (LLMs) be used to act in interactive environments? In this paper, we investigate the possibility of grounding high-level tasks, expressed in natural language (e.g. “make breakfast”), to a chosen set of actionable steps (e.g. “open fridge”). While…

2022

Legged Locomotion in Challenging Terrains using Egocentric Vision

CoRL 2022oral

Animals are capable of precise and agile locomotion using vision. Replicating this ability has been a long-standing goal in robotics. The traditional approach has been to decompose this problem into elevation mapping and foothold planning phases. The elevation mapping, however, is susceptible to fai…

Cited by 242SourcecodeScholar
2022

REvolveR: Continuous Evolutionary Models for Robot-to-robot Policy Transfer

ICML 2022oral

A popular paradigm in robotic learning is to train a policy from scratch for every new robot. This is not only inefficient but also often impractical for complex robots. In this work, we consider the problem of transferring a policy across two different robots with significantly different parameters…

2022

Understanding Collapse in Non-Contrastive Siamese Representation Learning

ECCV 2022poster

"Contrastive methods have led a recent surge in the performance of self-supervised representation learning (SSL). Recent methods like BYOL or SimSiam purportedly distill these contrastive methods down to their essence, removing bells and whistles, including the negative examples, that do not contrib…

2021

Accelerating Robotic Reinforcement Learning via Parameterized Action Primitives

NeurIPS 2021poster

Despite the potential of reinforcement learning (RL) for building general-purpose robotic systems, training RL agents to solve robotics tasks still remains challenging due to the difficulty of exploration in purely continuous action spaces. Addressing this problem is an active area of research with…

Cited by 115SourcePDFScholar
2021

Discovering and Achieving Goals via World Models

NeurIPS 2021poster

How can artificial agents learn to solve many diverse tasks in complex visual environments without any supervision? We decompose this question into two challenges: discovering new goals and learning to reliably achieve them. Our proposed agent, Latent Explorer Achiever (LEXA), addresses both challen…

2021

Functional Regularization for Reinforcement Learning via Learned Fourier Features

NeurIPS 2021poster

We propose a simple architecture for deep reinforcement learning by embedding inputs into a learned Fourier basis and show that it improves the sample efficiency of both state-based and image-based RL. We perform infinite-width analysis of our architecture using the Neural Tangent Kernel and theoret…

2021

Interesting Object, Curious Agent: Learning Task-Agnostic Exploration

NeurIPS 2021oral

Common approaches for task-agnostic exploration learn tabula-rasa --the agent assumes isolated environments and no prior knowledge or experience. However, in the real world, agents learn in many environments and always come with prior experiences as they explore new ones. Exploration is a lifelong p…

2021

Learning Long-term Visual Dynamics with Region Proposal Interaction Networks

ICLR 2021poster

Learning long-term dynamics models is the key to understanding physical common sense. Most existing approaches on learning dynamics from visual input sidestep long-term predictions by resorting to rapid re-planning with short-term models. This not only requires such models to be super accurate but a…

2021

Minimizing Energy Consumption Leads to the Emergence of Gaits in Legged Robots

CoRL 2021poster

Legged locomotion is commonly studied and expressed as a discrete set of gait patterns, like walk, trot, gallop, which are usually treated as given and pre-programmed in legged robots for efficient locomotion at different speeds. However, fixing a set of pre-programmed gaits limits the generality of…

Cited by 138SourcecodeScholar
2021

Planning in Learned Latent Action Spaces for Generalizable Legged Locomotion

RA-L 2021

Hierarchical learning has been successful at learning generalizable locomotion skills on walking robots in a sample-efficient manner. However, the low-dimensional “latent” action used to communicate between two layers of the hierarchy is typically user-designed. In this letter, we present a fully-le

Cited by 33SourceScholar
2021

RB2: Robotic Manipulation Benchmarking with a Twist

NeurIPS 2021poster

Benchmarks offer a scientific way to compare algorithms using objective performance metrics. Good benchmarks have two features: (a) they should be widely useful for many research groups; (b) and they should produce reproducible findings. In robotic manipulation research, there is a trade-off between…

Cited by 25SourceScholar
2021

The CLEAR Benchmark: Continual LEArning on Real-World Imagery

NeurIPS 2021poster

Continual learning (CL) is widely regarded as crucial challenge for lifelong AI. However, existing CL benchmarks, e.g. Permuted-MNIST and Split-CIFAR, make use of artificial temporal variation and do not align with or generalize to the real- world. In this paper, we introduce CLEAR, the first contin…

Cited by 112SourcecodeScholar
2021

Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis From a Single Image

ICCV 2021poster

We present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input image, consistent with the learned intermediate depth, captures underlying geometry sufficient to generate photorealisti…

Cited by 85PDFcodeScholar
2020

Neural Dynamic Policies for End-to-End Sensorimotor Learning

NeurIPS 2020spotlight

The current dominant paradigm in sensorimotor control, whether imitation or reinforcement learning, is to train policies directly in raw action spaces such as torque, joint angle, or end-effector position. This forces the agent to make decision at each point in training, and hence, limits the scalab…

2020

One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic Control

ICML 2020poster

Reinforcement learning is typically concerned with learning control policies tailored to a particular agent. We investigate whether there exists a single global policy that can generalize to control a wide variety of agent morphologies – ones in which even dimensionality of state and action spaces c…

2020

Planning to Explore via Self-Supervised World Models

ICML 2020poster

Reinforcement learning allows solving complex tasks, however, the learning tends to be task-specific and the sample efficiency remains a challenge. We present Plan2Explore, a self-supervised reinforcement learning agent that tackles both these challenges through a new approach to self-supervised exp…

2020

Sparse Graphical Memory for Robust Planning

NeurIPS 2020poster

To operate effectively in the real world, agents should be able to act from high-dimensional raw sensory input such as images and achieve diverse goals across long time-horizons. Current deep reinforcement and imitation learning methods can learn directly from high-dimensional inputs but do not sca…

2019

Large-Scale Study of Curiosity-Driven Learning

ICLR 2019poster

Reinforcement learning algorithms rely on carefully engineered rewards from the environment that are extrinsic to the agent. However, annotating each environment with hand-designed, dense rewards is difficult and not scalable, motivating the need for developing reward functions that are intrinsic to…

2019

Learning to Control Self-Assembling Morphologies: A Study of Generalization via Modularity

NeurIPS 2019spotlight

Contemporary sensorimotor learning approaches typically start with an existing complex agent (e.g., a robotic arm), which they learn to control. In contrast, this paper investigates a modular co-evolution strategy: a collection of primitive agents learns to dynamically self-assemble into composite b…

2019

Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller

NeurIPS 2019poster

We study a generalized setup for learning from demonstration to build an agent that can manipulate novel objects in unseen scenarios by looking at only a single video of human demonstration from a third-person perspective. To accomplish this goal, our agent should not only learn to understand the in…

2018

Investigating Human Priors for Playing Video Games

ICLR 2018workshop

What makes humans so good at solving seemingly complex video games? Unlike computers, humans bring in a great deal of prior knowledge about the world, enabling efficient decision making. This paper investigates the role of human priors for solving video games. Given a sample game, we conduct a seri…

Cited by 210SourceScholar
2018

Investigating Human Priors for Playing Video Games

ICML 2018oral

What makes humans so good at solving seemingly complex video games? Unlike computers, humans bring in a great deal of prior knowledge about the world, enabling efficient decision making. This paper investigates the role of human priors for solving video games. Given a sample game, we conduct a serie…

2018

Zero-Shot Visual Imitation

ICLR 2018oral

The current dominant paradigm for imitation learning relies on strong supervision of expert actions to learn both 'what' and 'how' to imitate. We pursue an alternative paradigm wherein an agent first explores the world without any expert supervision and then distills its experience into a goal-condi…

2017

Curiosity-driven Exploration by Self-supervised Prediction

ICML 2017poster

In many real-world scenarios, rewards extrinsic to the agent are extremely sparse, or absent altogether. In such cases, curiosity can serve as an intrinsic reward signal to enable the agent to explore its environment and learn skills that might be useful later in its life. We formulate curiosity as…

2017

Learning Features by Watching Objects Move

CVPR 2017poster

This paper presents a novel yet intuitive approach to unsupervised feature learning. Inspired by the human visual system, we explore whether low-level motion-based grouping cues can be used to learn an effective visual representation. Specifically, we use unsupervised motion-based segmentation on vi…

Cited by 640PDFcodeScholar
2017

Toward Multimodal Image-to-Image Translation

NeurIPS 2017poster

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a distribution of possible outputs in a conditional generative modeling setting. The ambiguity of the mapping is distilled in a low-dimensional l…

2016

Context Encoders: Feature Learning by Inpainting

CVPR 2016poster

We present an unsupervised visual feature learning algorithm driven by context-based pixel prediction. By analogy with auto-encoders, we propose Context Encoders -- a convolutional neural network trained to generate the contents of an arbitrary image region conditioned on its surroundings. In order…

Cited by 7142PDFcodeScholar
2015

Constrained Convolutional Neural Networks for Weakly Supervised Segmentation

ICCV 2015poster

We present an approach to learn a dense pixel-wise labeling from image-level tags. Each image-level tag imposes constraints on the output labeling of a Convolutional Neural Network (CNN) classifier. We propose Constrained CNN (CCNN), a method which uses a novel loss function to optimize for any set…

Cited by 791PDFcodeScholar
2015

Detector Discovery in the Wild: Joint Multiple Instance and Representation Learning

CVPR 2015poster

We develop methods for detector learning which exploit joint training over both weak and strong labels and which transfer learned perceptual representations from strongly-labeled auxiliary tasks. Previous methods for weak-label learning often learn detector models independently using latent variable…

Cited by 98SourcePDFScholar