← Search

Debidatta Dwibedi

14 accepted papers

2026

Long-Context Robot Imitation Learning by Focusing on Key History Frames

RSS 2026poster

Many useful robot tasks require attending to the history of past observations. For example, finding an item in a room requires remembering which places have already been searched. However, the best-performing robot policies typically condition only on the current observation, limiting their applicab…

Cited by 0SourceScholar
2025

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

CoRL 2025poster

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of…

Cited by 0SourceScholar
2024

FlexCap: Describe Anything in Images in Controllable Detail

NeurIPS 2024poster

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input boxes, enabling control over information density, with descriptions ranging from concise object labels to detailed caption…

2024

RT-H: Action Hierarchies using Language

RSS 2024poster

Language provides a way to break down complex concepts into digestible pieces. Recent works in robot imitation learning have proposed learning language-conditioned policies that predict actions given visual observations and the high-level task specified in language. These methods leverage the struct…

2024

RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

ICRA 2024poster

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down step-by-step collection. We collect realistic data by performing any u…

Cited by 67SourceScholar
2024

Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

RSS 2024poster

Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the robot must understand a person's intent and perform the inferred task despite differences in the embodiments and environme…

2023

Visuomotor Control in Multi-Object Scenes Using Object-Aware Representations

ICRA 2023poster

Perceptual understanding of the scene and the relationship between its different components is important for successful completion of robotic tasks. Representation learning has been shown to be a powerful technique for this, but most of the current methodologies learn task specific representations t…

Cited by 20SourceScholar
2021

With a Little Help From My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

ICCV 2021poster

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other ins…

Cited by 561PDFScholar
2021

XIRL: Cross-embodiment Inverse Reinforcement Learning

CoRL 2021oral

We investigate the visual cross-embodiment imitation setting, in which agents learn policies from videos of other agents (such as humans) demonstrating the same task, but with stark differences in their embodiments -- shape, actions, end-effector dynamics, etc. In this work, we demonstrate that it i…

Cited by 134SourcecodeScholar
2020

Counting Out Time: Class Agnostic Video Repetition Counting in the Wild

CVPR 2020poster

We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temporal self-similarity as an intermediate representation bottleneck that allows generalization to unseen repetitions in vi…

Cited by 159PDFcodeScholar
2019

Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning

ICLR 2019poster

We identify two issues with the family of algorithms based on the Adversarial Imitation Learning framework. The first problem is implicit bias present in the reward functions used in these algorithms. While these biases might work well for some environments, they can also lead to sub-optimal behavio…

Cited by 348SourcePDFScholar
2018

Learning Actionable Representations from Visual Observations

IROS 2018poster

In this work we explore a new approach for robots to teach themselves about the world simply by observing it. In particular we investigate the effectiveness of learning task-agnostic representations for continuous control tasks. We extend Time-Contrastive Networks (TCN) that learn from visual observ…

Cited by 101SourceScholar