← Search

Lerrel Pinto

71 accepted papers

2026

Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models

RSS 2026poster

The prevalent paradigm in robot learning attempts to generalize across environments, embodiments, and tasks with language prompts at runtime. A fundamental tension limits this approach: language is often too abstract to guide the concrete physical understanding required for robust manipulation. In t…

Cited by 0SourceScholar
2026

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations

ICRA 2026poster

Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on lab…

2026

Quality Over Quantity: Demonstration Curation Via Influence Functions for Data-Centric Robot Learning

ICRA 2026poster

Learning from demonstrations has emerged as a promising paradigm for end-to-end robot control, particularly when scaled to diverse and large datasets. However, the quality of demonstration data, often collected through human teleoperation, remains a critical bottleneck for effective data-driven robo…

2025

AnySkin: Plug-and-Play Skin Sensing for Robotic Touch

ICRA 2025

While tactile sensing is widely accepted as an important and useful sensing modality, its use pales in comparison to other sensory modalities like vision and proprioception. AnySkin addresses the critical challenges that impede the use of tactile sensing - versatility, replaceability, and data reusa

Cited by 40SourcecodeScholar
2025

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

ICLR 2025poster

Large Language Models (LLMs) and Vision Language Models (VLMs) possess extensive knowledge and exhibit promising reasoning abilities, however, they still struggle to perform well in complex, dynamic environments. Real-world tasks require handling intricate interactions, advanced spatial reasoning, l…

Cited by 9SourcePDFScholar
2025

Bridging the Human to Robot Dexterity Gap Through Object-Oriented Rewards

ICRA 2025

Training robots directly from human videos is an emerging area in robotics and computer vision. While there has been notable progress with two-fingered grippers, learning autonomous tasks without teleoperation remains a difficult problem for multi-fingered robot hands. A key reason for this difficul

Cited by 26SourcecodeScholar
2025

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

ICML 2025poster

The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remain challenging to learn and are typically developed for task-specific solutions with online policy learning. To unlock world models' tru…

Cited by 17SourcePDFScholar
2025

Dynamem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

ICRA 2025

Significant progress has been made in openvocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment, which limits the system's applicability in realworld scenarios

Cited by 34SourcecodeScholar
2025

P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies

ICRA 2025

Developing generalizable robot policies that can robustly handle varied environmental conditions and object instances remains a fundamental challenge in robot learning. While considerable efforts have focused on collecting large robot datasets and developing policy architectures to learn from such d

Cited by 19SourcecodeScholar
2025

Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation

CoRL 2025poster

Building robotic agents capable of operating across diverse environments and object types remains a significant challenge, often requiring extensive data collection. This is particularly restrictive in robotics, where each data point must be physically executed in the real world. Consequently, there…

Cited by 0SourcecodeScholar
2025

Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

ICRA 2025

Robot models, particularly those trained with large amounts of data, have recently shown a plethora of real-world manipulation and navigation capabilities. Several independent efforts have shown that given sufficient training data in an environment, robot policies can generalize to demonstrated vari

Cited by 49SourcecodeScholar
2025

Training Language Models on Synthetic Edit Sequences Improves Code Synthesis

ICLR 2025poster

Software engineers mainly write code by editing existing programs. In contrast, language models (LMs) autoregressively synthesize programs in a single pass. One explanation for this is the scarcity of sequential edit data. While high-quality instruction data for code synthesis is scarce, edit data f…

2024

Adaptive Sampling of k-Space in Magnetic Resonance for Rapid Pathology Prediction

ICML 2024poster

Magnetic Resonance (MR) imaging, despite its proven diagnostic utility, remains an inaccessible imaging modality for disease surveillance at the population level. A major factor rendering MR inaccessible is lengthy scan times. An MR scanner collects measurements associated with the underlying anatom…

Cited by 2SourcePDFScholar
2024

BAKU: An Efficient Transformer for Multi-Task Policy Learning

NeurIPS 2024poster

Training generalist agents capable of solving diverse tasks is challenging, often requiring large datasets of expert demonstrations. This is particularly problematic in robotics, where each data point requires physical execution of actions in the real world. Thus, there is a pressing need for archit…

2024

Behavior Generation with Latent Actions

ICML 2024spotlight

Generative modeling of complex behaviors from labeled datasets has been a longstanding problem in decision-making. Unlike language or image generation, decision-making requires modeling actions – continuous-valued vectors that are multimodal in their distribution, potentially drawn from uncurated so…

2024

Demonstrating OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

RSS 2024poster

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile systems, and grasping models that can handle a wide range of o…

2024

DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control

NeurIPS 2024poster

Imitation learning has proven to be a powerful tool for training complex visuo-motor policies. However, current methods often require hundreds to thousands of expert demonstrations to handle high-dimensional visual observations. A key reason for this poor data efficiency is that visual representatio…

2024

Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling

ICML 2024poster

Reasoning from sequences of raw sensory data is a ubiquitous problem across fields ranging from medical devices to robotics. These problems often involve using long sequences of raw sensor data (e.g. magnetometers, piezoresistors) to predict sequences of desirable physical quantities (e.g. force, in…

2024

OPEN TEACH: A Versatile Teleoperation System for Robotic Manipulation

CoRL 2024poster

Open-sourced, user-friendly tools form the bedrock of scientific advancement across disciplines. The widespread adoption of data-driven learning has led to remarkable progress in multi-fingered dexterity, bimanual manipulation, and applications ranging from logistics to home robotics. However, exist…

Cited by 56SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

See to Touch: Learning Tactile Dexterity through Visual Incentives

ICRA 2024poster

Equipping multi-fingered robots with tactile sensing is crucial for achieving the precise, contact-rich, and dexterous manipulation that humans excel at. However, relying solely on tactile sensing fails to provide adequate cues for reasoning about objects’ spatial configurations, limiting the abilit…

Cited by 39SourcecodeScholar
2023

CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

RSS 2023poster

We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this…

2023

Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play

CoRL 2023poster

Teaching dexterity to multi-fingered robots has been a longstanding challenge in robotics. Most prominent work in this area focuses on learning controllers or policies that either operate on visual observations or state estimates derived from vision. However, such methods perform poorly on fine-grai…

Cited by 64SourceScholar
2023

Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation

ICRA 2023poster

Optimizing behaviors for dexterous manipulation has been a longstanding challenge in robotics, with a variety of methods from model-based control to model-free reinforcement learning having been previously explored in literature. Such prior work often require extensive trial-and-error training along…

Cited by 124SourcecodeScholar
2023

From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data

ICLR 2023top-5%

While large-scale sequence modelling from offline data has led to impressive performance gains in natural language generation and image generation, directly translating such ideas to robotics has been challenging. One critical reason for this is that uncurated robot demonstration data, i.e. play dat…

2023

Holo-Dex: Teaching Dexterity with Immersive Mixed Reality

ICRA 2023poster

A fundamental challenge in teaching robots is to provide an effective interface for human teachers to demonstrate useful skills to a robot. This challenge is exacerbated in dexterous manipulation, where teaching high-dimensional, contact-rich behaviors often require esoteric teleoperation tools. In…

Cited by 71SourcecodeScholar
2023

Improving Long-Horizon Imitation through Instruction Prediction

AAAI 2023technical

Complex, long-horizon planning and its combinatorial nature pose steep challenges for learning-based agents. Difficulties in such settings are exacerbated in low data regimes where over-fitting stifles generalization and compounding errors hurt accuracy. In this work, we explore the use of an often…

2023

Learning Simultaneous Navigation and Construction in Grid Worlds

ICLR 2023poster

We propose to study a new learning task, mobile construction, to enable an agent to build designed structures in 1/2/3D grid worlds while navigating in the same evolving environments. Unlike existing robot learning tasks such as visual navigation and object manipulation, this task is challenging bec…

2023

Teach a Robot to FISH: Versatile Imitation from One Minute of Demonstrations

RSS 2023poster

While imitation learning provides us with an efficient toolkit to train robots, learning skills that are robust to environment variations remains a significant challenge. Current approaches address this challenge by relying either on large amounts of demonstrations that span environment variations o…

2023

Train Offline, Test Online: A Real Robot Learning Benchmark

ICRA 2023poster

Three challenges limit the progress of robot learning research: robots are expensive (few labs can participate), everyone uses different robots (findings do not generalize across labs), and we lack internet-scale robotics data. We take on these challenges via a new benchmark: Train Offline, Test Onl…

Cited by 20SourcecodeScholar
2022

Behavior Transformers: Cloning $k$ modes with one stone

NeurIPS 2022accept

While behavior learning has made impressive progress in recent times, it lags behind computer vision and natural language processing due to its inability to leverage large, human-generated datasets. Human behavior has a wide variance, multiple modes, and human demonstrations naturally do not come wi…

2022

Context is Everything: Implicit Identification for Dynamics Adaptation

ICRA 2022poster

Understanding environment dynamics is necessary for robots to act safely and optimally in the world. In realistic scenarios, dynamics are non-stationary and the causal variables such as environment parameters cannot necessarily be precisely measured or inferred, even during training. We propose Impl…

Cited by 25SourcecodeScholar
2022

Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation

IROS 2022poster

Recent advances in unsupervised representation learning significantly improved the sample efficiency of training Reinforcement Learning policies in simulated environments. However, similar gains have not yet been seen for real-robot reinforcement learning. In this work, we focus on enabling data-eff…

Cited by 0SourceScholar
2022

Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

ICLR 2022poster

We present DrQ-v2, a model-free reinforcement learning (RL) algorithm for visual continuous control. DrQ-v2 builds on DrQ, an off-policy actor-critic approach that uses data augmentation to learn directly from pixels. We introduce several improvements that yield state-of-the-art results on the DeepM…

2022

One After Another: Learning Incremental Skills for a Changing World

ICLR 2022poster

Reward-free, unsupervised discovery of skills is an attractive alternative to the bottleneck of hand-designing rewards in environments where task supervision is scarce or expensive. However, current skill pre-training methods, like many RL techniques, make a fundamental assumption -- stationary envi…

2022

Watch and Match: Supercharging Imitation with Regularized Optimal Transport

CoRL 2022oral

Imitation learning holds tremendous promise in learning policies efficiently for complex decision making problems. Current state-of-the-art algorithms often use inverse reinforcement learning (IRL), where given a set of expert demonstrations, an agent alternatively infers a reward function and the a…

Cited by 76SourceScholar
2021

Learning Cross-Domain Correspondence for Control with Dynamics Cycle-Consistency

ICLR 2021oral

At the heart of many robotics problems is the challenge of learning correspondences across domains. For instance, imitation learning requires obtaining correspondence between humans and robots; sim-to-real requires correspondence between physics simulators and real hardware; transfer learning requir…

Cited by 73SourcePDFScholar
2021

RB2: Robotic Manipulation Benchmarking with a Twist

NeurIPS 2021poster

Benchmarks offer a scientific way to compare algorithms using objective performance metrics. Good benchmarks have two features: (a) they should be widely useful for many research groups; (b) and they should produce reproducible findings. In robotic manipulation research, there is a trade-off between…

Cited by 25SourceScholar
2021

Reinforcement Learning with Prototypical Representations

ICML 2021spotlight

Learning effective representations in image-based environments is crucial for sample efficient Reinforcement Learning (RL). Unfortunately, in RL, representation learning is confounded with the exploratory experience of the agent – learning a useful representation requires diverse data, while effecti…

2021

Self-Supervised Policy Adaptation during Deployment

ICLR 2021spotlight

In most real world scenarios, a policy trained by reinforcement learning in one environment needs to be deployed in another, potentially quite different environment. However, generalization across different environments is known to be hard. A natural solution would be to keep training after deployme…

2021

State-Only Imitation Learning for Dexterous Manipulation

IROS 2021poster

Modern model-free reinforcement learning methods have recently demonstrated impressive results on a number of problems. However, complex domains like dexterous manipulation remain a challenge due to the high sample complexity. To address this, current approaches employ expert demonstrations in the f…

Cited by 135SourceScholar
2021

URLB: Unsupervised Reinforcement Learning Benchmark

NeurIPS 2021poster

Deep Reinforcement Learning (RL) has emerged as a powerful paradigm to solve a range of complex yet specific control tasks. Training generalist agents that can quickly adapt to new tasks remains an outstanding challenge. Recent advances in unsupervised RL have shown that pre-training RL agents with…

Cited by 181SourcecodeScholar
2020

Learning Predictive Representations for Deformable Objects Using Contrastive Estimation

CoRL 2020

Using visual model-based learning for deformable object manipulation is challenging due to difficulties in learning plannable visual representations along with complex dynamic models. In this work, we propose a new learning framework that jointly optimizes both the visual representation model and th

Cited by 0SourcePDFScholar
2020

Learning to Manipulate Deformable Objects without Demonstrations

RSS 2020poster

In this paper we tackle the problem of deformable object manipulation through model-free visual reinforcement learning (RL). In order to circumvent the sample inefficiency of RL, we propose two key ideas that accelerate learning. First, we propose an iterative pick-place action space that encodes th…

2020

Reinforcement Learning with Augmented Data

NeurIPS 2020spotlight

Learning from visual observations is a fundamental yet challenging problem in Reinforcement Learning (RL). Although algorithmic advances combined with convolutional neural networks have proved to be a recipe for success, current methods are still lacking on two fronts: (a) data-efficiency of learnin…

2020

Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation

CoRL 2020

Vision-based robotics often factors the control loop into separate components for perception and control. Conventional perception components usually extract hand-engineered features from the visual input that are then used by the control component in an explicit manner. In contrast, recent advances

Cited by 0SourcePDFScholar
2018

Asymmetric Actor Critic for Image-Based Robot Learning

RSS 2018poster

Deep reinforcement learning (RL) has proven a powerful technique in many sequential decision making domains. However, robotics poses many challenges for RL, most notably training on a physical system can be expensive and dangerous, which has sparked significant interest in learning control policies…

Cited by 452SourcePDFScholar
2018

CASSL: Curriculum Accelerated Self-Supervised Learning

ICRA 2018poster

Recent self-supervised learning approaches focus on using a few thousand data points to learn policies for high-level, low-dimensional action spaces. However, scaling this framework for higher-dimensional control requires either scaling up the data collection efforts or using a clever sampling strat…

Cited by 41SourceScholar
2018

Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation

CoRL 2018

In recent years, we have seen an emergence of data-driven approaches in robotics. However, most existing efforts and datasets are either in simulation or focus on a single task in isolation such as grasping, pushing or poking. In order to make progress and capture the space of manipulation, we would

Cited by 0SourcePDFScholar
2018

Robot Learning in Homes: Improving Generalization and Reducing Dataset Bias

NeurIPS 2018poster

Data-driven approaches to solving robotic tasks have gained a lot of traction in recent years. However, most existing policies are trained on large-scale datasets collected in curated lab settings. If we aim to deploy these models in unstructured visual environments like people's homes, they will be…

Cited by 167SourcePDFScholar
2017

Predictive-State Decoders: Encoding the Future into Recurrent Networks

NeurIPS 2017poster

Recurrent neural networks (RNNs) are a vital modeling technique that rely on internal states learned indirectly by optimization of a supervised, unsupervised, or reinforcement training loss. RNNs are used to model dynamic processes that are characterized by underlying latent states whose form is oft…

Cited by 46SourcePDFScholar