← Search

Abhishek Gupta

103 accepted papers

2026

Adapting to Evolving Graphs: A Scalable Framework for Dynamic Coarsening

ICML 2026poster

Graph coarsening is a fundamental dimensionality reduction technique for scaling large graphs while preserving structural and feature information. However, most existing coarsening methods are designed for static graphs and do not extend well to dynamic settings where nodes, edges, and connectivity …

Cited by 0SourceScholar
2026

Amortized Multi-Objective Optimization Across Tasks with Generative Solution Modeling

IJCAI 2026

Many real-world applications require solving families of expensive multi-objective optimization problems~(EMOPs) under varying operational conditions. This can be formulated as parametric expensive multi-objective optimization problems (P-EMOPs) where each task parameter defines a distinct optimizat

Cited by 0Scholar
2026

Difference-Aware Retrieval Polices for Imitation Learning

ICLR 2026poster

Behavior cloning suffers from poor generalization to out-of-distribution states due to compounding errors during deployment. We present Difference-Aware Retrieval Polices for Imitation Learning (DARP), a novel nearest-neighbor-based imitation learning approach that addresses this limitation by repar…

Cited by 0SourceScholar
2026

Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement Learning

ICLR 2026poster

Reinforcement learning in GPU-enabled physics simulation has been the driving force behind many of the breakthroughs in sim-to-real robot learning. However, current approaches for data generation in simulation are unwieldy and task-specific, requiring extensive human effort to engineer training curr…

Cited by 0SourceScholar
2026

LieStoNet: Learning Lie Symmetries from Spatiotemporal Data for Stochastic Dynamical Systems

ICML 2026poster

Symmetry is central to modern machine learning and physics: invariances and equivariances improve sample efficiency, robustness, and out-of-distribution generalization, while symmetry principles guide scientific modeling. Yet for stochastic dynamical systems, the relevant continuous symmetries are r…

Cited by 0SourceScholar
2026

PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies

ICRA 2026poster

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to vision-language models (VLMs), leaving policies to specialize in h…

2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

RFS: Reinforcement learning with Residual flow steering for dexterous manipulation

ICLR 2026poster

Imitation learning has been an effective tool for bootstrapping sequential decision making behavior, showing surprisingly strong results as methods are scaled up to high-dimensional, dexterous problems in robotics. These ``behavior cloning" methods have been further bolstered by the integration of g…

Cited by 0SourcecodeScholar
2026

Refinery: Active Fine-Tuning and Deployment-Time Optimization for Contact-Rich Policies

ICRA 2026poster

Simulation-based learning has enabled policies for precise, contact-rich tasks (e.g., robotic assembly) to reach high success rates (~80%) under high levels of observation noise and control error. Although such performance may be sufficient for research applications, it falls short of industry stand…

2026

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

RSS 2026poster

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large scale real-world robotics datasets where failed and s…

Cited by 0SourceScholar
2026

SPARR: Simulation-Based Policies with Asymmetric Real-World Residuals for Assembly

ICRA 2026poster

Robotic assembly presents a long-standing challenge due to its requirement for precise, contact-rich manipulation. While simulation-based learning has enabled the development of robust assembly policies, their performance often degrades when deployed in real-world settings due to the sim-to-real gap…

2026

Sample Efficient Full-Finetuning of Generative Control Policies

ICML 2026poster

Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning. Yet there remains substantial debate over how to sample efficiently fine-tune them via reinforcement learning. A prevailing view holds that fine-tun…

Cited by 0SourceScholar
2026

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation

RSS 2026poster

Simulation-to-real transfer remains a central challenge in robotics, as mismatches between simulated and real-world dynamics often lead to failures. While reinforcement learning offers a principled mechanism for adaptation, existing sim-to-real finetuning methods struggle with exploration and long-h…

Cited by 0SourceScholar
2026

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

RSS 2026poster

Efficient exploration remains a bottleneck in reinforcement learning (RL), particularly for long-horizon, high-dimensional tasks. While recent methods leverage pre-trained policies for guidance, they are often constrained by the base policy’s original behavior distribution. We introduce Timestep Mod…

Cited by 0SourceScholar
2026

Using Non-Expert Data to Robustify Imitation Learning Via Offline Reinforcement Learning

ICRA 2026poster

Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In …

2025

AMPS: ASR with Multimodal Paraphrase Supervision

NAACL 2025short

Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS, that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational…

Cited by 0SourcePDFScholar
2025

ATK: Automatic Task-driven Keypoint Selection for Robust Policy Learning

CoRL 2025poster

Learning visuamotor policy through imitation learning often suffers from perceptual challenges, where visual differences between training and evaluation environments degrade policy performance. Policies relying on state estimations like 6D pose, require task-specific tracking and are difficult to sc…

Cited by 0SourceScholar
2025

DRAWER: Digital Reconstruction and Articulation With Environment Realism

CVPR 2025poster

Creating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach cen…

2025

Duolingo: Dynamics Utilization for Online Translation of Actions

ICRA 2025

Robots in the real world experience wear and tear, leading to changing system dynamics. This challenge is particularly exacerbated for non-rigid systems such as soft robots or robotic systems made of metamaterials with hysteresis. This setting results in a challenging problem for most learning-based

Cited by 0SourceScholar
2025

Evolvable Conditional Diffusion

IJCAI 2025

This paper presents an evolvable conditional diffusion method such that black-box, non-differentiable multi-physics models, as are common in domains like computational fluid dynamics and electromagnetics, can be effectively used for guiding the generative process to facilitate autonomous scientific

Cited by 0SourcePDFScholar
2025

HAMSTER: Hierarchical Action Models for Open-World Robot Manipulation

ICLR 2025poster

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is the lack of robotic data, which are typically obtained through expensive on-robot ope…

2025

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

EMNLP 2025

Large language models (LLMs) have shown promise in robotic procedural planning, yet their human-centric reasoning often omits the low-level, grounded details needed for robotic execution. Vision-language models (VLMs) offer a path toward more perceptually grounded plans, but current methods either r

2025

Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-Tuning

ICLR 2025poster

Robot learning requires a considerable amount of high-quality data to realize the promise of generalization. However, large data sets are costly to collect in the real world. Physics simulators can cheaply generate vast data sets with broad coverage over states, actions, and environments. However, p…

Cited by 2SourcePDFScholar
2025

RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

CoRL 2025oral

Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standardization, either by specifying fixed evaluation tasks and environments, or by hosting centralized "robot challenges", an…

Cited by 0SourceScholar
2025

SRSA: Skill Retrieval and Adaptation for Robotic Assembly Tasks

ICLR 2025spotlight

Enabling robots to learn novel tasks in a data-efficient manner is a long-standing challenge. Common strategies involve carefully leveraging prior experiences, especially transition data collected on related tasks. Although much progress has been made for general pick-and-place manipulation, far few…

2025

STRAP: Robot Sub-Trajectory Retrieval for Augmented Policy Learning

ICLR 2025poster

Robot learning is witnessing a significant increase in the size, diversity, and complexity of pre-collected datasets, mirroring trends in domains such as natural language processing and computer vision. Many robot learning methods treat such datasets as multi-task expert data and learn a multi-task,…

Cited by 1SourcePDFScholar
2025

Steering Your Diffusion Policy with Latent Space Reinforcement Learning

CoRL 2025oral

Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications. However, in scenarios where initial performance is not satisfactory, as is often the case in novel open-world settings, such behavioral cloning (BC)-learned policies typically…

Cited by 0SourceScholar
2025

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

RSS 2025poster

Imitation learning has emerged as a promising approach towards building generalist robots. However, the reliance on high-quality expert demonstrations poses a challenge in scaling imitation learning for large-scale robot foundation models. On the other hand, large amounts of video data depicting a w…

Cited by 2PDFScholar
2024

ASID: Active Exploration for System Identification in Robotic Manipulation

ICLR 2024oral

Model-free control strategies such as reinforcement learning have shown the ability to learn control strategies without requiring an accurate model or simulator of the world. While this is appealing due to the lack of modeling requirements, such methods can be sample inefficient, making them impract…

Cited by 14SourcePDFScholar
2024

CCIL: Continuity-Based Data Augmentation for Corrective Imitation Learning

ICLR 2024poster

We present a new technique to enhance the robustness of imitation learning methods by generating corrective data to account for compounding error and disturbances. While existing methods rely on interactive expert labeling, additional offline datasets, or domain-specific invariances, our approach re…

Cited by 9SourcePDFScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

Data Efficient Behavior Cloning for Fine Manipulation via Continuity-based Corrective Labels

IROS 2024poster

We consider imitation learning with access only to expert demonstrations, whose real-world application is often limited by covariate shift due to compounding errors during execution. We investigate the effectiveness of the Continuity-based Corrective Labels for Imitation Learning (CCIL) framework in…

Cited by 2SourceScholar
2024

Distributional Successor Features Enable Zero-Shot Policy Optimization

NeurIPS 2024poster

Intelligent agents must be generalists, capable of quickly adapting to various tasks. In reinforcement learning (RL), model-based RL learns a dynamics model of the world, in principle enabling transfer to arbitrary reward functions through planning. However, autoregressive model rollouts suffer from…

2024

Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning

ICLR 2024poster

Off-policy dynamic programming (DP) techniques such as $Q$-learning have proven to be important in sequential decision-making problems. In the presence of function approximation, however, these techniques often diverge due to the absence of Bellman completeness in the function classes considered, a…

2024

Learning to Cooperate with Humans using Generative Agents

NeurIPS 2024poster

Training agents that can coordinate zero-shot with humans is a key mission in multi-agent reinforcement learning (MARL). Current algorithms focus on training simulated human partner policies which are then used to train a Cooperator agent. The simulated human is produced either through behavior clon…

2024

Learning to Grasp in Clutter with Interactive Visual Failure Prediction

ICRA 2024poster

Modern warehouses process millions of unique objects which are often stored in densely packed containers. To automate tasks in this environment, a robot must be able to pick diverse objects from highly cluttered scenes. Real-world learning is a promising approach, but executing picks in the real wor…

Cited by 2SourcecodeScholar
2024

Lifelong Robot Learning with Human Assisted Language Planners

ICRA 2024poster

Large Language Models (LLMs) have been shown to act like planners that can decompose high-level instructions into a sequence of executable instructions. However, current LLM-based planners are only able to operate with a fixed set of skills. We overcome this critical limitation and present a method…

Cited by 19SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL

NeurIPS 2024poster

In order to mitigate the sample complexity of real-world reinforcement learning, common practice is to first train a policy in a simulator where samples are cheap, and then deploy this policy in the real world, with the hope that it generalizes effectively. Such \emph{direct sim2real} transfer is no…

Cited by 1SourcePDFScholar
2024

Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

NeurIPS 2024spotlight

Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these dif…

Cited by 29SourcePDFScholar
2024

Rank2Reward: Learning Shaped Reward Functions from Passive Video

ICRA 2024poster

Teaching robots novel skills with demonstrations via human-in-the-loop data collection techniques like kinesthetic teaching or teleoperation puts a heavy burden on human supervisors. In contrast to this paradigm, it is often significantly easier to provide raw, action-free visual data of tasks being…

Cited by 5SourcecodeScholar
2024

Reconciling Reality through Simulation: A Real-To-Sim-to-Real Approach for Robust Manipulation

RSS 2024poster

Imitation learning methods need significant human supervision to learn policies robust to changes in object poses, physical disturbances, and visual distractors. Reinforcement learning, on the other hand, can explore the environment autonomously to learn robust behaviors but may require impractical…

Cited by 55SourcePDFScholar
2024

SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

ICRA 2024poster

In recent years, significant progress has been made in the field of robotic reinforcement learning (RL), enabling methods that handle complex image observations, train in the real world, and incorporate auxiliary data, such as demonstrations and prior experience. However, despite these advances, rob…

Cited by 48SourcecodeScholar
2024

Teaching Robots with Show and Tell: Using Foundation Models to Synthesize Robot Policies from Language and Visual Demonstration

CoRL 2024poster

We introduce a modular, neuro-symbolic framework for teaching robots new skills through language and visual demonstration. Our approach, ShowTell, composes a mixture of foundation models to synthesize robot manipulation programs that are easy to interpret and generalize across a wide range of tasks…

Cited by 2SourceScholar
2024

URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

RSS 2024poster

Constructing accurate and targeted simulation scenes that are both visually and physically realistic is a problem of significant practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large data-hungry learning metho…

Cited by 21SourcePDFScholar
2024

Universal Visual Decomposer: Long-Horizon Manipulation Made Easy

ICRA 2024poster

Real-world robotic tasks stretch over extended horizons and encompass multiple stages. Learning long-horizon manipulation tasks, however, is a long-standing challenge, and demands decomposing the overarching task into several manageable subtasks to facilitate policy learning and generalization to un…

Cited by 21SourceScholar
2023

Autonomous Robotic Reinforcement Learning with Asynchronous Human Feedback

CoRL 2023poster

Ideally, we would place a robot in a real-world environment and leave it there improving on its own by gathering more experience autonomously. However, algorithms for autonomous robotic learning have been challenging to realize in the real world. While this has often been attributed to the challenge…

Cited by 6SourcecodeScholar
2023

Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced Datasets

NeurIPS 2023poster

Offline reinforcement learning (RL) enables learning a decision-making policy without interaction with the environment. This makes it particularly beneficial in situations where such interactions are costly. However, a known challenge for offline RL algorithms is the distributional mismatch between…

2023

Breadcrumbs to the Goal: Goal-Conditioned Exploration from Human-in-the-Loop Feedback

NeurIPS 2023poster

Exploration and reward specification are fundamental and intertwined challenges for reinforcement learning. Solving sequential decision making tasks with a non-trivial element of exploration requires either specifying carefully designed reward functions or relying on indiscriminate, novelty seeking…

2023

Cherry-Picking with Reinforcement Learning

RSS 2023poster

Grasping small objects surrounded by unstable or non-rigid material plays a crucial role in applications such as surgery, harvesting, construction, disaster recovery, and assisted feeding. This task is especially difficult when fine manipulation is required in the presence of sensor noise and percep…

2023

Demonstration-Bootstrapped Autonomous Practicing via Multi-Task Reinforcement Learning

ICRA 2023poster

Reinforcement learning systems have the potential to enable continuous improvement in unstructured environments, leveraging data collected autonomously. However, in practice these systems require significant amounts of instrumentation or human intervention to learn in the real world. In this work, w…

Cited by 16SourceScholar
2023

Dexterous Manipulation from Images: Autonomous Real-World RL via Substep Guidance

ICRA 2023poster

Complex and contact-rich robotic manipulation tasks, particularly those that involve multi-fingered hands and underactuated object manipulation, present a significant challenge to any control method. Methods based on reinforcement learning offer an appealing choice for such settings, as they can ena…

Cited by 24SourceScholar
2023

Fighting Uncertainty with Gradients: Offline Reinforcement Learning via Diffusion Score Matching

CoRL 2023poster

Gradient-based methods enable efficient search capabilities in high dimensions. However, in order to apply them effectively in offline optimization paradigms such as offline Reinforcement Learning (RL) or Imitation Learning (IL), we require a more careful consideration of how uncertainty estimation…

Cited by 11SourceScholar
2023

GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

RSS 2023poster

Robot learning methods have the potential for widespread generalization across tasks, environments, and objects. However, these methods are severely limited by the amount of data that they are provided or are able to collect. Robots in the real world are likely to only be able to collect a small dat…

Cited by 88SourcePDFScholar
2023

Guiding Pretraining in Reinforcement Learning with Large Language Models

ICML 2023poster

Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivated exploration methods address this limitation by rewarding agents for visiting novel states or transitions, but these methods offer limited benefits in large environment…

2023

Learning to Extrapolate: A Transductive Approach

ICLR 2023poster

Machine learning systems, especially with overparameterized deep neural networks, can generalize to novel test instances drawn from the same distribution as the training data. However, they fare poorly when evaluated on out-of-support test points. In this work, we tackle the problem of developing ma…

2023

REBOOT: Reuse Data for Bootstrapping Efficient Real-World Dexterous Manipulation

CoRL 2023poster

Dexterous manipulation tasks involving contact-rich interactions pose a significant challenge for both model-based control systems and imitation learning algorithms. The complexity arises from the need for multi-fingered robotic hands to dynamically establish and break contacts, balance forces on th…

Cited by 11SourceScholar
2023

RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability

NeurIPS 2023spotlight

Visual model-based RL methods typically encode image observations into low-dimensional representations in a manner that does not eliminate redundant information. This leaves them susceptible to spurious variations -- changes in task-irrelevant components such as background distractors or lighting co…

Cited by 13SourcePDFScholar
2023

RoboHive: A Unified Framework for Robot Learning

NeurIPS 2023poster

We present RoboHive, a comprehensive software platform and ecosystem for research in the field of Robot Learning and Embodied Artificial Intelligence. Our platform encompasses a diverse range of pre-existing and novel environments, including dexterous manipulation with the Shadow Hand, whole-arm man…

2023

Self-Supervised Reinforcement Learning that Transfers using Random Features

NeurIPS 2023poster

Model-free reinforcement learning algorithms have exhibited great potential in solving single-task sequential decision-making problems with high-dimensional observations and long horizons, but are known to be hard to generalize across tasks. Model-based RL, on the other hand, learns task-agnostic mo…

Cited by 11SourcePDFScholar
2023

TactoFind: A Tactile Only System for Object Retrieval

ICRA 2023poster

We study the problem of object retrieval in scenarios where visual sensing is absent, object shapes are unknown beforehand and objects can move freely, like grabbing objects out of a drawer. Successful solutions require localizing free objects, identifying specific object instances, and then graspin…

Cited by 14SourceScholar
2022

Autonomous Reinforcement Learning: Formalism and Benchmarking

ICLR 2022poster

Reinforcement learning (RL) provides a naturalistic framing for learning through trial and error, which is appealing both because of its simplicity and effectiveness and because of its resemblance to how humans and animals acquire skills through experience. However, real-world embodied learning, suc…

2022

Autonomous Service Robots for Urban Waste Management - Multiagent Route Planning and Cooperative Operation

RA-L 2022

A sustainable approach towards smart waste management means reduced impact on the health of the workers, lower greenhouse gas emissions and low operational costs. Incorporating fleets of the robot MARBLE (Mobile Autonomous RoBot for Litter Emptying) effectuates these requirements for autonomously em

Cited by 21SourceScholar
2022

CounteRGAN: Generating counterfactuals for real-time recourse and interpretability using residual GANs

UAI 2022poster

Model interpretability, fairness, and recourse for end users have increased as machine learning models have become increasingly popular in areas including criminal justice, finance, healthcare, and job marketplaces. This work presents a novel method of addressing these issues by producing meaningful…

2022

Distributionally Adaptive Meta Reinforcement Learning

NeurIPS 2022accept

Meta-reinforcement learning algorithms provide a data-driven way to acquire policies that quickly adapt to many tasks with varying rewards or dynamics functions. However, learned meta-policies are often effective only on the exact task distribution on which they were trained and struggle in the pres…

Cited by 19SourcePDFScholar
2022

Graph Learning Assisted Multi-Objective Integer Programming

NeurIPS 2022accept

Objective-space decomposition algorithms (ODAs) are widely studied for solving multi-objective integer programs. However, they often encounter difficulties in handling scalarized problems, which could cause infeasibility or repetitive nondominated points and thus induce redundant runtime. To mitigat…

Cited by 11SourcePDFScholar
2022

Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation

CoRL 2022poster

Dexterous robotic hands have the capability to interact with a wide variety of household objects. However, learning robust real world grasping policies for arbitrary objects has proven challenging due to the difficulty of generating high quality training data. In this work, we propose a learning sys…

Cited by 32SourceScholar
2022

Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity

NeurIPS 2022accept

The success of reinforcement learning in a variety of challenging sequential decision-making problems has been much discussed, but often ignored in this discussion is the consideration of how the choice of reward function affects the behavior of these algorithms. Most practical RL algorithms require…

Cited by 83SourcePDFScholar
2022

Weighted Gaussian Process Bandits for Non-stationary Environments

AISTATS 2022poster

In this paper, we consider the Gaussian process (GP) bandit optimization problem in a non-stationary environment. To capture external changes, the black-box function is allowed to be time-varying within a reproducing kernel Hilbert space (RKHS). To this end, we develop WGP-UCB, a novel UCB-type algo…

Cited by 29SourcePDFScholar
2021

Adaptive Risk Minimization: Learning to Adapt to Domain Shift

NeurIPS 2021poster

A fundamental assumption of most machine learning algorithms is that the training and test data are drawn from the same underlying distribution. However, this assumption is violated in almost all practical applications: machine learning systems are regularly tested under distribution shift, due to c…

2021

Autonomous Reinforcement Learning via Subgoal Curricula

NeurIPS 2021poster

Reinforcement learning (RL) promises to enable autonomous acquisition of complex behaviors for diverse agents. However, the success of current reinforcement learning algorithms is predicated on an often under-emphasised requirement -- each trial needs to start from a fixed initial state distribution…

Cited by 34SourcePDFScholar
2021

Fully Autonomous Real-World Reinforcement Learning with Applications to Mobile Manipulation

CoRL 2021poster

In this paper, we study how robots can autonomously learn skills that require a combination of navigation and grasping. Learning robotic skills in the real world remains challenging without large scale data collection and supervision. Our aim is to devise a robotic reinforcement learning system for…

Cited by 58SourceScholar
2021

Learning to Reach Goals via Iterated Supervised Learning

ICLR 2021oral

Current reinforcement learning (RL) algorithms can be brittle and difficult to use, especially when learning goal-reaching behaviors from sparse rewards. Although supervised imitation learning provides a simple and stable alternative, it requires access to demonstrations from a human supervisor. In…

2021

MURAL: Meta-Learning Uncertainty-Aware Rewards for Outcome-Driven Reinforcement Learning

ICML 2021spotlight

Exploration in reinforcement learning is, in general, a challenging problem. A common technique to make learning easier is providing demonstrations from a human supervisor, but such demonstrations can be expensive and time-consuming to acquire. In this work, we study a more tractable class of reinfo…

Cited by 46SourcePDFScholar
2021

Reset-Free Reinforcement Learning via Multi-Task Learning: Learning Dexterous Manipulation Behaviors without Human Intervention

ICRA 2021poster

Reinforcement Learning (RL) algorithms can in principle acquire complex robotic skills by learning from large amounts of data in the real world, collected via trial and error. However, most RL algorithms use a carefully engineered setup in order to collect data, requiring human supervision and inter…

Cited by 118SourceScholar
2021

Teachable Reinforcement Learning via Advice Distillation

NeurIPS 2021poster

Training automated agents to complete complex tasks in interactive environments is challenging: reinforcement learning requires careful hand-engineering of reward functions, imitation learning requires specialized infrastructure and access to a human expert, and learning from intermediate forms of s…

2021

Which Mutual-Information Representation Learning Objectives are Sufficient for Control?

NeurIPS 2021poster

Mutual information (MI) maximization provides an appealing formalism for learning representations of data. In the context of reinforcement learning (RL), such representations can accelerate learning by discarding irrelevant and redundant information, while retaining the information necessary for con…

Cited by 41SourcePDFScholar
2020

DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction

NeurIPS 2020spotlight

Deep reinforcement learning can learn effective policies for a wide range of tasks, but is notoriously difficult to use due to instability and sensitivity to hyperparameters. The reasons for this remain unclear. In this paper, we study how RL methods based on bootstrapping-based Q-learning can suffe…

Cited by 131SourcePDFScholar
2020

Gradient Surgery for Multi-Task Learning

NeurIPS 2020poster

While deep learning and deep reinforcement learning (RL) systems have demonstrated impressive results in domains such as image classification, game playing, and robotic control, data efficiency remains a major challenge. Multi-task learning has emerged as a promising approach for sharing structure a…

2020

The Ingredients of Real World Robotic Reinforcement Learning

ICLR 2020spotlight

The success of reinforcement learning in the real world has been limited to instrumented laboratory scenarios, often requiring arduous human supervision to enable continuous learning. In this work, we discuss the required elements of a robotic system that can continually and autonomously improve wit…

Cited by 220SourceScholar
2019

Automatically Composing Representation Transformations as a Means for Generalization

ICLR 2019poster

A generally intelligent learner should generalize to more complex tasks than it has previously encountered, but the two common paradigms in machine learning -- either training a separate learner per task or training a single learner for all tasks -- both have difficulty with such generalization beca…

2019

Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost

ICRA 2019poster

Dexterous multi-fingered robotic hands can perform a wide range of manipulation skills, making them an appealing component for general-purpose robotic manipulators. However, such hands pose a major challenge for autonomous control, due to the high dimensionality of their configuration space and comp…

Cited by 274SourceScholar
2019

Diversity is All You Need: Learning Skills without a Reward Function

ICLR 2019poster

Intelligent creatures can explore their environments and learn useful skills without supervision. In this paper, we propose ``Diversity is All You Need''(DIAYN), a method for learning useful skills without a reward function. Our proposed method learns skills by maximizing an information theoretic ob…

Cited by 1341SourcePDFScholar
2019

Domain Randomization for Active Pose Estimation

ICRA 2019poster

Accurate state estimation is a fundamental component of robotic control. In robotic manipulation tasks, as is our focus in this work, state estimation is essential for identifying the positions of objects in the scene, forming the basis of the manipulation plan. However, pose estimation typically re…

Cited by 59SourceScholar
2019

Guided Meta-Policy Search

NeurIPS 2019spotlight

Reinforcement learning (RL) algorithms have demonstrated promising results on complex tasks, yet often require impractical numbers of samples because they learn from scratch. Meta-RL aims to address this challenge by leveraging experience from previous tasks so as to more quickly solve new tasks. Ho…

Cited by 87SourcePDFScholar
2019

Guiding Policies with Language via Meta-Learning

ICLR 2019poster

Behavioral skills or policies for autonomous agents are conventionally learned from reward functions, via reinforcement learning, or from demonstrations, via imitation learning. However, both modes of task specification have their disadvantages: reward functions require manual engineering, while dem…

Cited by 75SourcePDFScholar
2019

ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots

CoRL 2019

ROBEL is an open-source platform of cost-effective robots designed for reinforcement learning in the real world. ROBEL introduces two robots, each aimed to accelerate reinforcement learning research in different task domains: D’Claw is a three-fingered hand robot that facilitates learning dexterous

Cited by 0SourcePDFScholar
2019

Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

CoRL 2019

We present relay policy learning, a method for imitation and reinforcement learning that can solve multi-stage, long-horizon robotic tasks. This general and universally-applicable, two-phase approach consists of an imitation learning stage resulting in goal-conditioned hierarchical policies that can

2019

Unsupervised Curricula for Visual Meta-Reinforcement Learning

NeurIPS 2019spotlight

In principle, meta-reinforcement learning algorithms leverage experience across many tasks to learn fast and effective reinforcement learning (RL) strategies. However, current meta-RL approaches rely on manually-defined distributions of training tasks, and hand-crafting these task distributions can…

Cited by 79SourcePDFScholar
2018

Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation

ICRA 2018poster

Imitation learning is an effective approach for autonomous systems to acquire control policies when an explicit reward function is unavailable, using supervision provided as demonstrations from an expert, typically a human operator. However, standard imitation learning methods assume that the agent…

Cited by 455SourceScholar
2018

Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

RSS 2018poster

Dexterous multi-fingered hands are extremely versatile and provide a generic way to perform a multitude of tasks in human-centric environments. However, effectively controlling them remains challenging due to their high dimensionality and large number of potential contacts. Deep reinforcement learni…

Cited by 1325SourcePDFScholar
2018

Meta-Reinforcement Learning of Structured Exploration Strategies

NeurIPS 2018spotlight

Exploration is a fundamental challenge in reinforcement learning (RL). Many current exploration methods for deep RL use task-agnostic objectives, such as information gain or bonuses based on state visitation. However, many practical applications of RL involve learning more than a single task, and pr…

2018

Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings

ICML 2018oral

In this work, we take a representation learning perspective on hierarchical reinforcement learning, where the problem of learning lower layers in a hierarchy is transformed into the problem of learning trajectory-level generative models. We show that we can learn continuous latent representations of…

Cited by 193SourcePDFScholar
2017

Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning

ICLR 2017poster

People can learn a wide range of tasks from their own experience, but can also learn from observing other creatures. This can accelerate acquisition of new skills even when the observed agent differs substantially from the learning agent in terms of morphology. In this paper, we examine how reinforc…

Cited by 369SourceScholar
2017

Learning modular neural network policies for multi-task and multi-robot transfer

ICRA 2017poster

Reinforcement learning (RL) can automate a wide variety of robotic skills, but learning each new skill requires considerable real-world data collection and manual representation engineering to design policy classes or features. Using deep reinforcement learning to train general purpose neural networ…

Cited by 489SourceScholar
2016

Guided search for task and motion plans using learned heuristics

ICRA 2016

Tasks in mobile manipulation planning often require thousands of individual motions to complete. Such tasks require reasoning about complex goals as well as the feasibility of movements in configuration space. In discrete representations, planning complexity is exponential in the length of the plan.

Cited by 83SourceScholar
2016

Learning dexterous manipulation for a soft robotic hand from human demonstrations

IROS 2016poster

Dexterous multi-fingered hands can accomplish fine manipulation behaviors that are infeasible with simple robotic grippers. However, sophisticated multi-fingered hands are often expensive and fragile. Low-cost soft hands offer an appealing alternative to more conventional devices, but present consid…

Cited by 228SourceScholar
2015

Learning force-based manipulation of deformable objects from multiple demonstrations

ICRA 2015poster

Manipulation of deformable objects often requires a robot to apply specific forces to bring the object into the desired configuration. For instance, tightening a knot requires pulling on the ends, flattening an article of clothing requires smoothing out wrinkles, and erasing a whiteboard requires ap…

Cited by 185SourceScholar
2015

Learning from multiple demonstrations using trajectory-aware non-rigid registration with applications to deformable object manipulation

IROS 2015poster

Learning from demonstration by means of non-rigid point cloud registration is an effective tool for learning to manipulate a wide range of deformable objects. However, most methods that use non-rigid registration to transfer demonstrated trajectories assume that the test and demonstration scene are…

Cited by 49SourceScholar