← Search

Dorsa Sadigh

137 accepted papers

2026

A Taxonomy for Evaluating Generalist Robot Manipulation Policies

ICRA 2026poster

Machine learning for robot manipulation promises to unlock generalization to novel tasks and environments. But how should we measure the progress of these policies towards generalization? Evaluating and quantifying generalization is the Wild West of modern robotics, with each work proposing and meas…

2026

Cross-Embodiment Transfer Via Behavior-Aligned Representations

ICRA 2026poster

Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned rep…

Cited by 0codeScholar
2026

HoMeR: Learning In-The-Wild Mobile Manipulation Via Hybrid Imitation and Whole-Body Control

ICRA 2026poster

We introduce HoMeR, an imitation learning framework for mobile manipulation that combines whole-body control with hybrid action modes that handle both long-range and fine-grained motion, enabling effective performance on realistic in-the-wild tasks. At its core is a fast, kinematics-based whole-body…

2026

Polychromic Objectives for Reinforcement Learning

ICLR 2026poster

Reinforcement learning fine-tuning (RLFT) is a dominant paradigm for improving pretrained policies for downstream tasks. These pretrained policies, trained on large datasets, produce generations with a broad range of promising but unrefined behaviors. Often, a critical failure mode of RLFT arises wh…

Cited by 0SourceScholar
2026

Reinforcement Learning for Non-Verifiable Problems

ICML 2026poster

Many real-world tasks are non-verifiable—there is no objective ground truth, and quality must be judged subjectively—making reward design for RL difficult. Existing approaches based on scalar rubric scores or single comparisons are often noisy, poorly calibrated, or provide sparse learning signals. …

Cited by 0SourceScholar
2026

RoboCade: Gamifying Robot Data Collection

ICRA 2026poster

Imitation learning from human demonstrations has become a dominant approach for training autonomous robot policies. However, collecting demonstration datasets is costly: it often requires access to robots and needs sustained effort in a tedious, long process. These factors limit the scale of data av…

2026

TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse

ICML 2026poster

Despite scale driving substantial recent advancements in machine learning, reinforcement learning (RL) methods still primarily use small value functions. Naively scaling value functions -- including with a transformer architecture, which is known to be highly scalable -- often results in learning in…

Cited by 0SourceScholar
2026

What Matters for Batch Online Reinforcement Learning in Robotics?

ICLR 2026poster

The ability to learn from large batches of autonomously collected data for policy improvement---a paradigm we refer to as batch online reinforcement learning---holds the promise of enabling truly scalable robot learning by significantly reducing the need for human effort of data collection while get…

Cited by 0SourceScholar
2026

Will People Enjoy a Robot Trainer? a Case Study with Snoopie the Pacerbot

ICRA 2026poster

The physicality of exercise makes the role of athletic trainers unique. Their physical presence allows them to guide a student through a motion, demonstrate an exercise, and give intuitive feedback. Robot quadrupeds are also embodied agents with robust agility and athleticism. In our work, we invest…

2025

Action-Free Reasoning for Policy Generalization

CoRL 2025poster

End-to-end imitation learning offers a promising approach for training robot policies. However, generalizing to new settings—such as unseen scenes, tasks, and object instances—remains a significant challenge. Although large-scale robot demonstration datasets have shown potential for inducing general…

Cited by 0SourcecodeScholar
2025

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

RSS 2025poster

In this work, we investigate how spatially-grounded auxiliary representations can provide both broad, high-level grounding, as well as direct, actionable information, and help policy learning performance and generalization. We study these mid-level representations across three critical dimensions: o…

Cited by 0PDFScholar
2025

Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation

ICML 2025poster

Denoising Diffusion Probabilistic Models (DDPMs) have emerged as powerful tools for generative modeling. However, their sequential computation requirements lead to significant inference-time bottlenecks. In this work, we utilize the connection between DDPMs and Stochastic Localization to prove that,…

Cited by 0SourcePDFScholar
2025

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning

ICRA 2025

Expressive robotic behavior is essential for the widespread acceptance of robots in social environments. Recent advancements in learned legged locomotion controllers have enabled more dynamic and versatile robot behaviors. However, determining the optimal behavior for interactions with different use

Cited by 3SourceScholar
2025

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

CoRL 2025poster

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of…

Cited by 0SourceScholar
2025

Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning

ICRA 2025

Teaching robots to autonomously complete everyday tasks remains a challenge. Imitation Learning (IL) is a powerful approach that imbues robots with skills via demonstrations, but is limited by the labor-intensive process of collecting teleoperated robot data. Human videos offer a scalable alternativ

Cited by 65SourcecodeScholar
2025

ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration

RA-L 2025

Collaborative robots must quickly adapt to their partner's intent and preferences to proactively identify helpful actions. This is especially true in situated settings where human partners can continually teach robots new high-level behaviors, visual concepts, and physical skills (e.g., through demo

Cited by 4SourceScholar
2025

RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

ICRA 2025

We explore how intermediate policy representations can facilitate generalization by providing guidance on how to perform manipulation tasks. Existing representations such as language, goal images, and trajectory sketches have been shown to be helpful, but these representations either do not provide

Cited by 43SourceScholar
2025

RoboCrowd: Scaling Robot Data Collection Through Crowdsourcing

ICRA 2025

In recent years, imitation learning from large-scale human demonstrations has emerged as a promising paradigm for training robot policies. However, the burden of collecting large quantities of human demonstrations is significant in terms of collection time and the need for access to expert operators

Cited by 9SourcecodeScholar
2025

Robot Data Curation with Mutual Information Estimators

RSS 2025poster

The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, litt…

Cited by 3PDFScholar
2025

Scaffolding Dexterous Manipulation with Vision-Language Models

NeurIPS 2025poster

Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensional control. While reinforcement learning (RL) can alleviate the data bottleneck by generating experience in simulation, i…

Cited by 0SourceScholar
2025

Vision Language Models are In-Context Value Learners

ICLR 2025spotlight

Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal value function, across different tasks and domains requires both a large amount of diverse data and methods which can s…

Cited by 2SourcePDFScholar
2025

What's the Move? Hybrid Imitation Learning via Salient Points

ICLR 2025poster

While imitation learning (IL) offers a promising framework for teaching robots various behaviors, learning complex tasks remains challenging. Existing IL policies struggle to generalize effectively across visual and spatial variations even for simple tasks. In this work, we introduce **SPHINX**: **S…

2024

Chain of Code: Reasoning with a Language Model-Augmented Code Emulator

ICML 2024oral

Code provides a general syntactic structure to build complex programs and perform precise computations when paired with a code interpreter – we hypothesize that language models (LMs) can leverage code-writing to improve Chain of Thought reasoning not only for logic and arithmetic tasks, but also for…

Cited by 69SourcePDFScholar
2024

Contrastive Preference Learning: Learning from Human Feedback without Reinforcement Learning

ICLR 2024poster

Reinforcement Learning from Human Feedback (RLHF) has emerged as a popular paradigm for aligning models with human intent. Typically RLHF algorithms operate in two phases: first, use human preferences to learn a reward function and second, align the model by optimizing the learned reward via reinfor…

Cited by 27SourcePDFScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections

ICRA 2024poster

Today’s robot policies exhibit subpar performance when faced with the challenge of generalizing to novel environments. Human corrective feedback is a crucial form of guidance to enable such generalization. However, adapting to and learning from online human corrections is a non-trivial endeavor: not…

Cited by 43SourcecodeScholar
2024

Efficient Data Collection for Robotic Manipulation via Compositional Generalization

RSS 2024poster

Data collection has become an increasingly important problem in robotic manipulation, yet there still lacks much understanding of how to effectively collect data to facilitate broad generalization. Recent works on large-scale robotic data collection typically vary many environmental factors of varia…

Cited by 15SourcePDFScholar
2024

Explore until Confident: Efficient Exploration for Embodied Question Answering

RSS 2024poster

We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a question. In this work, we leverage the strong semantic reasoning…

Cited by 39SourcePDFScholar
2024

FLAIR: Feeding via Long-Horizon AcquIsition of Realistic dishes

RSS 2024poster

Robot-assisted feeding holds immense promise for improving the quality of life for individuals with mobility limitations who are unable to feed themselves independently. However, there exists a large gap between the kinds of homogeneous, curated plates existing assistive feeding systems can handle,…

Cited by 13SourcePDFScholar
2024

FlowRetrieval: Flow-Guided Data Retrieval for Few-Shot Imitation Learning

CoRL 2024poster

Imitation learning policies in robotics tend to require an extensive amount of demonstrations. It is critical to develop few-shot adaptation strategies that rely only on a small amount of task-specific human demonstrations. Prior works focus on learning general policies from large scale dataset with…

Cited by 9SourceScholar
2024

How to Prompt Your Robot: A PromptBook for Manipulation Skills with Code as Policies

ICRA 2024poster

Large Language Models (LLMs) have demonstrated the ability to perform semantic reasoning, planning and write code for robotics tasks. However, most methods rely on pre-existing primitives (i.e. pick, open drawer) or similar examples of robot code alone, which heavily limits their scalability to new…

Cited by 30SourceScholar
2024

Learning to Learn Faster from Human Feedback with Language Model Predictive Control

RSS 2024poster

Large language models (LLMs) have been shown to exhibit a wide range of capabilities, such as writing robot code from language commands -- enabling non-experts to direct robot behaviors, modify them based on feedback, or compose them to perform new tasks. However, these capabilities (driven by in-co…

2024

Octo: An Open-Source Generalist Robot Policy

RSS 2024poster

Large policies pretrained on diverse robot datasets have the potential to transform robotic learning: instead of training new policies from scratch, such generalist robot policies may be finetuned with only a little in-domain data, yet generalize broadly. However, to be widely applicable across a ra…

Cited by 327SourcePDFScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

OpenVLA: An Open-Source Vision-Language-Action Model

CoRL 2024poster

Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust,…

Cited by 437SourceScholar
2024

Physically Grounded Vision-Language Models for Robotic Manipulation

ICRA 2024poster

Recent advances in vision-language models (VLMs) have led to improved performance on tasks such as visual question answering and image captioning. Consequently, these models are now well-positioned to reason about the physical world, particularly within domains such as robotic manipulation. However,…

Cited by 131SourceScholar
2024

Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

ICML 2024poster

Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as LLaVa, InstructBLIP, and PaLI-3. Despite the volume of new releases, key design deci…

Cited by 102SourcePDFScholar
2024

Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation

RSS 2024poster

Recent years in robotics and imitation learning have shown remarkable progress in training large-scale foundation models by leveraging data across a multitude of embodiments. The success of such policies might lead us to wonder: just how diverse can the robots in the training set be while still faci…

2024

RT-H: Action Hierarchies using Language

RSS 2024poster

Language provides a way to break down complex concepts into digestible pieces. Recent works in robot imitation learning have proposed learning language-conditioned policies that predict actions given visual observations and the high-level task specified in language. These methods leverage the struct…

2024

RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches

CoRL 2024poster

Natural language and images are commonly used as goal representations in goal-conditioned imitation learning. However, language can be ambiguous and images can be over-specified. In this work, we study hand-drawn sketches as a modality for goal specification. Sketches can be easy to provide on the f…

Cited by 11SourcecodeScholar
2024

ReMix: Optimizing Data Mixtures for Large Scale Imitation Learning

CoRL 2024poster

Increasingly large robotics datasets are being collected to train larger foundation models in robotics. However, despite the fact that data selection has been of utmost importance to scaling in vision and natural language processing (NLP), little work in robotics has questioned what data such models…

Cited by 16SourcecodeScholar
2024

So You Think You Can Scale Up Autonomous Robot Data Collection?

CoRL 2024poster

A long-standing goal in robot learning is to develop methods for robots to acquire new skills autonomously. While reinforcement learning (RL) comes with the promise of enabling autonomous data collection, it remains challenging to scale in the real-world partly due to the significant effort required…

Cited by 4SourceScholar
2024

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

CVPR 2024poster

Understanding and reasoning about spatial relationships is crucial for Visual Question Answering (VQA) and robotics. Vision Language Models (VLMs) have shown impressive performance in some VQA benchmarks but struggle with 3D spatial reasoning such as recognizing distances or size differences between…

Cited by 198SourcePDFScholar
2024

Toward Grounded Commonsense Reasoning

ICRA 2024poster

Consider a robot tasked with tidying a desk with a meticulously constructed Lego sports car. A human may recognize that it is not appropriate to disassemble the sports car and put it away as part of the "tidying." How can a robot reach that conclusion? Although large language models (LLMs) have rece…

Cited by 23SourcecodeScholar
2024

Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration

CoRL 2024poster

We introduce Vocal Sandbox, a framework for enabling seamless human-robot collaboration in situated environments. Systems in our framework are characterized by their ability to *adapt and continually learn* at multiple levels of abstraction from diverse teaching modalities such as spoken dialogue, o…

Cited by 0SourcecodeScholar
2023

Behavior Retrieval: Few-Shot Imitation Learning by Querying Unlabeled Datasets

RSS 2023poster

Enabling robots to learn novel visuomotor skills in a data-efficient manner remains an unsolved problem with myriad challenges. A popular paradigm for tackling this problem is through leveraging large unlabeled datasets that have many behaviors in them and then adapting a policy to a specific task u…

Cited by 34SourcePDFScholar
2023

Generating Language Corrections for Teaching Physical Control Tasks

ICML 2023poster

AI assistance continues to help advance applications in education, from language learning to intelligent tutoring systems, yet current methods for providing students feedback are still quite limited. Most automatic feedback systems either provide binary correctness feedback, which may not help a stu…

2023

In-Mouth Robotic Bite Transfer with Visual and Haptic Sensing

ICRA 2023poster

Assistance during eating is essential for those with severe mobility issues or eating risks. However, dependence on traditional human caregivers is linked to malnutrition, weight loss, and low self-esteem. For those who require eating assistance, a semi-autonomous robotic platform can provide indepe…

Cited by 13SourceScholar
2023

KITE: Keypoint-Conditioned Policies for Semantic Manipulation

CoRL 2023poster

While natural language offers a convenient shared interface for humans and robots, enabling robots to interpret and follow language commands remains a longstanding challenge in manipulation. A crucial step to realizing a performant instruction-following robot is achieving semantic manipulation – whe…

Cited by 25SourceScholar
2023

Language to Rewards for Robotic Skill Synthesis

CoRL 2023oral

Large language models (LLMs) have demonstrated exciting progress in acquiring diverse new capabilities through in-context learning, ranging from logical reasoning to code-writing. Robotics researchers have also explored using LLMs to advance the capabilities of robotic control. However, since low-le…

Cited by 326SourceScholar
2023

Language-Driven Representation Learning for Robotics

RSS 2023poster

Recent work in visual representation learning for robotics demonstrates the viability of learning from large video datasets of humans performing everyday tasks. Leveraging methods such as masked autoencoding and contrastive learning, these representations exhibit strong transfer to policy learning f…

2023

Large Language Models as General Pattern Machines

CoRL 2023poster

We observe that pre-trained large language models (LLMs) are capable of autoregressively completing complex token sequences—from arbitrary ones procedurally generated by probabilistic context-free grammars (PCFG), to more rich spatial patterns found in the Abstraction and Reasoning Corpus (ARC), a g…

Cited by 220SourceScholar
2023

Learning Tool Morphology for Contact-Rich Manipulation Tasks with Differentiable Simulation

ICRA 2023poster

When humans perform contact-rich manipulation tasks, customized tools are often necessary to simplify the task. For instance, we use various utensils for handling food, such as knives, forks and spoons. Similarly, robots may benefit from specialized tools that enable them to more easily complete a v…

Cited by 10SourceScholar
2023

Masked Imitation Learning: Discovering Environment-Invariant Modalities in Multimodal Demonstrations

IROS 2023poster

Multimodal demonstrations provide robots with an abundance of information to make sense of the world. However, such abundance may not always lead to good performance when it comes to learning sensorimotor control policies from human demonstrations. Extraneous data modalities can lead to state over-s…

Cited by 2SourceScholar
2023

Parallel Sampling of Diffusion Models

NeurIPS 2023spotlight

Diffusion models are powerful generative models but suffer from slow sampling, often taking 1000 sequential denoising steps for one sample. As a result, considerable efforts have been directed toward reducing the number of denoising steps, but these methods hurt sample quality. Instead of reducing t…

2023

Polybot: Training One Policy Across Robots While Embracing Variability

CoRL 2023poster

Reusing large datasets is crucial to scale vision-based robotic manipulators to everyday scenarios due to the high cost of collecting robotic datasets. However, robotic platforms possess varying control schemes, camera viewpoints, kinematic configurations, and end-effector morphologies, posing signi…

Cited by 27SourceScholar
2023

RoboCLIP: One Demonstration is Enough to Learn Robot Policies

NeurIPS 2023poster

Reward specification is a notoriously difficult problem in reinforcement learning, requiring extensive expert supervision to design robust reward functions. Imitation learning (IL) methods attempt to circumvent these problems by utilizing expert demonstrations instead of using an extrinsic reward fu…

Cited by 74SourcePDFScholar
2023

Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners

CoRL 2023oral

Large language models (LLMs) exhibit a wide range of promising capabilities --- from step-by-step planning to commonsense reasoning --- that may provide utility for robots, but remain prone to confidently hallucinated predictions. In this work, we present KnowNo, a framework for measuring and aligni…

Cited by 248SourceScholar
2023

Robots That Can See: Leveraging Human Pose for Trajectory Prediction

RA-L 2023

Anticipating the motion of all humans in dynamic environments such as homes and offices is critical to enable safe and effective robot navigation. Such spaces remain challenging as humans do not follow strict rules of motion and there are often multiple occluded entry points such as corners and door

Cited by 39SourcecodeScholar
2022

Assistive Teaching of Motor Control Tasks to Humans

NeurIPS 2022accept

Recent works on shared autonomy and assistive-AI technologies, such as assistive robotic teleoperation, seek to model and help human users with limited ability in a fixed task. However, these approaches often fail to account for humans' ability to adapt and eventually learn how to execute a control…

2022

Balancing Efficiency and Comfort in Robot-Assisted Bite Transfer

ICRA 2022poster

Robot-assisted feeding in household environments is challenging because it requires robots to generate trajectories that effectively bring food items of varying shapes and sizes into the mouth while making sure the user is comfortable. Our key insight is that in order to solve this challenge, robots…

Cited by 26SourceScholar
2022

Eliciting Compatible Demonstrations for Multi-Human Imitation Learning

CoRL 2022poster

Imitation learning from human-provided demonstrations is a strong approach for learning policies for robot manipulation. While the ideal dataset for imitation learning is homogenous and low-variance - reflecting a single, optimal method for performing a task - natural human behavior has a great deal…

Cited by 22SourceScholar
2022

Imitation Learning by Estimating Expertise of Demonstrators

ICML 2022spotlight

Many existing imitation learning datasets are collected from multiple demonstrators, each with different expertise at different parts of the environment. Yet, standard imitation learning algorithms typically treat all demonstrators as homogeneous, regardless of their expertise, absorbing the weaknes…

2022

Leveraging Smooth Attention Prior for Multi-Agent Trajectory Prediction

ICRA 2022poster

Multi-agent interactions are important to model for forecasting other agents' behaviors and trajectories. At a certain time, to forecast a reasonable future trajectory, each agent needs to pay attention to the interactions with only a small group of most relevant agents instead of unnecessarily payi…

Cited by 11SourceScholar
2022

Partner-Aware Algorithms in Decentralized Cooperative Bandit Teams

AAAI 2022technical

When humans collaborate with each other, they often make decisions by observing others and considering the consequences that their actions may have on the entire team, instead of greedily doing what is best for just themselves. We would like our AI agents to effectively collaborate in a similar way…

Cited by 3SourcePDFScholar
2022

Training and Inference on Any-Order Autoregressive Models the Right Way

NeurIPS 2022accept

Conditional inference on arbitrary subsets of variables is a core problem in probabilistic inference with important applications such as masked language modeling and image inpainting. In recent years, the family of Any-Order Autoregressive Models (AO-ARMs) -- closely related to popular models such a…

2021

Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

NeurIPS 2021poster

Most existing imitation learning approaches assume the demonstrations are drawn from experts who are optimal, but relaxing this assumption enables us to use a wider range of data. Standard imitation learning may learn a suboptimal policy from demonstrations with varying optimality. Prior works use c…

2021

Cooperative Autonomous Vehicles that Sympathize with Human Drivers

IROS 2021poster

Widespread adoption of autonomous vehicles will not become a reality until solutions are developed that enable these intelligent agents to co-exist with humans. This includes safely and efficiently interacting with human-driven vehicles, especially in both conflictive and competitive scenarios. We b…

Cited by 57SourceScholar
2021

ELLA: Exploration through Learned Language Abstraction

NeurIPS 2021poster

Building agents capable of understanding language instructions is critical to effective and robust human-AI collaboration. Recent work focuses on training these agents via reinforcement learning in environments with synthetic language; however, instructions often define long-horizon, sparse-reward t…

2021

Emergent Prosociality in Multi-Agent Games Through Gifting

IJCAI 2021poster

Coordination is often critical to forming prosocial behaviors -- behaviors that increase the overall sum of rewards received by all agents in a multi-agent game. However, state of the art reinforcement learning algorithms often suffer from converging to socially less desirable equilibria when multip…

Cited by 35SourcePDFScholar
2021

Learning Feasibility to Imitate Demonstrators with Different Dynamics

CoRL 2021poster

The goal of learning from demonstrations is to learn a policy for an agent (imitator) by mimicking the behavior in the demonstrations. Prior works on learning from demonstrations assume that the demonstrations are collected by a demonstrator that has the same dynamics as the imitator. However, in m…

Cited by 14SourcecodeScholar
2021

Learning Human Objectives from Sequences of Physical Corrections

ICRA 2021poster

When personal, assistive, and interactive robots make mistakes, humans naturally and intuitively correct those mistakes through physical interaction. In simple situations, one correction is sufficient to convey what the human wants. But when humans are working with multiple robots or the robot is pe…

Cited by 44SourceScholar
2021

On the Critical Role of Conventions in Adaptive Human-AI Collaboration

ICLR 2021poster

Humans can quickly adapt to new partners in collaborative tasks (e.g. playing basketball), because they understand which fundamental skills of the task (e.g. how to dribble, how to shoot) carry over across new partners. Humans can also quickly adapt to similar tasks with the same partners by carryin…

2021

Open-domain clarification question generation without question examples

EMNLP 2021main

An overarching goal of natural language processing is to enable machines to communicate seamlessly with humans. However, natural language can be ambiguous or unclear. In cases of uncertainty, humans engage in an interactive process known as repair: asking questions and seeking clarification until th…

2021

ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes

ICRA 2021poster

Characterizing what types of exoskeleton gaits are comfortable for users, and understanding the science of walking more generally, require recovering a user’s utility landscape. Learning these landscapes is challenging, as walking trajectories are defined by numerous gait parameters, data collection…

Cited by 52SourcecodeScholar
2021

Targeted Data Acquisition for Evolving Negotiation Agents

ICML 2021spotlight

Successful negotiators must learn how to balance optimizing for self-interest and cooperation. Yet current artificial negotiation agents often heavily depend on the quality of the static datasets they were trained on, limiting their capacity to fashion an adaptive response balancing self-interest an…

Cited by 11SourcePDFScholar
2020

Active Preference-Based Gaussian Process Regression for Reward Learning

RSS 2020poster

Designing reward functions is a challenging problem in AI and robotics. Humans usually have a difficult time directly specifying all the desirable behaviors that a robot needs to optimize. One common approach is to learn reward functions from collected expert demonstrations. However, learning reward…

2020

Controlling Assistive Robots with Learned Latent Actions

ICRA 2020poster

Assistive robotic arms enable users with physical disabilities to perform everyday tasks without relying on a caregiver. Unfortunately, the very dexterity that makes these arms useful also makes them challenging to teleoperate: the robot has more degrees-of-freedom than the human can directly coordi…

Cited by 97SourceScholar
2020

Dynamic Multi-Robot Task Allocation under Uncertainty and Temporal Constraints

RSS 2020poster

We consider the problem of dynamically allocating tasks to multiple agents under time window constraints and task completion uncertainty. Our objective is to minimize the number of unsuccessful tasks at the end of the operation horizon. We present a multi-robot allocation algorithm that decouples t…

2020

Learning Latent Representations to Influence Multi-Agent Interaction

CoRL 2020

Seamlessly interacting with humans or robots is hard because these agents are non-stationary. They update their policy in response to the ego agent’s behavior, and the ego agent must anticipate these changes to co-adapt. Inspired by humans, we recognize that robots do not need to explicitly model ev

Cited by 0SourcePDFScholar
2020

Reinforcement Learning based Control of Imitative Policies for Near-Accident Driving

RSS 2020poster

Autonomous driving has achieved significant progress in recent years, but autonomous cars are still unable to tackle high-risk situations where a potential accident is likely. In such near-accident scenarios, even a minor change in the vehicle's actions may result in drastically different consequenc…

2019

Active Learning of Reward Dynamics from Hierarchical Queries

IROS 2019poster

Enabling robots to act according to human preferences across diverse environments is a crucial task, extensively studied by both roboticists and machine learning researchers. To achieve it, human preferences are often encoded by a reward function which the robot optimizes for. This reward function i…

Cited by 45SourceScholar
2019

Asking Easy Questions: A User-Friendly Approach to Active Reward Learning

CoRL 2019

Robots can learn the right reward function by querying a human expert. Existing approaches attempt to choose questions where the robot is most uncertain about the human’s response; however, they do not consider how easy it will be for the human to answer! In this paper we explore an information gain

2019

Deep Local Trajectory Replanning and Control for Robot Navigation

ICRA 2019poster

We present a navigation system that combines ideas from hierarchical planning and machine learning. The system uses a traditional global planner to compute optimal paths towards a goal, and a deep local trajectory planner and velocity controller to compute motion commands. The latter components of t…

Cited by 88SourceScholar
2019

Hierarchical Game-Theoretic Planning for Autonomous Vehicles

ICRA 2019poster

The actions of an autonomous vehicle on the road affect and are affected by those of other drivers, whether overtaking, negotiating a merge, or avoiding an accident. This mutual dependence, best captured by dynamic game theory, creates a strong coupling between the vehicle's planning and its predict…

Cited by 323SourceScholar
2019

Learning Reward Functions by Integrating Human Demonstrations and Preferences

RSS 2019poster

Our goal is to accurately and efficiently learn reward functions for autonomous robots. Current approaches to this problem include inverse reinforcement learning (IRL), which uses expert demonstrations, and preference-based learning, which iteratively queries the user for her preferences between tra…

2019

Learning from My Partner’s Actions: Roles in Decentralized Robot Teams

CoRL 2019

When teams of robots collaborate to complete a task, communication is often necessary. Like humans, robot teammates should implicitly communicate through their actions: but interpreting our partner’s actions is typically difficult, since a given action may have many different underlying reasons. Her

Cited by 0SourcePDFScholar
2019

Unsupervised Visuomotor Control through Distributional Planning Networks

RSS 2019poster

While reinforcement learning (RL) has the potential to enable robots to autonomously acquire a wide range of skills, in practice, RL usually requires manual, per-task engineering of reward functions, especially in real world settings where aspects of the environment needed to compute progress are no…

Cited by 48SourcePDFScholar
2018

Multi-Agent Generative Adversarial Imitation Learning

NeurIPS 2018poster

Imitation learning algorithms can be used to learn a policy from expert demonstrations without access to a reward signal. However, most existing approaches are not applicable in multi-agent settings due to the existence of multiple (Nash) equilibria and non-stationary environments. We propose a new…

2016

Information gathering actions over human internal state

IROS 2016poster

Much of estimation of human internal state (goal, intentions, activities, preferences, etc.) is passive: an algorithm observes human actions and updates its estimate of human state. In this work, we embrace the fact that robot actions affect what humans do, and leverage it to improve state estimatio…

Cited by 253SourceScholar
2016

Planning for Autonomous Cars that Leverage Effects on Human Actions

RSS 2016poster

Traditionally, autonomous cars make predic- tions about other drivers’ future trajectories, and plan to stay out of their way. This tends to result in defensive and opaque behaviors. Our key insight is that an autonomous car’s actions will actually affect what other cars will do in response, whe…

Cited by 661SourcePDFScholar