← Search

Daniel S. Brown

21 accepted papers

2025

Agreement Volatility: A Second-Order Metric for Uncertainty Quantification in Surgical Robot Learning

CoRL 2025poster

Autonomous surgical robots are a promising solution to the increasing demand for surgery amid a shortage of surgeons. Recent work has proposed learning-based approaches for the autonomous manipulation of soft tissue. However, due to variability in aspects such as tissue geometries and stiffnesses, t…

Cited by 0SourceScholar
2024

Bayesian Constraint Inference from User Demonstrations Based on Margin-Respecting Preference Models

ICRA 2024poster

It is crucial for robots to be aware of the presence of constraints in order to acquire safe policies. However, explicitly specifying all constraints in an environment can be a challenging task. State-of-the-art constraint inference algorithms learn constraints from demonstrations, but tend to be co…

Cited by 3SourceScholar
2023

Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

ICLR 2023poster

Learning policies via preference-based reward learning is an increasingly popular method for customizing agent behavior, but has been shown anecdotally to be prone to spurious correlations and reward hacking behaviors. While much prior work focuses on causal confusion in reinforcement learning and b…

Cited by 59SourcePDFScholar
2023

Contextual Reliability: When Different Features Matter in Different Contexts

ICML 2023poster

Deep neural networks often fail catastrophically by relying on spurious correlations. Most prior work assumes a clear dichotomy into spurious and reliable features; however, this is often unrealistic. For example, most of the time we do not want an autonomous car to simply copy the speed of surround…

Cited by 3SourcePDFScholar
2023

Efficient Preference-Based Reinforcement Learning Using Learned Dynamics Models

ICRA 2023poster

Preference-based reinforcement learning (PbRL) can enable robots to learn to perform tasks based on an individual's preferences without requiring a hand-crafted re-ward function. However, existing approaches either assume access to a high-fidelity simulator or analytic model or take a model-free app…

Cited by 24SourceScholar
2023

Quantifying Assistive Robustness Via the Natural-Adversarial Frontier

CoRL 2023poster

Our ultimate goal is to build robust policies for robots that assist people. What makes this hard is that people can behave unexpectedly at test time, potentially interacting with the robot outside its training distribution and leading to failures. Even just measuring robustness is a challenge. Adve…

Cited by 0SourceScholar
2023

The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types

AAAI 2023technical

When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a "rationality coefficient" capturing how much noise or entropy we expect to see in the human beha…

Cited by 37SourcePDFScholar
2022

LEGS: Learning Efficient Grasp Sets for Exploratory Grasping

ICRA 2022poster

While deep learning has enabled significant progress in designing general purpose robot grasping systems, there remain objects which still pose challenges for these systems. Recent work on Exploratory Grasping has formalized the problem of systematically exploring grasps on these adversarial objects…

Cited by 14SourceScholar
2022

Learning Representations that Enable Generalization in Assistive Tasks

CoRL 2022poster

Recent work in sim2real has successfully enabled robots to act in physical environments by training in simulation with a diverse ``population'' of environments (i.e. domain randomization). In this work, we focus on enabling generalization in \emph{assistive tasks}: tasks in which the robot is acting…

Cited by 35SourceScholar
2022

Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal Demonstrations

NeurIPS 2022accept

Providing densely shaped reward functions for RL algorithms is often exceedingly challenging, motivating the development of RL algorithms that can learn from easier-to-specify sparse reward functions. This sparsity poses new exploration challenges. One common way to address this problem is using dem…

Cited by 26SourcePDFScholar
2022

Teaching Robots to Span the Space of Functional Expressive Motion

IROS 2022poster

Our goal is to enable robots to perform functional tasks in emotive ways, be it in response to their users' emotional states, or expressive of their confidence levels. Prior work has proposed learning independent cost functions from user feedback for each target emotion, so that the robot may optimi…

Cited by 14SourceScholar
2021

Dynamically Switching Human Prediction Models for Efficient Planning

ICRA 2021poster

As environments involving both robots and humans become increasingly common, so does the need to account for people during planning. To plan effectively, robots must be able to respond to and sometimes influence what humans do. This requires a human model which predicts future human actions. A simpl…

Cited by 9SourceScholar
2021

Policy Gradient Bayesian Robust Optimization for Imitation Learning

ICML 2021spotlight

The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the tru…

Cited by 26SourcePDFScholar
2021

Situational Confidence Assistance for Lifelong Shared Autonomy

ICRA 2021poster

Shared autonomy enables robots to infer user intent and assist in accomplishing it. But when the user wants to do a new task that the robot does not know about, shared autonomy will hinder their performance by attempting to assist them with something that is not their intent. Our key idea is that th…

Cited by 34SourceScholar
2021

ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning

CoRL 2021oral

Effective robot learning often requires online human feedback and interventions that can cost significant human time, giving rise to the central challenge in interactive imitation learning: is it possible to control the timing and length of interventions to both facilitate learning and limit burden…

Cited by 87SourceScholar
2019

Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations

CoRL 2019

The performance of imitation learning is typically upper-bounded by the performance of the demonstrator. While recent empirical results demonstrate that ranked demonstrations allow for better-than-demonstrator performance, preferences over demonstrations may be difficult to obtain, and little is kno