← Search

J. Andrew Bagnell

11 accepted papers

2024

REBEL: Reinforcement Learning via Regressing Relative Rewards

NeurIPS 2024poster

While originally developed for continuous control problems, Proximal Policy Optimization (PPO) has emerged as the work-horse of a variety of reinforcement learning (RL) applications, including the fine-tuning of generative models. Unfortunately, PPO requires multiple heuristics to enable stable conv…

2021

CMAX++ : Leveraging Experience in Planning and Execution using Inaccurate Models

AAAI 2021technical

Given access to accurate dynamical models, modern planning approaches are effective in computing feasible and optimal plans for repetitive robotic tasks. However, it is difficult to model the true dynamics of the real world before execution, especially for tasks requiring interactions with objects w…

2021

Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation Gap

ICML 2021spotlight

We provide a unifying view of a large family of previous imitation learning algorithms through the lens of moment matching. At its core, our classification scheme is based on whether the learner attempts to match (1) reward or (2) action-value moments of the expert’s behavior, with each option leadi…

2018

TRUNCATED HORIZON POLICY SEARCH: COMBINING REINFORCEMENT LEARNING & IMITATION LEARNING

ICLR 2018poster

In this paper, we propose to combine imitation and reinforcement learning via the idea of reward shaping using an oracle. We study the effectiveness of the near- optimal cost-to-go oracle on the planning horizon and demonstrate that the cost- to-go oracle shortens the learner’s planning horizon as f…

Cited by 112SourcePDFScholar
2017

A Probabilistic Planning Framework for Planar Grasping Under Uncertainty

RA-L 2017

How can a robot design a sequence of grasping actions that will succeed despite the presence of bounded state uncertainty and an inherently stochastic system? In this letter, we propose a probabilistic algorithm that generates sequential actions to iteratively reduce uncertainty until object pose is

Cited by 39SourceScholar
2017

Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction

ICML 2017poster

Recently, researchers have demonstrated state-of-the-art performance on sequential prediction problems using deep neural networks and Reinforcement Learning (RL). For some of these problems, oracles that can demonstrate good performance may be available during training, but are not used by plain RL…

Cited by 299SourcePDFScholar
2016

A convex polynomial force-motion model for planar sliding: Identification and application

ICRA 2016poster

We propose a polynomial force-motion model for planar sliding. The set of generalized friction loads is the 1-sublevel set of a polynomial whose gradient directions correspond to generalized velocities. Additionally, the polynomial is confined to be convex even-degree homogeneous in order to obey th…

Cited by 112SourceScholar
2016

Introspective perception: Learning to predict failures in vision systems

IROS 2016poster

As robots aspire for long-term autonomous operations in complex dynamic environments, the ability to reliably take mission-critical decisions in ambiguous situations becomes critical. This motivates the need to build systems that have situational awareness to assess how quali ed they are at that mom…

Cited by 105SourceScholar
2015

Predicting Multiple Structured Visual Interpretations

ICCV 2015poster

We present a simple approach for producing a small number of structured visual outputs which have high recall, for a variety of tasks including monocular pose estimation and semantic scene segmentation. Current state-of-the-art approaches learn a single model and modify inference procedures to produ…

Cited by 36PDFScholar
2015

Visual chunking: A list prediction framework for region-based object detection

ICRA 2015poster

We consider detecting objects in an image by iteratively selecting from a set of arbitrarily shaped candidate regions. Our generic approach, which we term visual chunking, reasons about the locations of multiple object instances in an image while expressively describing object boundaries. We design…

Cited by 5SourceScholar