← Search

Anusha Nagabandi

14 accepted papers

2026

RFS: Reinforcement learning with Residual flow steering for dexterous manipulation

ICLR 2026poster

Imitation learning has been an effective tool for bootstrapping sequential decision making behavior, showing surprisingly strong results as methods are scaled up to high-dimensional, dexterous problems in robotics. These ``behavior cloning" methods have been further bolstered by the integration of g…

Cited by 0SourcecodeScholar
2026

Residual Off-Policy RL for Finetuning Behavior Cloning Policies

ICRA 2026poster

Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement lear…

2026

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

RSS 2026poster

Efficient exploration remains a bottleneck in reinforcement learning (RL), particularly for long-horizon, high-dimensional tasks. While recent methods leverage pre-trained policies for guidance, they are often constrained by the base policy’s original behavior distribution. We introduce Timestep Mod…

Cited by 0SourceScholar
2025

Steering Your Diffusion Policy with Latent Space Reinforcement Learning

CoRL 2025oral

Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications. However, in scenarios where initial performance is not satisfactory, as is often the case in novel open-world settings, such behavioral cloning (BC)-learned policies typically…

Cited by 0SourceScholar
2021

Model-Based Reinforcement Learning via Latent-Space Collocation

ICML 2021spotlight

The ability to plan into the future while utilizing only raw high-dimensional observations, such as images, can provide autonomous agents with broad and general capabilities. However, realistic tasks require performing temporally extended reasoning, and cannot be solved with only myopic, short-sight…

2020

MELD: Meta-Reinforcement Learning from Images via Latent State Models

CoRL 2020

Meta-reinforcement learning algorithms can enable autonomous agents, such as robots, to quickly acquire new behaviors by leveraging prior experience in a set of related training tasks. However, the onerous data requirements of meta-training compounded with the challenge of learning from sensory inpu

2020

Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model

NeurIPS 2020poster

Deep reinforcement learning (RL) algorithms can use high-capacity deep networks to learn directly from image observations. However, these high-dimensional observation spaces present a number of challenges in practice, since the policy must now solve two problems: representation learning and task le…

Cited by 484SourcePDFScholar
2019

Deep Online Learning Via Meta-Learning: Continual Adaptation for Model-Based RL

ICLR 2019poster

Humans and animals can learn complex predictive models that allow them to accurately and reliably reason about real-world phenomena, and they can adapt such models extremely quickly in the face of unexpected changes. Deep neural network models allow us to represent very complex functions, but lack t…

Cited by 236SourcePDFScholar
2019

Learning to Adapt in Dynamic, Real-World Environments through Meta-Reinforcement Learning

ICLR 2019poster

Although reinforcement learning methods can achieve impressive results in simulation, the real world presents two major challenges: generating samples is exceedingly expensive, and unexpected perturbations or unseen situations cause proficient but specialized policies to fail at test time. Given tha…

Cited by 740SourcePDFScholar
2018

Learning Image-Conditioned Dynamics Models for Control of Underactuated Legged Millirobots

IROS 2018poster

Millirobots are a promising robotic platform for many applications due to their small size and low manufacturing costs. Legged millirobots, in particular, can provide increased mobility in complex environments and improved scaling of obstacles. However, controlling these small, highly dynamic, and u…

Cited by 34SourceScholar
2018

Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning

ICRA 2018poster

Model-free deep reinforcement learning algorithms have been shown to be capable of learning a wide range of robotic skills, but typically require a very large number of samples to achieve good performance. Model-based algorithms, in principle, can provide for much more efficient learning, but have p…

Cited by 1376SourceScholar
2017

Cooperative inchworm localization with a low cost team

ICRA 2017poster

In this paper we address the problem of multi-robot localization with a heterogeneous team of low-cost mobile robots. The team consists of a single centralized observer with an inertial measurement unit (IMU) and monocular camera, and multiple picket robots with only IMUs and Red Green Blue (RGB) li…

Cited by 15SourceScholar
2016

A path planning algorithm for single-ended continuous planar robotic ribbon folding

IROS 2016poster

Ribbon folding is a new approach to structure formation that forms higher dimensional structures using a lower dimensional primitive, namely a ribbon. In this paper, we present a novel algorithm to address path planning for ribbon folding of multi-link planar structures. We first represent the desir…

Cited by 1SourceScholar