← Search

Rituraj Kaushik

5 accepted papers

2023

Imitation-Guided Multimodal Policy Generation from Behaviourally Diverse Demonstrations

IROS 2023poster

Learning policies from multiple demonstrators is often difficult because different individuals perform the same task differently due to hidden factors such as preferences. In the context of policy learning, this leads to multimodal policies. Existing policy learning methods often converge to a singl…

Cited by 0SourceScholar
2022

SafeAPT: Safe Simulation-to-Real Robot Learning Using Diverse Policies Learned in Simulation

RA-L 2022

The framework of sim-to-real learning, i.e., training policies in simulation and transferring them to real-world systems, is one of the most promising approaches towards data-efficient learning in robotics. However, due to the inevitable reality gap between the simulation and the real world, a polic

Cited by 13SourcecodeScholar
2020

Fast Online Adaptation in Robotics through Meta-Learning Embeddings of Simulated Priors

IROS 2020poster

Meta-learning algorithms can accelerate the model-based reinforcement learning (MBRL) algorithms by finding an initial set of parameters for the dynamical model such that the model can be trained to match the actual dynamics of the system with only a few data-points. However, in the real world, a ro…

Cited by 73SourcecodeScholar
2018

Multi-objective Model-based Policy Search for Data-efficient Learning with Sparse Rewards

CoRL 2018

The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. However, the current algo

2017

Black-box data-efficient policy search for robotics

IROS 2017poster

The most data-efficient algorithms for reinforcement learning (RL) in robotics are based on uncertain dynamical models: after each episode, they first learn a dynamical model of the robot, then they use an optimization algorithm to find a policy that maximizes the expected return given the model and…

Cited by 144SourcecodeScholar