← Search

Shambhuraj Sawant

2 accepted papers

2021

Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers

ICLR 2021poster

Solving high-dimensional, continuous robotic tasks is a challenging optimization problem. Model-based methods that rely on zero-order optimizers like the cross-entropy method (CEM) have so far shown strong performance and are considered state-of-the-art in the model-based reinforcement learning comm…

Cited by 13SourcePDFScholar
2020

Sample-efficient Cross-Entropy Method for Real-time Planning

CoRL 2020

Trajectory optimizers for model-based reinforcement learning, such as the Cross-Entropy Method (CEM), can yield compelling results even in high-dimensional control tasks and sparse-reward environments. However, their sampling inefficiency prevents them from being used for real-time planning and cont