← Search

Supratik Paul

6 accepted papers

2024

Rate-Informed Discovery via Bayesian Adaptive Multifidelity Sampling

CoRL 2024poster

Ensuring the safety of autonomous vehicles (AVs) requires both accurate estimation of their performance and efficient discovery of potential failure cases. This paper introduces Bayesian adaptive multifidelity sampling (BAMS), which leverages the power of adaptive Bayesian sampling to achieve effici…

Cited by 0SourceScholar
2022

Embedding Synthetic Off-Policy Experience for Autonomous Driving via Zero-Shot Curricula

CoRL 2022oral

ML-based motion planning is a promising approach to produce agents that exhibit complex behaviors, and automatically adapt to novel environments. In the context of autonomous driving, it is common to treat all available training data equally. However, this approach produces agents that do not perfor…

Cited by 21SourceScholar
2022

Hierarchical Model-Based Imitation Learning for Planning in Autonomous Driving

IROS 2022poster

We demonstrate the first large-scale application of model-based generative adversarial imitation learning (MGAIL) to the task of dense urban self-driving. We augment standard MGAIL using a hierarchical model to enable generalization to arbitrary goal routes, and measure performance using a closed-lo…

Cited by 60SourceScholar
2019

Fast Efficient Hyperparameter Tuning for Policy Gradient Methods

NeurIPS 2019poster

The performance of policy gradient methods is sensitive to hyperparameter settings that must be tuned for any new application. Widely used grid search methods for tuning hyperparameters are sample inefficient and computationally expensive. More advanced methods like Population Based Training that le…

2019

Fingerprint Policy Optimisation for Robust Reinforcement Learning

ICML 2019oral

Policy gradient methods ignore the potential value of adjusting environment variables: unobservable state features that are randomly determined by the environment in a physical setting, but are controllable in a simulator. This can lead to slow learning, or convergence to suboptimal policies, if the…

Cited by 27SourcePDFScholar
2019

Learning From Demonstration in the Wild

ICRA 2019poster

Learning from demonstration (LfD) is useful in settings where hand-coding behaviour or a reward function is impractical. It has succeeded in a wide range of problems but typically relies on manually generated demonstrations or specially deployed sensors and has not generally been able to leverage th…

Cited by 80SourceScholar