Backprop-MPDM: Faster Risk-Aware Policy Evaluation Through Efficient Gradient Optimization
In Multi-Policy Decision-Making (MPDM), many computationally-expensive forward simulations are performed in order to predict the performance of a set of candidate policies. In risk-aware formulations of MPDM, only the worst outcomes affect the decision making process, and efficiently finding these i…