← Search

Surya Bhupatiraju

2 accepted papers

2018

The Mirage of Action-Dependent Baselines in Reinforcement Learning

ICML 2018oral

Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces…

Cited by 164SourcePDFScholar
2017

RobustFill: Neural Program Learning under Noisy I/O

ICML 2017poster

The problem of automatically generating a computer program from some specification has been studied since the early days of AI. Recently, two competing approaches for `automatic program learning’ have received significant attention: (1) `neural program synthesis’, where a neural network is condition…

Cited by 483SourcePDFScholar