← Search

David Carter

1 accepted papers

2017

Batch Policy Gradient Methods for Improving Neural Conversation Models

ICLR 2017poster

We study reinforcement learning of chat-bots with recurrent neural network architectures when the rewards are noisy and expensive to obtain. For instance, a chat-bot used in automated customer service support can be scored by quality assurance agents, but this process can be expensive, time consumin…

Cited by 39SourceScholar