2017
Batch Policy Gradient Methods for Improving Neural Conversation Models
ICLR 2017poster
We study reinforcement learning of chat-bots with recurrent neural network architectures when the rewards are noisy and expensive to obtain. For instance, a chat-bot used in automated customer service support can be scored by quality assurance agents, but this process can be expensive, time consumin…