2020
LAVA: Latent Action Spaces via Variational Auto-encoding for Dialogue Policy Optimization
COLING 2020main
Reinforcement learning (RL) can enable task-oriented dialogue systems to steer the conversation towards successful task completion. In an end-to-end setting, a response can be constructed in a word-level sequential decision making process with the entire system vocabulary as action space. Policies t…