2024
Deep Reinforcement Learning with Hierarchical Action Exploration for Dialogue Generation
COLING 2024main
Traditionally, approximate dynamic programming is employed in dialogue generation with greedy policy improvement through action sampling, as the natural language action space is vast. However, this practice is inefficient for reinforcement learning (RL) due to the sparsity of eligible responses with…