2022
[CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue
ACL 2022long
The recent success of reinforcement learning (RL) in solving complex tasks is often attributed to its capacity to explore and exploit an environment. Sample efficiency is usually not an issue for tasks with cheap simulators to sample data online. On the other hand, Task-oriented Dialogues (ToD) are…