← Search

Yangyang Zhao

7 accepted papers

2025

An Efficient Dialogue Policy Agent with Model-Based Causal Reinforcement Learning

COLING 2025main

Dialogue policy trains an agent to select dialogue actions frequently implemented via deep reinforcement learning (DRL). The model-based reinforcement methods built a world model to generate simulated data to alleviate the sample inefficiency. However, traditional world model methods merely consider…

Cited by 0SourcePDFScholar
2025

An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals

ACL 2025long

Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the high dimensionality of state and action spaces. This challenge often results in local optima or poor convergence. Evolut…

Cited by 0SourcePDFScholar
2025

Semantic-Aware Action Space Compression via LLM-DRL Synergy for Efficient Task-oriented Dialogue Policy Exploration

EMNLP 2025

The flexibility of natural language significantly expands the action space in task-oriented dialogue systems, causing inefficient exploration and slow convergence in deep reinforcement learning (DRL)-based policy optimization. Pre-trained large language models (LLMs), with world knowledge and semant

Cited by 0SourcePDFScholar
2024

Bootstrapped Policy Learning for Task-oriented Dialogue through Goal Shaping

EMNLP 2024main

Reinforcement learning shows promise in optimizing dialogue policies, but addressing the challenge of reward sparsity remains crucial. While curriculum learning offers a practical solution by strategically training policies from simple to complex, it hinges on the assumption of a gradual increase in…

2024

TransLoc4D: Transformer-based 4D Radar Place Recognition

CVPR 2024poster

Place Recognition is crucial for unmanned vehicles in terms of localization and mapping. Recent years have witnessed numerous explorations in the field where 2D cameras and 3D LiDARs are mostly employed. Despite their admirable performance they may encounter challenges in adverse weather such as rai…

2021

Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy Learning

AAAI 2021technical

Dialogue policy learning based on reinforcement learning is difficult to be applied to real users to train dialogue agents from scratch because of the high cost. User simulators, which choose random user goals for the dialogue agent to train on, have been considered as an affordable substitute for r…

Cited by 20SourcePDFScholar
2021

Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy

EMNLP 2021main

Deep reinforcement learning has shown great potential in training dialogue policies. However, its favorable performance comes at the cost of many rounds of interaction. Most of the existing dialogue policy methods rely on a single learning system, while the human brain has two specialized learning a…

Cited by 13SourcePDFScholar