← Search

Huanhuan Yang

5 accepted papers

2026

Learning Reward–Cost Balance in Safe RL via Score-Based World Models

ICML 2026poster

Safe reinforcement learning (Safe RL) seeks to optimize long-term performance while ensuring adherence to safety constraints. However, most existing approaches address safety in a simplified manner, typically by linearly combining rewards and costs, which provides limited guidance when safety and pe…

Cited by 0SourceScholar
2025

Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios

AAAI 2025technical

Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a nove…

Cited by 0SourcePDFScholar
2025

Multi-Agent Hierarchical Graph Attention Actor-Critic Reinforcement Learning

ICASSP 2025accepted

Multi-agent systems often face challenges such as elevated communication demands and intricate interactions. We propose an innovative hierarchical graph attention actor-critic reinforcement learning method to address the issues, which uses the hierarchical graph attention to capture the relationship…

Cited by 0SourceScholar
2024

Crowd Perception Communication-Based Multi- Agent Path Finding With Imitation Learning

RA-L 2024

Deep reinforcement learning-based Multi-Agent Path Finding (MAPF) has gained significant attention due to its remarkable adaptability to environments. Existing methods primarily leverage multi-agent communication in a fully-decentralized framework to maintain scalability while enhancing information

Cited by 2SourceScholar
2022

Self-supervised representations for multi-view reinforcement learning

UAI 2022poster

Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, suc…