← Search

Yitian Hong

2 accepted papers

2026

HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement Learning

AAAI 2026technical

In cooperative Multi-Agent Reinforcement Learning (MARL), efficient exploration is crucial for optimizing the performance of joint policy. However, existing methods often update joint policies via independent agent exploration, without coordination among agents, which inherently constrains the expre

Cited by 0SourcePDFScholar
2022

Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2022accept

In cooperative multi-agent reinforcement learning, centralized training and decentralized execution (CTDE) has achieved remarkable success. Individual Global Max (IGM) decomposition, which is an important element of CTDE, measures the consistency between local and joint policies. The majority of IGM…

Cited by 34SourcePDFScholar