ICRA 2026poster0 citations

Transformer-Based Hierarchical Reinforcement Learning for Sequential Decision-Making in Swarm Confrontation

Ruozhai Sun, Qizhen Wu, Lei Chen

Abstract

Hierarchical Reinforcement Learning (HRL) is a potent paradigm for addressing long-horizon sequential decision-making in swarm confrontation. However, its strategic capabilities are often bottlenecked by high-level policies that struggle to reason over the dynamic, variable-sized observations of other agents. To address this, we introduce a novel decentralized HRL framework featuring a Transformer-based strategic policy. The Transformer's self-attention mechanism is uniquely suited to capture complex spatio-temporal relationships among a varying number of entities, enabling robust long-horizon task allocation. This high-level strategy is then translated by a low-level policy into collision-free navigation. In complex swarm confrontation scenarios, our method significantly outperforms established baselines, achieving win rates of up to 81%. Beyond this performance, the learned policies exhibit strong zero-shot generalization to larger swarms, offer decision-making interpretability via the attention mechanism, and foster the autonomous emergence of complex cooperative tactics. This work provides a blueprint for scalable, strategically sophisticated, and interpretable multi-agent systems.

Reinforcement LearningMulti-Robot SystemsTask and Motion Planning
Transformer-Based Hierarchical Reinforcement Learning for Sequential Decision-Making in Swarm Confrontation · ICRA 2026