← Search

Tien Anh Mai

9 accepted papers

2025

ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization

ICLR 2025poster

Offline reinforcement learning (RL) has garnered significant attention for its ability to learn effective policies from pre-collected datasets without the need for further environmental interactions. While promising results have been demonstrated in single-agent settings, offline multi-agent reinfor…

Cited by 0SourcePDFScholar
2025

IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy Matching

NeurIPS 2025poster

Offline Learning from Observations (LfO) focuses on enabling agents to imitate expert behavior using datasets that contain only expert state trajectories and separate transition data with suboptimal actions. This setting is both practical and critical in real-world scenarios where direct environment…

Cited by 0SourceScholar
2024

Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation Learning

NeurIPS 2024poster

This paper concerns imitation learning (IL) in cooperative multi-agent systems. The learning problem under consideration poses several challenges, characterized by high-dimensional state and action spaces and intricate inter-agent dependencies. In a single-agent setting, IL was shown to be done effi…

Cited by 1SourcePDFScholar
2024

Mimicking To Dominate: Imitation Learning Strategies for Success in Multiagent Games

NeurIPS 2024poster

Training agents in multi-agent games presents significant challenges due to their intricate nature. These challenges are exacerbated by dynamics influenced not only by the environment but also by strategies of opponents. Existing methods often struggle with slow convergence and instability. To addre…

Cited by 0SourcePDFScholar
2024

SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning

NeurIPS 2024poster

We focus on offline imitation learning (IL), which aims to mimic an expert's behavior using demonstrations without any interaction with the environment. One of the main challenges in offline IL is the limited support of expert demonstrations, which typically cover only a small fraction of the state-…

2022

Scalable Distributional Robustness in a Class of Non-Convex Optimization with Guarantees

NeurIPS 2022accept

Distributionally robust optimization (DRO) has shown a lot of promise in providing robustness in learning as well as sample-based optimization problems. We endeavor to provide DRO solutions for a class of sum of fractionals, non-convex optimization which is used for decision making in prominent area…

Cited by 2SourcePDFScholar