← Search

Thanh Hong Nguyen

8 accepted papers

2026

CTPD: Cross Tokenizer Preference Distillation

AAAI 2026technical

While knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplored, particularly in the more realistic cross-tokenizer setting. The incompatibility of tokenization schemes between teach

Cited by 4SourcePDFScholar
2025

ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization

ICLR 2025poster

Offline reinforcement learning (RL) has garnered significant attention for its ability to learn effective policies from pre-collected datasets without the need for further environmental interactions. While promising results have been demonstrated in single-agent settings, offline multi-agent reinfor…

Cited by 0SourcePDFScholar
2024

Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation Learning

NeurIPS 2024poster

This paper concerns imitation learning (IL) in cooperative multi-agent systems. The learning problem under consideration poses several challenges, characterized by high-dimensional state and action spaces and intricate inter-agent dependencies. In a single-agent setting, IL was shown to be done effi…

Cited by 1SourcePDFScholar
2024

Mimicking To Dominate: Imitation Learning Strategies for Success in Multiagent Games

NeurIPS 2024poster

Training agents in multi-agent games presents significant challenges due to their intricate nature. These challenges are exacerbated by dynamics influenced not only by the environment but also by strategies of opponents. Existing methods often struggle with slow convergence and instability. To addre…

Cited by 0SourcePDFScholar
2023

Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement Learning

NeurIPS 2023poster

Many problems in Reinforcement Learning (RL) seek an optimal policy with large discrete multidimensional yet unordered action spaces; these include problems in randomized allocation of resources such as placements of multiple security resources and emergency response units, etc. A challenge in this…

2022

Information design for multiple independent and self-interested defenders: Work less, pay off more

UAI 2022poster

This paper studies the problem of information design in a general security game setting in which multiple independent self-interested defenders attempt to provide protection simultaneously on the same set of important targets against an unknown attacker. A principal, who can be one of the defenders,…

Cited by 0SourcePDFScholar