← Search

Long Tran-Thanh

26 accepted papers

2026

BRIDGE: Bi-level Reinforcement Learning for Dynamic Group Structure in Coalition Formation Games

ICLR 2026poster

The challenge of coalition formation games lies in efficiently navigating the exponentially large space of possible coalitions to identify the optimal partition. While existing approaches to solve coalition formation games either provide optimal solutions with limited scalability or approximate solu…

Cited by 0SourceScholar
2026

Investigating the Role of Implicit Signals in Adaptive User-Aware Human-Robot Interactions

ICRA 2026poster

Our work investigates how social robots can act in a user-aware manner by adapting their behaviour to users' personal characteristics and preferences without unnecessarily exposing them to frustration through the robot's actions. In particular, we investigate how implicit social signals inadvertentl…

Cited by 0Scholar
2026

Pruning at Initialisation through the lens of Graphon Limit: Convergence, Expressivity, and Generalisation

ICML 2026poster

Pruning at Initialisation methods discover sparse, trainable subnetworks before training, but their theoretical mechanisms remain elusive. Existing analyses are often limited to finite-width statistics, lacking a rigorous characterisation of the global sparsity patterns that emerge as networks grow …

Cited by 0SourceScholar
2025

DPaI: Differentiable Pruning at Initialization with Node-Path Balance Principle

ICLR 2025poster

Pruning at Initialization (PaI) is a technique in neural network optimization characterized by the proactive elimination of weights before the network's training on designated tasks. This innovative strategy potentially reduces the costs for training and inference, significantly advancing computatio…

2025

Non-stochastic Budgeted Online Pricing with Semi-Bandit Feedback

AAAI 2025technical

We consider a general non-stochastic online pricing bandit setting in a procurement scenario where a buyer with a budget wants to procure items from a fixed set of sellers to maximize the buyer's reward by dynamically offering purchasing prices to the sellers, where the sellers' costs and values at…

Cited by 0SourcePDFScholar
2025

Provably Improving Generalization of Few-shot models with Synthetic Data

ICML 2025poster

Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between r…

Cited by 4SourcePDFScholar
2025

The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis

NeurIPS 2025spotlight

Sparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, understanding why some sparse structures are better trainable than others with the same level of sparsity remains poorly und…

Cited by 0SourceScholar
2025

User-Aware Collaborative Learning in Human-Robot Interactions

ICRA 2025

Our work investigates how social robots can efficiently collaborate with human users in a user-aware manner, minimising the generated frustration in human colleagues, thus enhancing their experience. As part of this, we develop a useraware framework for human-robot collaborative learning. We model u

Cited by 0SourceScholar
2024

Learning the Expected Core of Strictly Convex Stochastic Cooperative Games

NeurIPS 2024poster

Reward allocation, also known as the credit assignment problem, has been an important topic in economics, engineering, and machine learning. An important concept in reward allocation is the core, which is the set of stable allocations where no agent has the motivation to deviate from the grand coali…

2024

Symmetric Linear Bandits with Hidden Symmetry

NeurIPS 2024poster

High-dimensional linear bandits with low-dimensional structure have received considerable attention in recent studies due to their practical significance. The most common structure in the literature is sparsity. However, it may not be available in practice. Symmetry, where the reward is invariant un…

2023

Towards Data-Agnostic Pruning At Initialization: What Makes a Good Sparse Mask?

NeurIPS 2023poster

Pruning at initialization (PaI) aims to remove weights of neural networks before training in pursuit of training efficiency besides the inference. While off-the-shelf PaI methods manage to find trainable subnetworks that outperform random pruning, their performance in terms of both accuracy and com…

2022

Class Similarity Weighted Knowledge Distillation for Continual Semantic Segmentation

CVPR 2022poster

Deep learning models are known to suffer from the problem of catastrophic forgetting when they incrementally learn new classes. Continual learning for semantic segmentation (CSS) is an emerging field in computer vision. We identify a problem in CSS: A model tends to be confused between old and new c…

Cited by 65PDFScholar
2022

Expected Improvement for Contextual Bandits

NeurIPS 2022accept

The expected improvement (EI) is a popular technique to handle the tradeoff between exploration and exploitation under uncertainty. This technique has been widely used in Bayesian optimization but it is not applicable for the contextual bandit problem which is a generalization of the standard bandit…

Cited by 0SourcePDFScholar
2022

Saving Stochastic Bandits from Poisoning Attacks via Limited Data Verification

AAAI 2022technical

This paper studies bandit algorithms under data poisoning attacks in a bounded reward setting. We consider a strong attacker model in which the attacker can observe both the selected actions and their corresponding rewards, and can contaminate the rewards with additive noise. We show that any bandit…

Cited by 16SourcePDFScholar
2022

Sequential Vaccine Allocation with Delayed Feedback

IJCAI 2022poster

In this work we consider the problem of how to best allocate a limited supply of vaccines in the aftermath of an infectious disease outbreak by viewing the problem as a sequential game between a learner and an environment (specifically, a bandit problem). The difficulty of this problem lies in the f…

Cited by 0SourcePDFScholar
2022

Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning

IJCAI 2022poster

To understand the security threats to reinforcement learning (RL) algorithms, this paper studies poisoning attacks to manipulate any order-optimal learning algorithm towards a targeted policy in episodic RL and examines the potential damage of two natural types of poisoning attacks, i.e., the manip…

Cited by 23SourcePDFScholar
2020

Fighting Wildfires under Uncertainty - A Sequential Resource Allocation Approach

IJCAI 2020poster

Standard disaster response involves using drones (or helicopters) for reconnaissance and using people on the ground to mitigate the damage. In this paper, we look at the problem of wildfires and propose an efficient resource allocation strategy to cope with both dynamically changing environment and…

Cited by 0SourcePDFScholar
2020

Learning Optimal Temperature Region for Solving Mixed Integer Functional DCOPs

IJCAI 2020poster

Distributed Constraint Optimization Problems (DCOPs) are an important framework for modeling coordinated decision-making problems in multi-agent systems with a set of discrete variables. Later works have extended DCOPs to model problems with a set of continuous variables, named Functional DCOPs (F-D…

Cited by 0SourcePDFScholar
2020

To Ask or Not to Ask: A User Annoyance Aware Preference Elicitation Framework for Social Robots

IROS 2020poster

In this paper we investigate how social robots can efficiently gather user preferences without exceeding the allowed user annoyance threshold. To do so, we use a Gazebo based simulated office environment with a TIAGo Steel robot. We then formulate the user annoyance aware preference elicitation prob…

Cited by 6SourceScholar
2019

Manipulating a Learning Defender and Ways to Counteract

NeurIPS 2019poster

In Stackelberg security games when information about the attacker's payoffs is uncertain, algorithms have been proposed to learn the optimal defender commitment by interacting with the attacker and observing their best responses. In this paper, we show that, however, these algorithms can be easily m…

Cited by 22SourcePDFScholar
2019

Streaming Bayesian Inference for Crowdsourced Classification

NeurIPS 2019poster

A key challenge in crowdsourcing is inferring the ground truth from noisy and unreliable data. To do so, existing approaches rely on collecting redundant information from the crowd, and aggregating it with some probabilistic method. However, oftentimes such methods are computationally inefficient, a…

Cited by 5SourcePDFScholar
2015

Efficient Thompson Sampling for Online Matrix-Factorization Recommendation

NeurIPS 2015poster

Matrix factorization (MF) collaborative filtering is an effective and widely used method in recommendation systems. However, the problem of finding an optimal trade-off between exploration and exploitation (otherwise known as the bandit problem), a crucial problem in collaborative filtering from col…

Cited by 231SourcePDFScholar