← Search

Hao-Lun Hsu

7 accepted papers

2026

NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces

ICML 2026poster

Generative modeling of neural network parameters is often tied to architectures because standard parameter representations rely on known weight-matrix dimensions. Generation is further complicated by permutation symmetries that allow networks to model similar input-output functions while having wide…

Cited by 0SourceScholar
2025

Variational Adversarial Training Towards Policies with Improved Robustness

AISTATS 2025poster

Reinforcement learning (RL), while being the benchmark for policy formulation, often struggles to deliver robust solutions across varying scenarios, leading to marked performance drops under environmental perturbations.~Traditional adversarial training, based on a two-player max-min game, is known t…

Cited by 0SourceScholar
2024

Finite-Time Frequentist Regret Bounds of Multi-Agent Thompson Sampling on Sparse Hypergraphs

AAAI 2024technical

We study the multi-agent multi-armed bandit (MAMAB) problem, where agents are factored into overlapping groups. Each group represents a hyperedge, forming a hypergraph over the agents. At each round of interaction, the learner pulls a joint arm (composed of individual arms for each agent) and receiv…

2024

REFORMA: Robust REinFORceMent Learning via Adaptive Adversary for Drones Flying under Disturbances

ICRA 2024poster

In this work, we introduce REFORMA, a novel robust reinforcement learning (RL) approach to design controllers for unmanned aerial vehicles (UAVs) robust to unknown disturbances during flights. These disturbances, typically due to wind turbulence, electromagnetic interference, temperature extremes an…

Cited by 6SourceScholar
2024

Randomized Exploration in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2024poster

We present the first study on provably efficient randomized exploration in cooperative multi-agent reinforcement learning (MARL). We propose a unified algorithm framework for randomized exploration in parallel Markov Decision Processes (MDPs), and two Thompson Sampling (TS)-type algorithms, CoopTS-P…

Cited by 7SourcePDFScholar
2024

Steering Decision Transformers via Temporal Difference Learning

IROS 2024poster

Decision Transformers (DTs) have been highly effective for offline reinforcement learning (RL) tasks, successfully modeling the sequences of actions in a given set of demonstrations. However, DTs may perform poorly in stochastic environments, which are prevalent in robotics scenarios. In this paper,…

Cited by 0SourceScholar
2022

Improving Safety in Deep Reinforcement Learning using Unsupervised Action Planning

ICRA 2022poster

One of the key challenges to deep reinforcement learning (deep RL) is to ensure safety at both training and testing phases. In this work, we propose a novel technique of unsupervised action planning to improve the safety of on-policy reinforcement learning algorithms, such as trust region policy opt…

Cited by 19SourceScholar