← Search

Haitong Ma

9 accepted papers

2026

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces

ICML 2026poster

Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a novel framework for training discrete diffusion models as highly effective policies in these complex settings. Our key innovation is an efficient online tr…

Cited by 0SourceScholar
2025

Learning Preferences without Interaction for Cooperative AI: A Hybrid Offline-Online Approach

NeurIPS 2025poster

Reinforcement learning (RL) for collaborative agents capable of cooperating with humans to accomplish tasks has long been a central goal in the RL community. While prior approaches have made progress in adapting collaborative agents to diverse human partners, they often focus solely on optimizing ta…

Cited by 0SourceScholar
2025

Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations

IROS 2025

Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducing a pre-training stage that learns dynamics representations, derived from factorizations of the transition dynamics. We

Cited by 4SourceScholar
2024

Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based Viewpoint

IROS 2024poster

We study sim-to-real skill transfer and discovery in the context of robotics control using representation learning. We draw inspiration from spectral decomposition of Markov decision processes. The spectral decomposition brings about representation that can linearly represent the state-action value…

Cited by 2SourceScholar
2024

Synthesize Efficient Safety Certificates for Learning-Based Safe Control using Magnitude Regularization

ICRA 2024poster

Safety certificates based on energy functions can provide demonstrable safety for complex robotic systems. However, all recent studies on learning-based energy function synthesis only consider the feasibility of the control policy, which might cause over-conservativeness and even fail to achieve the…

Cited by 2SourceScholar
2023

Gaussian Max-Value Entropy Search for Multi-Agent Bayesian Optimization

IROS 2023poster

We study the multi-agent Bayesian optimization (BO) problem, where multiple agents maximize a black-box function via iterative queries. We focus on Entropy Search (ES), a sample-efficient BO algorithm that selects queries to maximize the mutual information about the maximum of the black-box function…

Cited by 15SourcecodeScholar
2021

Model-based Constrained Reinforcement Learning using Generalized Control Barrier Function

IROS 2021poster

Model information can be used to predict future trajectories, so it has huge potential to avoid dangerous regions when applying reinforcement learning (RL) on real-world tasks, like autonomous driving. However, existing studies mostly use model-free constrained RL, which causes inevitable constraint…

Cited by 86SourcecodeScholar