← Search

Siyi Hu

8 accepted papers

2026

SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models

ICLR 2026poster

Language models have shown strong performance on mathematical reasoning tasks. Post-training with outcome-based reinforcement learning (RL) can further enhance reasoning but is inefficient because it relies solely on final rewards. Recent credit assignment–based RL methods provide intermediate feedb…

Cited by 0SourceScholar
2024

Maximum Entropy Heterogeneous-Agent Reinforcement Learning

ICLR 2024spotlight

*Multi-agent reinforcement learning* (MARL) has been shown effective for cooperative games in recent years. However, existing state-of-the-art methods face challenges related to sample complexity, training instability, and the risk of converging to a suboptimal Nash Equilibrium. In this paper, we pr…

Cited by 18SourcePDFScholar
2024

ProAgent: Building Proactive Cooperative Agents with Large Language Models

AAAI 2024technical

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they int…

2022

Policy Diagnosis via Measuring Role Diversity in Cooperative Multi-agent RL

ICML 2022spotlight

Cooperative multi-agent reinforcement learning (MARL) is making rapid progress for solving tasks in a grid world and real-world scenarios, in which agents are given different attributes and goals, resulting in different behavior through the whole multi-agent task. In this study, we quantify the agen…

Cited by 34SourcePDFScholar
2021

UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers

ICLR 2021spotlight

Recent advances in multi-agent reinforcement learning have been largely limited in training one model from scratch for every new task. The limitation is due to the restricted model architecture related to fixed input and output dimensions. This hinders the experience accumulation and transfer of the…

Cited by 0SourcePDFScholar
2019

Modeling Perceptual Aliasing in SLAM via Discrete-Continuous Graphical Models

RA-L 2019

Perceptual aliasing is one of the main causes of the failure for simultaneous localization and mapping (SLAM) systems operating in the wild. Perceptual aliasing is a phenomenon where different places generate a similar visual (or, in general, perceptual) footprint. This causes spurious measurements

Cited by 95SourceScholar