← Search

Xingzhou Lou

5 accepted papers

2025

Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward Modeling

AAAI 2025technical

Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessness) or struggle with the complexity of managing multiple reward models. To addre…

2025

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

NeurIPS 2025poster

Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically employ reasoning chains at a coarse-grained level, which struggles to perform fi…

Cited by 0SourcecodeScholar
2024

Position: Foundation Agents as the Paradigm Shift for Decision Making

ICML 2024poster

Decision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor generalization. In contrast, foundation models in language and vision have showcased…

2024

TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy Gradient

AAAI 2024technical

Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. Whi…

2023

An Efficient End-to-End Training Approach for Zero-Shot Human-AI Coordination

NeurIPS 2023poster

The goal of zero-shot human-AI coordination is to develop an agent that can collaborate with humans without relying on human data. Prevailing two-stage population-based methods require a diverse population of mutually distinct policies to simulate diverse human behaviors. The necessity of such popul…

Cited by 13SourcePDFScholar