← Search

Xiaowei Chen

5 accepted papers

2026

Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations

ICLR 2026poster

Large Language Models (LLMs) have achieved impressive performance on complex tasks by generating human-like, step-by-step rationales, referred to as \textit{reasoning trajectory}, before arriving at final answers. However, the length of these reasoning trajectories often far exceeds that of the fina…

Cited by 0SourcecodeScholar
2025

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

COLING 2025main

Aligning Large Language Models (LLMs) with human feedback is crucial for their development. Existing preference optimization methods such as DPO and KTO, while improved based on Reinforcement Learning from Human Feedback (RLHF), are inherently derived from PPO, requiring a reference model that adds…

Cited by 0SourcePDFScholar
2024

SuDA: Support-based Domain Adaptation for Sim2Real Hinge Joint Tracking with Flexible Sensors

ICML 2024poster

Flexible sensors hold promise for human motion capture (MoCap), offering advantages such as wearability, privacy preservation, and minimal constraints on natural movement. However, existing flexible sensor-based MoCap methods rely on deep learning and necessitate large and diverse labeled datasets f…

Cited by 1SourcePDFScholar
2021

Multi-layered Network Exploration via Random Walks: From Offline Optimization to Online Learning

ICML 2021oral

Multi-layered network exploration (MuLaNE) problem is an important problem abstracted from many applications. In MuLaNE, there are multiple network layers where each node has an importance weight and each layer is explored by a random walk. The MuLaNE task is to allocate total random walk budget $B$…

Cited by 18SourcePDFScholar
2018

Community Exploration: From Offline Optimization to Online Learning

NeurIPS 2018poster

We introduce the community exploration problem that has various real-world applications such as online advertising. In the problem, an explorer allocates limited budget to explore communities so as to maximize the number of members he could meet. We provide a systematic study of the community explor…

Cited by 7SourcePDFScholar