← Search

Hong Xie

13 accepted papers

2026

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

ICML 2026poster

Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges. Systematic studies reveal pervasive feature splitting that fragments coherent concept…

Cited by 0SourceScholar
2026

GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series data

ICLR 2026poster

Despite recent progress in time-series foundation models, challenges persist in improving representation learning and adapting to diverse downstream tasks. We introduce a General Time-series Model (GTM), which advances representation learning via a novel frequency-domain attention mechanism that cap…

Cited by 0SourcecodeScholar
2026

Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing

AAAI 2026technical

This paper proposes a variant of multiple-play stochastic bandits tailored to resource allocation problems arising from LLM applications, edge intelligence, etc. The model is composed of finite number of arms and plays. Each arm has a stochastic number of capacities, and each unit of capacity is ass

Cited by 0SourcePDFScholar
2026

Revisiting Fairness-aware Interactive Recommendation: Item Lifecycle as a Control Knob

AAAI 2026technical

This paper revisits fairness-aware interactive recommendation (e.g., TikTok, KuaiShou) by introducing a novel control knob, i.e., the lifecycle of items. We make threefold contributions. First, we conduct a comprehensive empirical analysis and uncover that item lifecycles in short-video platforms fo

Cited by 0SourcePDFScholar
2025

Exploring the Choice Behavior of Large Language Models

ACL 2025finding

Large Language Models (LLMs) are increasingly deployed as human assistants across various domains where they help to make choices. However, the mechanisms behind LLMs’ choice behavior remain unclear, posing risks in safety-critical situations. Inspired by the intrinsic and extrinsic motivation frame…

Cited by 0SourcePDFScholar
2024

Adaptive Order Q-learning

IJCAI 2024poster

This paper revisits the estimation bias control problem of Q-learning, motivated by the fact that the estimation bias is not always evil, i.e., some environments benefit from overestimation bias or underestimation bias, while others suffer from these biases. Different from previous coarse-grained b…

Cited by 0SourcePDFScholar
2024

Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users

AAAI 2024technical

We study the problem of federated contextual combinatorial cascading bandits, where agents collaborate under the coordination of a central server to provide tailored recommendations to users. Existing works consider either a synchronous framework, necessitating full agent participation and global sy…

Cited by 6SourcePDFScholar
2024

Understanding and Patching Compositional Reasoning in LLMs

ACL 2024findings

LLMs have marked a revolutonary shift, yet they falter when faced with compositional reasoning tasks. Our research embarks on a quest to uncover the root causes of compositional reasoning failures of LLMs, uncovering that most of them stem from the improperly generated or leveraged implicit reasonin…

2023

Uncertainty-Aware Instance Reweighting for Off-Policy Learning

NeurIPS 2023poster

Off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has shown importance in various important real-world applications, such as search engines and recommender systems. While the ground-truth logging policy is usually unknown, previous work…

Cited by 7SourcePDFScholar
2022

Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications

IJCAI 2022poster

Multi-player multi-armed bandits (MMAB) study how decentralized players cooperatively play the same multi-armed bandit so as to maximize their total cumulative rewards. Existing MMAB models mostly assume when more than one player pulls the same arm, they either have a collision and obtain zero rewar…

Cited by 10SourcePDFScholar
2022

Multiple-Play Stochastic Bandits with Shareable Finite-Capacity Arms

ICML 2022spotlight

We generalize the multiple-play multi-armed bandits (MP-MAB) problem with a shareable arms setting, in which several plays can share the same arm. Furthermore, each shareable arm has a finite reward capacity and a “per-load” reward distribution, both of which are unknown to the learner. The reward f…

Cited by 10SourcePDFScholar