← Search

Zhixuan Fang

14 accepted papers

2026

Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMs

ICML 2026poster

The rapid advancement of large language models (LLMs) has led to remarkable performance across diverse domains, making them indispensable assistants in daily life and work. Currently, LLM services are primarily accessed in two ways: (i) paid access to cloud-hosted LLMs, which are powerful but introd…

Cited by 0SourceScholar
2026

The Pareto-optimal Trade-off between Regret and Statistical Inference in Linear Stochastic Bandits under Safety Constraints

ICML 2026poster

Linear bandits traditionally prioritize regret minimization, often overlooking statistical inference of the underlying parameter as a critical objective. In high-stakes settings such as healthcare, precise parameter estimation is indispensable, as it provides fundamental insights into system mechani…

Cited by 0SourceScholar
2025

Linear Streaming Bandit: Regret Minimization and Fixed-Budget Epsilon-Best Arm Identification

AAAI 2025technical

Recently, there has been a focus on the streaming setting in a line of works on the Multi-Armed Bandit (MAB). In this scenario, a large number of arms arrive in a streaming manner, and the algorithm scans through the stream and stores some arms in its limited processing memory. We advance this line…

Cited by 0SourcePDFScholar
2025

User-side Model Consistency Monitoring for Open Source Large Language Models Inference Services

ACL 2025long

With the continuous advancement in the performance of open-source large language models (LLMs), their inference services have attracted a substantial user base by offering quality comparable to closed-source models at a significantly lower cost. However, it has also given rise to trust issues regard…

2024

Decentralized Scheduling with QoS Constraints: Achieving O(1) QoS Regret of Multi-Player Bandits

AAAI 2024technical

We consider a decentralized multi-player multi-armed bandit (MP-MAB) problem where players cannot observe the actions and rewards of other players and no explicit communication or coordination between players is possible. Prior studies mostly focus on maximizing the sum of rewards of the players o…

Cited by 3SourcePDFScholar
2024

RL-CFR: Improving Action Abstraction for Imperfect Information Extensive-Form Games with Reinforcement Learning

ICML 2024poster

Effective action abstraction is crucial in tackling challenges associated with large action spaces in Imperfect Information Extensive-Form Games (IIEFGs). However, due to the vast state space and computational complexity in IIEFGs, existing methods often rely on fixed abstractions, resulting in sub-…

Cited by 1SourcePDFScholar
2024

The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation

ACL 2024long

Large language models (LLMs) encapsulate vast amounts of knowledge but still remain vulnerable to external misinformation. Existing research mainly studied this susceptibility behavior in a single-turn setting. However, belief can change during a multi-turn conversation, especially a persuasive one.…

Cited by 61SourcePDFScholar
2022

Combinatorial Bandits with Linear Constraints: Beyond Knapsacks and Fairness

NeurIPS 2022accept

This paper proposes and studies for the first time the problem of combinatorial multi-armed bandits with linear long-term constraints. Our model generalizes and unifies several prominent lines of work, including bandits with fairness constraints, bandits with knapsacks (BwK), etc. We propose an upp…

Cited by 25SourcePDFScholar