← Search

Haoran Gu

2 accepted papers

2026

Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing

AAAI 2026technical

This paper proposes a variant of multiple-play stochastic bandits tailored to resource allocation problems arising from LLM applications, edge intelligence, etc. The model is composed of finite number of arms and plays. Each arm has a stochastic number of capacities, and each unit of capacity is ass

Cited by 0SourcePDFScholar
2026

ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models Using Pareto High-Quality Data

AAAI 2026technical

Aligning large language models with multiple human expectations and values is crucial for ensuring that they adequately serve a variety of user needs. To this end, offline multiobjective alignment algorithms such as the Rewards-in-Context algorithm have shown strong performance and efficiency. Howev

Cited by 0SourcePDFScholar