← Search

Haoyuan Hu

9 accepted papers

2025

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

IROS 2025

Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel

Cited by 0SourcecodeScholar
2024

Dr3: Ask Large Language Models Not to Give Off-Topic Answers in Open Domain Multi-Hop Question Answering

COLING 2024main

Open Domain Multi-Hop Question Answering (ODMHQA) plays a crucial role in Natural Language Processing (NLP) by aiming to answer complex questions through multi-step reasoning over retrieved information from external knowledge sources. Recently, Large Language Models (LLMs) have demonstrated remarkab…

2023

An Online Algorithm for Chance Constrained Resource Allocation

ICASSP 2023accepted

This paper studies the online stochastic resource allocation problem (RAP) with chance constraints. The online RAP is a 0-1 integer linear programming problem where the resource consumption coefficients are revealed column by column along with the corresponding revenue coefficients. When a column is…

Cited by 0SourceScholar
2023

Nearly Optimal Competitive Ratio for Online Allocation Problems with Two-sided Resource Constraints and Finite Requests

ICML 2023poster

In this paper, we investigate the online allocation problem of maximizing the overall revenue subject to both lower and upper bound constraints. Compared to the extensively studied online problems with only resource upper bounds, the two-sided constraints affect the prospects of resource consumption…

Cited by 2SourcePDFScholar
2023

OFCOURSE: A Multi-Agent Reinforcement Learning Environment for Order Fulfillment

NeurIPS 2023poster

The dramatic growth of global e-commerce has led to a surge in demand for efficient and cost-effective order fulfillment which can increase customers' service levels and sellers' competitiveness. However, managing order fulfillment is challenging due to a series of interdependent online sequential d…

2023

Online Learning for Non-monotone DR-Submodular Maximization: From Full Information to Bandit Feedback

AISTATS 2023poster

In this paper, we revisit the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, which finds wide real-world applications in the domain of machine learning, economics, and operations research. At first, we present the Meta-MFW algorithm achieving a $1/e$…

Cited by 13SourcePDFScholar
2022

Stochastic Continuous Submodular Maximization: Boosting via Non-oblivious Function

ICML 2022spotlight

In this paper, we revisit Stochastic Continuous Submodular Maximization in both offline and online settings, which can benefit wide applications in machine learning and operations research areas. We present a boosting framework covering gradient ascent and online gradient ascent. The fundamental ing…

Cited by 22SourcePDFScholar
2019

Katalyst: Boosting Convex Katayusha for Non-Convex Problems with a Large Condition Number

ICML 2019oral

An important class of non-convex objectives that has wide applications in machine learning consists of a sum of $n$ smooth functions and a non-smooth convex function. Tremendous studies have been devoted to conquering these problems by leveraging one of the two types of variance reduction techniques…

Cited by 4SourcePDFScholar