← Search

Hengquan Guo

7 accepted papers

2026

Towards Safe and Optimal Online Bidding: A Modular Look-ahead Lyapunov Framework

ICLR 2026poster

This paper studies online bidding subject to simultaneous budget and return-on-investment (ROI) constraints, which encodes the goal of balancing high volume and profitability. We formulate the problem as a general constrained online learning problem that can be applied to diverse bidding settings (e…

Cited by 0SourceScholar
2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

NeurIPS 2025poster

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov D…

Cited by 0SourcecodeScholar
2025

No Regret Reinforcement Learning Algorithms for Online Scheduling with Multi-Stage Tasks

IJCAI 2025

We study online task scheduling problems where tasks arrive sequentially and are processed by the platform or server. The service processes for tasks are multi-stage and are modeled as episodic Markov Decision Processes (MDPs). While processing a task, the system acquires rewards by consuming resour

Cited by 0SourcePDFScholar
2025

Triple-Optimistic Learning for Stochastic Contextual Bandits with General Constraints

ICML 2025poster

We study contextual bandits with general constraints, where a learner observes contexts and aims to maximize cumulative rewards while satisfying a wide range of general constraints. We introduce the Optimistic$^3$ framework, a novel learning and decision-making approach that integrates optimistic de…

Cited by 0SourcePDFScholar
2022

Online Convex Optimization with Hard Constraints: Towards the Best of Two Worlds and Beyond

NeurIPS 2022accept

This paper considers online convex optimization with hard constraints and analyzes achievable regret and cumulative hard constraint violation (violation for short). The problem distinguishes itself from online convex optimization with soft constraints, where a violation at one round can be compensat…

Cited by 40SourcePDFScholar