← Search

Ziyu Shao

5 accepted papers

2026

Towards Achieving Optimal Strong Regret and Constraint Violation via Computational Efficient Model-free RL

ICML 2026poster

We study episodic constrained Markov decision processes (CMDPs) with linear function approximation, where the goal is to achieve strong regret and constraint violation guarantees without allowing error cancellations. Unlike the existing work, which focuses on either tabular CMDP or model-based reinf…

Cited by 0SourceScholar
2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

NeurIPS 2025poster

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov D…

Cited by 0SourcecodeScholar
2025

No Regret Reinforcement Learning Algorithms for Online Scheduling with Multi-Stage Tasks

IJCAI 2025

We study online task scheduling problems where tasks arrive sequentially and are processed by the platform or server. The service processes for tasks are multi-stage and are modeled as episodic Markov Decision Processes (MDPs). While processing a task, the system acquires rewards by consuming resour

Cited by 0SourcePDFScholar
2022

Anisotropic Fourier Features for Neural Image-Based Rendering and Relighting

AAAI 2022technical

Recent neural rendering techniques have greatly benefited image-based modeling and relighting tasks. They provide a continuous, compact, and parallelable representation by modeling the plenoptic function as multilayer perceptrons (MLPs). However, vanilla MLPs suffer from spectral biases on multidime…

Cited by 7SourcePDFScholar
2020

Learning-Aided Content Placement in Caching-Enabled fog Computing Systems Using Thompson Sampling

ICASSP 2020accepted

In this paper, we focus on the problem of online content placement with unknown content popularity in caching-enabled fog computing systems, i.e., how to decide and update cached content on resourcelimited edge fog nodes to maximize cache hit rate and minimize switching costs of content update. Face…

Cited by 0SourceScholar