← Search

Han Shen

12 accepted papers

2026

On Entropy Control in LLM-RL Algorithms

ICLR 2026poster

For RL algorithms, appropriate entropy control is crucial to their effectiveness. To control the policy entropy, a commonly used method is entropy regularization, which is adopted in various popular RL algorithms including PPO, SAC and A3C. Although entropy regularization proves effective in robotic…

Cited by 0SourceScholar
2026

REVIS: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models

ICML 2026poster

Despite the advanced capabilities of Large Vision-Language Models (LVLMs), they frequently suffer from object hallucination. One reason is that visual features and pretrained textual representations often become intertwined in the deeper network layers. To address this, we propose REVIS, a training-…

Cited by 0SourceScholar
2025

SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

ICLR 2025poster

Fine-tuning on task-specific data to boost downstream performance is a crucial step for leveraging Large Language Models (LLMs). However, though fine-tuning enhances the model performance for specialized applications, previous studies have demonstrated that fine-tuning the models on several adversar…

2024

A Method for Bilevel Optimization with Convex Lower-Level Problem

ICASSP 2024accepted

Gradient-based bilevel optimization methods have been applied to a wide range of applications including hyper-parameter optimization, meta-learning, and model pruning. However, it is known that the bilevel optimization problem is difficult to solve, and the finite-time guarantee has only been establ…

Cited by 0SourceScholar
2024

Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization

ICASSP 2024accepted

In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term bi-level joint unsupervised and supervised training (BL-JUST). BL-JUST employs a lower and upper level optimization with an unsupervise…

Cited by 0SourceScholar
2024

Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF

ICML 2024poster

Bilevel optimization has been recently applied to many machine learning tasks. However, their applications have been restricted to the supervised learning setting, where static objective functions with benign structures are considered. But bilevel problems such as incentive design, inverse reinforce…

Cited by 13SourcePDFScholar
2023

Alternating Projected SGD for Equality-constrained Bilevel Optimization

AISTATS 2023poster

Bilevel optimization, which captures the inherent nested structure of machine learning problems, is gaining popularity in many recent applications. Existing works on bilevel optimization mostly consider either the unconstrained problems or the constrained upper-level problems. In this context, this…

2023

Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Approach

ICLR 2023top-5%

Many machine learning problems today have multiple objective functions. They appear either in learning with multiple criteria where learning has to make a trade-off between multiple performance metrics such as fairness, safety and accuracy; or, in multi-task learning where multiple tasks are optimiz…

Cited by 55SourcePDFScholar
2022

A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled Sequences

NeurIPS 2022accept

Stochastic approximation (SA) with multiple coupled sequences has found broad applications in machine learning such as bilevel learning and reinforcement learning (RL). In this paper, we study the finite-time convergence of nonlinear SA with multiple coupled sequences. Different from existing multi-…

Cited by 19SourcePDFScholar
2021

Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation

ICASSP 2021accepted

This paper considers the policy evaluation problem in reinforcement learning with agents of a decentralized and directed network. The focus is on decentralized temporal-difference (TD) learning with linear function approximation in the presence of unreliable or even malicious agents, termed as Byzan…

Cited by 0SourceScholar