← Search

Wenjie Qiu

9 accepted papers

2026

Preference-based Policy Optimization from Sparse-reward Offline Dataset

ICLR 2026poster

Offline reinforcement learning (RL) holds the promise of training effective policies from static datasets without the need for costly online interactions. However, offline RL faces key limitations, most notably the challenge of generalizing to unseen or infrequently encountered state-action pairs. W…

Cited by 0SourceScholar
2025

Explainable Reinforcement Learning from Human Feedback to Improve Alignment

NeurIPS 2025poster

A common and effective strategy for humans to improve an unsatisfactory outcome in daily life is to find a cause of this outcome and correct the cause. In this paper, we investigate whether this human improvement strategy can be applied to improving reinforcement learning from human feedback (RLHF)…

Cited by 0SourceScholar
2025

MetaBox-v2: A Unified Benchmark Platform for Meta-Black-Box Optimization

NeurIPS 2025poster

Meta-Black-Box Optimization (MetaBBO) streamlines the automation of optimization algorithm design through meta-learning. It typically employs a bi-level structure: the meta-level policy undergoes meta-training to reduce the manual effort required in developing algorithms for low-level optimization t…

Cited by 0SourcecodeScholar
2025

Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation

ICLR 2025poster

Large Language Models (LLMs), trained on a large amount of corpus, have demonstrated remarkable abilities. However, it may not be sufficient to directly apply open-source LLMs like Llama to certain real-world scenarios, since most of them are trained for \emph{general} purposes. Thus, the demands fo…

2024

Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics

ICML 2024poster

Developing policies that can adapt to non-stationary environments is essential for real-world reinforcement learning applications. Nevertheless, learning such adaptable policies in offline settings, with only a limited set of pre-collected trajectories, presents significant challenges. A key difficu…

2023

Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic Objectives

NeurIPS 2023poster

Goal-conditioned reinforcement learning (RL) is a powerful approach for learning general-purpose skills by reaching diverse goals. However, it has limitations when it comes to task-conditioned policies, where goals are specified by temporally extended instructions written in the Linear Temporal Logi…

Cited by 13SourcePDFScholar
2021

T-IK: An Efficient Multi-Objective Evolutionary Algorithm for Analytical Inverse Kinematics of Redundant Manipulator

RA-L 2021

This letter proposes a new method combined by the parameterization method and T-IK to solve the inverse kinematics problem of redundant manipulators in the position domain. T-IK is an improved multi-objective optimization algorithm based on NSGA-II. By adding population migration strategy and adapti

Cited by 29SourceScholar