← Search

Yupei Yang

3 accepted papers

2026

Factored Causal Representation Learning for Robust Reward Modeling in RLHF

ICML 2026poster

A reliable reward model is essential for aligning large language models (LLMs) with human preferences through reinforcement learning from human feedback (RLHF). However, standard reward models are susceptible to spurious features that are not causally related to human labels. This can lead to *rewar…

Cited by 0SourceScholar
2025

Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations

ICLR 2025poster

General intelligence requires quick adaptation across tasks. While existing reinforcement learning (RL) methods have made progress in generalization, they typically assume only distribution changes between source and target domains. In this paper, we explore a wider range of scenarios where not only…

Cited by 1SourcePDFScholar
2024

Boosting Efficiency in Task-Agnostic Exploration through Causal Knowledge

IJCAI 2024poster

The effectiveness of model training heavily relies on the quality of available training resources. However, budget constraints often impose limitations on data collection efforts. To tackle this challenge, we introduce causal exploration in this paper, a strategy that leverages the underlying causal…