← Search

Kunhe Yang

11 accepted papers

2025

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

NeurIPS 2025poster

After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have…

Cited by 0SourceScholar
2024

Is Knowledge Power? On the (Im)possibility of Learning from Strategic Interactions

NeurIPS 2024poster

When learning in strategic environments, a key question is whether agents can overcome uncertainty about their preferences to achieve outcomes they could have achieved absent any uncertainty. Can they do this solely through interactions with each other? We focus this question on the ability of agent…

Cited by 4SourcePDFScholar
2024

Strategic Littlestone Dimension: Improved Bounds on Online Strategic Classification

NeurIPS 2024poster

We study the problem of online binary classification in settings where strategic agents can modify their observable features to receive a positive classification. We model the set of feasible manipulations by a directed graph over the feature space, and assume the learner only observes the manipulat…

Cited by 1SourcePDFScholar
2023

Calibrated Stackelberg Games: Learning Optimal Commitments Against Calibrated Agents

NeurIPS 2023spotlight

In this paper, we introduce a generalization of the standard Stackelberg Games (SGs) framework: _Calibrated Stackelberg Games_. In CSGs, a principal repeatedly interacts with an agent who (contrary to standard SGs) does not have direct access to the principal's action but instead best responds to _c…

Cited by 34SourcePDFScholar
2023

Optimal Conservative Offline RL with General Function Approximation via Augmented Lagrangian

ICLR 2023top-25%

Offline reinforcement learning (RL), which aims at learning good policies from historical data, has received significant attention over the past years. Much effort has focused on improving offline RL practicality by addressing the prevalent issue of partial data coverage through various forms of con…

Cited by 47SourcePDFScholar
2022

Oracle-Efficient Online Learning for Smoothed Adversaries

NeurIPS 2022accept

We study the design of computationally efficient online learning algorithms under smoothed analysis. In this setting, at every step, an adversary generates a sample from an adaptively chosen distribution whose density is upper bounded by $1/\sigma$ times the uniform density. Given access to an offli…

Cited by 14SourcePDFScholar