← Search

Andi Nika

6 accepted papers

2025

Corruption Robust Offline Reinforcement Learning with Human Feedback

AISTATS 2025oral

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilon$-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory f…

Cited by 0SourceScholar
2025

Policy Teaching via Data Poisoning in Learning from Human Preferences

AISTATS 2025poster

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy $\pi^\dagger$ by synthesizing preference data. We seek to understand the susceptibility of different preference-based learning paradigms to poisoned pr…

Cited by 0SourceScholar
2024

Corruption-Robust Offline Two-Player Zero-Sum Markov Games

AISTATS 2024poster

We study data corruption robustness in offline two-player zero-sum Markov games. Given a dataset of realized trajectories of two players, an adversary is allowed to modify an $\epsilon$-fraction of it. The learner’s goal is to identify an approximate Nash Equilibrium policy pair from the corrupted d…

Cited by 3SourcePDFScholar
2024

Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

ICML 2024poster

In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the cla…

Cited by 10SourcePDFScholar
2023

Online Defense Strategies for Reinforcement Learning Against Adaptive Reward Poisoning

AISTATS 2023poster

We consider the problem of defense against reward-poisoning attacks in reinforcement learning and formulate it as a game in $T$ rounds between a defender and an adaptive attacker in an adversarial environment. To address this problem, we design two novel defense algorithms. First, we propose Exp3-DA…

Cited by 4SourcePDFScholar
2020

Contextual Combinatorial Volatile Multi-armed Bandit with Adaptive Discretization

AISTATS 2020poster

We consider contextual combinatorial volatile multi-armed bandit (CCV-MAB), in which at each round, the learner observes a set of available base arms and their contexts, and then, selects a super arm that contains $K$ base arms in order to maximize its cumulative reward. Under the semi-bandit feedba…

Cited by 26SourcePDFScholar