← Search

Goran Radanovic

29 accepted papers

2026

Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks

ICML 2026poster

We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of reward poisoning attacks targeting the DPT, we propose a novel adversarial training framework, called Adversarially Tra…

Cited by 0SourceScholar
2025

Corruption Robust Offline Reinforcement Learning with Human Feedback

AISTATS 2025oral

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilon$-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory f…

Cited by 0SourceScholar
2025

Counterfactual Effect Decomposition in Multi-Agent Sequential Decision Making

ICML 2025poster

We address the challenge of explaining counterfactual outcomes in multi-agent Markov decision processes. In particular, we aim to explain the total counterfactual effect of an agent's action on the outcome of a realized scenario through its influence on the environment dynamics and the agents' behav…

2025

Independent Learning in Performative Markov Potential Games

AISTATS 2025poster

Performative Reinforcement Learning (PRL) refers to a scenario in which the deployed policy changes the reward and transition dynamics of the underlying environment. In this work, we study multi-agent PRL by incorporating performative effects into Markov Potential Games (MPGs). We introduce the not…

Cited by 0SourcecodeScholar
2025

On Corruption-Robustness in Performative Reinforcement Learning

AAAI 2025technical

In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining approaches to a performatively stable policy. In the finite s…

Cited by 1SourcePDFScholar
2025

Policy Teaching via Data Poisoning in Learning from Human Preferences

AISTATS 2025poster

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy $\pi^\dagger$ by synthesizing preference data. We seek to understand the susceptibility of different preference-based learning paradigms to poisoned pr…

Cited by 0SourceScholar
2025

Stochastic Principal-Agent Problems: Computing and Learning Optimal History-Dependent Policies

NeurIPS 2025poster

We study a stochastic principal-agent model. A principal and an agent interact in a stochastic environment, each privy to observations about the state not available to the other. The principal has the power of commitment, both to elicit information from the agent and to signal her own information. T…

Cited by 0SourceScholar
2025

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints

AAAI 2025technical

Recent work has proposed automated red-teaming methods for testing the vulnerabilities of a given target large language model (LLM). These methods use red-teaming LLMs to uncover inputs that induce harmful behavior in a target LLM. In this paper, we study red-teaming strategies that enable a targe…

Cited by 7SourcePDFScholar
2024

Agent-Specific Effects: A Causal Effect Propagation Analysis in Multi-Agent MDPs

ICML 2024poster

Establishing causal relationships between actions and outcomes is fundamental for accountable multi-agent decision-making. However, interpreting and quantifying agents' contributions to such relationships pose significant challenges. These challenges are particularly prominent in the context of mult…

2024

Corruption-Robust Offline Two-Player Zero-Sum Markov Games

AISTATS 2024poster

We study data corruption robustness in offline two-player zero-sum Markov games. Given a dataset of realized trajectories of two players, an adversary is allowed to modify an $\epsilon$-fraction of it. The learner’s goal is to identify an approximate Nash Equilibrium policy pair from the corrupted d…

Cited by 3SourcePDFScholar
2024

Learning Embeddings for Sequential Tasks Using Population of Agents

IJCAI 2024poster

We present an information-theoretic framework to learn fixed-dimensional embeddings for tasks in reinforcement learning. We leverage the idea that two tasks are similar if observing an agent's performance on one task reduces our uncertainty about its performance on the other. This intuition is captu…

Cited by 0SourcePDFScholar
2024

Performative Reinforcement Learning in Gradually Shifting Environments

UAI 2024poster

When Reinforcement Learning (RL) agents are deployed in practice, they might impact their environment and change its dynamics. We propose a new framework to model this phenomenon, where the current environment depends on the deployed policy as well as its previous dynamics. This is a generalization…

Cited by 6SourcePDFScholar
2024

Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

ICML 2024poster

In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the cla…

Cited by 10SourcePDFScholar
2023

Markov Decision Processes with Time-Varying Geometric Discounting

AAAI 2023technical

Canonical models of Markov decision processes (MDPs) usually consider geometric discounting based on a constant discount factor. While this standard modeling approach has led to many elegant results, some recent studies indicate the necessity of modeling time-varying discounting in certain applicati…

Cited by 2SourcePDFScholar
2023

Online Defense Strategies for Reinforcement Learning Against Adaptive Reward Poisoning

AISTATS 2023poster

We consider the problem of defense against reward-poisoning attacks in reinforcement learning and formulate it as a game in $T$ rounds between a defender and an adaptive attacker in an adversarial environment. To address this problem, we design two novel defense algorithms. First, we propose Exp3-DA…

Cited by 4SourcePDFScholar
2023

Online Reinforcement Learning with Uncertain Episode Lengths

AAAI 2023technical

Existing episodic reinforcement algorithms assume that the length of an episode is fixed across time and known a priori. In this paper, we consider a general framework of episodic reinforcement learning when the length of each episode is drawn from a distribution. We first establish that this prob…

Cited by 7SourcePDFScholar
2022

Admissible Policy Teaching through Reward Design

AAAI 2022technical

We study reward design strategies for incentivizing a reinforcement learning agent to adopt a policy from a set of admissible policies. The goal of the reward designer is to modify the underlying reward function cost-efficiently while ensuring that any approximately optimal deterministic policy unde…

Cited by 17SourcePDFScholar
2021

Explicable Reward Design for Reinforcement Learning Agents

NeurIPS 2021poster

We study the design of explicable reward functions for a reinforcement learning agent while guaranteeing that an optimal policy induced by the function belongs to a set of target policies. By being explicable, we seek to capture two properties: (a) informativeness so that the rewards speed up the ag…

2021

On Blame Attribution for Accountable Multi-Agent Sequential Decision Making

NeurIPS 2021poster

Blame attribution is one of the key aspects of accountable decision making, as it provides means to quantify the responsibility of an agent for a decision making outcome. In this paper, we study blame attribution in the context of cooperative multi-agent sequential decision making. As a particular s…

Cited by 10SourcePDFScholar
2020

Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning

ICML 2020poster

We study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes average reward in undiscounted infin…

2017

Multi-View Decision Processes: The Helper-AI Problem

NeurIPS 2017poster

We consider a two-player sequential game in which agents have the same reward function but may disagree on the transition probabilities of an underlying Markovian model of the world. By committing to play a specific policy, the agent with the correct model can steer the behavior of the other agent,…

Cited by 32SourcePDFScholar