← Search

Ilija Bogunovic

38 accepted papers

2026

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

ICML 2026poster

RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-task adaptation of GRPO often leads to imbalanced outcomes, with some tasks domin…

Cited by 0SourceScholar
2026

RSPO: Regularized Self-Play Alignment of Large Language Models

ICML 2026poster

Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. However, the regularization with respect to the reference policy, which is crucial for mitigating over-optimization, has be…

Cited by 0SourceScholar
2026

Robust Multi-Objective Controlled Decoding of Large Language Models

ICLR 2026poster

We introduce Robust Multi-Objective Decoding (RMOD), a novel inference-time algorithm that robustly aligns Large Language Models (LLMs) to multiple human objectives (e.g., instruction-following, helpfulness, safety) by maximizing the worst-case rewards. RMOD formulates the robust decoding problem as…

Cited by 0SourcecodeScholar
2026

wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models

ICLR 2026poster

Improving the reasoning capabilities of diffusion-based large language models (dLLMs) through reinforcement learning (RL) remains an open problem. The intractability of dLLMs likelihood function necessitates approximating the current, old, and reference policy likelihoods at each policy optimization…

Cited by 62SourceScholar
2025

Imagined Autocurricula

NeurIPS 2025poster

Training agents to act in embodied environments typically requires vast training data or access to accurate simulation, neither of which exists for many cases in the real world. Instead, world models are emerging as an alternative–leveraging offline, passively collected data, they make it possible t…

Cited by 0SourceScholar
2025

PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks

AAAI 2025technical

It is widely known that state-of-the-art machine learning models, including vision and language models, can be seriously compromised by adversarial perturbations. It is therefore increasingly relevant to develop capabilities to certify their performance in the presence of the most effective adversar…

Cited by 0SourcePDFScholar
2025

Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift

ICML 2025poster

Current Large Language Model (LLM) preference optimization algorithms do not account for temporal preference drift, which can lead to severe misalignment. To address this limitation, we propose **Non-Stationary Direct Preference Optimisation (NS-DPO)** that models time-dependent reward functions wit…

Cited by 0SourcePDFScholar
2024

Adversarially Robust Decision Transformer

NeurIPS 2024poster

Decision Transformer (DT), as one of the representative Reinforcement Learning via Supervised Learning (RvS) methods, has achieved strong performance in offline learning tasks by leveraging the powerful Transformer architecture for sequential decision-making. However, in adversarial environments, th…

2024

Distributionally Robust Model-based Reinforcement Learning with Large State Spaces

AISTATS 2024poster

Three major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment. To overcome these issues, we study distributionally robust Markov deci…

2024

Group Robust Preference Optimization in Reward-free RLHF

NeurIPS 2024poster

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional…

Cited by 20SourcePDFScholar
2024

REDUCR: Robust Data Downsampling using Class Priority Reweighting

NeurIPS 2024poster

Modern machine learning models are becoming increasingly expensive to train for real-world image and text classification tasks, where massive web-scale data is collected in a streaming fashion. To reduce the training cost, online batch selection techniques have been developed to choose the most info…

2024

Sample-efficient Bayesian Optimisation Using Known Invariances

NeurIPS 2024poster

Bayesian optimisation (BO) is a powerful framework for global optimisation of costly functions, using predictions from Gaussian process models (GPs). In this work, we apply BO to functions that exhibit invariance to a known group of transformations. We show that vanilla and constrained BO algorithm…

Cited by 2SourcePDFScholar
2023

Efficient Planning in Combinatorial Action Spaces with Applications to Cooperative Multi-Agent Reinforcement Learning

AISTATS 2023poster

A practical challenge in reinforcement learning are combinatorial action spaces that make planning computationally demanding. For example, in cooperative multi-agent reinforcement learning, a potentially large number of agents jointly optimize a global reward function, which leads to a combinatorial…

Cited by 5SourcePDFScholar
2023

Near-optimal Policy Identification in Active Reinforcement Learning

ICLR 2023top-5%

Many real-world reinforcement learning tasks require control of complex dynamical systems that involve both costly data acquisition processes and large state spaces. In cases where the expensive transition dynamics can be readily evaluated at specified states (e.g., via a simulator), agents can oper…

Cited by 8SourcePDFScholar
2022

A Robust Phased Elimination Algorithm for Corruption-Tolerant Gaussian Process Bandits

NeurIPS 2022accept

We consider the sequential optimization of an unknown, continuous, and expensive to evaluate reward function, from noisy and adversarially corrupted observed rewards. When the corruption attacks are subject to a suitable budget $C$ and the function lives in a Reproducing Kernel Hilbert Space (RKHS),…

Cited by 11SourcePDFScholar
2022

Movement Penalized Bayesian Optimization with Application to Wind Energy Systems

NeurIPS 2022accept

Contextual Bayesian optimization (CBO) is a powerful framework for sequential decision-making given side information, with important applications, e.g., in wind energy systems. In this setting, the learner receives context (e.g., weather conditions) at each round, and has to choose an action (e.g.,…

Cited by 13SourcePDFScholar
2021

Combining Pessimism with Optimism for Robust and Efficient Model-Based Deep Reinforcement Learning

ICML 2021spotlight

In real-world tasks, reinforcement learning (RL) agents frequently encounter situations that are not present during training time. To ensure reliable performance, the RL agents need to exhibit robustness to such worst-case situations. The robust-RL framework addresses this challenge via a minimax op…

Cited by 16SourcePDFScholar
2021

Online Submodular Resource Allocation with Applications to Rebalancing Shared Mobility Systems

ICML 2021spotlight

Motivated by applications in shared mobility, we address the problem of allocating a group of agents to a set of resources to maximize a cumulative welfare objective. We model the welfare obtainable from each resource as a monotone DR-submodular function which is a-priori unknown and can only be lea…

Cited by 3SourcePDFScholar
2021

Risk-averse Heteroscedastic Bayesian Optimization

NeurIPS 2021poster

Many black-box optimization tasks arising in high-stakes applications require risk-averse decisions. The standard Bayesian optimization (BO) paradigm, however, optimizes the expected value only. We generalize BO to trade mean and input-dependent variance of the objective, both of which we assume to…

2021

Stochastic Linear Bandits Robust to Adversarial Attacks

AISTATS 2021poster

We consider a stochastic linear bandit problem in which the rewards are not only subject to random noise, but also adversarial attacks subject to a suitable budget $C$ (i.e., an upper bound on the sum of corruption magnitudes across the time horizon). We provide two variants of a Robust Phased Elimi…

Cited by 91SourcePDFScholar
2020

Contextual Games: Multi-Agent Learning with Side Information

NeurIPS 2020poster

We formulate the novel class of contextual games, a type of repeated games driven by contextual information at each round. By means of kernel-based regularity assumptions, we model the correlation between different contexts and game outcomes and propose a novel online (meta) algorithm that exploits…

Cited by 23SourcePDFScholar
2020

Distributionally Robust Bayesian Optimization

AISTATS 2020poster

Robustness to distributional shift is one of the key challenges of contemporary machine learning. Attaining such robustness is the goal of distributionally robust optimization, which seeks a solution to an optimization problem that is worst-case robust under a specified distributional shift of an un…

Cited by 106SourcePDFScholar
2020

Learning to Play Sequential Games versus Unknown Opponents

NeurIPS 2020poster

We consider a repeated sequential game between a learner, who plays first, and an opponent who responds to the chosen action. We seek to design strategies for the learner to successfully interact with the opponent. While most previous approaches consider known opponent models, we focus on the settin…

Cited by 27SourcePDFScholar
2020

Mixed Strategies for Robust Optimization of Unknown Objectives

AISTATS 2020poster

We consider robust optimization problems, where the goal is to optimize an unknown objective function against the worst-case realization of an uncertain parameter. For this setting, we design a novel sample-efficient algorithm GP-MRO, which sequentially learns about the unknown objective from noisy…

Cited by 18SourcePDFScholar
2019

No-Regret Learning in Unknown Games with Correlated Payoffs

NeurIPS 2019poster

We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback…

Cited by 48SourcePDFScholar
2018

Adversarially Robust Optimization with Gaussian Processes

NeurIPS 2018spotlight

In this paper, we consider the problem of Gaussian process (GP) optimization with an added robustness requirement: The returned point may be perturbed by an adversary, and we require the function value to remain as high as possible even after this perturbation. This problem is motivated by settings…

2018

High-Dimensional Bayesian Optimization via Additive Models with Overlapping Groups

AISTATS 2018poster

Bayesian optimization (BO) is a popular technique for sequential black-box function optimization, with applications including parameter tuning, robotics, environmental monitoring, and more. One of the most important challenges in BO is the development of algorithms that scale to high dimensions, wh…

Cited by 0SourcePDFScholar
2017

Robust Submodular Maximization: A Non-Uniform Partitioning Approach

ICML 2017poster

We study the problem of maximizing a monotone submodular function subject to a cardinality constraint $k$, with the added twist that a number of items $\tau$ from the returned set may be removed. We focus on the worst-case setting considered by Orlin et al.\ (2016), in which a constant-factor approx…

Cited by 77SourcePDFScholar
2017

Streaming Robust Submodular Maximization: A Partitioned Thresholding Approach

NeurIPS 2017poster

We study the classical problem of maximizing a monotone submodular function subject to a cardinality constraint k, with two additional twists: (i) elements arrive in a streaming fashion, and (ii) m items from the algorithm’s memory are removed after the stream is finished. We develop a robust submod…

Cited by 63SourcePDFScholar
2016

An Efficient Streaming Algorithm for the Submodular Cover Problem

NeurIPS 2016poster

We initiate the study of the classical Submodular Cover (SC) problem in the data streaming model which we refer to as the Streaming Submodular Cover (SSC). We show that any single pass streaming algorithm using sublinear memory in the size of the stream will fail to provide any non-trivial approxima…

Cited by 27SourcePDFScholar
2016

Truncated Variance Reduction: A Unified Approach to Bayesian Optimization and Level-Set Estimation

NeurIPS 2016poster

We present a new algorithm, truncated variance reduction (TruVaR), that treats Bayesian optimization (BO) and level-set estimation (LSE) with Gaussian processes in a unified fashion. The algorithm greedily shrinks a sum of truncated variances within a set of potential maximizers (BO) or unclassified…

2015

Active learning of self-concordant like multi-index functions

ICASSP 2015accepted

We study the problem of actively learning a multi-index function of the form f(x) = g <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> (A <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> x) fr…

Cited by 0SourceScholar